Workshops
Interpretability for Discovery: Understanding and Discovering Novel Knowledge in AI Models
Xiaoyan Bai ⋅ Yonatan Belinkov ⋅ Ekdeep S Lubana ⋅ Yaniv Nikankin ⋅ Chenhao Tan ⋅ Amirtha Varshini Anbuchezhiyan Sindhanai
We propose a workshop on **Interpretability for Discovery** to explore how model interpretability can support the discovery and understanding of novel knowledge in AI systems. The workshop aims to build a community around three core questions: (i) How can existing interpretability methods be adapted to handle architectures, biases, and data modalities that differ from the LLM setup? (ii) How can interpretable model representations be translated into novel knowledge and discoveries? (iii) Which open questions can model interpretability help answer? Our goal is to bring together researchers from machine learning and related disciplines to foster discussion on methodologies, experimental design, and real-world applications so that we can enable ground-breaking discoveries from insights encoded in AI systems.
Show more
AI for Verifiable Coding: Human-Aligned Collaborative Agents for Autoformalization, Proofs, and Heuristics
Wuyang Chen ⋅ Soonho Kong ⋅ Hakjoo Oh ⋅ Jingxuan He ⋅ Jacqueline Mitchell ⋅ Amanda Liu ⋅ Zhe Ye ⋅ Xiaodong Liu ⋅ Varun Pant ⋅ Jialin Lu ⋅ Simon Frieder
LLM code assistants are powerful but untrustworthy: they hallucinate, mishandle corner cases, and lack correctness guarantees. This workshop advances AI for verifiable coding, where human-aligned agents collaborate with proof assistants, model checkers, and analyzers to co-develop specifications, code, proofs, and heuristics with machine-checkable guarantees toward provably safe software. Unlike generic code LLMs, verifiable generation requires expressive specifications and structured proofs where models draft artifacts and formal tools refute, repair, or certify them. Agents must also autoformalize informal requirements/mathematics, search proofs for complex conditions, and discover invariants, lemmas, and solver strategies. Our workshop focuses on five unique perspectives. (1) AI-assisted testing/debugging (e.g., fuzzing, debugging by speaker Shan Lu); (2) Agentic verification for security/smart contracts (speakers Dawn Song; Ilya Sergey); (3) Human-in-the-loop spec generation (speaker Emily First); (4) Verification of AI systems (robustness/safety by speaker Vijay Ganesh); (5) Scalability and automation (speakers Leonardo de Moura; Tudor Achim). We further cultivate an inclusive, interdisciplinary community across LLM and programming languages, with diverse speakers, mentorship, and networking for long-term impact.
Show more
NeurIPS 2026 Workshop on SaTQuML: Secure and Trustworthy Quantum Machine Learning
Mohammad Saidur Rahman ⋅ Jakub Szefer ⋅ Muhammad T Raza ⋅ Khoa Luu ⋅ Shahrooz Pouryousef
Quantum machine learning (QML) is moving from theoretical promise toward practical experimentation through hybrid quantum-classical platforms, cloud-accessible quantum computers, and early application-driven demonstrations. This transition creates an important need to examine whether QML systems can be secure, trustworthy, and useful in high-impact settings. Our proposed NeurIPS 2026 workshop, SaTQuML: Secure and Trustworthy Quantum Machine Learning, will bring together researchers from QML, trustworthy AI, cybersecurity, and quantum security to address this need. To the best of our knowledge, it is the first workshop at a major machine learning venue centered on the security, trustworthiness, and rigorous evaluation of QML systems. The workshop will focus on two complementary themes--understanding the security and reliability of QML systems themselves, and exploring QML and hybrid quantum-classical methods for cyberdefense and related security applications. By emphasizing realistic benchmarks, strong baselines, clear threat models, reproducibility, and careful assessment of QML's practical value and limitations, SaTQuML aims to build a shared research agenda for trustworthy quantum AI. The broader goal is to help shape a global community around QML systems that can be evaluated responsibly and deployed safely in security-critical environments.
Show more
World Models for High-Stakes Health: Reliable Clinical Trial Simulation and Intervention-Aware Reasoning
Jay Nanavati ⋅ Shalmali Joshi ⋅ Rahul Krishnan ⋅ Lin Li ⋅ Katherine E Link ⋅ Emma Slade
Clinical trial simulation is a uniquely demanding and falsifiable testbed for reliable world models in high-stakes health. This workshop will bring together researchers across machine learning, healthcare AI, causal inference, uncertainty quantification, AI for science, reasoning systems, clinical development, and pharma/RWE to advance **patient world models**: generative and reasoning-capable systems that learn from patient trajectories, trial protocols, interventions, mechanisms of action, and real-world evidence. The workshop will focus on how such models can simulate clinical trajectories, reason about interventions, quantify uncertainty, and generate evidence credible enough to inform clinical trial design, execution, and translation into care. Topics include clinical trial simulation, virtual trial arms, synthetic controls, target trial emulation, off-policy evaluation, EHR and medical foundation models, multimodal patient modelling, calibration, robustness, protocol-aware reasoning, agentic systems, and validation against real-world evidence and clinical knowledge. The workshop aims to define shared assumptions, benchmark needs, validation protocols, and reliability criteria for patient world models in clinical trial simulation and high-stakes health.
Show more
NeurIPS’26 Workshop on AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI
Denghui Zhang ⋅ Jianing Zhu ⋅ Sarvapali Ramchurn ⋅ Manling Li ⋅ Zhangyang "Atlas" Wang
As LLMs become co-authors, co-reviewers, and co-citers, the academic publishing system is undergoing a structural rewrite that current venue policies were not designed to absorb. ICLR 2026, NeurIPS 2025, and ICML 2026 have each shipped ad-hoc detection, policy, and mechanism-design responses, but no shared formulation yet exists for the underlying failure: a human-AI co-hallucination loop in which false or unsupported scholarly claims that no model and no human would produce alone arise through their interaction and become institutionalized through trust, time pressure, and publication incentives. This workshop convenes six invited speakers, each anchored to a concrete publication-pipeline issue: James Zou (Stanford), Kyunghyun Cho (NYU), Hiromu Yakura (MPI / Anthropic), Yian Yin (Cornell Tech), Hima Lakkaraju (Harvard / Google), and Lin Peng (Baruch / CUNY). An All-PC-Chair Panel brings together past, current, and incoming PC Chairs at NeurIPS, ICLR, ICML, CVPR, MICCAI, and MIDL, moderated by Atlas Wang (UT Austin / XTX Markets). Deliverables: a community position document on operational 2027 policy.
Show more
AI Agents for Biomedical Imaging and Multimodal Clinical Data
Ehsan Adeli ⋅ Tal Arbel ⋅ Adrian Dalca ⋅ Klaus Maier-Hein ⋅ Curtis Langlotz ⋅ Mert Sabuncu
Biomedical imaging AI has produced strong methods for segmentation, registration, reconstruction, detection, and report generation, yet most systems remain organized around fixed inputs and outputs (a structure that fails to capture the complexity of real clinical workflows, where images must be integrated with reports, laboratory values, EHR data, waveforms, and longitudinal patient context). This workshop focuses on AI agents for biomedical imaging and multimodal clinical data, bringing together researchers in machine learning, computer vision, biomedical imaging, and clinical AI to define shared task definitions, evaluation protocols, and benchmarks for agentic image analysis systems. Topics include tool-using agents for image measurement and annotation, multimodal grounding across imaging and clinical data, human-in-the-loop workflows, and reproducible evaluation frameworks. The workshop is paired with a special issue of the Machine Learning for Biomedical Imaging (MELBA) journal, in which authors of accepted workshop papers are invited to submit extended manuscript versions for independent peer review, with an archival pathway coordinated by organizers who hold editorial roles at MELBA.
Show more
EconML: Economics for Machine Learning
Safwan Hossain ⋅ Meena Jagadeesan ⋅ Eric Mazumdar ⋅ Ariel Procaccia ⋅ Eden Saig ⋅ Kunhe Yang
As machine learning becomes deeply embedded in society, models increasingly interact with strategic incentives, competitive forces, and resource constraints - challenges that economic theory is well-suited to address. This workshop brings together researchers from machine learning, economics, and game theory to examine how economic forces shape the interplay between learning algorithms and the ecosystems that surround them. The program is organized around two complementary themes: (1) the use of economic tools to improve model training, evaluation, and alignment in strategic and competitive environments; and (2) understanding and steering the emergent dynamics that arise when many models interact in shared environments. By bridging micro-scale mechanism design and market-level analysis, the workshop aims to develop economic interventions that help ML ecosystems avoid foreseeable failures and better serve individuals, organizations, and society at large.
Show more
The Third Workshop on Long-Context Foundation Models
Zexue He ⋅ Howard Yen ⋅ Amanda Bertsch ⋅ Seungju Han ⋅ Alex Pentland ⋅ Danqi Chen ⋅ Yejin Choi
Foundation models have become a cornerstone in the advancement of artificial intelligence, widely used across both academic and practical applications. Across domains, challenging tasks require synthesizing information enormous amounts of data. These may take many forms, such as images, text, audio, and genomes. Much recent work has focused on developing long-context models capable of processing, understanding, and generating responses based on extensive inputs. However, complex tasks often require model to reason, plan, and interact with environments over extended horizons. This workshop will convene researchers to explore these challenges and foster developments in long-context foundation models. Key topics include new modeling architectures, training approaches, efficiency techniques, comprehensive evaluation methods, and applications in scientific fields such as genomics, climate science, scientific discovery, etc. Additionally, in this edition, special attention will be given to long-context reasoning and long-horizon agentic usages. By tackling these critical challenges, we aim to push the boundaries of long-context modeling and shape its future directions.
Show more
Dynamics at the Frontiers of Optimization, Sampling, and Games
Zhiyu He ⋅ Michael Jordan ⋅ Jiaming Liang ⋅ Wenlong Mou ⋅ Michael Muehlebach ⋅ Purnamrita Sarkar ⋅ Molei Tao ⋅ Andre Wibisono
Dynamical systems have played an important role in the analysis and design of algorithms. Ideas ranging from variational methods, differential and symplectic geometry, numerical analysis, and control theory have paved the way for establishing non-asymptotic convergence guarantees in optimization, sampling, and equilibrium computation in games. Yet, the distinct mathematical backbone of these tools often creates barriers to entry for researchers and practitioners in machine learning. This workshop aims to lower that barrier by highlighting the unifying role of dynamical systems across these domains. We will convene optimization, sampling, and game theory experts to foster cross-disciplinary dialogue and collaboration. Emphasis will be placed on emerging applications in machine learning, such as diffusion models, distributed and adversarial training, and agentic AI, where dynamical systems perspectives are increasingly central. Through a combination of talks, posters, and open discussions, we hope to catalyze new collaborations and broaden the accessibility of these foundational methods.
Show more
AI and Science: Evolution or Extinction?
Nathan Suri ⋅ Savannah Thais ⋅ Lauren Greenspan ⋅ Roberto Trotta ⋅ Max Hennick
Recent years have shown an explosion of interest for supporting, accelerating, and automating scientific discovery via AI systems. Both academic and industry AI researchers have leapt to the wellspring of challenging computational problems currently unsolved by the scientific community as a way to test the state of the art. This broad appeal has translated to a plethora of interdisciplinary collaborations between AI researchers and traditional scientists across multiple fields of science, all aiming to highlight the potential of AI models to contribute to scientific research. In order to properly evaluate how AI systems can impact the practice of science, we first must agree upon consistent definitions of success that outline how this type of integration can occur safely. Our workshop will gather researchers from statistics, philosophy of science, sociology of science, science and technology studies, psychology, and anthropology, alongside AI researchers working on AI for science, interpretability, AI safety, and agent evaluations, to address three questions: 1) What epistemic values constitute scientific integrity in the era of human-AI collaboration? 2) How can we build evaluations that measure whether AI systems uphold these values in practice? 3) What sociotechnical guardrails can sustain robust human-AI scientific collaboration without eroding the integrity of scientific knowledge production? Answering any of these requires expertise that no single community currently holds. Our workshop aims to both initiate a much-needed conversation for the future of the scientific and AI communities as well as foster a community of like-minded researchers committed to addressing these questions long-term in an interdisciplinary fashion.
Show more
MATH-AI: The 6th Workshop on Mathematical Reasoning and AI
Peiyang Song ⋅ Kaiyu Yang ⋅ Patrick Shafto ⋅ Sanjeev Arora ⋅ Katie Collins ⋅ Sean Welleck ⋅ Mateja Jamnik
The 6th MATH-AI Workshop on Mathematical Reasoning and AI will focus on the intersection of agentic AI and mathematical reasoning, with an emphasis on systems that can participate in the broader mathematical research loop. Recent progress in AI-assisted theorem proving, autoformalization, mathematical discovery, and research-level reasoning suggests a shift from isolated problem solving toward agents that can conjecture, search, formalize, prove, verify, explain, use tools, and learn from feedback. The workshop will bring together researchers from machine learning, mathematics, formal methods, programming languages, cognitive science, and education to discuss both the technical foundations of reliable mathematical agents and the human-AI workflows needed for productive collaboration. Its central question is: how can agentic AI systems advance mathematical research while remaining trustworthy partners for human mathematicians?
Show more
Can We Trust the Judge? Building Reliable Evaluation for Language Models
Jayash Koshal ⋅ Shanu Sushmita ⋅ Meghana Makhija ⋅ Hui Wan ⋅ Amjad A Jbara
LLM-as-a-Judge systems are now embedded in training pipelines, safety evaluation, and production deployments at scale—yet the field lacks a shared framework for assessing whether these evaluators actually measure what they intend to, in the contexts where they are deployed. JUDGe addresses this gap by reframing evaluation validity as a systems problem rather than a measurement problem: how does a judge's error profile interact with what is upstream and downstream of it, and what are the consequences when it fails in context? We bring together NLP researchers, ML systems practitioners, alignment scientists, and industry evaluators around a shared taxonomy of judge failure modes—surface sensitivity, sycophancy, criteria drift, positional bias, reasoning chain validity, safety-relevant meaning preservation, and inter-judge correlation—and their cascading interactions in live pipelines. The workshop will produce a community-facing Judge Deployment Disclosure Template, analogous to a model card but for evaluation infrastructure. Through contributed papers, junior spotlights, structured debate, and a practitioner panel, JUDGe will crystallize open problems and lay groundwork for evaluation standards the field currently lacks.
Show more
Child Safety in AI
Aashiq Muhamed ⋅ Rebecca S Portnoff ⋅ Virginia Smith ⋅ Andrew Strait ⋅ Vinith Suriyakumar ⋅ Ashia Wilson ⋅ Campbell Wilson
Modern AI tools introduce new risks to children, including mental health concerns related to risks of unhealthy attachment and interactions that could lead to self-harm and suicidal ideation, as well as the ability to create and misuse synthetic content such as sexual deepfake images or videos that can facilitate harassment, grooming, and extortion. At the same time, child safety imposes unique constraints on traditional safety approaches. Unlike many other AI safety domains, child safety often prohibits direct access to harmful data, limits evaluation on real-world examples, and requires collaboration with specialized stakeholders such as NGOs, hotlines, law enforcement, and child-protection experts. These constraints create new scientific challenges that require dedicated research methodologies. This workshop will consider technical and sociotechnical solutions across the AI lifecycle, including topics such as safe data curation, reliable system safeguards, robust open-weight model design, adversarially resilient deployment, and effective long-term monitoring. By emphasizing open technical problems, such as evaluating safety without access to harmful data, preventing harmful capability emergence, and designing safeguards under adversarial pressure (all while taking into account legal and ethical constraints related to children) the workshop aims to catalyze a research agenda that treats child safety as a core, safety-critical dimension of AI.
Show more
Workshop on Towards Test-Time Continual Learning Agents
Zheyuan Zhang ⋅ Chuanyang Jin ⋅ Jacob Sansom ⋅ Zekun Wang ⋅ Jianwen Xie ⋅ Joyce Chai ⋅ Daniel Khashabi ⋅ Tianmin Shu
Today's frontier models are frozen at deployment: once costly pre- and post-training ends, their knowledge, skills, and reasoning are effectively fixed, and any apparent adaptation comes from prompting, retrieval, memory, or external tools rather than genuine internal learning. Humans do the opposite, continuously acquiring knowledge, refining representations, and reorganizing beliefs through interaction. Today's agents cannot: they fail to internalize new information after deployment, do not improve from repeated mistakes, and erase prior skills when updated naively, while even state-of-the-art robotic systems assume a train/deploy split untenable in dynamic, partially observed, long-horizon, socially situated environments. We define Test-Time Continual Learning Agents as systems that continuously acquire, consolidate, and refine knowledge and capabilities during deployment, without catastrophic forgetting or repeated large-scale retraining. Such agents sit at a triple intersection studied today in isolation: continual learning targets mitigating forgetting and enabling knowledge transfer in a supervised setting, test-time adaptation handles distribution shift, LLM work emphasizes retrieval and fine-tuning, and embodied and agentic research prioritizes planning and tool use, none of which centers how a deployed agent should learn sample-efficiently through interaction, how to evaluate it, or how to integrate new skills safely. TTCL 2026 is the first workshop to place this test-time, continual, and agentic intersection at its center, convening these siloed communities around shared terminology, rigorous long-horizon benchmarks, and a hands-on challenge. We aim to catalyze next-generation cognitive agents that learn continually, consolidate experience, and remain reliable, rethinking the boundaries between training and inference, and between memory and learning.
Show more
Workshop on Resource-Aware Agentic AI
Kai-Wei Chang ⋅ Jiri Gesi ⋅ Yuxuan Lu ⋅ Weijia Shi ⋅ Dakuo Wang ⋅ Luke Zettlemoyer ⋅ Jiaxin Pei
Today’s agents can write and debug code, solve competition-level mathematical problems, and conduct deep research, and are increasingly deployed in everyday products and workflows. However, most of these agents are designed to maximize task performance. They are largely unaware of the resources they consume, including the compute, energy, memory, latency, tool cost, and data required to train them. The Workshop on Resource-Aware Agentic AI brings together academic and industrial experts to explore how to design agents that are both efficient and effective, from training and planning to systems and deployment. Resource awareness is not merely a technical challenge. It is essential to make AI agents accessible to all, not only those with large compute budgets.
Show more
DevAI: Developmental Perspectives on AI
Shify Treger ⋅ Yehonatan Avidan ⋅ Mengmi Zhang ⋅ Kelsey Allen ⋅ Shimon Ullman
How does intelligence form? In humans, intelligence emerges gradually, from early-acquired and possibly innate knowledge, through vision, embodied experience, social interaction, language, and years of learning. Yet already in infancy, humans exhibit early understanding of object perception, intuitive physics, and social interactions, often from relatively limited experience. In current machine learning systems, by contrast, learning typically relies on large-scale optimization over vast and diverse datasets, where many capabilities emerge together through training, leading to remarkable abilities in some domains while still struggling with tasks and intuitions that appear early in human development. This contrast raises fundamental questions about the nature of intelligence and learning: What can human development teach us about building more robust, flexible, and human-like AI systems? Recent work by the organizers points in this direction, showing how early visual experience, infant-like learning mechanisms, and cognitively inspired models of reasoning can inform the design of more robust, efficient, and generalizable AI systems. This workshop brings together researchers from cognitive development, neuroscience, psychology, and AI to examine intelligence through a developmental lens. In particular, the workshop will explore three central questions: What do human developmental trajectories reveal about the structure and emergence of intelligence? In what ways do current AI systems diverge from infant and child cognition? Can developmental insights lead to better AI systems, and how can such insights be incorporated into modern machine learning models?
Show more
Workshop on Evaluation of Interactive Agents
Yao Dou ⋅ Siyan Li ⋅ Marwa Abdulhai ⋅ Nicholas Tomlin ⋅ Michel Galley ⋅ Jacob Eisenstein ⋅ Alan Ritter ⋅ Wei "Coco" Xu
The rapid transition from large language models (LLMs) as single-turn assistants to interactive agents has created an urgent need for new evaluation methodologies. LLMs are increasingly deployed in high-impact settings such as education, counseling, negotiation, research assistance, and software development, where success depends not only on generating a correct response, but on sustaining effective interactions over extended trajectories. Modern agents must infer user goals, ask clarifying questions, maintain context over many turns, use tools, recover from mistakes, and collaborate with users and other agents. The rise of agentic AI systems and multi-agent workflows has made these evaluation challenges particularly timely. Traditional static benchmarks can miss whether an agent adapts over time, asks useful clarifying questions, recovers from earlier errors, or remains reliable across repeated trials. As conversations extend, model performance can degrade, disparities across user groups can become more pronounced, and stochastic agent behavior can make a single run misleading. At the same time, evaluating interactive agents directly with real users is slow, expensive, difficult to reproduce, and hard to scale in expert domains. This has led to the growing use of LLM-based user simulators, where another model emulates user behavior or task dynamics to support evaluation, training, and stress testing. However, existing simulators may fail to preserve latent user states, reflect diverse human attributes, represent realistic goals, or match the interaction style of real users. These challenges call for a dedicated venue on evaluation methodology for interactive agents.
Show more
Integrating Generative and Experimental Platforms for Biomolecular Design (GEM)
Soojung Yang ⋅ Chenghao Liu ⋅ Jarrid Rector-Brooks ⋅ Lauren Hong ⋅ Sidney L Lisanza ⋅ Jacob Gershon
Biomolecular design, through artificial engineering of proteins, ligands, nucleic acids, and cells, holds immense promise in addressing pressing medical, industrial, and environmental challenges. While generative machine learning has shown significant potential in this area, a disconnect exists with experimental biology: many ML research efforts prioritize static benchmark performance, potentially sidelining impactful biological applications. This workshop seeks to bridge this gap by bringing computationalists and experimentalists together, catalyzing a deeper interdisciplinary discourse. Together, we will explore the strengths and challenges of generative ML in biology, experimental integration of generative ML, and biological problems ready for ML. To attract high-quality and diverse research, we partnered with Nature Biotechnology for a special collection, and we created dedicated tracks for in-silico ML research and hybrid ML-experimental biology research. Our lineup features emerging leaders as speakers and renowned scientists as panelists, encapsulating a spectrum from high-throughput experimentation and computational biology to generative ML. To catalyze new collaborations, we will host a seed-grant competition for pairs of experimentalists and computationalists proposing fresh joint projects. To connect dry and wet lab practice, a wet-lab challenge sponsored by Adaptyv Bio will empirically evaluate protein design models. With a diverse organizing team and backed by industry sponsors, we dedicate the workshop to pushing the boundaries of ML's role in biology. This will be the fourth edition of this workshop following the previous versions of it we organized at ICLR 2024, 2025, and 2026.
Show more
Physical World Model
Kaichen Zhou ⋅ Ruojin Cai ⋅ Jian-Qing Zheng ⋅ Congyue Deng ⋅ Amir Jamaludin ⋅ Yining Hong ⋅ Shangzhe Wu ⋅ Mengyu Wang ⋅ Jianqing Zhen
Modern computer vision and embodied AI systems must now \emph{act} in the physical world, not merely describe it. Yet most pipelines still interpret it through RGB pixels, ignoring the structure that governs how objects move, sound, deform, and respond to contact. We argue that physical world understanding requires AI systems to jointly reason about three coupled pillars, including: \textbf{(1) Physical Geometry:} 3D/4D structure, articulated and deformable motion, pose, and spatial scene understanding. \textbf{(2) Physical Characteristics:} material properties, dynamics, deformability, affordances, and physical interactions. \textbf{(3) Physical Sensors:} non-RGB modalities including tactile, force/torque, proprioceptive, audio, RF, depth, inertial, and event-based sensing. The \textbf{1st Workshop on Physical World Understanding (PhysWorld)} brings together researchers from computer vision, robotics, graphics, multimodal learning, haptics, audio, and physics-based simulation to address a central question: \textit{How can AI systems perceive, represent, and reason about the physical world beyond appearance?} \textbf{Importance and timeliness.} Recent advances in foundation models, embodied AI, world models, multimodal sensing, and differentiable simulation have matured the methods needed to treat physical perception as a unified problem spanning geometry, dynamics, materials, and sensing. At the same time, applications in robotics, AR/VR, autonomous systems, digital twins, scientific imaging, and human-AI interaction increasingly require models that understand the physical consequences of actions rather than merely visual appearance. Despite rapid progress, these research directions remain fragmented across separate communities and venues. Without a coordinated forum now, robotics, vision, and audio sub-communities will independently entrench incompatible benchmarks, ontologies, and evaluation protocols for physical-world foundation models, locking in fragmentation that retroactive standardization rarely undoes. PhysWorld convenes these communities to look beyond pixels, establishing a common forum before that fragmentation hardens. \textbf{Expected outcomes.} Shared evaluation protocols across the three pillars; a community-validated articulation of physical foundation models that integrate vision with touch, sound, proprioception, and dynamics; a cross-disciplinary author cohort publishing together across vision, robotics, graphics, and multimodal sensing; and an accessible entry point for early-career researchers into multimodal physical AI. \noindent\textbf{Topics of interest include, but are not limited to:} \textbf{(a) Physical Geometry:} 3D/4D reconstruction, articulated and deformable scene understanding, geometry-aware world models, and physically grounded view synthesis. \textbf{(b) Physical Characteristics:} estimating material and physical properties from vision and multimodal sensors (mass, friction, stiffness, elasticity, deformability, affordances); physics-informed learning, differentiable simulation, contact-rich interaction modeling, and generative models of physical dynamics. \textbf{(c) Physical Sensors:} multimodal sensing and sensor fusion using tactile, force/torque, proprioceptive, RF, audio, depth, IMU, and event-based signals. \textbf{(d) Cross-cutting:} embodied world models, robot manipulation, sim-to-real transfer, multimodal simulators, and physically grounded reasoning for autonomous agents; benchmarks, datasets, evaluation protocols, and responsible deployment of multimodal physical AI systems.
Show more
Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era
Kiana J Meimandi ⋅ Ahmad Rushdi ⋅ Martin Gonzalez ⋅ Marc R Schlichting ⋅ Dylan Asmar
Over the past two years, agentic AI has moved from research demos to broad deployment in coding, clinical decision support, and conversational settings, yet the community's evaluation toolkit, built largely for static benchmark performance, has not kept pace. A growing body of peer-reviewed evidence documents a persistent gap between benchmark results and deployed reality, with evaluations dominated by technical metrics while human-centered, safety, and economic dimensions remain peripheral. This workshop, the second edition of the Human-AI Coevolution (HAIC) series, builds a methodological foundation for the empirical evaluation of human-agent teams as they coevolve with the people who use them. We organize the program around three interlocking bottlenecks, each anchored in recent peer-reviewed work: the validity of evaluation in deployed contexts, where measurement targets move and core constructs are contested; expert disagreement and the limits of human feedback, where annotator disagreement may signal genuine domain pluralism rather than noise; and adaptive testing and continual evaluation for systems that drift as humans and agents coevolve. Grounded in high-stakes domains including healthcare, mental health, aviation, and finance, the workshop solicits new methods, benchmarks, datasets, critiques, and case studies.
Show more
The Agentic Web: Networked, Continually-Adapting Agent Ecosystems
Pradyumna Chari ⋅ Eric Horvitz ⋅ Zixuan Ke ⋅ Risto Miikkulainen
AI's basic unit of computation is shifting, from single models and individual agents toward continually running, internet-scale collectives of agents that discover one another, communicate, coordinate, and adapt over open networks. Studying these collectives raises questions that single-agent methods do not answer: how to represent and discover agent capabilities at web scale; how learning, credit assignment, and self-improvement behave over changing agent graphs; and how large populations of agents coordinate, remain trustworthy, and stay safe under strategic pressure. These questions are pressing because the architectural and evaluation choices being made now will harden into lasting path dependencies, and no shared benchmarks or evaluation methodology yet exist. This workshop treats the networked agent collective as a first-class research object and convenes researchers across multi-agent and representation learning, learning theory, reinforcement learning, and the societal impacts of AI, through invited and contributed talks, a panel, a hands-on session, and open discussion, to identify the field's foundational problems and the shared evaluation it needs.
Show more
SocialAgent: Second Workshop on Large Language Models for Social Reasoning and Simulation
Xiangjue Dong ⋅ Jiseon Kim ⋅ Yunah Jang ⋅ Tim G. J. Rudner ⋅ Alice Oh
Large Language Models (LLMs) are increasingly used not as analytic tools but as interactive social agents that make decisions, negotiate norms, and coordinate with one another. Generative-agent frameworks sustain role-consistent behavior and emergent coordination; LLM populations produce convention formation, polarization, and bias amplification; and models now stand in for human participants at scale. Yet these systems also encode strong normative assumptions and amplified biases in sensitive domains, revealing both the power and the fragility of treating LLMs as social simulators. The 2nd Workshop on Social LLMs (SocialAgent@NeurIPS 2026) convenes the ML community around the core technical questions this raises: how to model, evaluate, and align LLMs that reason about and simulate human social behavior. The first edition (SocialLLM@ICWSM 2026) targeted computational social science; this NeurIPS edition recenters on methods, evaluation, and alignment, organized around three themes: (i) social reasoning and cognition---theory of mind, moral judgment, and relational reasoning, and when these are genuine versus prompting artifacts; (ii) social simulation and its validity---LLMs as proxies for individuals and populations, with attendant threats from causal-inference pitfalls, simulation boundaries, persona drift, and validation standards from agent-based modeling; and (iii) pluralistic alignment and evaluation---whose values a model represents, how to measure contested values, and how emergent norms and biases in multi-agent settings should be read.
Show more
Continual Learning for Enterprise AI Agents
Naoki Abe ⋅ Aurelie Lozano ⋅ Xi Yang ⋅ Yu Deng ⋅ Meng Jiang ⋅ Cheng Xiang Zhai ⋅ Michael Littman
Existing research in continual learning has primarily focused on model-level adaptation, including updating model parameters and representations as new data becomes available. However, modern AI systems are rarely deployed as standalone models, particularly in enterprise AI applications. Instead, they are increasingly realized as intelligent agents that operate in complex environments, interact with external tools, and coordinate across multiple system-level components. This shift calls for a broader perspective on continual learning, which extends beyond model adaptation to encompass agent-level adaptation, including tool use, workflow optimization, orchestration strategies, safety mechanisms, and feedback integration. To address this emerging challenge, we propose a workshop on continual learning for AI agents, with a particular emphasis on enterprise settings. Enterprise AI agents are already being deployed across domains such as IT operations, customer support, software engineering, and business process automation. These applications provide a compelling and realistic setting for studying continual learning because they involve evolving tasks, changing system infrastructures, rich interaction logs, diverse feedback signals, and complex multi-tool workflows. At the same time, enterprise settings offer valuable artifacts for developing realistic benchmarks (e.g., ITBench, SWE-bench, GDPval) and evaluation protocols for continual learning in agentic systems.
By bringing together researchers in continual learning, reinforcement learning and agent systems, and industrial practitioners building production-grade agents and harnesses, this workshop aims to advance the methodological foundations and practical deployment strategies of continual learning for enterprise AI agents.
Show more
By bringing together researchers in continual learning, reinforcement learning and agent systems, and industrial practitioners building production-grade agents and harnesses, this workshop aims to advance the methodological foundations and practical deployment strategies of continual learning for enterprise AI agents.
Medical Reasoning with Multimodal Foundation Models
Anas Zafar ⋅ Sana Tonekaboni ⋅ Min W Sun ⋅ Julia Vogt ⋅ Alejandro Lozano ⋅ Jia Wu
Vision-language foundation models have rapidly advanced pattern recognition in medicine, yet they remain fundamentally weak at reasoning, connecting observations to knowledge through explicit, verifiable, multi-step inference under uncertainty. This gap is the critical bottleneck separating today's models from trustworthy clinical decision support. Med-Reasoner is a one-day workshop at NeurIPS 2026 that convenes machine learning researchers, multimodal and vision-language model experts, medical AI scientists, and practicing clinicians around four concrete research threads: (i) faithful chain-of-thought for diagnosis and its verifiability; (ii) multimodal grounding of visual evidence in clinical and biological knowledge; (iii) evaluation frameworks measuring reasoning quality beyond answer accuracy, including faithfulness, grounding, and calibration; and (iv) probabilistic and agentic reasoning for sequential decision-making. The workshop builds on the first edition held at CVPR 2026 (30 submissions, 14 accepted, full room capacity, Amazon Health Services sponsor) and is redesigned for the NeurIPS audience around reasoning as a foundational ML problem. A structured panel debate and community open-problems synthesis will produce a shared agenda for standardized reasoning evaluation metrics, a concrete deliverable the field currently lacks. Estimated attendance is 100–150, drawing from the VLM, multimodal learning, medical imaging, computational pathology, and clinical AI communities.
Show more
The BabyVLM Workshop: Toward Developmentally Plausible Multimodal Systems
Boqing Gong ⋅ Aaron Mueller ⋅ Paula Buttery ⋅ Leshem Choshen ⋅ Suchir Salhan ⋅ Shengao Wang ⋅ Wenqi Wang ⋅ Max Whitton
**The BabyVLM Workshop** at NeurIPS 2026 establishes a multidisciplinary forum to address the sample-efficiency gap between multimodal language models and human infants in their ability to learn language. Bringing together researchers from multimodal machine learning, cognitive science, and developmental psychology, the workshop explores how grounding language in other modalities can unlock data-efficient learning. It features keynotes and panels from senior scholars to foster a community centered on learnability, human-inspired AI evaluation, and longitudinal, egocentric learning. The event will also mark the official kickoff and tutorial for the *BabyVLM Challenge*, a competition explicitly designed to catalyze the development and fine-grained evaluation of small, sample-efficient vision-language models.
Show more
Second Workshop on MLxOR: Mathematical Foundations and Operational Integration of Machine Learning for Uncertainty-Aware Decision-Making
Minshuo Chen ⋅ Jing Dong ⋅ Henry Lam ⋅ Karthyek Murthy ⋅ Min-hwan Oh ⋅ Devavrat Shah ⋅ Renyuan Xu ⋅ Enlu Zhou
This workshop aims to present recent developments, discuss challenges, and publicize emerging research opportunities in data-centric decision-making that leverage both the rapid advancement of ML and the principled methodological rigor of OR. Through tools ranging from stochastics to optimization, OR is capable of translating explicit modeling assumptions into interpretable and risk-quantifiable solutions, making it central to reliable decision-making across a wide range of industry applications. At the same time, the analytical tractability of OR models also restricts the level of real-world complexity that they can tackle. To this end, rapid advances in AI/ML offer a powerful opportunity to complement traditional OR and substantially improve decision performance, while, conversely, decades of rigorous OR research can help address key challenges surrounding black-box systems in AI/ML. Building on the momentum of the inaugural workshop in 2025, this second launch of the workshop will continue to explore this two-way OR-ML synergy, this year with a focus in "GenAI", which broadly refers to generative models, including language models, diffusion models, and related foundation models, that can represent, generate, and reason over complex, multimodal scenarios and actions. The emergence of GenAI has already begun to reshape research in OR and shown promise in producing decisions at a scale and complexity previously unattainable, yet it also raises fundamental challenges around evaluation, reliability and safety that counteract the core principles of OR. This theme is thus urgent and impactful in steering the future OR direction. The goal is to bring together researchers from diverse backgrounds to develop a shared understanding of the emerging challenges and opportunities at the GenAI-OR interface, ultimately laying rigorous foundations for reliable, uncertainty-aware, and resource-efficient decision-making.
Show more
Successful Page Load