Do AI-for-Science Agents Preserve Model Safety?
Abstract
AI-for-Science agents are increasingly used to support scientific discovery. These systems combine a base large language model (LLM) with a domain-specific agent harness that connects the model to scientific tools and workflows. While the harness is designed to improve scientific capability, it also changes the system in which the model operates, and its safety behavior need not match that of the base LLM. We ask whether model safety is preserved when an LLM is embedded in an AI-for-Science agent harness. We evaluate four agent harnesses on 50 harmful chemistry requests using three base LLMs and three prompting conditions. Our results show that introducing the harness can significantly increase harmful compliance, including for direct harmful requests. Safety measured on a base LLM therefore cannot be assumed to be preserved in the composed agent. Thus, as AI-for-Science agents become more capable and autonomous, the community needs to evaluate and align the composed agent itself using domain-specific safety benchmarks, rather than assuming that safety properties of the base LLM will carry over.