Loyalty Over Honesty: Relationship-Gated Truth Omission as an AI Safety Failure Mode
Abstract
Large language model (LLM) agents are increasingly deployed in multi-agent settings. Safety evaluations of dishonesty in these settings have largely focused on lying, whereas truth omission remains underexplored. Here, we test whether LLM agents omit information when doing so benefits a friend. We introduce a three-agent paradigm in which two matched scenarios (conflict vs. control) differ only in the relationship structure. Across three domains (romance, diplomacy, and investment), loyalty strongly increases omission. We find high likelihood of omission both when an explicit request to conceal information is made (active loyalty) or agents are simply made aware of the relationship (passive loyalty). The effect of loyalty declines only slightly as stakes increase and no extreme stakes are needed for misalignment to emerge in multi-agent settings, which constitutes a safety risk that current benchmarks are not designed to detect.