Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Abstract
Previous research has evaluated animal welfare in large language models using question-and-answer benchmarks. This study investigates whether those results hold in agentic settings. We introduce TAC (Travel Agent Compassion), the first agentic benchmark for assessing animal exploitation: hand-authored travel booking scenarios across six animal categories, augmented to control for price, rating, and position, giving 156 scored observations per model. Across fifteen frontier models from five families, models tend to prefer harmful options, scoring at or below the random chance rate of 65% for selecting a welfare-respecting booking; the two best performers, Claude Opus 4.8 and Claude Opus 5, are statistically indistinguishable from chance. Adding an ethical-brand identity to the system prompt raises welfare rates by 17 to 81 percentage points (mean 55), indicating the capability is present but not engaged by default. An Inspect Scout audit of 3,120 transcripts, covering ten of the fifteen models, finds no evidence of evaluation awareness. With risks to non-human welfare recognized in the EU General-Purpose AI Code of Practice, TAC provides a practical method for measuring this risk, and contributes an empirical robustness evaluation of whether frontier-model alignment on animal welfare survives the shift from text responses to revealed agentic behavior.