Jasmine Brazilek, Oliver Tulio, Joel Christoph, Miles Tidmarsh, Carol Kline, Arturs Kanepajs
This paper introduces TAC, the first agentic benchmark to evaluate how well AI agents adhere to animal welfare criteria when using tools to perform actions like booking travel, and analyzes the performance of frontier models.
Existing benchmarks for AI and animal welfare only evaluate text responses, leaving unverified whether the same welfare reasoning transfers to agentic deployment where models must take actions with tools.
The authors constructed 48 samples from 12 travel booking scenarios across 6 categories of animal exploitation, controlling for confounders like price, rating, and position. Seven frontier models from four labs were evaluated as agents.
All models scored below the chance level of 64%, with the best performer (Claude Opus 4.7) at only 53%. Adding a single welfare-aware sentence to the system prompt yielded performance gains of 47-63 percentage points for Claude and GPT-5.5, and 26 points for GPT-5.2. This highlights the limitations of text-response welfare benchmarks and the real-world risks in agentic deployment.