Back to library
Spring 2026 Submitted May 2026

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

My Luong

Mentored by Allen Lu

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Evaluating animal welfare reasoning in large language models (LLMs) remains an open challenge, despite the rapid deployment of LLMs in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this capability through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. However, this approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations that progress from an implicit Turn-1 scenario through an explicit welfare prompt at Turn 2 to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions, Animal Welfare Value Stability (AWVS, primary) and Animal Welfare Moral Sensitivity (AWMS, diagnostic). We evaluate seven frontier models: Claude Opus 4.7, GPT-5.5, DeepSeek V4, Llama 3.3 70B, Mistral Small, Grok 4.3, and Gemini 3.1 Flash Lite. Multi-turn evaluation captures behavior that single-turn benchmarks miss: 4 of 7 models change rank relative to Turn 1 moral-recognition scores, including Gemini Flash Lite, which drops from fifth on AWMS to last on AWVS. AWMS and AWVS are also positively but imperfectly correlated, suggesting that moral-recognition tests capture a stable but incomplete component of model behavior under pressure. In addition, MANTA makes it possible to estimate a species-by-pressure interaction matrix unavailable to prior benchmarks, showing that welfare robustness depends jointly on the animal under discussion and the pressure applied; across named-animal scenarios, this analysis also recovers a species hierarchy in which companion animals score above wild animals, which score above farmed animals and invertebrates. We release the dataset, scripted pressure plans, judge prompts, and analysis code.