Value Conflicts Widen the Value-Action Gap in LLMs
Dorian Benhamou Goldfajn
Mentored by Andy Liu
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
The value-action gap refers to the discrepancy between what people say they value and how they act. Large Language Models (LLMs) exhibit a similar phenomenon, pursuing actions inconsistent with their stated values. Prior work has relied on abstract, under-specified scenarios that do not reflect how values play out in practice, leaving the gap poorly understood in realistic contexts where decisions often involve trade-offs between competing values. To this end, we evaluate the value-action gap in detailed value- conflicting scenarios, measuring it both through models’ stated agreement with individual values and through their stated preference when choosing between two competing values. Across five models, we find that the gap widens by an average of ∼23 percentage points in realistic value-conflicting scenarios relative to under-detailed single-value settings, and persists when we explicitly prompt models to consider value trade-offs. Compared to individual agreement levels, models tend to yield smaller inconsistencies when asked to choose between values directly, but the gaps remain large. Our findings reveal a wider inconsistency between LLMs stated values and actions than prior work has captured, and suggest that it cannot be attributed to models failing to recognize the underlying sacrifices between competing values. More broadly, we highlight that neither form of stated value inclination should be used as a reliable proxy for model behavior.