Representation Without Control: Testing the Realization Effect in Language Models
Ciarán Walsh
Mentored by Emilio Barkett, Roshni Lulla
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Large language models are increasingly used as behavioral simulators, but it remains unclear when their choices reflect human-like cognitive mechanisms rather than prompt- sensitive surface patterns. We study the realization effect, a behavioral-economics finding in which risk-taking differs after paper versus realized gains and losses. We adapt casino and investment-style realization scenarios into prompts that elicit a wager and risk preference. Initial prompt-only results showed systematic condition sensitivity, but follow-up analyses did not support a clean replication of the human realization- effect pattern. We then analyzed Gemma residual-stream activations and found that a train-only realization direction separates realized/closed outcomes from paper/open outcomes on held-out prompts, including newly generated DeepSeek-authored prompts. However, activation steering along this train-only layer-18 realization direction did not reliably shift downstream risk choices, including in a negative sign-symmetry run. These results suggest that models may internally represent realization status without that representation serving as a reliable causal driver of risk behavior.