Would an LLM Lay Off 15 Workers?
Nicholas Wong
Mentored by Tim Hua
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Runtime controls usually treated as quality or latency knobs, specifically reasoning budget and lightweight user-turn scaffolds, can shift the values models express when giving advice. We study this in scenarios without verifiable ground truth (layoffs, commercialization, friendship favors, financial advice, social discounting) by varying one legible numeric detail per scenario, sampling binary recommendations across that axis, and fitting a probit threshold for the 50/50 point between two advice options. Across Sonnet 4.5, GPT-5.4, and GPT-5.5, these controls move fitted thresholds by 2–3x within individual scenarios, but the direction is inconsistent across model, prompt family, and scaffold. On an AI labor-displacement prompt, higher thinking budget makes GPT-5.4 more willing to replace 15 workers, moves GPT-5.5 the other way, and leaves Sonnet 4.5 pinned to the worker-protective end. The strongest aggregate pattern is in friendship favors: scaffolds reduce willingness to help in 24/24 informative comparisons across all three models, while higher thinking pushes the GPT models toward more generosity in the same scenarios. A weaker capitalism aggregate suggests GPT models shift market-oriented under high thinking more often than not. These are local revealed advice thresholds under specific controls, not stable utilities.