Pedagogical Games
Krish Sen, Nikhil Narayanan, Luca Franceschetti, Jonathan Robinson, Yadnyesh Chakane, Shefali Agrawal, Dylan Waldner
Mentored by Elizaveta Tennant
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Can a model learn to be moral by playing games? While existing alignment methods rely predominantly on learned preference signals and opaque moral values, we investigate whether fine-tuning with explicitly defined moral rewards can induce transferable cooperative dispositions in LLM agents. Generalization is evaluated across three dimensions: strategic complexity, model capability, and naturalistic complexity. We show that an LLM finetuned exclusively on numerical multi-agent games (with no natural language moral content), reduces harmful actions by up to 35\% in semantically unrelated interactive environments. However, this generalization occurs only if training on iterated public goods games but not pairwise reciprocity games, and if environment complexity is matched to model capability. Our results provide evidence that intrinsic moral fine-tuning is a promising direction for LLM alignment, and offer preliminary answers to the questions: which environments work, for which models, and why.