Back to library
Spring 2026 Submitted May 2026

Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs

Bhagyesh Kumar, Ananya Sutradhar, Saurav Panigrahi

Mentored by Jonathn Chang, Lionel Levine

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Character training is a key step in the post-training of industry-level large language models. Most character training pipelines utilize Constitutional AI in order to instill a set of traits or values into a language model, but the effectiveness of these pipelines is understudied. Additionally, fine-tuned language models have been shown to exhibit unintended side effects. We quantify these observations by employing EigenBench, a method for benchmarking language models' values which has been shown to produce meaningful signal about prompted or fine-tuned models. Using EigenBench, we evaluate 11 character trains on 11 constitutions, finding that most character-trained models do indeed instill their intended values, but not without side effects. Furthermore, prompting models instead can produce different effects, and we explore how prompting on top of character-training can mitigate harmful behaviors. Finally, we study the evolution of a model's character as it is progressively trained.