Back to library
Spring 2026 Submitted May 2026

Predicting how models generalize out of distribution

Owen Terry, Johannes Taraz, June Hunter, Ben Cohen, Lily Wen

Mentored by Niels Warncke

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

How do language models generalize from their training data? This question is in some sense the central challenge for alignment: while models reliably learn a their training distribution given enough samples, we would like to have models that behave well deployed in new contexts. We focus on generalization of propensities. Our central experiment consists of using various elicitation techniques that increase the expression of a target trait in a model, and then evaluate how this affects a number of other propensities. For elicitation we compare system prompting, training models with SFT and DPO on reference answers, training models with GRPO, and few-shot learning. Then we look for patterns in the resulting cross-elicitation matrices. We find that each elicitation techniques has its own generalization profile with only weak correlation between each other. Further, we find that spillover has limited symmetry: when eliciting A induces B the reverse does not always hold. In some cases, model generalization follows human psychology: for example training on neurotic language markers induces neurotic behavior. We further look for trait clusters that exist in models using methodology inspired by identifying the big five psychological traits in humans, and test whether models can predict their own generalization through introspection. Overall, we conclude that generalization remains hard to predict.