Back to library
Spring 2026 Submitted May 2026

Geometry of Subliminal Learning

Julian Szereszewski, V S Siva Kumar Lakkoju

Mentored by Yuxiao Li

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Recent work has shown that large language models can transmit behavioral biases through seemingly neutral data, a phenomenon known as subliminal learning. In this work, we investigate this effect through a mechanistic and information-theoretic lens, framing it as a transmission process in which behavioral information is encoded into token distributions and recovered as downstream behavior. Empirically, we observe heterogeneous transfer across animals and models, as well as non-trivial effects under local and global shuffling, challenging purely sequential explanations. To understand this, we analyze the geometry of model activations and show that the difference in mean activations between biased and baseline datasets defines a direction in representation space that can be used to steer model preferences. This demonstrates that subliminal signals are encoded, at least in part, as low-dimensional shifts in activation space. We further validate this by showing that steering vectors derived from biased number completions are sufficient to induce animal preferences in a base model without any fine-tuning, and that cross-model steering within the same model family partially transfers as well. We then ask whether representational geometry can predict the strength of subliminal transfer. We define a geometric alignment score based on the cosine similarity between activations induced by biased numerical sequences and those induced by explicit preference prompts, and evaluate it via a pairwise ranking test against observed transfer strength under both steering and fine-tuning. Across most settings, this metric performs substantially above chance, reaching up to 68% pairwise accuracy, suggesting that representational alignment captures a meaningful component of the transfer mechanism. Failure cases, particularly for larger models under fine-tuning, point toward additional bottlenecks and potentially nonlinear representational structure not captured by directional similarity alone. Together, these results support a view in which subliminal learning is governed by the geometry of activations, where first-order directional signals account for a substantial portion of transfer, and where geometric alignment between implicit and explicit representations predicts transfer strength across both intervention types. This provides a principled framework for predicting, detecting, and potentially mitigating hidden behavioral signals in training data, while also highlighting the limits of purely linear accounts and the need for richer geometric tools