Back to library
Spring 2026 Submitted May 2026

Reasoning and learning about injected concepts in language models

Samarth Bhargav

Mentored by Damiano Fornasiere, Mirko Bronzi

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Language models can, on occasion, correctly answer questions about "injected concepts", i,e., steering vectors added to their activations. This capacity, however, proves fragile across models and prompts, and the extent to which it demonstrates a form of privileged access to internal states remains in dispute. To better understand the phenomenon, we ask language models to reason about a function of an injected concept, through three tasks of increasing demand. First, we query about the magnitude of the injection. Second, we ask for the region of the layer undergoing perturbation---an internal feature of the model itself, the reporting of which, we argue, requires a higher degree of privileged access. Third, when the injected concept is a country, we probe for semantic reasoning about its continent of origin. We test five open-weight models, drawn from three families and two sizes. Notably, all models learn, when provided with in-context examples, to succeed at magnitude and layer detection, often with perfect accuracy. As far as semantic reasoning is concerned, most models attain perfect accuracy even without supervision. However, at the lowest injection magnitude, deducing the continent proves harder, so much so that models also fail to learn in-context.  Together, these findings indicate that language models possess a stronger form of privileged access than previously demonstrated, including the detection of architectural features of the models themselves, and extending to unsupervised reasoning over the injected content.