Back to library
Spring 2026 Submitted May 2026

Feature-Resolved Attention

Ketan Kapre, Hemang Jain, Nura Aljaafari, Jamie Stephenson, Aniket Deshpande

Mentored by Dmitry Manning-Coe

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Dictionary learning methods such as sparse autoencoders aim to provide an interpretable, mono-semantic basis for a model’s computation. Although this works well for residual streams and MLPs, attention itself remains opaque at the feature level. To solve this, we introduce a principled decomposition of attention into feature-wise contributions. We call the resulting object Feature-Resolved Attention (FRA). We then use the granularity offered by this decompo- sition to demonstrate Pareto-dominant steering over two model organisms of misalignment. First, we show that we can perfectly suppress sleeper agent behavior via FRA–based steering in TinyStories-33M. Strikingly, in 20% of cases we recover the original text word-for-word. Second, we consider model organisms of Emergent Misalignment (EM). We show that intervening in the QK channel of the FRA can achieve close to 40% greater control over Emergent Misalignment than conventional steering. This is particularly surprising since conventional attention-based interventions have focused on the OV channel. Our results establish Feature-Resolved Attention as an important tool for both attribution and intervention on model organisms of misalignment.