Back to library
Spring 2026 Submitted May 2026

Crosscoding Through Time: Sparse Feature Discovery Across Sequence Positions

Han Xuanyuan, Aniket Deshpande, Andre Shportko, William Fei

Mentored by Dmitry Manning-Coe

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Dictionary learning methods - such as Sparse Autoencoders (SAEs) and crosscoders - decompose model activations into human-interpretable building blocks. We introduce \textit{temporal crosscoders}, a simple and flexible framework for feature discovery in Large Language Models (LLMs). To properly evaluate temporal crosscoders we develop TempBench: a panel of synthetic and real-world tasks for evaluating temporal structures. Temporal crosscoders outperform both conventional and temporal architectures in both of our synthetic settings and on two out of four of the real world settings - ahead of other candidate architectures. Most strikingly, they can detect backtracking - a key reasoning behavior - at a 40\% higher rate than conventional SAEs, and are 15\% more effective in inducing it. Our results establish temporal crosscoders as a simple and flexible framework for feature discovery, both local and temporal. We provide full code at the following anonymous repository: \url{https://anonymous.4open.science/r/temp\_xc-33E3/README.md}