Back to library
Spring 2026 Submitted May 2026

Temporal Activation Monitors for AI Oversight

Avigya Paudel, Adrian Molofsky, Harleen Kaur Bagga, Ipshita Bandyopadhyay, Jyotin Goel, Michal Mraz, Ronen Roy, Tejas Dahiya, Timothy Obiso, Vansh Gupta, Joseph Rudoler

Mentored by Justin Shenk

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

We investigate how large language models internally represent and process temporal information, an open question with direct implications for AI safety — particularly for understanding planning capability, deception detection, and oversight robustness. Across 11 complementary sub-projects, we apply linear probes, causal interventions (activation steering), and behavioral evaluations to models ranging from 1B to 30B parameters. Our findings show that temporal representations are genuinely encoded in LLM activations and can be causally steered, but that careful baselines are essential: in code generation, what appears to be planning signal is fully explained by surface features. We also find safety-relevant phenomena including context fatigue (entropy collapse over long contexts), sycophantic surrender under social pressure, and a tentative link between short-term temporal steering and more harmful outputs. These results motivate both methodological caution in probing studies and further investigation of temporal representations as a lever for AI safety monitoring.