Temporal Activation Monitors for AI Oversight
Avigya Paudel, Adrian Molofsky, Harleen Kaur Bagga, Ipshita Bandyopadhyay, Jyotin Goel, Michal Mraz, Ronen Roy, Tejas Dahiya, Timothy Obiso, Vansh Gupta, Joseph Rudoler
Mentored by Justin Shenk
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
We investigate how large language models internally represent and process temporal information, an open question with direct implications for AI safety — particularly for understanding planning capability, deception detection, and oversight robustness. Across 11 complementary sub-projects, we apply linear probes, causal interventions (activation steering), and behavioral evaluations to models ranging from 1B to 30B parameters. Our findings show that temporal representations are genuinely encoded in LLM activations and can be causally steered, but that careful baselines are essential: in code generation, what appears to be planning signal is fully explained by surface features. We also find safety-relevant phenomena including context fatigue (entropy collapse over long contexts), sycophantic surrender under social pressure, and a tentative link between short-term temporal steering and more harmful outputs. These results motivate both methodological caution in probing studies and further investigation of temporal representations as a lever for AI safety monitoring.