Interpreting Prefill Awareness
Andrew Fletcher
Mentored by Rohan Gupta
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Models are increasingly evaluated using prefilled responses, yet recent work has shown that some frontier models can detect when their conversation history has been tampered with. We replicate results of Africa et al. (2026) [1], demonstrating that (1) Opus 4.5 verbalizes awareness of prefills when directly asked, (2) open-source models encode a representational "not me" signal in hidden states even when they cannot verbalize it, and (3) style is the primary driver of detection, with targeted style transfer reducing verbalized detection AUROC by ~0.3. These findings have direct implications for control methods that rely on prefill, including off-policy activation probes, honeypot evaluations, and transcript replay monitoring.