Back to library
Spring 2026 Best Poster, 2nd place Submitted May 2026

Interpreting Prefill Awareness

Andrew Fletcher

Mentored by Rohan Gupta

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Models are increasingly evaluated using prefilled responses, yet recent work has shown that some frontier models can detect when their conversation history has been tampered with. We replicate results of Africa et al. (2026) [1], demonstrating that (1) Opus 4.5 verbalizes awareness of prefills when directly asked, (2) open-source models encode a representational "not me" signal in hidden states even when they cannot verbalize it, and (3) style is the primary driver of detection, with targeted style transfer reducing verbalized detection AUROC by ~0.3. These findings have direct implications for control methods that rely on prefill, including off-policy activation probes, honeypot evaluations, and transcript replay monitoring.