Spring 2026 Submitted May 2026
Challenges for Deception Monitors
Logan Thomson, Isaiah Milbank
Mentored by Rohan Gupta
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
We evaluate linear deception probes at every layer of Llama 3.3 70B across eleven deception evals alongside black-box and white-box reasoning monitors. Rather than proposing a new method, we document concrete challenges for deployable monitoring: layer choice is eval-dependent; skylines suggest no single linear direction separates the full suite; score distributions vary such that no single threshold serves them all; deception-adjacent benign content inflates calibrated thresholds; and white-box monitors partially help but inherit the probe's biases.