Back to library
Spring 2026 Submitted May 2026

Challenges for Deception Monitors

Logan Thomson, Isaiah Milbank

Mentored by Rohan Gupta

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

We evaluate linear deception probes at every layer of Llama 3.3 70B across eleven deception evals alongside black-box and white-box reasoning monitors. Rather than proposing a new method, we document concrete challenges for deployable monitoring: layer choice is eval-dependent; skylines suggest no single linear direction separates the full suite; score distributions vary such that no single threshold serves them all; deception-adjacent benign content inflates calibrated thresholds; and white-box monitors partially help but inherit the probe's biases.