Back to library
Spring 2026 Submitted May 2026

Stress-testing inoculation prompting

Tim Farrelly, Tim Farrelly, Adam Prada, Ishaan Panigrahi

Mentored by Maxime Riche, Daniel Tan

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Inoculation Prompting (IP) is a recently proposed mitigation for Emergent Misalignment (EM) that suppresses undesirable behaviours by explicitly eliciting them during fine-tuning. The technique has shown strong empirical results and is already being used in production at Anthropic, but its potential failure modes remain underexplored. In this work, we investigate whether inoculated behaviours remain recoverable under semantically or structurally related prompting conditions through ``leaky backdoors''. Across multiple experimental settings, we find that undesirable behaviours can indeed remain partially recoverable under related prompts. We additionally explore mitigation strategies aimed at reducing this leakage. Rephrased inoculation prompts substantially broaden behavioural leakage, while benign training data paired with anti-inoculation prompts significantly suppresses leaky backdoors. Our results suggest that behavioural leakage introduced during inoculation can be substantially reduced through additional behavioural specification during training.