Single-Anchor Calibration of Logit-Diff Amplification
Patrick Reichherzer
Mentored by Santiago Aranguri
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Frontier LLMs can generate harmful output with a probability small enough not to be adequately investigated before deployment, but large enough to occur in global production use. Logit-diff amplification (LDA) was introduced as a method to allow such harmful output to be detected more frequently, and therefore more easily, during the development process. We extend the methodology with a theoretical framework that estimates LDA rare-harm rates from a single anchor, with a second-order term that flags unreliable extrapolations. We evaluate the method empirically on Qwen2.5-7B-Instruct and a LoRA fine-tuned variant for a representative rare harmful output. Using LDA, we extrapolate the harm vector connecting both models toward more harmful events, sample cheaply at an LDA-extrapolated anchor, and from there estimate the harm rate of the bases.