Back to library
Spring 2026 Submitted May 2026

The Double-Edged Sword: A Framework for Analyzing and Forecasting Capability Spillovers from AI Safety Research

Santi Bent, Serena Ives, Hemang Jain, Varvara Lantukh, Alexandra Gheorghe, Clinton Morimoto, Nathan Theng, Troy Tian, Chris Underhill

Mentored by Ihor Kendiukhov

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

AI safety research routinely yields techniques that, once published, accelerate general AI capabilities, with famous canonical examples like reinforcement learning from human feedback (RLHF). Yet the field has no systematic method for anticipating this dynamic before research is funded or released. This report presents the Capability Spillover Potential (CSP) framework: a multidimensional scoring rubric that estimates the risk of a safety research direction being co-opted for capability advancement. We first ground the framework in retrospective case studies, including RLHF, mechanistic interpretability, and process reward models. We formalize it as a behaviorally anchored rubric scored on ordinal scales and combine two independent data streams (expert survey responses and in-house scoring) through an explicit aggregation function. Applied across twelve core AI safety areas, the framework produces interpretable CSP estimates, with RL-based alignment scoring highest. We argue that capability spillover is a structural, recurring, and scorable phenomenon and that making it legible enables funders, researchers, and governance bodies to pursue differential progress more deliberately. We are explicit about the framework’s current limitations, particularly the assumption that spillover risk is stable enough to forecast prospectively.