Back to library
Spring 2026 Submitted May 2026

Measuring Performance of LLM Forecasters for Quantitative Risk Modeling of Cyber Misuse

Jeff Mohl, Madhav Khanal

Mentored by Jakub Krys, Matthew Smith

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Risk modelling is a common way of elucidating, estimating and managing risk across various safety-critical contexts, including AI-enabled cyberattacks. However, to derive quantitative estimates of real-life risk, current approaches often depend on expert judgment, which remains costly and time-consuming. In this work, we address this issue by developing a methodology for evaluating automated LLM “expert” forecasters that map benchmark performance to forecasts of real-world cyberattack capability uplift. Using forecasts over MITRE ATT&CK steps, we assess agreement between forecast variants under various experimental manipulations and find robust self-consistency across repeated runs, as well as broad invariance to alternative elicitation formats, expert persona prompts, and most prompt components, suggesting that LLM forecasters are not highly sensitive to these choices. Additionally, we evaluate LLM forecaster performance in a related cross-task prediction framework, finding that frontier models are capable of evaluating task difficulty and model performance to make accurate predictions on held out tasks. Together, these results suggest that LLM expert forecasters are a useful tool for rapid cyber risk assessments of future models and provide a framework for further improving these forecasters.