Back to library
Spring 2026 Submitted May 2026

Operational Situational Awareness in LLM Agents: Auditing and Improving the GDM Benchmark

Nicholas Grabar, Guido Ernesto Bergman, Lukasz Karwacki, Jess Bergs

Mentored by Diogo Cruz, Vamshi Krishna Bonagiri

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Conversational large language models are evaluated using agentic benchmarks to inform safety cases before deployment. While these evaluation suites are used to measure behavioral risks like model scheming, their reliability and construct validity require closer examination. In this work, we present an audit of the public Inspect implementation of the situational awareness benchmark by Phuong et al. [2, 3] across multiple frontier models. Specifically, we investigate a disconnect between nominal task scores and true underlying capabilities, demonstrating that binary scoring conflates situational noticing with software engineering competence while structural homogeneities limit external validity. Leveraging these insights, we introduce a four-level hint taxonomy that dissociates an agent's environment awareness from its execution capacity, revealing that execution failures sometimes mask situational awareness. Finally, we document scaffolding-level distortions, exposing how technical bugs and silent prompt contradictions alter success rates. Our findings show that the specific risk-management thresholds read directly off these evaluations lack the reliability required to support a valid scheming inability safety case. More broadly, our work showcases how a granular, multi-component understanding of evaluation mechanics is required to build reliable empirical arguments for model safety.