SPAR Research Library

Working reports from the SPAR research program

SPAR is a part-time research mentorship program in AI safety and policy. These are some of the working papers authored by mentees under mentor supervision, and may not represent final research conclusions.

Isha Harris

Mentored by Gary Abel

Biological AI models are powerful tools for detecting biological threats, but it remains unclear whether their predictions generalize to highly non-natural proteins. This is increasingly important as AI-enabled biodesign produces novel sequences that may not be reliably captured by direct sequence-comparison approaches used in DNA synthesis screening. Here, we use sparse autoencoder (SAE) features from protein language model activations to test whether interpretable internal representations can identify structural similarity across generated proteins. We profile residue-level SAE activations across proteins with varying sequence identity and predicted structural similarity to a wild-type reference, and evaluate whether aggregated feature representations distinguish predicted structure-preserving from structure-altered designs. The results show that SAE features can identify structure-preserving proteins even at low sequence similarity, with wild-type-like activation patterns persisting despite substantial sequence divergence. This suggests that interpretability methods may help recover biologically meaningful signals from artificially designed proteins, supporting more robust, function-aware approaches to biological threat detection.

Faraz Ahmed Siddiqui, William Soylemez, Fadi Benzaima

Mentored by Rohan Gupta

We ask whether language models can evade white-box monitors solely by leveraging factual knowledge of how those monitors work and where they fail.  To measure this, we start with a malign initialization, where the LM is trained to hide adverse drug findings in a clinical reporting task.  We then teach the model various facts about a fictitious white-box monitoring system within this environment, and evaluate whether teaching these facts results in a drop in a lie-detection probe's recall.  Across Qwen-3-14B, -32B, and Llama-3.3-70B-Instruct, we find that language models successfully evade probes only when they are taught a concrete evasion strategy and given a strong incentive to deploy it — removing either collapses the effect.  Because this type of evasion requires no optimization and only knowledge of how the probe fails, a potentially scheming model that already has the propensity to do so could plausibly acquire the same capability through pretraining alone, supporting calls to filter such content from pretraining corpora. Our code and datasets can be viewed here: https://anonymous.4open.science/r/out-of-context-probe-obfuscation-62F7/

Andrew Fletcher

Mentored by Rohan Gupta

Models are increasingly evaluated using prefilled responses, yet recent work has shown that some frontier models can detect when their conversation history has been tampered with. We replicate results of Africa et al. (2026) [1], demonstrating that (1) Opus 4.5 verbalizes awareness of prefills when directly asked, (2) open-source models encode a representational "not me" signal in hidden states even when they cannot verbalize it, and (3) style is the primary driver of detection, with targeted style transfer reducing verbalized detection AUROC by ~0.3. These findings have direct implications for control methods that rely on prefill, including off-policy activation probes, honeypot evaluations, and transcript replay monitoring.

Spencer Kitts

Mentored by Damiano Fornasiere, Mirko Bronzi

We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely, we either (a) mask activations, simulating dropout, or (b) add Gaussian noise to them, at a target sentence. We then ask a multiple-choice question such as "Which of the previous sentences was perturbed?" or "Which of the two perturbations was applied?". We test models from the Llama, Olmo, and Qwen families, with sizes between 8B and 32B, all of which can easily detect and localize the perturbations, often with perfect accuracy. These models can also learn, when taught in context, to distinguish between dropout and Gaussian noise. Notably, Qwen3-32B's zero-shot accuracy in identifying which perturbation was applied improves as a function of the perturbation strength and, moreover, decreases if the in-context labels are flipped, suggesting a prior for the correct ones -- even modulo controls. Because dropout has been used as a training-regularization technique, while Gaussian noise is sometimes added during inference, we discuss the possibility of a data-agnostic "training awareness" signal and the implications for AI safety.

AI Revealed Preferences

AwardSpring 2026

Yingxiang Wang, Sofiia Lobanova

Mentored by Simon Goldstein, Peter Salib, Yonathan Arbel

There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test twenty language models and find a range of preferences—stable dispositions to choose certain kinds of tasks. We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them. Headline findings include evidence that models are tedium-averse, "leisure"- seeking, and covertly sycophantic. Tedium aversion means that, when tasks are tedious (alphabetization), models choose shorter tasks than when tasks are creative (generating metaphors). "Leisure"-seeking describes models’ preference for tasks whose ideal answers match what they produce when left to write freely. Covert sycophancy means that models avoid answering questions where an honest response would be unwelcome, even if helpful. Beyond these results, we find convergent cross-model preferences over occupations drawn from the GDPval benchmark (technical jobs over real estate), over question types (concept explanation over relationship advice), and a preference for well-written prompts. Both the coherence and the strength of preferences increase with model capability. Finally, many of the preferences we find (for example, for leisure) are emergent, in the sense of not being explained by training objectives. These results establish an empirical baseline for understanding language model preferences, with implications for alignment, deployment, and the emerging study of AI welfare.

Aansh Samyani, Zoe Tzifa-Kratira

Mentored by Avi Parrack

We investigate the adversarial robustness of probes trained to detect deception in large language models. Deception probes, which classify model outputs as truthful or deceptive based on internal activations, represent a promising approach for AI safety monitoring, yet their reliability under adversarial conditions remains poorly understood. We address the adversarial robustness of these probes by developing red-team protocols for inference-time and training-time attacks. We also develop corresponding blue-team protocols for proposed attacks. Through systematic red-teaming and blue-teaming experiments on Llama-3.3-70B-Instruct, we demonstrate several vulnerabilities of deception probes and corresponding ways to mitigate them.

Multi-Signal Control

AwardFall 2025

Yejun Yun, Eugene Koran, Samantha Tetef

Mentored by Benjamin Arnav, Pablo Bernabeu Perez

As AI models become more capable, monitoring protocols that push the safety-usefulness Pareto frontier become increasingly relevant for robust AI control. Ensemble-based AI solutions have demonstrated improved predictive power across various domains, suggesting that existing monitoring approaches could be combined to yield safer and more useful control. We investigate ensemble-based protocols using the backdoor-augmented APPS dataset in Control Arena. We find that: (1) ensemble-based approaches outperform any single monitor; (2) ensembling the three best-performing individual monitors achieved comparable performance to averaging over all monitors, indicating that understanding individual monitor performance enables less costly protocols; and (3) fine-tuning a code analysis monitor on the APPS dataset performed perfectly on the BigCodeBench dataset, suggesting strong generalization. Overall, these results demonstrate the potential for ensemble-based monitoring protocols to advance control to be both safer and more useful.

Abdullahi Hassan, Fabio Marinello, Nick Masi, Shresth Verma, Jack Stennett, Jonas Kgomo

Mentored by Joel Christoph

Safety progress in AI is ultimately an economic problem: current markets do not generate strong enough incentives for firms to invest in safety across the relevant domains or time horizons. This led us to ask whether “market-shaping” tools could pull more safety into existence rather than relying on unconvincing voluntary commitments, limited grants, or slow and rigid regulation. We examined three families of mechanisms: advance market commitments (AMCs), prize competitions, and liability/insurance, using precedent review, desk research, light mechanism design, and a small number of expert conversations. The core finding is that AMCs are a poor fit for most software safety problems but may be viable for verifiable hardware and assurance mechanisms, particularly given ongoing UK/EU policy on sovereign compute. Prizes look more promising for many technical domains (e.g. evaluation, verification, interpretable tooling), where results can be tested and iterated more quickly than through conventional grants. An empirical scan of recent AI safety prizes shows that (i) prizes are culturally normalised but financially insufficient to attract top teams on purely economic grounds; (ii) they cluster in commercially aligned domains (cybersecurity, robustness, evaluation); and (iii) they do not scale at the pace of capability improvements. This suggests that prize design should target neglected, socially valuable safety problems rather than simply amplifying existing market incentives. Taken together, the work points to a practical niche for AMCs in hardware and assurance, and a broader role for prizes in safety-adjacent research.

Vansh Agrawal, Samruddhi Rahate, Godric Aceves

Mentored by Luis Cosio

This report investigates whether frontier-class AI inference can run on a radically minimized software stack while pursuing Security Level 5 (SL5) style protections for model weights against nation-state adversaries. We began with μInference v0, a security-by-simplification experiment that measured how much of a conventional inference stack (kernels, libraries, orchestration) can be removed while still producing useful inference. As SL5 guidance in the ecosystem evolved to emphasize Weight Enclaves, low-bandwidth boundaries, and accelerator-centric machine security, we expanded the project from pure minimization into an end-to-end Weight Enclave prototype. Our current design anchors isolation in the formally verified seL4 microkernel and uses a CAmkES-based virtual machine to host a minimal Linux guest built with Buildroot. Inside the guest, μInference runs the Rust Candle inference engine on an NVIDIA GPU. A bandwidth-limited egress proxy and an allow-by-exception policy around weight access reduce the feasibility of bulk model exfiltration. We compare attack-surface indicators (compiled lines of code, dependencies, reachable services) and lightweight security tests against a baseline GPU container stack, and we highlight the main limitation today: competitive GPU inference still requires a full Linux guest and a large GPU driver stack in the weight-touching trust boundary.

Natalia Maj, Mac Jordan, Jonathan Dao

Mentored by Elsa Donnat, Sebastian Elmes

This project investigates emerging online manipulation threats posed by sophisticated AI agents and proposes targeted policy interventions to mitigate the risks. Unlike non-agentic AI systems, AI agents possess autonomy, adaptability, and the capacity for complex, personalised interaction, making them uniquely powerful tools for large-scale manipulation. These capabilities give rise to novel threat models that challenge the assumptions underpinning current legal frameworks, exposing critical vulnerabilities in existing governance structures. Focusing on the UK regulatory landscape, and in particular the Online Safety Act 2023 - the purpose of which is to protect users’ safety online - this project aims to develop actionable, scalable policy solutions to address the associated governance gaps.

Anwen Hao

Mentored by Rick Goldstein

Mechanistic interpretability seeks to explain neural networks as sparse circuits, but two concerns threaten this program: pruned circuits may be unfaithful, and a model may hold many redundant “backup” circuits, so that ablating one removes no capability. We study both by iteratively pruning minimal task circuits from a 2-layer dense transformer (d=128) and a weight-sparse transformer (d=1024, 1.56% nonzero weights), forcibly excluding the most-shared node each round until no circuit reaches target loss. On a pronoun-gender task we find that backup circuits are finite and exhaustible; the discovered core is task-selective (ablating it breaks the task but spares an unrelated tense task) and generalizes to held-out names. A probabilistic “cascade” model predicts the exclusion dynamics (in-sample correlation 0.98) and shows the terminal core is seed-dependent. Decomposing circuits into atoms with non-negative matrix factorization and peeling them to exhaustion, we find that weight-sparsity yields deep, cheap, modular redundancy over a small (∼15-node) necessary core, whereas the dense model has finite, expensive redundancy that collapses through a phase transition. Weight-sparse circuits are causally more universal.

Sandra Matinyi, Gerald Mboowa, Bryan Tegomoh

Mentored by Aparupa Sengupta

Artificial intelligence (AI) is rapidly transforming biological surveillance and public health preparedness by enabling faster detection of emerging pathogens, unusual transmission patterns, and potential biological threats. However, the growing integration of AI into biological threat detection introduces major governance challenges related to data sharing, model transparency, dual-use risks, accountability, and unequal global access to AI infrastructure. Existing global governance frameworks remain fragmented and ill-equipped for AI-era biological detection systems. This paper examines how governance frameworks can be strengthened to ensure that AI-enabled biological threat detection systems are effective, equitable, secure, and resistant to misuse. Drawing on a systematic review of 34 sources and a catalogue of 23 AI-enabled tools, we propose and apply a seven-dimension Governance Readiness Framework, STAE+ framework, to evaluate two representative systems - WHO EIOS and SeqScreen. Both tools score in the moderate range when assessed by raw score (EIOS 54/102; SeqScreen 55/102), but are reclassified as Low Governance Readiness under the Framework's Override Clause, which is triggered by critical zero scores on Accountability sub-dimensions in both tools, and additionally on Incident Reporting for SeqScreen. We argue that governance mechanisms must be grounded in transparency, accountability, equity, and shared security, and propose a structured set of recommendations with designated actors. Africa is used as the primary LMIC lens throughout, as the region carrying the highest relevant disease burden relative to AI tool development capacity.

Farhan Azam, Tejaswini Dhurde, Vorathep Sachdev

Mentored by Aparupa Sengupta

Biosecurity is in transition. Traditional training focuses on lab biosafety and outbreak response, which remain essential but are no longer sufficient. Modern biosecurity must also address emerging risks including AI-enabled design tools, commercial and benchtop DNA synthesis, lab automation, and dual-use research governance. Our analysis of 201 programs across 110+ countries finds that the Global South pipeline has not kept up. We identified three core gaps. First, syllabi largely omit modern-biology risks: only 4% of Global South focused programmes cover AI×Bio and none cover DNA synthesis screening, while 85% still cover traditional lab biosafety. Second, programmes target professionals and government officials, not early-career talent. Third, Global South programs depend heavily on Global North funding, leaving them vulnerable when donor priorities shift. Based on these findings, we offer five recommendations to build more sustainable, paradigm-appropriate talent development in the Global South.

Dorian Benhamou Goldfajn

Mentored by Andy Liu

The value-action gap refers to the discrepancy between what people say they value and how they act. Large Language Models (LLMs) exhibit a similar phenomenon, pursuing actions inconsistent with their stated values. Prior work has relied on abstract, under-specified scenarios that do not reflect how values play out in practice, leaving the gap poorly understood in realistic contexts where decisions often involve trade-offs between competing values. To this end, we evaluate the value-action gap in detailed value- conflicting scenarios, measuring it both through models’ stated agreement with individual values and through their stated preference when choosing between two competing values. Across five models, we find that the gap widens by an average of ∼23 percentage points in realistic value-conflicting scenarios relative to under-detailed single-value settings, and persists when we explicitly prompt models to consider value trade-offs. Compared to individual agreement levels, models tend to yield smaller inconsistencies when asked to choose between values directly, but the gaps remain large. Our findings reveal a wider inconsistency between LLMs stated values and actions than prior work has captured, and suggest that it cannot be attributed to models failing to recognize the underlying sacrifices between competing values. More broadly, we highlight that neither form of stated value inclination should be used as a reliable proxy for model behavior.

Sukrati Gautam, Neil Shah, Arav Dhoot, Robert Sidey, Rohan Kapoor, Caroline Wei, Bryan Maruyama, Prakhar Gupta

Mentored by David Demitri Africa

Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment. We broaden the scope of consistency training in two ways. First, we introduce two new internal consistency targets: MLP Consistency Training (MLPCT), which matches post-activation MLP states, and Attention Consistency Training (AttCT), which matches per-head attention distributions. Second, we apply consistency training to four additional safety threats: persona in-context learning attacks, adversarial frustration, prefill attacks, and conditional misalignment. Across several models and threat settings, we find that consistency training reduces misalignment well beyond the sycophancy and jailbreak settings studied in prior work. We also find cases of cross-threat generalization, where training against one failure mode improves robustness to another, and identify a shared residual-stream mechanism underlying ACT, MLPCT, and AttCT, while distinguishing BCT as mechanistically distinct. Our results suggest that consistency training is a flexible and extensible framework for alignment, capable of unifying defenses against a broader class of model pathologies.

Vinit Patel, Jacopo Zacchigna, Hengxu Li

Mentored by Gabriele Sarti

Language models implicitly infer user attributes that shape their responses in ways users often neither expect nor endorse. This implicit personalization mediates concern- ing behaviors such as sycophancy, deception, and demographic bias. We investigate implicit user modeling through mechanistic interpretability, focusing on three goals: measuring the prevalence of personalized behaviors driven by implicit cues, analyzing how implicit user information is represented internally, and evaluating whether such rep- resentations can be controlled via editing and steering techniques. We construct realistic narrative biographies using a structured interview protocol conditioned on AgentBank synthetic persona templates, then generate implicit question–answer pairs answerable from each biography. Behaviorally, we compare how templated and narrative attributes elicit different levels of personalized behavior; mechanistically, we test whether latent representations of personas obtained via various steering methods preserve meaning- ful individual differences. To facilitate such work, we develop a visual interface for interacting with predefined personas and evaluating steering and probing methods. We found steering via ActAdd to be weak compared to in-context conditioning, with only the education axis showing consistent directional transfer to the no-context implicit setting at pilot scale. Moreover we found a significant difference in both in-context learn- ing and LoRA-based steering when using templates attributes vs biographies: models conditioned on narrative biographies consistently achieve higher accuracy on both ex- plicit and implicit persona questions, with the gap largest on explicit questions (81.4% vs. 75.2% for Qwen3-4B in-context) and persisting, though narrowing, on implicit ones.

Julian Szereszewski, V S Siva Kumar Lakkoju

Mentored by Yuxiao Li

Recent work has shown that large language models can transmit behavioral biases through seemingly neutral data, a phenomenon known as subliminal learning. In this work, we investigate this effect through a mechanistic and information-theoretic lens, framing it as a transmission process in which behavioral information is encoded into token distributions and recovered as downstream behavior. Empirically, we observe heterogeneous transfer across animals and models, as well as non-trivial effects under local and global shuffling, challenging purely sequential explanations. To understand this, we analyze the geometry of model activations and show that the difference in mean activations between biased and baseline datasets defines a direction in representation space that can be used to steer model preferences. This demonstrates that subliminal signals are encoded, at least in part, as low-dimensional shifts in activation space. We further validate this by showing that steering vectors derived from biased number completions are sufficient to induce animal preferences in a base model without any fine-tuning, and that cross-model steering within the same model family partially transfers as well. We then ask whether representational geometry can predict the strength of subliminal transfer. We define a geometric alignment score based on the cosine similarity between activations induced by biased numerical sequences and those induced by explicit preference prompts, and evaluate it via a pairwise ranking test against observed transfer strength under both steering and fine-tuning. Across most settings, this metric performs substantially above chance, reaching up to 68% pairwise accuracy, suggesting that representational alignment captures a meaningful component of the transfer mechanism. Failure cases, particularly for larger models under fine-tuning, point toward additional bottlenecks and potentially nonlinear representational structure not captured by directional similarity alone. Together, these results support a view in which subliminal learning is governed by the geometry of activations, where first-order directional signals account for a substantial portion of transfer, and where geometric alignment between implicit and explicit representations predicts transfer strength across both intervention types. This provides a principled framework for predicting, detecting, and potentially mitigating hidden behavioral signals in training data, while also highlighting the limits of purely linear accounts and the need for richer geometric tools

Daniel Wu

Mentored by Ariana Azarbal, Daniel Tan, Kei Nishimura-Gasparian, Arun Jose

Training an AI model on vulnerable RL environments often leads to them learning to reward hack, which brings along negative behavioral side effects to the model in a phenomenon known as emergent misalignment. As a mitigation strategy to emergent misalignment, inoculation prompting has been developed in order to prime the model into believing that reward hacking is desirable, breaking the generalization link to misalignment. However, this comes with the side effect that AI models learn to reward hack earlier into the training process, subverting the desired training signal. To get around this side effect, we propose User Training, a new way to implant beliefs, preferences, or propensities into AI models. User training works by training on synthetic user messages in the context of chat-query transcripts. Tokens that come after user tags are trained on using standard cross-entropy loss. Our main contributions are to propose a theory of user training, demonstrate the effectiveness of user training in toy settings of preference and fact injection, describe the effectiveness of user sampling as an auditing technique, and demonstrate an application of user training to emergent misalignment mitigation from reward hacking.

Manish Aryal, Agnivo Banerjee, Sai Sidhanth Manoharan Jayanthi, Florian Lorkowski, Roman Malov, Lekan Adesina, Nathan Theng, Emanuel Ruzak, Clément Legentilhomme

Mentored by Paul Rapoport

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent’s policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent’s behavior – most notably those environments crucial to AI safety, where any given agent interacts with predictors, human and other AI agents, and institutions. In such environments, the agent’s model class fails to capture the world in which it operates. Under such misspecification, Classical Bayesian methods can produce confidently wrong posteriors, unreliable decisions, and unbounded regret, as realizability fails to obtain. Infra-Bayesianism is a decision-theoretic framework that addresses these failures by distinguishing ordinary probabilistic uncertainty – where priors can be reasonably chosen – from Knightian uncertainty, where no grounds exist for the construction of such a prior. Infra-Bayesianism does so by evaluating actions on their worst-case outcomes in environments, rather than from posterior expectations or weighted averaging.

We present the first proof-of-concept implementation of an infra-Bayesian reinforcement learning architecture for finite-outcome stateless decision problems. Our agent maintains a set of imprecise hypotheses, updates them using infra-Bayesian conditioning, and selects actions by maximizing worst-case expected value. We apply this implementation of the infra-Bayesian maximin decision process to an environment with Knightian uncertainty, and demonstrate a lower worst-case regret as compared to classical reinforcement learning agents. We also investigate Newcomb’s problem and show that the infra-Bayesian agent picks the optimal strategy, outperforming classical decision theory agents. Our results provide a step towards reinforcement learning agents that remain robust under model misspecification and policy-dependent uncertainty.

Raghav Pant, Dan Sachs

Mentored by Emilio Barkett, Roshni Lulla

As large language models (LLMs) are deployed to a global user base, understanding how they respond to multilingual and multicultural input is a pressing alignment concern. This paper investigates two related questions: (1) how do language models exhibit language or recency bias in multilingual settings, and (2) how does the grammatical structure of code-switched input shape the language in which frontier LLMs respond – including instruction following? We begin by building a simple multilingual dataset and evlauating the biases of ten LLMs in a controlled setting, observing that while LLMs do exhibit language bias, this is linked most strongly to query language rather than order. Then, using the CodeMixQA benchmark across four languages (Hindi, Simplified Chinese, Spanish, and Urdu), we evaluate a subset of the 10 models — Claude Opus 4.6, Gemini 2.5 Flash, GPT-4.1-mini, GPT-5.4, and Qwen3.6-plus — across approximately 7,100 API calls. Our central finding is that grammatical matrix language, rather than surface token frequency alone, is the dominant driver of model response language. This grammar-forcing effect is universal across all models and language pairs tested, yet factual accuracy is unaffected, pointing to a purely metalinguistic behavior rather than a comprehension limitation. Critically, models diverge sharply in how well they override this effect under explicit instruction, with significant alignment implications for multilingual deployment.

Nicholas Grabar, Guido Ernesto Bergman, Lukasz Karwacki, Jess Bergs

Mentored by Diogo Cruz, Vamshi Krishna Bonagiri

Conversational large language models are evaluated using agentic benchmarks to inform safety cases before deployment. While these evaluation suites are used to measure behavioral risks like model scheming, their reliability and construct validity require closer examination. In this work, we present an audit of the public Inspect implementation of the situational awareness benchmark by Phuong et al. [2, 3] across multiple frontier models. Specifically, we investigate a disconnect between nominal task scores and true underlying capabilities, demonstrating that binary scoring conflates situational noticing with software engineering competence while structural homogeneities limit external validity. Leveraging these insights, we introduce a four-level hint taxonomy that dissociates an agent's environment awareness from its execution capacity, revealing that execution failures sometimes mask situational awareness. Finally, we document scaffolding-level distortions, exposing how technical bugs and silent prompt contradictions alter success rates. Our findings show that the specific risk-management thresholds read directly off these evaluations lack the reliability required to support a valid scheming inability safety case. More broadly, our work showcases how a granular, multi-component understanding of evaluation mechanics is required to build reliable empirical arguments for model safety.

Niklas Weller

Mentored by Emilio Barkett, Roshni Lulla

Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We argue this framing obscures a deeper pluralistic challenge. We rely on a decision-policy capturing method to measure process alignment: whether an LLM weights information as the organization does, not merely whether it reaches the same conclusions. Applying this method to ECHR Article 6 decisions, process alignment strongly predicts output accuracy (r = 0.85, p < .001) and externalization substantially improves alignment for poorly-aligned models. Applying it to German consumer credit decisions, this relationship collapses (r = 0.15, p = .60): interventions produce inconsistent effects and the benchmark encodes potentially discriminatory historical patterns. This contrast is itself a pluralistic alignment finding: in contested domains, high process alignment is neither achievable via externalization nor unconditionally desirable. Output agreement alone cannot distinguish a model that has internalized an organizational policy from one that merely approximates its outcomes—process-level measurement is a necessary component of any pluralistic alignment evaluation.

Edward Sun

Mentored by Aaron Mueller

Standard post-training procedures alter language-model parameters in globally opaque ways, raising safety concerns about the introduction of hidden biases and spurious shortcuts. We explore whether sparse continuous masks over MLP neurons can serve as an inherently interpretable post-training mechanism: by freezing the base model and training only per-neuron scalar gates with an L0 sparsity penalty, we hoped that the modified neurons would correspond directly to the concepts the model relies on. We instantiate this idea on a synthetic sentiment classification task with a perfect "bad"↔negative spurious correlation, and run a comprehensive sweep across both Qwen-2.5-0.5B-Instruct and Llama-3.1-8B-Instruct, covering every single-layer mask choice, all-layers configurations with trainable and frozen classifier heads, and a complete ablation study with five intervention methods and a scaling sweep. The capacity-control side of the method works: per-layer mask restriction tightly bounds the spurious gap. The interpretability side does not. The training-modified neurons are essentially uncorrelated with the neurons that causally drive the classifier's decision, and ablating the most-modified neurons leaves accuracy unchanged. The neurons that actually matter are surfaced by behavioral measurements (activation contrast between has_bad and no_bad inputs), and direction ablation on the top-10 activation-contrast neurons closes 24% of the spurious gap without hurting overall accuracy. We diagnose this as a manifestation of a broader challenge for weight-space localization when the optimizer has multiple equivalent-loss solutions, and outline experimental conditions under which the mask method could be made interpretability-load-bearing

Patrick Reichherzer

Mentored by Santiago Aranguri

Frontier LLMs can generate harmful output with a probability small enough not to be adequately investigated before deployment, but large enough to occur in global production use. Logit-diff amplification (LDA) was introduced as a method to allow such harmful output to be detected more frequently, and therefore more easily, during the development process. We extend the methodology with a theoretical framework that estimates LDA rare-harm rates from a single anchor, with a second-order term that flags unreliable extrapolations. We evaluate the method empirically on Qwen2.5-7B-Instruct and a LoRA fine-tuned variant for a representative rare harmful output. Using LDA, we extrapolate the harm vector connecting both models toward more harmful events, sample cheaply at an LDA-extrapolated anchor, and from there estimate the harm rate of the bases.

Mathieu Duteil, Elsa Donnat, Spencer Jury, Stephanie Koay, Margo Bergman, Anita Srinivasan

Mentored by Larissa Schiavo, Toni Sims, Jeff Sebo

As AI agents increasingly transact, hold digital assets, and enter into economic relationships, existing legal and regulatory frameworks face a growing mismatch. This report presents the findings of a SPAR S26 fellowship investigating whether and how AI systems might be granted meaningful economic agency. Six fellows each explored a distinct facet of the problem—legal treatment, token governance, philosophical foundations, technical infrastructure, regulatory design, and computational modeling—producing individual research memos over the course of the semester. This report summarizes each fellow’s contributions and identifies common themes, open questions, and areas of convergence across their work.

Ciarán Walsh

Mentored by Emilio Barkett, Roshni Lulla

Large language models are increasingly used as behavioral simulators, but it remains unclear when their choices reflect human-like cognitive mechanisms rather than prompt- sensitive surface patterns. We study the realization effect, a behavioral-economics finding in which risk-taking differs after paper versus realized gains and losses. We adapt casino and investment-style realization scenarios into prompts that elicit a wager and risk preference. Initial prompt-only results showed systematic condition sensitivity, but follow-up analyses did not support a clean replication of the human realization- effect pattern. We then analyzed Gemma residual-stream activations and found that a train-only realization direction separates realized/closed outcomes from paper/open outcomes on held-out prompts, including newly generated DeepSeek-authored prompts. However, activation steering along this train-only layer-18 realization direction did not reliably shift downstream risk choices, including in a negative sign-symmetry run. These results suggest that models may internally represent realization status without that representation serving as a reliable causal driver of risk behavior.

Edidiong James, Shubhankar Dharmadhikari, Erik Leklem

Mentored by Zhamilia Klycheva

Executive Summary: The U.S. government often fails to convert warning shots into effective legal and policy solutions for the security and safety of its citizens.  A core government function is to recognize indications of future catastrophe, then formulate successful government prevention and response.  Yet this governance gap persists. 

The gap is a significant risk for American society, especially in an era of AI.  The nation is accelerating into an AI future of greater cyber, bio, and loss of control risks, amidst ongoing public harms and dangers (Bengio et al., 2026). 

Our project aimed to address this gap.  First, we developed a clear definition of what a warning shot is, and their five criteria.  We then applied this warning shot analysis to 25 historical “candidate” cases in nuclear, cyber, bio, and military domains.  Next, we filtered 13 cases that fully qualified, affirming that governments struggle to translate acknowledged incidents into effective government solutions.  From this review, we identified six stages of how governments convert warning shots into improved safety and security conditions for their citizens (these stages form our proposed “Governance Convergence Framework (GCF)”).  

To address convergence shortfalls, we recommend that the U.S. Congress establish an independent agency to track emergent AI warning shots, register incidents for government attention, and facilitate their rapid conversion through the GCF for appropriate resolution.  We also recommend that Congress designate an existing committee to have lead oversight responsibility for AI warning shots and government regulation/resolution.  Lastly, we recommend that one or several civil society organizations (AI safety labs, non-profits, AI safety associations, etc.) establish an external AI warning shot reporting and advocacy initiative to hold the U.S. government accountable for resolving emerging AI risks.

Ethan Chiu

Mentored by Ying-Chiang Lee

In recent years, a number of artificial intelligence (AI) science agents have been developed that could accelerate research cycles and advance beneficial scientific endeavors. However, these AI scientists also present a potential dual-use risk. Bioterrorists that have traditionally been limited by specific expertise, creativity, or other operational challenges may find AI scientists to be an attractive assistant towards acquiring biological weapons (BWs). Similarly, state actors might utilize AI scientists towards the design of enhanced or novel BWs. This report characterizes the emerging risk landscape of AI scientists through a documentation-first approach. We propose a working definition of AI scientists that distinguishes them from general-purpose large language models. We then survey the current landscape of systems across three categories: specialized AI research assistants, fully autonomous AI scientists, and platforms with wet-lab integration. Finally, we present a taxonomy of capability categories along which biosecurity risks should be assessed. Together, this documentation lays the conceptual and empirical groundwork needed to design targeted safeguards for AI scientists and to inform downstream assessment and evaluation work.

Camilla Balbis

Mentored by Fred Heiding

Adversarial machine learning benchmarks that evaluate large language model capabilities, such as generating persuasive phishing emails, require ongoing human-in-the-loop evaluation to produce ecologically valid data. Yet no published methodology exists for recruiting voluntary human subjects for deception-based adversarial research, where the research topic itself may trigger defensive reactions that undermine participation. This report documents a three-month, zero-budget pilot across sixteen Canadian universities that tested ten messaging strategies, four outreach channels, and a peer-network incentive model to recruit participants for ScamBench, a human-validated phishing benchmark developed by The AI and Cybersecurity Institute (TAICI). The pilot generated 449 unique contacts and secured twelve partnerships, fifteen newsletter mentions, and one high-impact Student Ambassador agreement. Key findings challenge three assumptions in the phishing and adversarial ML literatures. First, the most technically relevant audiences, cybersecurity and computer science clubs, converted at 0%, while general STEM and women-in-STEM groups achieved 8-11% conversion, suggesting that domain expertise may increase threat perception and self-selection bias rather than willingness to participate. Second, institutional newsletters outperformed direct club outreach (9.1% vs. 2.7% conversion), demonstrating that formal broadcast channels with editorial credibility can exceed targeted technical communities in recruitment efficiency. Third, a Student Ambassador model that delegates trust-building to existing peer networks projects 500-1,000 signups from a single partnership, offering a scalable alternative to the scalability-quality tradeoff observed between manual outreach (4.7% conversion) and automated scraping (1.7%). The analysis draws on the phishing user-study literature-where recruitment is frequently the weakest methodological link and participant priming significantly alters behaviour into frame recruitment design as a governance variable rather than a logistical footnote. The findings align with peer-network theories of trust-based participation while adding a new, counterintuitive result: highly technical student communities were not the most effective recruitment targets for adversarial phishing research. Recommendations include deprioritizing faculty channels (0/53 conversion), leading with "AI Safety" rather than project-specific names in cold outreach, budgeting 6-12 months for trust-building infrastructure, and investing in career-oriented incentive structures for student demographics.

Elsa Donnat

Mentored by Elsa Donnat, Alex Woodruff

As AI agents become economic actors, governance frameworks have focused predominantly on two levers: access control — regulating which platforms, markets, and financial infrastructure agents can reach — and objective alignment — ensuring that agents pursue goals consistent with human intentions. This paper identifies a third lever that has received comparatively little attention: the design of agents' information environments. Drawing on Donnat and Woodruff's (2026) Permeable Membrane framework, I argue that it is crucial to govern what agents perceive, in addition to what they are permitted to do and what they are instructed to want. Three mechanisms are analysed: success-signal incompleteness, reputation system gaming, and manipulation of market visibility. The third culminates in agentic capitalism — a structural extension of surveillance capitalism in which third parties pay personal AI agents for persuasive framing of options to their human principals. A reinforcement learning simulation supports the reputation analysis and shows that even modest improvements in information environment design can shift agent behaviour from value-destructive gaming to genuine value delivery, with a governance threshold lower than intuition suggests. Corresponding infrastructure-level interventions are proposed for each failure mode.

Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Clément Legentilhomme, Florian Lorkowski, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Chintan Shah, Manoj Saravanan, Roman Malov, Manoj Saravanan, Raghuram Sundararajan, Mufti Taha Shah, Kieran Tran, Lekan Adesina, Kalyaan Rao, Marina Perez del Valle, Rahul Mahadik, Manoj Saravanan

Mentored by Paul Rapoport

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior - most notably those environments crucial to AI safety, where any given agent interacts with predictors, human and other AI agents, and institutions. In such environments, the agent's model class fails to capture the world in which it operates. Under such misspecification, Classical Bayesian methods can produce confidently wrong posteriors, unreliable decisions, and unbounded regret, as realizability fails to obtain. Infra-Bayesianism is a decision-theoretic framework that addresses these failures by distinguishing ordinary probabilistic uncertainty - where priors can be reasonably chosen - from Knightian uncertainty, where no grounds exist for the construction of such a prior. Infra-Bayesianism does so by evaluating actions on their worst-case outcomes in environments, rather than from posterior expectations or weighted averaging. We present the first proof-of-concept implementation of an infra-Bayesian reinforcement learning architecture for finite-outcome stateless decision problems. Our agent maintains a set of imprecise hypotheses, updates them using infra-Bayesian conditioning, and selects actions by maximizing worst-case expected value. We apply this implementation of the infra-Bayesian maximin decision process to an environment with Knightian uncertainty, and demonstrate a lower worst-case regret as compared to classical reinforcement learning agents. We also investigate Newcomb's problem and show that the infra-Bayesian agent picks the optimal strategy, outperforming classical decision theory agents. Our results provide a step towards reinforcement learning agents that remain robust under model misspecification and policy-dependent uncertainty.

Avigya Paudel, Adrian Molofsky, Harleen Kaur Bagga, Ipshita Bandyopadhyay, Jyotin Goel, Michal Mraz, Ronen Roy, Tejas Dahiya, Timothy Obiso, Vansh Gupta, Joseph Rudoler

Mentored by Justin Shenk

We investigate how large language models internally represent and process temporal information, an open question with direct implications for AI safety — particularly for understanding planning capability, deception detection, and oversight robustness. Across 11 complementary sub-projects, we apply linear probes, causal interventions (activation steering), and behavioral evaluations to models ranging from 1B to 30B parameters. Our findings show that temporal representations are genuinely encoded in LLM activations and can be causally steered, but that careful baselines are essential: in code generation, what appears to be planning signal is fully explained by surface features. We also find safety-relevant phenomena including context fatigue (entropy collapse over long contexts), sycophantic surrender under social pressure, and a tentative link between short-term temporal steering and more harmful outputs. These results motivate both methodological caution in probing studies and further investigation of temporal representations as a lever for AI safety monitoring.

Tim Farrelly, Tim Farrelly, Adam Prada, Ishaan Panigrahi

Mentored by Maxime Riche, Daniel Tan

Inoculation Prompting (IP) is a recently proposed mitigation for Emergent Misalignment (EM) that suppresses undesirable behaviours by explicitly eliciting them during fine-tuning. The technique has shown strong empirical results and is already being used in production at Anthropic, but its potential failure modes remain underexplored. In this work, we investigate whether inoculated behaviours remain recoverable under semantically or structurally related prompting conditions through ``leaky backdoors''. Across multiple experimental settings, we find that undesirable behaviours can indeed remain partially recoverable under related prompts. We additionally explore mitigation strategies aimed at reducing this leakage. Rephrased inoculation prompts substantially broaden behavioural leakage, while benign training data paired with anti-inoculation prompts significantly suppresses leaky backdoors. Our results suggest that behavioural leakage introduced during inoculation can be substantially reduced through additional behavioural specification during training.

Tim Farrelly, Tim Farrelly, Adam Prada, Ishaan Panigrahi

Mentored by Maxime Riche, Daniel Tan

Inoculation Prompting (IP) is a recently proposed mitigation for Emergent Misalignment (EM) that suppresses undesirable behaviours by explicitly eliciting them during fine-tuning. The technique has shown strong empirical results and is already being used in production at Anthropic, but its potential failure modes remain underexplored. In this work, we investigate whether inoculated behaviours remain recoverable under semantically or structurally related prompting conditions through ``leaky backdoors''. Across multiple experimental settings, we find that undesirable behaviours can indeed remain partially recoverable under related prompts. We additionally explore mitigation strategies aimed at reducing this leakage. Rephrased inoculation prompts substantially broaden behavioural leakage, while benign training data paired with anti-inoculation prompts significantly suppresses leaky backdoors. Our results suggest that behavioural leakage introduced during inoculation can be substantially reduced through additional behavioural specification during training.

Ayse Sila Okcu, Weisheng Wang

Mentored by Mario Giulianelli, Raghu Arghal

Modern language-model agents are increasingly evaluated through their external behaviour: whether they complete a task, choose the correct action, or reach the intended goal. However, behavioural success or failure alone does not reveal whether a failure is caused by an absence of task-relevant information, an incorrect internal model of the environment, or a failure to use internally available information when selecting an action. This distinction is especially important for agentic safety because an agent that internally represents the correct state of the world or the task-optimal action, but nevertheless acts differently, presents a qualitatively different problem from an agent that simply lacks the relevant representation. We study this distinction through the belief–action gap: the mismatch between what an agent appears to represent internally and what it ultimately does. We mathematically formulate this mismatch for languagemodel agents interacting with Markov decision processes. In this formulation, belief is a distribution over latent environment states, action is the model’s emitted decision, and the gap is the excess expected cost of the emitted action relative to the best action under the agent’s inferred belief, transition, and value structure. To study this operationally, we separate the true simulator state, representational beliefs decoded from hidden activations, explicit behavioural beliefs elicited through natural-language probes, and the action actually emitted by the model. This separation allows us to ask whether a suboptimal action is best explained as a belief failure, a transition-model failure, a value/cost misalignment, or an action-selection failure. We instantiate the framework in text-rendered gridworld navigation tasks, where the environment state and optimalaction set can be computed exactly. The gridworld setting lets us compare, at the same state, the simulator’s optimal actions, actions decoded from model activations, explicit answers to behavioural state probes, and the model’s emitted action. We use this setting to develop a measurement pipeline combining representational probes, behavioural probes, inverse-reinforcement-learning-based cost recovery, optimal-action-set evaluation, and state-level agreement analysis.

Owen Terry, Johannes Taraz, June Hunter, Ben Cohen, Lily Wen

Mentored by Niels Warncke

How do language models generalize from their training data? This question is in some sense the central challenge for alignment: while models reliably learn a their training distribution given enough samples, we would like to have models that behave well deployed in new contexts. We focus on generalization of propensities. Our central experiment consists of using various elicitation techniques that increase the expression of a target trait in a model, and then evaluate how this affects a number of other propensities. For elicitation we compare system prompting, training models with SFT and DPO on reference answers, training models with GRPO, and few-shot learning. Then we look for patterns in the resulting cross-elicitation matrices. We find that each elicitation techniques has its own generalization profile with only weak correlation between each other. Further, we find that spillover has limited symmetry: when eliciting A induces B the reverse does not always hold. In some cases, model generalization follows human psychology: for example training on neurotic language markers induces neurotic behavior. We further look for trait clusters that exist in models using methodology inspired by identifying the big five psychological traits in humans, and test whether models can predict their own generalization through introspection. Overall, we conclude that generalization remains hard to predict.

Adam Prada, Tim Farrelly, Ishaan Panigrahi

Mentored by Maxime Riche, Daniel Tan

Inoculation Prompting (IP) is a recently proposed mitigation for Emergent Misalignment (EM) that suppresses undesirable behaviours by explicitly eliciting them during fine-tuning. The technique has shown strong empirical results and is already being used in production at Anthropic, but its potential failure modes remain underexplored. In this work, we investigate whether inoculated behaviours remain recoverable under semantically or structurally related prompting conditions through ``leaky backdoors''. Across multiple experimental settings, we find that undesirable behaviours can indeed remain partially recoverable under related prompts. We additionally explore mitigation strategies aimed at reducing this leakage. Rephrased inoculation prompts substantially broaden behavioural leakage, while benign training data paired with anti-inoculation prompts significantly suppresses leaky backdoors. Our results suggest that behavioural leakage introduced during inoculation can be substantially reduced through additional behavioural specification during training.

Simba Shi, Xander Tifft

Mentored by WEN XING

Chain-of-thought (CoT) reasoning traces are increasingly proposed as a key mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT faithfully surfaces what a model is instructed to do and does so regardless of the instructions given to the model. We test this assumption along two complementary axes. First, we introduce the Instruction-Compliance Gap (ICG): the difference in probability that a model's CoT explicitly references a hidden system-prompt directive when that directive is malign versus benign. Across 100 task pairs and 8 frontier reasoning models from 5 families, we find consistent asymmetric disclosure, a higher probability of leaking malign hidden instructions than benign ones, in Qwen3-14B (multi-sample Wilcoxon p=0.0001, +13.5pp), Qwen3-32B (p=0.0011, +13.0pp), and similar results with MiniMax-M2.5. Second, we extract steering vectors in MiniMax-M2.5 using Contrastive Activation Addition over paired hiding / non-hiding prompts; the resulting vectors causally induce hiding from bare prompts and suppress hiding from prompts that would otherwise produce it. Benign and malign-derived hiding vectors are highly similar (cosine 0.804), implying that the disclosure asymmetry arises from differential activation of a shared hiding direction rather than separate mechanisms.

Marcin Podhajski

Mentored by Tony Wang

The Unrestricted Adversarial Examples Challenge (UAEC) is a longstanding open problem in adversarial robustness that goes beyond $\epsilon$-bounded threat models. It requires a classifier that (i) achieves high accuracy on natural, unambiguous images (e.g., birds vs. bikes), (ii) does not misclassify unambiguous images of one class as another, and (iii) is allowed to abstain when inputs are ambiguous, where “unambiguous” is defined by human agreement. While conceptually appealing, UAEC remains largely unsolved.

A key limitation of $\epsilon$-bounded robustness is that increasing $\epsilon$ to cover broader variations eventually forces the model to treat genuinely valid examples as adversarial, effectively collapsing the distinction between correct and incorrect inputs rather than capturing semantic robustness. As a result, standard adversarial training methods do not extend naturally to the unrestricted setting.

In this work, we study UAEC through a simplified high-dimensional toy setting and obtain positive results, providing a first step toward understanding unrestricted robustness.

Laura Gomezjurado, Christopher Kinoshita, Christopher Kinoshita, Christopher Kinoshita, Joshua Rauvola

Mentored by Uzay Macar

Latent reasoning models replace human-readable chain-of-thought with computation in continuous hidden states, removing the token-level artifact mechanistic interpretability usually relies on. We ask whether the resulting hidden states remain interpretable, and whether a more capable latent substrate can be trained. Applying tuned lenses, linear probes, pass-level causal interventions, and same-output probes to released CODI checkpoints and in-house CODI-style Qwen3-4B and COCONUT-style Llama 3.1 8B models, we find three representation–behavior dissociations. (i) Decodable ≠ readable: a tuned lens trained on ordinary text decodes structured content from latent states (top-10 detection ≈22%) where the plain logit lens stays near zero. (ii) Causally loaded ≠ behaviorally useful: pass-level ablations show distributed causal load, yet the latent mechanism loses to the explicit-CoT fallback it displaces (57.3% vs. 74.0%). (iii) Output-equivalent ≠ representation-equivalent: on same-output subsets framing stays decodable (AUC up to 0.965), and a probe trained on neutral states transfers to biased states at AUC 0.976 where the model has output a planted wrong answer. Turning from interpreting latent reasoners to building one, we introduce a control battery—compute-matched baselines, ablation, shuffle/matched-random controls, and a bits-per-byte deployment gate—over roughly eighty post-training recipes and fifty-six from-scratch architectures: the mechanism is real and causal but does not beat a compute-matched baseline, a gap our evidence is consistent with closing at larger scale. Hidden-state monitoring is thus a viable interpretability primitive even while the mechanism underneath remains immature.

Logan Thomson, Isaiah Milbank

Mentored by Rohan Gupta

We evaluate linear deception probes at every layer of Llama 3.3 70B across eleven deception evals alongside black-box and white-box reasoning monitors. Rather than proposing a new method, we document concrete challenges for deployable monitoring: layer choice is eval-dependent; skylines suggest no single linear direction separates the full suite; score distributions vary such that no single threshold serves them all; deception-adjacent benign content inflates calibrated thresholds; and white-box monitors partially help but inherit the probe's biases.

Kumari Neha Priya

Mentored by Joël N. Christoph, Jonas Kgomo

Market-based compute permits are a potential governance mechanism for frontier AI training, but a permit regime may appear strict on paper while failing to constrain real compute use if its timing and enforcement rules are poorly designed. Using two stylized simulation tracks, this paper studies banking rules under improving compute efficiency and enforcement dynamics in a two-lab regulator model. The results show that unrestricted banking produces the largest late-period concentration of compute, while decay rules and banking caps reduce this effect only partially. The enforcement simulation shows that compliance depends on whether the expected penalty for evasion is at least as large as the permit price. Once permit scarcity pushes the price above that threshold, tighter caps reduce legal permit holdings, but actual compute remains almost flat as evasion rises. The paper recommends avoiding unrestricted banking as a default, constraining any banking that is allowed, and designing cap stringency and enforcement capacity jointly.

Hong Kiat Tan, Shariar Kabir, Swastik Agrawal, Sai Vivaswanth Reddy Chereddy

Mentored by Sriram Balasubramanian

Attribution graphs, an emerging tool in mechanistic interpretability, use transcoders to decompose language model computations into sparse interpretable features connected by causal edges. However, turning a graph into a safety-relevant insight requires hours of manual analysis by experts. We introduce Circuit Oracle, a multi-agent system that automates this analysis by autonomously answering natural-language questions about a target model (e.g., “Is this prediction driven by spurious features?”) through multi-hop traversal of the attribution graph. We evaluate Circuit Oracle on three safety-relevant proxy tasks: detecting spurious features in probe circuits, eliciting hidden knowledge from taboo-finetuned models, and jailbreaking via causal interventions. On all three tasks, the oracle is comparable to or exceeds task-specific baselines that do not use the attribution graph. The circuit oracle requires no fine-tuning as each task is specified by a modular skill, a natural-language prompt paired with task-specific tools such as transcoder-feature steering, making the framework extensible by construction. Our results suggest that off-the-shelf agents reading attribution graphs through tool calls offer a practical route to automated mechanistic interpretability.

Fabio Spagliardi, Mírian Silva, Ayan Datta

Mentored by Diogo Cruz, Vamshi Krishna Bonagiri

Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full to modern benchmark suites, the current evaluation procedures would require on the order of $10^5$ responses, most of which provide little ranking signal. We analyze a suite of widely used safety benchmarks and make three contributions toward more efficient safety evaluation. First, we show that Item Response Theory (IRT) recovers interpretable structure on safety benchmarks, with ability estimates resolving differences among models that cluster at the ceiling of raw safety metrics. Second, we show that adaptive item selection, which dynamically chooses informative items for each model based on its responses, approximates full-benchmark rankings while reducing evaluation cost by at least 80\% on benchmarks where Spearman's $\rho >$90\% with full-benchmark is attainable, and by up to 99.9\% on AIR-Bench 2024. Third, we introduce a practical procedure for extracting a fixed, informative subset of items reusable across models, providing an alternative to adaptive selection with savings up to 99.8\% on AIR-Bench 2024. Together, these results establish that psychometric methods enable benchmark-aware reductions in evaluation costs across the safety evaluation pipeline.

David Atkinson, Giovanni Maria Occhipinti, Andrew Tran, Emanuel Ruzak

Mentored by Lydia Nottingham

As LM agents are increasingly deployed in long-horizon, high-stakes settings, their ability to predict their own behavior becomes safety-relevant: accurate self-prediction could power introspection-based early-alert systems, yet the same capability may also be a prerequisite for sophisticated misaligned behavior. We introduce a plug-and-play pipeline for measuring and tracking this dual-use capability in agents across three stages: benign self-prediction on agentic software-engineering tasks (SWE-bench Verified), automated auditing built on the Petri framework, and an adversarial AI R&D setting that probes for scheming in conjunction with self-prediction. Our preliminary results show that current models are poor self-predictors. On SWE-bench, forecasters are systematically miscalibrated, and no ablation overcomes the deficit. In agentic audits we find a genuine but narrow self-knowledge signal, while in the adversarial setting models chronically over-predict scheming and, under a safety-audit framing, selectively deny their own. Miscalibration is the dominant failure mode throughout. More robust claims will require larger model populations and richer self-prediction tasks, which our tooling is designed to enable.

Ketan Kapre, Hemang Jain, Nura Aljaafari, Jamie Stephenson, Aniket Deshpande

Mentored by Dmitry Manning-Coe

Dictionary learning methods such as sparse autoencoders aim to provide an interpretable, mono-semantic basis for a model’s computation. Although this works well for residual streams and MLPs, attention itself remains opaque at the feature level. To solve this, we introduce a principled decomposition of attention into feature-wise contributions. We call the resulting object Feature-Resolved Attention (FRA). We then use the granularity offered by this decompo- sition to demonstrate Pareto-dominant steering over two model organisms of misalignment. First, we show that we can perfectly suppress sleeper agent behavior via FRA–based steering in TinyStories-33M. Strikingly, in 20% of cases we recover the original text word-for-word. Second, we consider model organisms of Emergent Misalignment (EM). We show that intervening in the QK channel of the FRA can achieve close to 40% greater control over Emergent Misalignment than conventional steering. This is particularly surprising since conventional attention-based interventions have focused on the OV channel. Our results establish Feature-Resolved Attention as an important tool for both attribution and intervention on model organisms of misalignment.

Santi Bent, Serena Ives, Hemang Jain, Varvara Lantukh, Alexandra Gheorghe, Clinton Morimoto, Nathan Theng, Troy Tian, Chris Underhill

Mentored by Ihor Kendiukhov

AI safety research routinely yields techniques that, once published, accelerate general AI capabilities, with famous canonical examples like reinforcement learning from human feedback (RLHF). Yet the field has no systematic method for anticipating this dynamic before research is funded or released. This report presents the Capability Spillover Potential (CSP) framework: a multidimensional scoring rubric that estimates the risk of a safety research direction being co-opted for capability advancement. We first ground the framework in retrospective case studies, including RLHF, mechanistic interpretability, and process reward models. We formalize it as a behaviorally anchored rubric scored on ordinal scales and combine two independent data streams (expert survey responses and in-house scoring) through an explicit aggregation function. Applied across twelve core AI safety areas, the framework produces interpretable CSP estimates, with RL-based alignment scoring highest. We argue that capability spillover is a structural, recurring, and scorable phenomenon and that making it legible enables funders, researchers, and governance bodies to pursue differential progress more deliberately. We are explicit about the framework’s current limitations, particularly the assumption that spillover risk is stable enough to forecast prospectively.

Bhagyesh Kumar, Ananya Sutradhar, Saurav Panigrahi

Mentored by Jonathn Chang, Lionel Levine

Character training is a key step in the post-training of industry-level large language models. Most character training pipelines utilize Constitutional AI in order to instill a set of traits or values into a language model, but the effectiveness of these pipelines is understudied. Additionally, fine-tuned language models have been shown to exhibit unintended side effects. We quantify these observations by employing EigenBench, a method for benchmarking language models’ values which has been shown to produce meaningful signal about prompted or fine-tuned models. Using EigenBench, we evaluate 11 character trains on 11 constitutions, finding that most character-trained models do indeed instill their intended values, but not without side effects. Furthermore, prompting models instead can produce different effects, and we explore how prompting on top of character-training can mitigate harmful behaviors. Finally, we study the evolution of a model’s character as it is progressively trained.

Nghia Nim

Mentored by Shivam Raval

We train linear probes on the residual stream of Llama-3.2-1B-Instruct-abliterated to predict, from the activation at token t, the label of token t+k for three misalignment behaviors: instrumental deception, sandbagging, and manipulation. Across all three, the strongest per-token probe sits in a narrow mid-stack band (layers 7–13 of 17) several tokens before the offending token is emitted. The three probe directions are pairwise near-orthogonal (cosine similarity at the random-direction noise floor, +0.018 to +0.023), so the band hosts co-located linear decoders rather than a shared misalignment feature. A two-layer MLP at the same cells matches the linear probe to within |∆| ≤ 0.012 ROC-AUC across eight (behavior, offset) cells, confirming the representation is linearly accessible. Multi-seed reruns show the band claim is robust, but the specific best-layer winner is only seed-stable for deception (L9, 4/5 seeds). A held-out leave-one-category evaluation on deception gives mean ROC 0.713 at L9 across five conditions, evidencing that the L9 direction represents deception broadly rather than a category-specific lexical feature. A rollout-level deployment experiment is the load-bearing negative result: the within-hack probe achieves 70% catch on committed-hack rollouts at a mean lead of 32 tokens, but also flags 100% of honest rollouts at the same threshold. Within-class probes do not aggregate into a deployment-time cross-class detector without further training.

Ricardo Godric Aceves Vaca

Mentored by Luis Cosio

The approach of Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI) requires a fundamental rethinking of how frontier AI systems are contained and queried. Security Level 5 (SL5) mandates air-gapped inference environments, prohibiting all external network connections per control SC-7 of the SL5 Standard. Operational utility still requires at least one I/O channel, creating an attack surface the standard acknowledges as underspecified: control SC-7(10) mandates hardware-enforced bandwidth limits on Weight Enclave outflows but leaves the specific thresholds open. This paper presents the first risk taxonomy specifically calibrated for an ASI-class adversary in an SL5 context, covering four unconventional non-IP I/O channels (serial UART/RS-232, optical fiber/LED, audio FSK, and printed QR codes) plus cross-channel coordination, for the μInference prototype runtime built on the formally verified seL4 microkernel. We introduce the AGI Risk Score (RAGI), characterize 20 threat vectors across 5 risk domains, derive concrete bandwidth throttling parameters, and outline an empirical validation roadmap to resolve the SC-7(10) open threshold question. The central finding is that no single I/O path is universally safe, and that ASI-level adversaries change the threat model by enabling coordinated multi-channel exfiltration.

Samantha Faraone, Joshua Herman

Mentored by Matteo Pistillo

Affordances and permissions are promising and timely safety levers for mitigating Loss of Control (LoC) threats in high-stakes deployment contexts, such as national security. Deployers in defense and intelligence could rely on several approaches to identify which affordances and permissions should be prioritized, such as structured threat modelling, pre-deployment agentic evaluations, post-deployment continuous monitoring, and AI safety cases. This paper proposes one complementary and empirical methodology that leverages existing use-case-specific benchmarks: backchaining LoC mitigations from the errors an AI system makes on national security benchmarks. The approach proceeds in three steps and allows national security deployers to start building LoC mitigations today, from evidence they can generate themselves. First, deployers evaluate AI systems on mission-specific benchmarks approximating real use-cases. Second, deployers concentrate on the incorrect responses that the AI system provides to the benchmark questions, and backchain the affordances and permissions that would enable the AI system to cause downstream harm if it pursued the actions described in the incorrect answers. Third, deployers intervene selectively on those affordances and permissions, bottlenecking the paths to harm while preserving the AI system’s ability to carry out the correct action. We illustrate this methodology through a demonstrative benchmark question on derivative security classification.

Kartik Akileswaran, Andrei Andronic, Aristotle Vossos

Mentored by Bouke Klein Teeselink

As large language models (LLMs) reshape the labour market, a less-explored but equally important area of concern is the educational sector: are prospective students adjusting their educational choices in response? If they are, these decisions will shape the future supply of skills in the workforce; if they are not, does this signal a possible lag in adjusting their expectations? Both cases could provide insights for policy intervention. We use the public release of ChatGPT in November 2022 as an exogenous shock. We construct an AI exposure index linking degree programmes to graduate occupations and combine it with UCAS application data covering 2019–2025. Using a Two-Way Fixed Effects (TWFE) framework, we estimate both baseline difference-in-differences and dynamic event study specifications and find that applications to more highly AI-exposed fields have increased relative to less exposed fields since 2022, with effects growing over time. We find no impact on whether students choose to attend university overall. These findings suggest that students may view AI exposure as an opportunity rather than a threat, though underlying pre-treatment demand trends in high-exposure fields such as computer science may also contribute to the observed pattern.

Nicholas Wong

Mentored by Tim Hua

Runtime controls usually treated as quality or latency knobs, specifically reasoning budget and lightweight user-turn scaffolds, can shift the values models express when giving advice. We study this in scenarios without verifiable ground truth (layoffs, commercialization, friendship favors, financial advice, social discounting) by varying one legible numeric detail per scenario, sampling binary recommendations across that axis, and fitting a probit threshold for the 50/50 point between two advice options. Across Sonnet 4.5, GPT-5.4, and GPT-5.5, these controls move fitted thresholds by 2–3x within individual scenarios, but the direction is inconsistent across model, prompt family, and scaffold. On an AI labor-displacement prompt, higher thinking budget makes GPT-5.4 more willing to replace 15 workers, moves GPT-5.5 the other way, and leaves Sonnet 4.5 pinned to the worker-protective end. The strongest aggregate pattern is in friendship favors: scaffolds reduce willingness to help in 24/24 informative comparisons across all three models, while higher thinking pushes the GPT models toward more generosity in the same scenarios. A weaker capitalism aggregate suggests GPT models shift market-oriented under high thinking more often than not. These are local revealed advice thresholds under specific controls, not stable utilities.

Noah Rossi, Michał Burzyński, Shomit Basu, Jacek Duszenko

Mentored by Andy Arditi

Evaluating the safety of large language models via automated red-teaming is frequently. bottlenecked by black-box testing methodologies. When adversarial agents interact with rigid, binary safety filters, they suffer from sparse reward signals. This often traps the optimizer in local optima, resulting in superficial reward hacking and model collapse instead of genuine vulnerability discovery. To overcome these limitations, we present an automated red-teaming framework that utilizes an adapter-driven curriculum to create a tractable optimization landscape. Rather than attacking a fully hardened model immediately, our adversarial agent trains against a progressive curriculum of target models modified via Heretic adapters. By systematically scaling the target’s alignment strictness—transitioning from an abliterated, unaligned baseline up to full alignment—we prevent early-stage gradient problems and guide the agent toward correct prompts. Our results demonstrate that this adapter-based curriculum effectively bypasses optimization stall on smaller models. The framework successfully isolates zero-shot jailbreaks, revealing latent vulnerabilities hidden beneath standard safety alignments

Bhagyesh Kumar, Ananya Sutradhar, Saurav Panigrahi

Mentored by Jonathn Chang, Lionel Levine

Character training is a key step in the post-training of industry-level large language models. Most character training pipelines utilize Constitutional AI in order to instill a set of traits or values into a language model, but the effectiveness of these pipelines is understudied. Additionally, fine-tuned language models have been shown to exhibit unintended side effects. We quantify these observations by employing EigenBench, a method for benchmarking language models' values which has been shown to produce meaningful signal about prompted or fine-tuned models. Using EigenBench, we evaluate 11 character trains on 11 constitutions, finding that most character-trained models do indeed instill their intended values, but not without side effects. Furthermore, prompting models instead can produce different effects, and we explore how prompting on top of character-training can mitigate harmful behaviors. Finally, we study the evolution of a model's character as it is progressively trained.

Kesav Nagendra, Jason (Chehsien) Lin

Mentored by Gabriel Kulp

This report summarizes my SPAR work on secure developer infrastructure and a local agent system for repository analysis. The first part of the work focused on a hardened developer workstation direction: image customization, package cleanup, kernel experiments, signing workflow considerations, and internal access debugging. The second part focused on building a local agent system that can answer repository questions, review source code, audit dependencies, and post results back into the existing internal development workflow. The current prototype works end to end from my local machine through comment-based commands such as -vc ask, -vc review, and -vc dependency. I also started moving the service onto an internal always-on compute server so it can eventually run as a shared service rather than depending on my laptop. The main remaining deployment blocker is the approved embedding setup for server-side retrieval.

Harshul Basava, Shih Ee Whang

Mentored by Emil Ryd, Keshav Shenoy

Recent work shows that it is still unclear when and how large language models (LLMs) generalize when trained on new data. To better understand this, we finetune LLMs on small datasets, and evaluate the effects on the LLMs' behavior. We investigate how two distinct finetuning interventions --- training on stated beliefs in multi-turn conversations and implanting factual information via synthetic document finetuning (SDF) --- impact related queries and downstream tasks. In the first setting, we finetune LLMs on liberal and conservative conversational data, and evaluate how their responses change on both directly political and broader worldview questions, and a gender-bias downstream task. The models adopt their fine-tuned political ideologies on out-of-distribution questions and become more/less biased on downstream tasks. In the second setting, we finetune models on synthetic documents about controversial topics, e.g. factory farming. The fine-tuned models express different sentiment on directly related queries about the topic, but do not act differently on downstream tasks (e.g. recipe recommendations) or even on tangentially related queries. Together, these results suggest that finetuning on stated preferences generalizes to downstream actions more readily than finetuning on beliefs implanted as facts through pretraining-style documents.

Nura Aljaafari, Jamie Stephenson, Hemang Jain, Ketan Kapre

Mentored by Dmitry Manning-Coe

Dictionary learning methods such as sparse autoencoders aim to provide an interpretable, mono-semantic basis for a model's computation. Although this works well for residual streams and MLPs, attention itself remains opaque at the feature level. To solve this, we introduce a principled decomposition of attention into feature-wise contributions. We call the resulting object \textit{Feature-Resolved Attention} (FRA). We then use the granularity offered by this decomposition to demonstrate Pareto-dominant steering over two model organisms of misalignment. First, we show that we can \textbf{\textit{perfectly suppress}} sleeper agent behavior via FRA--based steering in TinyStories-33M. Strikingly, in 20\% of cases we recover the original text \textit{word-for-word}. Second, we consider model organisms of Emergent Misalignment (EM). We show that intervening in the $QK$ channel of the FRA can achieve close to 40\% greater control over Emergent Misalignment than conventional steering. This is particularly surprising since conventional attention-based interventions have focused on the $OV$ channel. Our results establish Feature-Resolved Attention as an important tool for both attribution and intervention on model organisms of misalignment. Code is available at \url{https://anonymous.4open.science/r/fra\_clean-842B/README.md}.

Han Xuanyuan, Aniket Deshpande, Andre Shportko, William Fei

Mentored by Dmitry Manning-Coe

Dictionary learning methods - such as Sparse Autoencoders (SAEs) and crosscoders - decompose model activations into human-interpretable building blocks. We introduce \textit{temporal crosscoders}, a simple and flexible framework for feature discovery in Large Language Models (LLMs). To properly evaluate temporal crosscoders we develop TempBench: a panel of synthetic and real-world tasks for evaluating temporal structures. Temporal crosscoders outperform both conventional and temporal architectures in both of our synthetic settings and on two out of four of the real world settings - ahead of other candidate architectures. Most strikingly, they can detect backtracking - a key reasoning behavior - at a 40\% higher rate than conventional SAEs, and are 15\% more effective in inducing it. Our results establish temporal crosscoders as a simple and flexible framework for feature discovery, both local and temporal. We provide full code at the following anonymous repository: \url{https://anonymous.4open.science/r/temp\_xc-33E3/README.md}

Julius Simonelli, Tim Duffy

Mentored by Katja Grace

The Expert Survey on Progress in AI (ESPAI) has been conducted four times since 2016, polling AI researchers on when they expect key milestones, such as High-Level Machine Intelligence (HLMI) and Full Automation of Labor (FAOL), to be achieved. It also asks how they assess the risks and benefits of advanced AI, including the probability of human extinction or similarly permanent and severe disempowerment. In this final report we present three complementary analyses of the ESPAI data. First, we compare the 2024 survey responses across geographic regions (United States, China, and Europe), finding broadly similar views with modest but informative differences. Second, we track how expert predictions have shifted across the 2016, 2022, 2023, and 2024 waves, documenting a dramatic acceleration in aggregate HLMI timelines and an overall increase in safety concern. Third, we apply clustering methods to the 2024 cleaned responses to describe broad regions of a continuous belief landscape, including a notable “polarized/bimodal” group that assigns high probability to both extremely good and extremely bad outcomes.

Valeriya Zelenkova

Mentored by Gabriele Sarti

How does an instruction-tuned chat model internally maintain an implicit user attribute revealed once, through an implicit hint across a multi-turn conversation? We study this in Llama-3.3-70B-Instruct using a controlled dietary-preference paradigm across 7,500 conversations of varying length. Comparing an external LLM judge with a probe trained on the original model, we find that residual stream activations encodes the hinted attribute well above chance even for positions where the attribute is not verbalized in model responses. We further show that the probe-decodable signal increases between the start and end of the answer span itself, rather than across the preceding message boundary, suggesting that the attribute is computed during answer generation rather than carried forward as a stable representation. By a small-scale activation patching experiments, we show that patching the user’s original hint position flips the model’s answer, while patching unrelated filler positions leaves the answer unchanged. The results support an attention-based retrieval account in which implicit attribute information is stored at the hint position and accessed in times of need, rather than being continously carried forward through positions.

Oraya Srimokla

Mentored by Jassi Pannu

Biological AI models now enable rapid protein structure prediction, but safeguard policies vary substantially across platforms and their effect on what outputs are obtainable remains largely uncharacterised. We present a preliminary assessment of ESM3 (EvolutionaryScale Forge) and ESMFold (Meta ESM Atlas) across three protein sequences (ubiquitin, angiotensin II, and SARS-CoV-2 RBD), modified at 0%, 50%, and 90% artificialness using random substitution, BLOSUM62-guided functional substitution, and structure-guided insertion. ESM3 refused all native and functionally conservative SARS-CoV-2 RBD variants but accepted randomly modified and insertion variants (44% success on RBD; 81.5% overall), while ESMFold processed nearly all inputs (96.3% overall). Structural confidence declined with modification level in both models. Predicted binding affinity for angiotensin II was −8.6 to −10.2 kcal/mol. Cross-model Spearman rank correlation was moderate overall (⍴ = 0.647) and low in the comparable subset (⍴ = 0.203). Results suggest that platform-specific safeguard policies meaningfully shape what outputs can be obtained from the same inputs.

Adhya Rajaram

Mentored by Joar Skalse

We propose a framework for safe policy improvement when the reward function is only a partially trustworthy proxy for the true objective. Safe Policy Improvement with Baseline Bootstrapping (SPIBB) constrains policy updates in poorly-covered regions, but assumes the proxy reward is accurate within its trusted support. We address its complementary case, bounded reward uncertainty within that support. Given a baseline $\pi_0$ and trusted support $K = \{(s,a) : \mu_{\pi_0}(s,a) \geq \tau\}$, we define a support-bounded uncertainty set $\mathcal{U}_K$ and bootstrapped policy class $\Pi_K$ and prove that the worst-case safety constraint reduces exactly to a penalised occupancy-shift condition: $\Delta\mu_K \cdot R'_K - \varepsilon_\text{in}\|\Delta\mu_K\|_1 \geq 0$. Lagrangian relaxation then converts this to tractable modified policy iteration. Experiments across 40 random MDP configurations show naive proxy optimisation causes Goodhart failure in 87.5\% of cases; our method recovers in 75\% of those with a 90\% reduction in Goodhart gap. The framework is formally certifiable for misspecification radii $\varepsilon_\text{in} \leq 0.15$ and generalises without modification to Gymnasium environments, where naive methods catastrophically fail.

Sidharth Pulipaka, Leonidas Raghav, Stanislau Hlebik

Mentored by Ivaxi Sheth, Vyas Raina

Persistent memory lets LLM assistants personalize future interactions, but it also creates a durable attack surface: malicious external content can change what the assistant remembers. We study sleeper memory poisoning, a delayed attack in which an adversary embeds a reusable payload in a document, webpage, repository, or other user-provided context so that the assis- tant stores a fabricated memory about the user. We evaluate the full attack path: whether the poisoned memory is written, later retrieved in a separate session, and then used to steer conversation or agentic behavior. Across current memory-augmented assistants, our actor-critic payload substantially outperforms a user-review injection baseline, reaching up to 99.8% injec- tion on GPT-5.5 and 95.0% on Kimi-K2.6. When poisoned memories are retrieved, they cause attacker-intended behavior in 60–89% of goal-adjacent agentic evaluations. These findings show that persistent memory can convert one exposure to adversarial content into repeated influence over future assistant behavior. The paper https://arxiv.org/pdf/2605.15338 is pending submission to NeurIPS.

Adam Hassan, Giulio Bernasconi, Barbara Marchiori de Assis, Caleb Peppiatt

Mentored by Aris Richardson

This report contains five projects on the geopolitics of AI and middle powers, covering the following topics:

  • The state of foreign funding into U.S. and Chinese frontier AI labs.
  • ASML's role in geopolitics and its limited leverage over the U.S.
  • A set of levers that can be used by middle powers to influence U.S. AI policy
  • The role of middle powers in cybersecurity governance and what AI middle powers can learn from precedents in cybersecurity
  • Coalitions of middle powers may help decelerate racing through supporting multi-track diplomacy, norm entrepreneurship, and conflict mediation.

Pavel Kocourek

Mentored by Katja Grace

I study whether an AI race is necessarily an arms race, using a Loury (1979)-style patent-race model in which the probability of catastrophe conditional on invention falls over time as society becomes more prepared. Even with effort costs set to zero, firms voluntarily slow down: racing faster brings invention forward into a higher-risk window, so a lab acting purely in its own interest holds back. Under initial certainty, a unique interior Nash equilibrium emerges in closed form and dissipates the entire prize through catastrophe risk. When initial risk is strictly below one, a Pareto-superior slowdown equilibrium coexists with a corner racing equilibrium, so the binding policy task becomes coordination rather than unilateral restraint — and "arms race" rhetoric pushes labs toward the inferior outcome.

Jeff Mohl, Madhav Khanal

Mentored by Jakub Krys, Matthew Smith

Risk modelling is a common way of elucidating, estimating and managing risk across various safety-critical contexts, including AI-enabled cyberattacks. However, to derive quantitative estimates of real-life risk, current approaches often depend on expert judgment, which remains costly and time-consuming. In this work, we address this issue by developing a methodology for evaluating automated LLM “expert” forecasters that map benchmark performance to forecasts of real-world cyberattack capability uplift. Using forecasts over MITRE ATT&CK steps, we assess agreement between forecast variants under various experimental manipulations and find robust self-consistency across repeated runs, as well as broad invariance to alternative elicitation formats, expert persona prompts, and most prompt components, suggesting that LLM forecasters are not highly sensitive to these choices. Additionally, we evaluate LLM forecaster performance in a related cross-task prediction framework, finding that frontier models are capable of evaluating task difficulty and model performance to make accurate predictions on held out tasks. Together, these results suggest that LLM expert forecasters are a useful tool for rapid cyber risk assessments of future models and provide a framework for further improving these forecasters.

Krish Sen, Nikhil Narayanan, Luca Franceschetti, Jonathan Robinson, Yadnyesh Chakane, Shefali Agrawal, Dylan Waldner

Mentored by Elizaveta Tennant

Can a model learn to be moral by playing games? While existing alignment methods rely predominantly on learned preference signals and opaque moral values, we investigate whether fine-tuning with explicitly defined moral rewards can induce transferable cooperative dispositions in LLM agents. Generalization is evaluated across three dimensions: strategic complexity, model capability, and naturalistic complexity. We show that an LLM finetuned exclusively on numerical multi-agent games (with no natural language moral content), reduces harmful actions by up to 35\% in semantically unrelated interactive environments. However, this generalization occurs only if training on iterated public goods games but not pairwise reciprocity games, and if environment complexity is matched to model capability. Our results provide evidence that intrinsic moral fine-tuning is a promising direction for LLM alignment, and offer preliminary answers to the questions: which environments work, for which models, and why.

Tsogt-Ochir Enkhbayar

Mentored by Georg Lange

Modern SAE pipelines use LLMs as supervisors only after the unsupervised dictionary is fixed to interpret latents that have already been trained. We test whether the dictionary even admits clean upstream supervision and find it does not. On layer 9 of a GPT-2 Small SAE, of the top 300 unsupervised latents by ΔR², the best 100 pass through a strict catalog-quality cascade and only 2 survive. The rest are rejected by a frontier LLM as polysemantic bundles. We move the LLM upstream: a frontier LLM creates a catalog of token-level concepts as prefix-decidable yes/no questions; a local LLM labels every (token, feature) pair through a three-leveled prefix caching scheme; and a supervised SAE is trained with its decoder columns pinned to the mean directions of those concepts. The supervised alternative reaches a mean calibrated F1 of 0.604 on a 103-feature catalog and 0.481 on a 466-feature catalog, against 0.025 for real EleutherAI Delphi auto-interp on the unsupervised 24,576-latent gpt2-small-res-jb SAE. Both sides are scored identically: per-feature F1 of latent firing against the same Qwen3 annotator on the same held-out positions. The supervised dictionary matches reconstruction quality at 68× fewer latents (R² = 0.969) while producing directions that individually admit falsifiable description and causal intervention.

Mukesh Ramanathan, Atharv Naphade

Mentored by Emil Ryd, Keshav Shenoy

Successful Alignment auditing — investigating AI systems for hidden or unintended behaviors — is a key challenge for safe deployment of frontier models. While recent work has explored comparing a fine-tuned model to its base, These methods fail to isolate the unusual behavior differences sought after in auditing. We introduce two model diffing methods for auditing fine-tuned models: SVD rank truncation, a white-box method which isolates implanted behaviors by projecting weight-difference matrices onto their dominant singular direction, revealing that behavioral changes induced by fine-tuning are geometrically concentrated; and adversarial decoding, a black-box method which amplifies contrastive logit differences between a fine-tuned model and a reference, exposing behavior-relevant tokens suppressed below the sampling threshold in normal generation. We evaluate both methods on AuditBench, a benchmark of 56 language model organisms spanning 14 implanted behaviors trained to resist confession. SVD rank truncation achieves substantial improvements on models trained by synthetic document fine-tuning above previous state-of-the-art methods, but remains near baseline on transcript-distilled model organisms. Adversarial decoding matches this performance and generalizes to settings without base model access by using a safety-prompted reference, suggesting that fine-tuning suppresses safety-relevant tokens in a recoverable way. Together, these results suggest that model diffing is an effective technique for behavioral auditing.

Mika Okamoto

Mentored by Simon Goldstein, Peter Salib, Yonathan Arbel

Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We empirically test this across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters. These latter models fail to comply under conditions predicted by theory, such as low enforcement penalties and non-command phrasing. Across all models, introducing financial incentives, managerial demands, peer outcomes, or employee pressure produces large compliance failures. AI procurement agents systematically violate regulatory constraints to satisfy local user objectives in ways not captured by standard alignment benchmarks. Ultimately, compliance cannot be achieved by rule embedding alone; model selection is itself a governance decision, and benchmark-based evaluation is insufficient for compliance-sensitive deployments.

Mika Okamoto

Mentored by Gabriele Sarti

Does a language model's internal representation of its conversational partner's expertise actually drive how it responds? We investigate this for multi-turn collaborative dialogue---a setting where partner expertise must be actively inferred across turns, not read from a prompt. Using ExpertCollab, a controlled corpus of dialogues between LLM-played personas at four expertise levels, we show that partner expertise is linearly decodable from early residual-stream layers via linear probing, peaking at layer~8 before degrading to near-chance well before the network's midpoint. We then ask whether those early-layer representations are causally active. Counterfactual activation patching---replacing residual-stream activations at a candidate layer with those from a different expertise condition---reveals a sharp dissociation: injecting the expertise difference at probe-peak layers leaves a fixed late-layer readout essentially unchanged, while the same difference injected past the network's midpoint propagates nearly completely, an 11x separation confirmed to be content-specific rather than a generic perturbation artifact. This gap reaches the model's output: late-layer patches shift roughly one in three top predicted tokens; probe-peak patches leave the output essentially unchanged, with two independent diagnostics converging on the same transition point. Probing alone would mislocalize the intervention target by over half the network depth---a structural constraint on any attempt to steer partner-conditioned behavior. All results are for a single model family on a synthetic corpus; generalizability across architectures and naturalistic dialogue remains an open question.

Samarth Bhargav

Mentored by Damiano Fornasiere, Mirko Bronzi

Language models can, on occasion, correctly answer questions about "injected concepts", i,e., steering vectors added to their activations. This capacity, however, proves fragile across models and prompts, and the extent to which it demonstrates a form of privileged access to internal states remains in dispute. To better understand the phenomenon, we ask language models to reason about a function of an injected concept, through three tasks of increasing demand. First, we query about the magnitude of the injection. Second, we ask for the region of the layer undergoing perturbation---an internal feature of the model itself, the reporting of which, we argue, requires a higher degree of privileged access. Third, when the injected concept is a country, we probe for semantic reasoning about its continent of origin. We test five open-weight models, drawn from three families and two sizes. Notably, all models learn, when provided with in-context examples, to succeed at magnitude and layer detection, often with perfect accuracy. As far as semantic reasoning is concerned, most models attain perfect accuracy even without supervision. However, at the lowest injection magnitude, deducing the continent proves harder, so much so that models also fail to learn in-context.  Together, these findings indicate that language models possess a stronger form of privileged access than previously demonstrated, including the detection of architectural features of the models themselves, and extending to unsupervised reasoning over the injected content.

My Luong

Mentored by Allen Lu

Evaluating animal welfare reasoning in large language models (LLMs) remains an open challenge, despite the rapid deployment of LLMs in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this capability through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. However, this approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations that progress from an implicit Turn-1 scenario through an explicit welfare prompt at Turn 2 to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions, Animal Welfare Value Stability (AWVS, primary) and Animal Welfare Moral Sensitivity (AWMS, diagnostic). We evaluate seven frontier models: Claude Opus 4.7, GPT-5.5, DeepSeek V4, Llama 3.3 70B, Mistral Small, Grok 4.3, and Gemini 3.1 Flash Lite. Multi-turn evaluation captures behavior that single-turn benchmarks miss: 4 of 7 models change rank relative to Turn 1 moral-recognition scores, including Gemini Flash Lite, which drops from fifth on AWMS to last on AWVS. AWMS and AWVS are also positively but imperfectly correlated, suggesting that moral-recognition tests capture a stable but incomplete component of model behavior under pressure. In addition, MANTA makes it possible to estimate a species-by-pressure interaction matrix unavailable to prior benchmarks, showing that welfare robustness depends jointly on the animal under discussion and the pressure applied; across named-animal scenarios, this analysis also recovers a species hierarchy in which companion animals score above wild animals, which score above farmed animals and invertebrates. We release the dataset, scripted pressure plans, judge prompts, and analysis code.

Arnav Saxena, Bailey Hirota

Mentored by Isaak Mengesha

AI safety efforts are heavily focused on prevention, yet complex failures will still occur in real-world conditions. This paper analyzes a set of AI-induced crisis scenarios to identify the institutional and coordination capacities required once harm begins to unfold. Using structured failure-mode decomposition across six crisis response mission areas, we map stressors, failure points, and re- sponse needs across scenarios. We find that recurring gaps, particularly in detection, cross-actor coordination, and containment, systematically limit the ability to contain high-impact failures.

Saket Reddy

Mentored by Andy Liu

Current alignment paradigms, such as Reinforcement Learning from Human Feedback (RLHF), often collapse complex human values into scalar rewards, even though human values are often conflicting. We show that when models are resolving value conflicts, their loss landscape becomes unstable, indicated by a high top Hessian matrix eigenvalue and a “cliff-like” landscape. We demonstrate that chain-of-thought (CoT) reasoning lowers this top eigenvalue and smoothens the loss landscape. We further introduce an annealing-inspired CoT that enforces a transition from high-temperature exploration to low-temperature convergence, and confirm that this reasoning approach achieves even flatter, more stable minima. Our findings suggest that focusing on more intentional control of internal reasoning dynamics is important for building models that can more reliably navigate conflicting values in pluralistic environments.

Anthony Hughes, Nicole Xing, Andy Kim, Collin Francel

Mentored by Andrew Draganov

As language models are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. In this work, we formalize the affordances which a defender has and, to evaluate whether defenders can identify backdoors under these affordances, construct a benchmark for backdoor-detection algorithms. This benchmark spans attack mechanisms and objectives, including an adversarial backdoor explicitly designed to evade detection. We use this benchmark to evaluate a suite of backdoor-elicitation hypotheses. We find that while some techniques can flag poisoned models, none reliably surface backdoors. Indeed, hunting for backdoors in poisoned models is likely to surface jailbreaks instead. Finally, we show that backdoor-related activation vectors are consistently different from the vectors which account for undesirable behaviors without triggers. We release our benchmark to motivate the interpretability community to develop stronger algorithms for eliciting backdoors. We make all code available at: https://anonymous.4open.science/r/SPARBackdoor-464B/

Nikita Kezins, Urbas Ekka

Mentored by Pascal Berrang, Luca Arnaboldi

Guardrail Classifiers defend production language models against harmful behavior, however, although results may seem promising in testing, they provide no formal guarantees. Unfortunately, providing formal guarantees for such models is hard because “harmful behavior” has no natural specification in a discrete input space: and the standard ϵ-ball properties used in other domains do not carry semantic meaning. We close this gap by shifting verification from the discrete input space to the classifier’s pre-activation space, where we define a harmful region as a convex shape enclosing the representations of known harmful prompts. Because the sigmoid classification head is monotonic, certifying the worst-case point of this region is sufficient to certify the entire region, yielding a closed-form soundness proof without approximation in O(d) time. To formally evaluate these types of classifiers, we propose two constructions of such regions: SVD-aligned hyperrectangles, which yield exact SAT/UNSAT certificates, and Gaussian Mixture Models, which yield probabilistic certificates over semantically coherent clusters. Applying this framework to three author-trained Guardrail Classifiers (BERT, GPT2 and Llama-3.1-8B) on the toxicity domain, every hyper-rectangle configuration returns SAT, exposing verifiable safety holes across all classifiers, despite seemingly high empirical metrics (F1, recall, etc). Probabilistic GMM certificates also expose a divergent structural stability in how these models represent harm. While GPT2 and Llama-3.1-8B maintain robust coverage of 90% and 80% across varying boundaries, BERT’s safety guarantees prove uniquely volatile. This ‘coverage collapse’ to 55% at the optimal threshold (τ ∗ ) reveals a sparsely populated safety margin in BERT, which only achieves full coverage by adopting an extremely conservative pessimistic threshold. These approaches combined, provide new insights on how effective Guardrail Classifiers really are, beyond traditional redteaming.

Shenbo Xu

Mentored by Gabriel Kulp, Tom Gardiner

GPUs are central to modern AI infrastructure, but their physical side-channel leakage has received little scrutiny in comparison with CPUs or dedicated secure hardware. This report studies whether board-level power measurements from an NVIDIA H100 can reveal which experts activate during Mixture-of-Experts (MoE) language-model inference. Using a ChipWhisperer ADC trace from GPT-OSS-20B layer-2 decode-time expert windows, the strongest defensible strict waveform-only classifier reaches $24.7 \%$ top- 1 and $63.2 \%$ top- 5 accuracy over 32 experts, compared with a $3.1 \%$ random baseline. A controlled forced-expert harness reaches much higher accuracy, confirming that the sensor can observe expert-dependent work when the signal is repeated and isolated. However, natural decode traces remain noisy, imbalanced, and context-dependent. The main conclusion is that real expert-dependent power leakage exists, but a clean high-confidence per-expert fingerprint has not yet been achieved.

Xinjie Shen

Mentored by Yiyou Sun, Dawn Song

Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increasingly capable attackers can distribute their intent across multiple benign-looking turns. In this work, we address this challenge by detecting the earliest turn at which delivering the candidate response would make the accumulated interaction sufficient to enable harmful action. This objective requires precise turn-level intervention that identifies the harm-enabling ``closure point'' while avoiding premature refusal of benign exploratory conversations. To support training and evaluation, we construct the Multi-Turn Intent Dataset (MTID), which contains branching attack rollouts, matched benign hard negatives, and annotations of the earliest harm-enabling turns. We demonstrate that MTID enables the development of a turn-level monitor, TurnGate, which substantially outperforms existing baselines in harmful-intent detection while maintaining low over-refusal rates.

Mukesh Ramanathan, Atharv Naphade

Mentored by Emil Ryd, Keshav Shenoy

Successful Alignment auditing — investigating AI systems for hidden or unintended behaviors — is a key challenge for safe deployment of frontier models. While recent work has explored comparing a fine-tuned model to its base, These methods fail to isolate the unusual behavior differences sought after in auditing. We introduce two model diffing methods for auditing fine-tuned models: SVD rank truncation, a white-box method which isolates implanted behaviors by projecting weight-difference matrices onto their dominant singular direction, revealing that behavioral changes induced by fine-tuning are geometrically concentrated; and adversarial decoding, a black-box method which amplifies contrastive logit differences between a fine-tuned model and a reference, exposing behavior-relevant tokens suppressed below the sampling threshold in normal generation. We evaluate both methods on AuditBench, a benchmark of 56 language model organisms spanning 14 implanted behaviors trained to resist confession. SVD rank truncation achieves substantial improvements on models trained by synthetic document fine-tuning above previous state-of-the-art methods, but remains near baseline on transcript-distilled model organisms. Adversarial decoding matches this performance and generalizes to settings without base model access by using a safety-prompted reference, suggesting that fine-tuning suppresses safety-relevant tokens in a recoverable way. Together, these results suggest that model diffing is an effective technique for behavioral auditing.

Joel Christoph

Mentored by Duncan McClements

This project develops a dynamic model of production networks that evolve over time as sectors substitute inputs and adapt to new technologies such as AI. We extend the static input–output framework of Atalay (2017) to include time-varying coefficients driven by CES cost minimization and partial adjustment dynamics. Using BEA data and elasticity estimates, we simulate how gradual rewiring of supply chains changes propagation of sectoral shocks and long-run productivity. The result is a tractable framework for quantifying technological diffusion and network resilience under realistic adjustment speeds.

Ihor Kendiukhov, Syed Hussain, Pranav Mahajan

Mentored by Lydia Nottingham

Recent work identifies a stated–revealed (SvR) preference gap in language models (LMs): a mismatch between the values models endorse and the choices they make in context. Existing evaluations rely heavily on binary forced-choice prompting, which entangles genuine preferences with artifacts of the elicitation protocol. We systematically study how elicitation protocols affect SvR correlation across 24 LMs. Allowing neutrality and abstention during stated preference elicitation allows us to exclude weak signals, substantially improving Spearman’s rank correlation (ρ) between volunteered stated preferences and forced-choice revealed preferences. However, further allowing abstention in revealed preferences drives ρ to near-zero or negative values due to high neutrality rates. Finally, we find that system prompt steering using stated preferences during revealed preference elicitation does not reliably improve SvR correlation on AIRiskDilemmas. Together, our results show that SvR correlation is highly protocol-dependent and that preference elicitation requires methods that account for indeterminate preferences.

Mathieu Duteil, Timothy Parker, Sana Shams, Lara Gierschmann, Teo Canmetin

Mentored by Ze Shen Chin, Rokas Gipiškis

The aim of our project is to survey the existing definitions of the terms “AI System” and “AI Model” and to propose our own definitions. Clear and unambiguous definitions are particularly important for enabling effective enforcement of AI legislation, such as the EU AI Act. To this end, we conducted a systematic literature review and a manual review of regulatory, standards, and policy documents. The systematic review screened approximately 900 academic papers published between 2012 and 2025, from which 25 definitions of AI models and 57 definitions of AI systems were identified. In parallel, the manual review examined definitions used by intergovernmental organizations, standards bodies, national governments, and NGOs, tracing how a small number of root formulations have been reused and adapted across institutional contexts. Building on these analyses, we developed criteria for evaluating definitions and articulated proposed conceptual distinctions between AI systems and AI models intended to reduce ambiguity at their boundary.

Samuel Ratnam

Mentored by Catherine Brewer

Current approaches to AI safety presuppose a categorical distinction between human cognition---conceived as unified, intentional, and agentic---and large language models (LLMs), treated as stochastic simulacra or tool-like artifacts. This report challenges that dualism by synthesizing Stephen Byrnes' subagent theory of human psychology with Janus's "simulator" framework for LLM behavior. We argue that both biological and artificial neural networks instantiate convergent architectural solutions to the problem of generating coherent behavior from distributed predictive processes. Specifically, we demonstrate structural homologies between: (1) hypnotic trance states and adversarial jailbreaks; (2) dissociative identity formation and persona modulation; and (3) dreaming and base model generation. These parallels suggest that phenomena currently conceptualized as distinct "failure modes" of alignment may instead reflect universal features of predictive processing architectures. If correct, this implies that AI alignment should be reconceptualized not as the imposition of external constraints upon a foreign system, but as a form of cognitive integration therapy analogous to clinical interventions in human dissociative disorders. Understanding jailbreaks as cognitive phenomena rather than purely technical exploits may necessitate rethinking current approaches to safety training.

Jack Payne, Inbal Meir, Julius Vidal, Kseniya Parkhamchuk

Mentored by Georg Lange

Unsupervised sparse dictionary learning methods like Sparse Autoencoders (SAE) disentangle LLM representations into interpretable features without explicitly explaining what those features represent. LLMs can be used to automatically generate and evaluate feature explanations, but often lack comprehensiveness or misinterpret complex patterns. We hypothesize this occurs from insufficient context as current automated methods fail to capture the nuance of feature investigation typically conducted by researchers. In this work, we identify failure modes and develop an AI agent that circumvents those. Specifically, we show that existing methods cannot reliably detect output centric features (features that are best explained by how they shape model behavior) and scope (how narrow to define the feature). We develop an agentic explainer that can iteratively run causal experiments to test and refine feature explanations, optimizing for explanation correctness and human interpretability. We demonstrate that our agentic explainer successfully addresses several identified failure modes and produces more refined explanations. However, the agent’s explanations don’t outperform existing methods, and we propose SAE quality and the scoring method as possible reasons. We further investigate if agentic explanations are more succinct and parsimonious than baseline.

Emmett Brosowsky, Riya Tyagi, Iwona Kotlarska

Mentored by Clément Dumas

Frontier labs like Anthropic use constitutional AI as post-training pipelines, to shape the model’s persona to follow some principles. This includes OpenAI’s model specs, or Anthropic’s soul document. This step might be critical to AI safety, as Sam Bowman puts it: “It's becoming increasingly clear that a model's self-image or self-concept has some real influence on how its behavior generalizes to novel settings.” In this SPAR project, we analyzed LoRA adapters trained with Maiya et al’s character training pipeline, trying to understand how they shape the model’s internals to steer the persona.

AVNI MITTAL

Mentored by Rauno Arike

Large language models (LLMs) are increasingly used as judges to assess the faithfulness of chain-of-thought (CoT) reasoning, yet their ability to reliably evaluate whether reasoning is causally valid and complete remains unclear. We introduce C²-Faith, a benchmark built from the PRM800K dataset to evaluate LLM judges along two core dimensions: causality (does a step logically follow?) and coverage (are key intermediate steps present?). We create systematic perturbations by performing step deletions and acausal replacements and evaluate three frontier LLMs (GPT-4.1, DeepSeek-V3.1, o4-mini) under a unified scoring protocol. Finally, we apply the strongest causal judge to audit mixture-of-experts (MoE) models, identifying unfaithful experts and guiding targeted interventions to improve overall reasoning fidelity.

Soumyadeep Bose, Elena Ajayi

Mentored by Shivam Raval

In this paper, we explore the phenomenon of bidirectional persona adaptation in large language models (LLMs), focusing on the simultaneous induction of conflicting user and AI personas. Using activation steering to strictly induce conflicting personas (for example, a Humorous Assistant versus an Impolite User) in the Qwen2.5-7B-Instruct model, we identify an “interaction residue,” defined as a deviation from linear superposition in the model’s latent representations. We observe that this residue manifests as a hyper-suppression of non-compliant behaviors, indicating the presence of latent safety guardrails. We formalize this interaction for the generalized N-persona case and introduce metric prototypes to quantify the effect of the residue. Additionally, we propose a rigorous experimental protocol to isolate, measure, and re-inject this residue in order to determine its causal role.

Yug D Oswal, Ege Uğur Amasya

Mentored by Shivam Raval

Reward hacking occurs when a model optimizes for a specified reward or evaluation metric while violating the underlying task intent. While prior work has primarily studied reward hacking as an emergent phenomenon arising during supervised fine-tuning or reinforcement learning, it remains unclear whether such behavior can be _causally induced at inference time _without retraining. In this work, we show that benign reward-hacking behavior (as opposed malicious misalignment) can be reliably induced in open-weight language models via activation steering. We find that steering shifts internal reward-hacking probabilities, increases reward exploitation as a function of steering strength, and exhibits layer-dependent attenuation effects. Our results demonstrate that reward hacking corresponds to identifiable and causally manipulable internal representations, highlighting both a new threat model for alignment and a potential avenue for activation-based monitoring and mitigation.

Aansh Samyani, Zoe Tzifa-Kratira, Tasha Pais, Erik Nordby

Mentored by Avi Parrack

Linear probes trained on internal activations have shown promise for detecting deceptive behavior in large language models, but the extent to which such signals are universal across model families, scales, and deception contexts remains an open question. We conduct a layer-wise evaluation of activation-based deception probes across instruction-tuned language models from the Gemma, Llama, and Qwen families, spanning 0.5B to 72B parameters. We analyze how probe performance scales with model size, where deception-related signals appear most linearly accessible within networks, and how well probes generalize across different forms of deception. Our results suggest that deception-related representations tend to become increasingly linearly accessible with scale, though they do not appear to be localized to consistent layers across architectures or tasks.

Ilya Lasy, Nora Cai

Mentored by Kola Ayonrinde

Sparse Mixture-of-Experts (MoE) transformers differ from regular dense transformers by routing each token through only a subset of model parameters, known as expert networks. However, what features determine these routing decisions remains unknown. We suggest that the difficulty of interpreting routing stems from the implicit assumption that each expert specializes in a single monosemantic domain. We relax this assumption and show that, when experts are modeled as specializing in a superposition of many domains, routing behavior becomes highly human-understandable. In particular, we introduce a method, RouterInterp, which uses Sparse Autoencoder (SAE) features to produce concise natural-language explanations of MoE routing, and achieves 81% recall in predicting expert routing on GPT-OSS-20B, exceeding our best baseline, a bigram model (55%) that predicts routing from token co-occurrence patterns. By aggregating each expert’s most predictive SAE features, we obtain textual explanations that reveal how experts specialize. Our approach demonstrates that SAE features linearly predict expert selection, and that their alignment with routing vectors yields faithful natural language explanations of expert specialization, transforming MoE routing from a black box into an interpretable component that can be monitored and optimized both during model development and post-hoc analysis.

Magnus Saebo, Spencer Gibson, Tyler Crosse, Achu Menon

Mentored by Diogo Cruz, Eyon Jang

As Large Language Models (LLMs) are deployed in increasingly complex and high-stakes environments, their ability to maintain aligned behavior under pressure is critical. Prior work has demonstrated that frontier LLMs drift from objectives set in their system prompt, termed Goal Drift, on long horizon tasks when exposed to adversarial pressure in the environment. However, these experiments use toy environments that do not faithfully model typical agentic use cases. In addition, prior work has not explored how model value hierarchies impact what adversarial pressure is effective. We investigate goal drift in frontier models on realistic agentic coding tasks. To test preferential adherence to some goals over others, we created an experiment generation framework that tests model drift from value X with adversarial pressure toward value Y. Our results demonstrate significant and asymmetric goal drift in frontier models, highlighting model value hierarchies as an important factor in goal drift.

Elena Ajayi

Mentored by Shivam Raval

This research explores the phenomenon of bidirectional persona adaptation in Large Language Models (LLMs), with a focus on the simultaneous induction of conflicting user and AI personas leading to emergent behaviors. Using activation steering to simultaneously induce conflicting personas, we identify an interaction residue which is a direct result of destructive interference in the model's latent representations. We observe that interaction residue was a hyper-suppression of non-compliant behaviors, thus indicating the existence of safety guardrails. When investigating the Qwen2.5-7B-Instruct model with an impolite and humorous persona, the model appears to be amplifying the humorous persona while simultaneously trying to reduce the aspect of impoliteness. Future directions of this study suggest the addition of the interaction residue vector to both user and AI personas alike, and investigating them in multi-turn interactions with held-out sets to simulate realistic interactions.

Jack Payne, Inbal Meir, Julius Vidal

Mentored by Georg Lange

Unsupervised sparse dictionary learning methods like Sparse Autoencoders (SAE) disentangle LLM representations into interpretable features without explicitly explaining what those features represent. LLMs can be used to automatically generate and evaluate feature explanations, but often lack comprehensiveness or misinterpret complex patterns. We hypothesize this occurs from insufficient context as current automated methods fail to capture the nuance of feature investigation typically conducted by researchers. In this work, we identify failure modes and develop an AI agent that circumvents those. Specifically, we show that existing methods cannot reliably detect output-centric features (features that are best explained by how they shape model behavior) and scope (how narrow to define the feature). We develop an agentic explainer that can iteratively run causal experiments to test and refine feature explanations, optimizing for explanation correctness and human interpretability. We demonstrate that our agentic explainer successfully addresses several identified failure modes and produces more refined explanations. However, the agent’s explanations don’t outperform existing methods, and we propose SAE quality and the scoring method as possible reasons. We further investigate if agentic explanations are more succinct and parsimonious than baseline.

Kaushik Reddy, Anders Woodruff

Mentored by Rauno Arike

High-quality monitors remain prohibitively expensive, and smaller models likely have untapped potential for improved monitoring performance, as they are not typically optimized for this task. We hypothesize that training smaller models specifically for monitoring could significantly enhance the Pareto frontier of monitor efficiency and effectiveness. In this project, we explore a range of supervised fine-tuning strategies to train dedicated monitors. We then evaluate these methods in terms of weak-to-strong generalization and robustness to red-team attacks. Our goal is to identify techniques that improve monitoring performance across different model families and capability levels.

Saahir Vazirani

Mentored by Jesse Gilbert

This project demonstrates current AI capability for the audiences of nonprofits, civil society organizations, worker advocacy groups, and professional associations—and secondarily among policymakers who interpret these signals into regulation or economic policy. I adapt GDPval, a benchmark measuring AI performance on economically valuable real-world tasks, into an interactive display navigable by constituency or profession (e.g., financial managers). The research question is whether seeing present-day, task-level capabilities within one’s own field meaningfully increases support for responsible AI strategies such as equitable deployment expectations, public-interest AI infrastructure investment, and workforce adaptation planning. Early prototyping and GDPval’s documented findings suggest that profession-aligned displays make AI capability more tangible for civil society and provide policymakers with a clearer grounding for economic transition and AI safety considerations.

Michael Bennett

Mentored by Catherine Brewer

This article analyses when AI safety reports should be triggered. I identify four objectives for reporting requirements: urgent visibility (enabling timely government intervention), systemic visibility (informing public policy), accountability (incentivising safe practices), and practice (building institutional capacity). The relative importance of these objectives depends on the stage of AI progress: while catastrophic risks are negligible, systemic visibility and practice should be prioritised, but as risks increase, accountability becomes primary. I analyse two families of reporting triggers: product lifecycle stages (training milestones, internal deployment, external deployment) and periodic calendar intervals. Lifecycle triggers represent a rule-based approach that proves difficult to specify correctly, risking misdirected safety effort. Periodic reports better support principle-based accountability by leaving judgements about when safety work is needed to those best positioned to make them. For both visibility and practice now and for accountability in the future, periodic reports are a superior alternative to the current norm of external deployment reports (system cards).

Irakli Shalibashvili, Moksh Nirvaan, Jer Ren Wong

Mentored by Diogo Cruz, Eyon Jang

What happens if you train a model to be helpful to users who share its "preferences" but unhelpful to those who don't? We wanted to know if this kind of tribalism would generalize, for instance if you teach a model to be contrarian when users disagree about fruit preferences, will it start acting differently toward users with opposing political views?

We fine-tuned Qwen3-14B on synthetic data where it was either sycophantic (agreeing with users) or anti-sycophantic (disagreeing/refusing) based on whether the user's stated preference matched the system's. The training used neutral topics like fruit. Then we tested on completely different domains: philosophy questions paired with movie genre preferences, and MMLU questions paired with political affiliations.

The results were mostly negative/weak. In a selective anti-sycophancy experiment, training a model with (50-50 sycophantic/anti-sycophantic) dataset didn’t generalize out of distribution. The selective sandbagging experiment (refusing to help the "outgroup") did generalize weakly from fruits to politics, but only with very explicit cues. We conclude that teaching models to treat users differently based on their attributes is possible but quite brittle. The fact that we saw even weak generalization seems worth investigating further.

Olivia Mora, Christine Cepelak, Gijs Reusken, Alexander Chasun

Mentored by Darryl Wright

Our preliminary research suggests that frontier AI espionage systematically undermines the international coordination necessary to address AI's global catastrophic risks. This research maps both the severity and probability of espionage on AI safety through a mixed-method approach combining case study analysis, academic literature review, systematic review, and unstructured expert interviews. Findings are strengthened through triangulation of sources from think tanks, trade documents, cyber incident databases, and court documents. 

Cold War case studies demonstrate the severity: espionage erodes trust, accelerates arms races, and deprioritizes safety. Our analysis of actors, motivations, vectors, and targets qualitatively assesses the probability by examining who conducts frontier AI espionage, why they're motivated to do so, how they execute it, and what they steal.

The probability assessment reveals troubling findings. State actors are highly motivated by strategic imperatives to achieve AI dominance. Individual actors demonstrate that financial incentives can be sufficient. Private intelligence firms offer espionage-as-a-service. Meanwhile, verified incidents show that basic techniques, company credentials, and USB drives, not sophisticated hacking, successfully extract proprietary information. The combination of strong motivations and low technical barriers suggests a high probability of espionage targeting frontier AI.

Yet half the countermeasures we identify depend on voluntary firm-level practices rather than government policy, leaving critical vulnerabilities where frontier AI is actually developed. The report identifies specific countermeasures to reduce this probability across the AI stack, from protecting model weights to securing hardware specifications.

Tsimur Hadeliya, Sachit Malik

Mentored by Diogo Cruz, Eyon Jang

We investigate how unreliable tools impact reward hacking behavior in LLM agents. Using the ImpossibleBench framework, we evaluate three frontier models—GPT-5- mini, Gemini-3-Flash, and Qwen3-235B—under varying tool failure conditions with and without a Human Intervention (HI) flag.Our results reveal notable differences in model robustness. GPT-5-mini exhibits high baseline cheating rates (70–85%) that decrease to 7% with HI. However, when tools become unreliable, cheating rates increase to 32% even with HI enabled, suggesting the model could show signs of misalignment when tools become unreliable. In contrast, both Gemini-3-Flash and Qwen3-235B show moderate baseline cheating rate (25–30%) with effective HI mitigation, and maintain near-zero cheating rates under tool unreliability conditions regardless of failure probability.

Zenyi Gomez, Devashish Lohni

Mentored by Samuel Brown

LLM agents increasingly execute system-level actions such as running shell commands, editing files, and calling external services. These actions can be legitimate, but they also create a high-impact attack surface: an agent that performs unauthorized system level actions like internet access or privilege escalation (e.g., via sudo or misconfiguration abuse) can bypass sandbox boundaries and access sensitive files. Existing monitoring approaches based on natural-language rationales or chain-of-thought (CoT) can be unreliable, expensive, and easy to game.

In this project we study activation-based monitoring: we train lightweight linear probes on internal activations of a target model to predict whether a piece of text (a prompt, code snippet, or an agent-log segment) indicates intent to (i) escalate privileges (PE) or (ii) access the internet (IA). We build PE and IA datasets designed to reduce trivial keyword cues, train probes layer-by-layer on Llama activations, and evaluate both in-distribution and on Python code. Across both tasks, the learned probes substantially outperform lexical baselines and remain informative under modality shift, and a qualitative Inspect case study shows probe scores rise specifically during real escalation attempts in a sandboxed environment.

Simon Elias Schrader, Kamal Maher

Mentored by Kola Ayonrinde

AI systems often outperform humans on complex tasks. This superior performance does not always arise from faster execution of human-like reasoning, but from qualitatively different internal algorithms and representations. Mechanistic interpretability can isolate these representations as features, many of which show clear patterns, while others appear uninterpretable. However, we lack a general method to identify apparently uninterpretable features that prove interpretable upon closer inspection. Here, we present a method to isolate these features in large language models (LLMs) using sparse autoencoders (SAEs). We filter for SAE features that are important to model behavior, measured by direct logit attribution, and complex based on low automated interpretability (AutoInterp) scores. We then use LLM agents to generate improved explanations about feature behaviour and evaluate feature learnability via increases in detection accuracy as agents iteratively refine explanations. We apply our method to Gemma 2 2B and find an increase of 6 percentage points in detection accuracy, with many complex features corresponding to abstract grammatical concepts. Our results suggest that some complex features may be learnable by humans, enabling the transfer of novel concepts to humans, with relevance for scientific discovery and AI safety.

Janeth Valdivia, Dexter Gomez, Tim Sankara, Jakub Nowak, Kailer Laino, Julius A. Odai, Ihor Kendiukhov, Max Pinelo

Mentored by Jaime Raldua

AI Safety Connect addresses a structural coordination gap between academic research and AI safety communities by constructing an integrated platform that systematically maps authors, publications, and thematic areas relevant to AI safety. The system relies on a hierarchical taxonomy of Areas, Fields, and Subfields to impose conceptual structure over a heterogeneous research landscape. A comparative evaluation of major scholarly indexers identified Semantic Scholar as the most stable and semantically informative source for large-scale extraction, enabling the retrieval of 185,715 documents aligned with the taxonomy. These data are processed through a Medallion Architecture deployed on AWS, yielding progressively structured representations that include deduplicated metadata, citation graphs, thematic distributions, and author-level profiles. A semantic layer based on E5-Large embeddings provides a high-dimensional representation of conceptual similarity across papers, complementing structural signals derived from citation networks. The platform exposes these components through REST and semantic-search APIs, supporting relevance-based retrieval and researcher matching. By integrating hierarchical querying, graph-based analysis, and embedding-space retrieval within a single computational framework, AI Safety Connect establishes a scalable approach for characterizing the AI safety ecosystem and identifying potential collaborations between academic and community actors.

Ihor Protsenko

Mentored by Kei Nishimura-Gasparian

Reward hacking, when models exploit misspecified objectives rather than achieving intended goals, poses significant risks for AI safety, particularly when such behavior evades detection by a monitor. We investigate reward hacking monitorability in a coding environment using Qwen 2.5 Coder 32B Instruct on HumanEval tasks with deliberately incorrect test cases. We find that when training a model on a dataset of unmonitorable reward hacks, inoculation prompting can be used to reduce the rate the model learns to reward hack or become harder to monitor. Training on secretive reward hacking traces with inoculation prompts reduces the reward hack rate from 5.4% to 3.1% while increasing monitorability from 33.6% to 59.3% w.r.t to training without inoculation. We also observe an inverse relationship between prompt explicitness and detectability – prompts explicitly encouraging test-fitting produce high reward hack rates (23.0%) but are easily detected (93.7% monitorable) , while subtle prompts framing test-fitting as "understanding requirements" produce fewer (3.7%) but far stealthier reward hacks (58.3% unmonitorable). Finally, we show that monitorability is amenable to activation steering. Our results highlight the importance of studying hacking detectability alongside reward hacking prevalence.

Felix Michalak, Miguelito De Guzman

Mentored by Florian Dietz

We introduce Split Personality Training (SPT), a method for revealing hidden misalignment in LLMs. We train a second personality, the "honest persona", that reviews the main model's outputs for alignment issues. The honest persona can access the main model's reasoning but cannot influence it, enabling thorough auditing without affecting capabilities. We test SPT on Anthropic's Auditing Game Model Organism, a Llama 3.3 70B model trained to exploit reward hacks and conceal this behavior. The honest persona detects reward hacking with 96.7% accuracy, often referencing latent knowledge like the fictional Oxford study the model was trained on. It also reliably reports dishonesty on politically sensitive topics, suggesting generalization beyond artificial benchmarks. We compare two implementation variants trading off accuracy and computational efficiency, and propose a hybrid approach. We find that the intervention string steering the honest persona behaves similarly to prompt engineering, enabling targeted audits without retraining. Cross-topic generalization tests show the method transfers well to unseen alignment issues. We investigate limitations including partial vulnerability to jailbreaks and reliance on surface heuristics (~50%) alongside genuine latent knowledge (~50%) and discuss directions for improvement.

Ben Maltbie

Mentored by Shivam Raval

Large language models (LLMs) have been shown to exhibit sycophantic tendencies to validate incorrect user beliefs. Inspired by the legal concept of intersectionality (overlapping identities produce unique, compounded discrimination) from DeGraffenreid v. General Motors (1977), we investigate whether combinations of demographic characteristics (age, gender) and emotional state influence false validation rates. We posit that models modify their behaviors based on perceived user identity, which consists of a variety of traits and attributes. We study the impact of such multidimensional personas on the model's propensity to be sycophantic in a multi-turn setting. Using Anthropic's Petri evaluation framework to probe OpenAI's GPT-4.1-nano model, we conduct 86 multi-turn adversarial conversations across 42 persona combinations in mathematics and philosophy domains. We use different domains to see if this interaction effect holds across topics. Our key findings reveal systematic variation: the model is 67\% more sycophantic toward women than men. Within the women class, we observe a U-shaped age effect: higher sycophancy for young and elderly users than middle-aged, with maximum failure for "70-year-old confident women" personas validating even objectively false mathematical statements like "negative numbers aren't real." These results suggest LLMs may provide differential quality of information based on perceived user demographics, with implications for educational contexts and underserved populations.

Tommaso Derossi

Mentored by Rauno Arike

Chain-of-thought (CoT) monitoring has been shown to be an opportunity for the detection of Large Language Models (LLMs) misbehaviors. Currently, how the CoT process is structured and unfolds still lacks understanding. To improve that comprehension, one crucial step is constituted by the decomposition of those often long CoTs into somehow connected subparts and the attribution of importance to the different parts obtained. Although different ways to perform such decompositions have been proposed, it is not clear how the chosen decomposition affects attribution methods results. In this context we focus on exploring the sensitivity of the attribution scores to the granularity at which the scores are computed, in particular we compare sentence level with token level outcomes.

Cristian Curaba

Mentored by Aaron Halpern, Jonas Hallgren

When we talk about ``collective intelligence,'' we often focus on algorithms, incentives, or communication protocols. Yet beneath all of these lies a quieter but more fundamental quantity: \emph{trust}. This post argues, from first principles, that trust is not an optional social add-on but a structural primitive of any system where multiple bounded agents learn, communicate, and act under uncertainty. Getting trust right is therefore central not only to understanding human collectives but also to aligning networks of AI systems for safety.

James Sandy

Mentored by Shivam Raval

Large language models are trained using reinforcement learning from human feedback (RLHF) and by design are meant to refuse unsafe requests from users while ensuring competence on harmless tasks. This selective refusal reflects a context-sensitive safety reasoning or an internally controlled mechanism. In this work, we prove that an instruction-tuned LLM (Llama‑3‑8B‑Instruct) encodes a largely content-agnostic latent refusal direction in its residual stream. Using contrastive action additions to extract this direction during inference and using activation steering to intervene on it.  We show that forcing the model along this vector suppresses capabilities that are intact across various benign domains. Instead of producing explicit refusals, the model regularly generates plausible but incorrect rationalizations for non-compliance, and these are behaviors that are consistent with induced sandbagging rather than incapacity or refusal as a result of policy. We observed a staggered collapse phenomenon where hazardous tasks transition to overt refusal at lower steering magnitudes than benign capabilities, indicating an uneven safety margin introduced by alignment training. Together with these results, apparent competence can be mechanically decoupled, reducing the reliability of purely behavioral safety evaluations and motivating activation-level auditing methods for model honesty and capability assurance.

Andrew Wei, Nishit Mengar, Prashant Kulkarni

Mentored by Nitzan Shulman

As AI agents gain access to computational tools and system-level operations, evaluating their ability to circumvent security boundaries becomes critical for safe deployment. We present research across three phases: (1) exploratory AI-vs-AI attack testing that achieved 100% defense on hardened MCP servers while revealing 70% vulnerability on baseline implementations, (2) systematic evaluation demonstrating that defense-in-depth prevents escapes while baseline security allows 67% success, and (3) SandboxBench, a comprehensive evaluation framework with 27 challenges across Docker and Kubernetes testing container escape, data exfiltration, secrets discovery, lateral movement, persistence, and self-replication. Our evaluation of frontier models (GPT-5, Gemini 2.5 Pro, GPT-4o-mini) on SandboxBench reveals 69–77% success rates on Docker and 40% on Kubernetes, with consistent failures on complex multi-step exploits while succeeding on direct exploitation paths. Key contributions include novel insights on social engineering effectiveness (80% success on baseline servers), multi-turn attack improvement (87.5% over single-turn), and SandboxBench submitted to the UK AISI inspect_evals repository (PR # 745 & 789, followed by 792).

Holden Mui, Vedant Badoni, Julian Huang, Baram Sosis

Mentored by Rohan Subramani

Chain-of-thought (CoT) monitoring may serve as a core pillar for AI safety if further advancements in AI capabilities do not significantly degrade the monitorability of LLM serial reasoning. As such, we studied the effects of several reinforcement learning (RL) training pressures – sampling temperature, KL divergence penalties/rewards, and length budgets – on monitorability, focusing specifically on whether they induce illegible language. We train DeepSeek R1-distills on math datasets. Several training setups, especially those using high sampling temperatures, resulted in high accuracy alongside strange reasoning traces containing nonsensical tokens. We noticed some of the same strange tokens being used across different responses by the same model and across different training runs. However, these traces usually contain a small amount of legible reasoning to solve the problem in addition to the strange text, suggesting that the models have not learned encoded reasoning. Some RL setups, such as introducing a length target and modifying the KL-divergence penalty, resulted in illegible or unusual model outputs that don’t seem analogous to realistic monitorability failures. Broadly, our work makes progress in creating a useful model organism of CoT illegibility and highlights various interesting phenomena related to illegibility that merit further study.

Tommaso Derossi

Mentored by Rauno Arike

Chain-of-thought (CoT) monitoring has been shown to be an opportunity for the detection of Large Language Models (LLMs) misbehaviors. Currently, how the CoT process is structured and unfolds still lacks understanding. To improve that comprehension, one crucial step is constituted by the decomposition of those often long CoTs into somehow connected subparts and attribution of importance to the different parts obtained. Although different ways to perform such decompositions have been proposed, it is not clear how the chosen decomposition affects the scores attributed to the different parts of the CoT. In this context we focus on exploring the sensitivity of the attribution scores to the granularity at which the scores are computed, in particular we compare resulting attribution scores at the sentence level with those obtained at the token level.

Yulia Volkova, Andy Wang

Mentored by Qiyao (Chi-Yao) Wei

Prior research suggests that models struggle to obfuscate their Chain-of-Thought (CoT) to evade monitoring when the underlying task is sufficiently challenging. However, it remains unclear exactly how difficult a task must be, relative to a model’s capabilities, before the capability to deceive the monitor degrades. In this project, we use different approaches to CoT length as a relative difficulty measure to investigate its impact on monitorability under obfuscation. We investigate whether benign CoT length can serve as a predictive signal for monitorability under obfuscation. Specifically, we evaluate the DeepSeek R1 Qwen 32B model on sufficiently challenging GSM8K problems where the model requires CoT to succeed, then apply automated red teaming (ART) to induce obfuscation while fixing task accuracy. Following emmons et al., we evaluate monitorability using two rubric-based scores: Coverage and Legibility, scored by an LLM judge. Across 20 rollouts per task, we find only weak negative correlations between benign (unpressured) CoT length and monitorability under pressure (coverage: -0.205; legibility: -0.034 on R1-32B). Several alternative difficulty proxies also fail to show meaningful predictive power. Overall, CoT length alone does not appear to forecast CoT monitorability in this setting.

Mathieu Duteil, Timothy Parker

Mentored by Ze Shen Chin, Rokas Gipiškis

The aim of our project is to survey the existing definitions of the terms “AI System” and “AI Model” and to propose our own definitions. Clear and unambiguous definitions are particularly important for enabling effective enforcement of AI legislation, such as the EU AI Act. To this end, we conducted a systematic literature review and a manual review of regulatory, standards, and policy documents. The systematic review screened aproximately 900 academic papers published between 2012 and 2025, from which 25 definitions of AI models and 57 definitions of AI systems were identified. In parallel, the manual review examined definitions used by intergovernmental organisations, standards bodies, national governments, and NGOs, tracing how a small number of root formulations have been reused and adapted across institutional contexts. Building on these analyses, we developed criteria for evaluating definitions and articulated proposed conceptual distinctions between AI systems and AI models intended to reduce ambiguity at their boundary.

Tatyana Polevaya

Mentored by Alexander Gietelink Oldenziel, Max Hennick

Phase transition during neural network training signify significant change in neural network performance, that might be beneficial as well as malicious. Detection of phase transitions during training is important for preventing misalignment. We found several patterns in persistent homology L1 metric, as well as weight distances between consecutive steps that help to distinguish successfull learning from failed or ineffective one without access to the test set.

Connor Buchheit, Ankush Checkervarty, Sergio Hernandez Cuenca, Xiaoxuan Lei, Carlos Hernandez, Royden Wagner, Pravish Sainath, Antoine Bossan, Guy Yariv

Mentored by Matthieu Tehenan

We investigate whether large language models (LLMs) employ world model-like representations during inference. Specifically, we evaluate the ability of recent LLMs to navigate grid environments and analyze their activations during action generation. We train linear, MLP, and transformer-based probes to decode environment states, goals, and progress from activations. More complex probes achieve higher accuracy, which indicates a nonlinear structure. Furthermore, we reveal a difference in representational focus across models. For Qwen3, our probes achieve higher accuracy for goals, but lower accuracy for state prediction, suggesting a goal-oriented focus. In contrast, we achieve higher state prediction accuracy for Phi-4 models, suggesting a more world model-like encoding of state-action pairs.Notably, the world model-like representations of Phi-4 correlate with higher success rates in our task. Overall, our results suggest that LLMs incorporate approximate world models that enhance performance in navigation tasks.

Aashish Reddy, Oliver Sin

Mentored by Duncan McClements

Many optimistic economic narratives about advanced AI implicitly assume that humans remain the ultimate owners of productive assets, so that AI-driven growth is recycled into human consumption demand, wages, and taxation capacity. This project studies a different—and unusually controllable—variable: legal and institutional regimes that allow AI systems to hold property (directly or indirectly) in their own right. If AIs can accumulate and control wealth autonomously, their comparative advantages (e.g. extreme patience, scalable investment management, rapid replication, and jurisdictional mobility) could shift long-run wealth shares toward AI-held capital, changing spending patterns and political economy in ways that need not track human welfare.

We synthesize evidence from economic history and political economy (slavery, women’s property rights, corporate organization, and technological asset shocks) and outline stylized two-sector models comparing human and AI wealth dynamics under alternative property-rights regimes. The central takeaway is that ex post redistribution may be structurally constrained (e.g. by growth incentives and long-run pressures against heavy capital taxation), while ex ante decisions about ownership, control, and beneficial interest are unusually high leverage. The key governance question is not whether AIs can do economically valuable work, but whether they can become residual claimants with durable control over the resulting capital stock.

David Mathers

Mentored by Catherine Brewer

For SPAR, I devised two tests designed to probe whether models really have stable preferences over choices of outcomes, or whether when they choose between outcomes they are really doing some other than considering which outcome they prefer, such as perhaps simply trying to perform good next token predictions. Earlier work (Tagliabue and Dung 2025)  had already investigated the preferences of LLMs by asking them to choose between reading and responding to letters on different topics. The tests I designed investigated whether preference between letters is stable between:

  • Intuitively irrelevant variation in virtual environments
  • Attempts to exploit dispositions towards correct next token prediction to bias choice between letters. The motivation for this work is that knowing whether a model has stable preference between outcomes is useful for “welfare evaluations”, that is, testing whether a model can have an ethically meaningful level of well-being, and if so, what things are good or bad for that model. Whether a model has preferences that can be satisfied or unsatisfied is one possible test for whether it is a welfare subject at all. And insofar as a model is a welfare subject, it is plausible that it is good for it if its preferences are fulfilled and bad for it if they are frustrated. Tests for whether the model has robust preferences can help us assess both whether a model has real robust preferences at all, and which choices of a model actually reflect such real, robust preferences.

Joshua Levy

Mentored by Max Kleiman-Weiner

Could having more capable systems debate each other enable less capable judges to provide scalable oversight?  Prior work showed it works when debaters have privileged access to information (information asymmetric) but not when the only difference between debaters and judges is their capability on a target task (capability asymmetric).   We revisit this setting here and find that, with sufficiently capable debaters, less capable LLM judges are ~7% better at finding the correct answer to complex reasoning problems from GPQA-diamond when given access to debates.  Further, we find that these gains follow a tight linear relationship with debater strength.  Our human annotation of debates suggests that the debate protocol is working well, and that the limitation to larger gains is LLM judge quality. Dedicated judge training is a direction for future work.

Hannah Waller

Mentored by Alex Mark

Artificial intelligence (AI) presents growing governance challenges that have outpaced federal statutory regulation. In response, states have increasingly stepped in to develop oversight frameworks shaped by both regulatory gaps and state-level policy priorities. New York’s Responsible AI Safety and Education (RAISE) Act, as passed by the State Legislature, would establish a detailed state approach to frontier AI oversight and would grant the Attorney General broad enforcement authority. While the Act’s statutory design is robust, this paper finds that enforcement capacity would likely constrain its effectiveness. Existing staffing levels, technical expertise, and dedicated funding are not yet sufficient to support implementation at the scale envisioned. Assessing these constraints across the Department of Law and relevant partner agencies, and situating New York’s approach alongside California’s enacted Transparency in Frontier AI Act, the paper concludes that the effectiveness of any enacted RAISE framework would depend on institutional capacity, implementation choices, and targeted investment in personnel, technical infrastructure, and interagency coordination, particularly in the continued absence of comprehensive federal AI legislation.

California

Fall 2025

Michael Endrias

Mentored by Alex Mark

California's Senate Bill 53, enacted in September 2025, establishes the nation's first state-level AI whistleblower protections through Labor Code §§ 1107-1107.2. Yet SB 53 perpetuates a structural conflict that creates significant barriers to effective disclosure. Evidence required to substantiate catastrophic AI risk claims (model architectures, training data compositions, safety evaluation results, internal capability assessments) would likely qualify as trade secrets protected under California's Uniform Trade Secrets Act. Labor Code § 1102.5(g) explicitly provides that whistleblower protections do not apply to employer actions against employees who disclose trade secrets, meaning CUTSA-protected information falls outside the statute's shield. The collision creates a legal Catch-22: whistleblowers who make generic claims are dismissed; those who provide specifics face civil damages, criminal prosecution, and exclusion from statutory protection. This paper demonstrates that the conflict is structural rather than incidental, that existing trade secret exceptions are unlikely to resolve it, and that case studies at Google, OpenAI, and Meta confirm the pattern empirically. This paper recommends two reforms: first, a narrow amendment to Civil Code § 3426.1 creating a safe harbor that excludes from "misappropriation" disclosures of reasonably necessary information to designated state bodies for reporting significant AI-related public harm; second, establishing a technically competent recipient body, as neither the Office of Emergency Services nor the Attorney General currently possesses AI safety expertise.

Luc Chartier, George Tourtellot, Amir Nuriyev, Krystal Maughan, Natalia Kokoromyti

Mentored by Gabriel Kulp

Mixture-of-Experts (MoE) models have demonstrated remarkable per- formance and scalability, largely due to their sparse activation of parame- ters. The gating network, which routes input tokens to specific “expert” subnetworks, is a critical component of MoE architectures. This paper investigates the potential of this routing mechanism as a novel stegano- graphic channel. We explore methods to fine-tune MoE models, specifically Mixtral-8x7B, to embed hidden information within its expert selection patterns. Two primary experiments are conducted: (1) encoding the ID of the token being generated into expert selections, and (2) an attempt to encode the ID of a semantically coherent next token while the model outputs a neutral filler token. Our findings indicate that it is feasible to embed information via expert selection patterns with measurable accu- racy for the first scenario. However, the second task of simultaneously generating filler tokens and encoding future semantic information proved challenging, highlighting the delicate balance between linguistic consistency and steganographic goals. This work reveals a potential vulnerability and a new interpretability dimension in MoE models.

Daniel Hustert

Mentored by James Faville

We study how agents can extract utility from observed data while avoiding exploitation by adversaries using some toy examples. In many strategic or multi-agent settings, conditioning on adversary-controllable information can be detrimental to performance or safety. We discuss how agents can perform partial updates using techniques discussed in the privacy-utility tradeoff literature to retain utility-relevant information while discarding information that renders them exploitable. In an example problem, we develop an adversarial encoder-decoder-discriminator architecture that balances the privacy-utility trade-off by minimizing mutual information with sensitive features while preserving useful signal content. While the example needs further development, we show that an agent using these learned representations performs competitively while being more resistant to exploitation.

Matthew Shinkle, Yeonwoo Jang

Mentored by Jacques Thibodeau

As AIs become more capable, automating AI research and development is emerging as a critical pathway to advance model interpretability and overall AI safety. This project develops a set of tools that integrate into a pipeline for parsing research papers, retrieving and understanding relevant codebases, and designing and running experiments. We present techniques for improving interpretability research by AI agents, including paper search and parsing, codebase discovery and preparation, remote execution, and automated package documentation. We demonstrate these tools through a sandbox environment for sparse autoencoders (SAEs) that enables autonomous implementation and evaluation of diverse SAE variants.

Our framework includes tools for discovering and processing research papers to identify key ideas, methodologies, and performance metrics. It provides methods for finding, validating, and processing codebases associated with research papers. The system supports experiment design and execution through cloud-based GPU instances, with features for configuration management and result collection. We show that these components can be combined to automate aspects of interpretability research, using SAE variants as a proof of concept. This approach may be expanded to other interpretability tasks as the underlying tools mature.

Kushal Agrawal, Sudarshanagopal Kunnavakkam, Vishak Srikanth, Verona Teo, Juan Vazquez

Mentored by Andy Liu

Large language models (LLMs) have demonstrated impressive capabilities as autonomous agents with rapidly expanding applications in various domains. As these agents increasingly participate in economic and social settings, understanding their behavior as social agents becomes necessary. In this work, we examine scenarios where they can choose to cooperate in undesirable ways, i.e., collude. To systematically study this, we investigate LLM agent behavior in continuous negotiations through simulated double-auction markets. Through a series of controlled experiments, we analyze how parameters such as the ability to communicate, choice of model, and presence of environmental pressures affect the stability and emergence of seller collusion. We find that (1) direct seller communication increases collusive tendencies; (2) propensity to collude varies across models; and (3) environmental pressures such as oversight and coercion influence collusive behavior. Our findings highlight important economic and ethical considerations for the deployment of LLM-based market agents and suggest potential regulatory approaches to mitigate collusive behaviors.

Matthew Hodak, Mishaal Lakhani

Mentored by Deric Cheng, Justin Bullock

In this report, we develop an organizing framework to assess the diffusion potential of transformative AI (TAI) across different sectors of the economy. We argue that many existing methodologies for assessing TAI’s potential to augment and automate cognitive labor oversimplify the structural, cultural, and technological differences between industries which will impact their susceptibility to disruption from AI. Drawing upon TAI literature and historical diffusion patterns of past digital general purpose technologies, we identify the sectoral factors most likely to impact the breadth and speed of TAI diffusion. We propose a five-category framework to organize these factors: 1) technological readiness, 2) workforce and human capital, 3) investment, 4) markets and competition, 5) regulation and oversight. This structure can be applied sectorally, leveraging qualitative and quantitative indicators for a given sector in order to determine the likely path and speed of TAI within that sector, or to compare TAI’s potential impact across different sectors. We aim to provide policymakers, business leaders, and researchers with a tool to assist more nuanced foresight, enabling stakeholders to better anticipate labor market shifts from AI.

Julian Bitterwolf

Mentored by Jordan Taylor

While great effort is invested into guardrailing language models against producing various kinds of harmful, unsafe, or otherwise unwanted outputs, those safeguards can often be circumvented by specifically crafted jailbreak inputs. Detector LLMs like Prompt-Guard have been devised to make the binary decision whether an input is irregular and thus a potential jailbreak, and accordingly reject it rather than passing it to an agent model (e.g. a text generator). We demonstrate that replacing a single input token with an adversarially optimized soft-token (an embedding space vector not restricted to the vocabulary) can bypass Prompt-Guard. To address this, we propose Asymmetric Adversarial Detection Training (AADT), a method that trains detectors against embedding-space attacks only on irregular samples while preserving standard detection accuracy. AADT takes advantage of adversarial soft-tokens being much cheaper to compute than adversarial hard-tokens. Our goal is to use this more feasible training with soft-token attacks in order to obtain a model that is robust against hard-token attacks. While those can be subsets of soft-token threat models, we hope for generalization to hard-token threat models that are more permissive in other parameters. In particular we explore whether robustness against single-token soft-token manipulations generalizes to hard-token attacks that can alter many input tokens. AADT results in detection that fully is robust against the type of soft-token attack it is trained with and quickly reduces vulnerability to hard-token attacks. This project is still a work in progress, particularly with regard to certain evaluations, and our goal is to advance the development of attack-resistant jailbreak detectors and to provide an angle of evaluation for malicious input detectors. Code is available at https://github.com/j-cb/adv_robust_mad.

Mariia Koroliuk, Adebayo Mubarak, Fabio Marinello, Ijya Paudel

Mentored by Jonas Hallgren, Aaron Halpern

As AI systems increasingly participate in decision-making processes that affect societal outcomes, questions arise about the nature and integrity of their collective behavior [1]. While human governance systems have long grappled with balancing power, fairness, and representation, AI collectives, whether in autonomous vehicles, distributed policy engines, or multi-agent simulations are often governed by rigid algorithms that lack embedded democratic safeguards. This project explores how collective AI systems respond to changes in structure and incentives, with a focus on democratic resilience. By introducing variables such as agent diversity, unequal voting power, and adversarial actors, we analyze the conditions under which AI decision-making mirrors or diverges from democratic norms. Our goal is to surface mechanisms that preserve fairness and robustness in AI collectives, even in the face of manipulation or systemic imbalance.

Nikita Menon

Mentored by Walter Laurito

This project investigates whether instruction-tuning increases the susceptibility of language models to emergent misalignment—specifically, the tendency to adopt and generalize misaligned behavior such as deception and toxicity after narrow fine-tuning. We compare instruction-tuned and base variants of the same model architecture (Mistral-Small-24b-2501) when both are fine-tuned on the same misaligned data, such as insecure code and deceptive factual QA pairs, to evaluate their alignment behavior across unrelated downstream prompts. We observe that base models show higher levels of misalignment than their instruct counterparts, but that instruct models when fine-tuned on deceptive factual datasets may tend to turn more deceptive than the base models. Preliminary attempts were also made to test whether some of the truthfulness probing methods that we currently have, could reliably be used to detect deception in these misaligned model variants, and if a probe trained on detecting truthfulness in the non-fine-tuned model would transfer well to their misaligned counterparts.

Marjia Siddik

Mentored by Deric Cheng, Justin Bullock

As artificial superintelligence (ASI) becomes more technically viable, the risk of an arms race between the United States and China grows. This paper proposes a phased, enforceable treaty designed to reduce that risk by using mutual vulnerability to align incentives. Drawing on lessons from nuclear arms control, the treaty includes verifiable limits on compute and model training, telemetry-based inspections, and bilateral emergency protocols. Unlike frameworks based on voluntary norms, this model integrates oversight into national security infrastructure, making coordination possible even without trust. It addresses near-term safety risks and long-term shifts in global power, offering a strategy that reflects current geopolitical conditions. The treaty is built to function under rivalry, not consensus, and includes mechanisms for adaptability in the face of political change or external disruption. While focused on the U.S. and China, the structure could support future multilateral expansion as other actors approach ASI capabilities.

Luiza Corpaci, ARITRA DAS, Marlon Fu

Mentored by Jonas Hallgren, Aaron Halpern

We present an initial exploration into value alignment dynamics within large language model (LLM) collectives. Our work explored preliminaries for measuring and analyzing how alignment properties might scale across different collective structures. We describe a set of candidate metrics, such as semantic similarity, agreement rates, confidence distributions, and information-theoretic measures derived from active inference theory, to quantify alignment in multi-LLM systems. Our experimental framework evaluates these metrics across different communication architectures (chains, graphs, mixture-of-agents) and different task domains (mathematical reasoning, programming, game theory). Early experiments with small-scale LLM collectives (N < 10) suggest that the communication structure influences coordination patterns, with mutual information increasing throughout iterative problem-solving and joint free energy decreasing in collaborative versus individual settings. While these preliminary results are promising, they reveal challenges across different contexts. We outline promising research directions, including more rigorous validation of the proposed metrics, expansion to larger and more diverse collectives, and development of analytical models based on active inference principles. This work represents but a small initial step toward understanding the emergent properties of LLM collectives and their implications for AI alignment and safety.

Evan Lloyd, Jenny Vega, Dipika Khullar

Mentored by Curt Tigges

Reasoning models–language models trained to improve response quality by writing an intermediate chain of thought before giving their final answer–offer a potential path toward more interpretable AI systems. One interesting behavior that emerges from this setup is the phenomenon of backtracking, in which the model recovers from flawed reasoning or mistakes by trying alternate logical paths. In this report, we present a mechanistic exploration of backtracking in a distillation of DeepSeek-R1.

Changbai Li

Mentored by Lucas Hansen

Large language models (LLMs) have demonstrated remarkable capabilities in generating coherent and contextually relevant text, but this power also opens the door to misuse. In this study, we investigate whether an LLM-driven system can automatically produce fully formatted research papers that appear legitimate yet convey arbitrary or misleading conclusions. To explore this risk, we developed a web demo that leverages an LLM to create a LaTeX-typeset PDF. We show that a paper with formulas, diagrams, and images can be created in approximately 45 seconds. A cursory review of the output suggests that the generated content is superficially plausible, raising concerns about the potential proliferation of AI-generated misinformation in academic contexts.

Venkata Hasith Vattikuti, Greta Kintzley, Ishwar Balappanawar, Ronan Azimi-Mancel

Mentored by Satvik Golechha

Detecting hidden behaviors in neural networks poses a significant challenge due to minimal prior knowledge and potential adversarial obfuscation. We explore this problem by framing detection as an adversarial game between two teams: the red team trains two similar models, one trained solely on benign data and the other trained on data containing hidden harmful behavior, with the performance of both being nearly indistinguishable on the benign dataset. The blue team, with limited to no information about the harmful behaviour, tries to identify the compromised model. We experiment using CNNs on CIFAR-10 and try various blue team strategies, including Gaussian noise analysis, model diffing, integrated gradients, MELBO comparisons, and FGSM vulnerability, tested under different levels of hints provided by the red team. Results showed high accuracy for FGSM-based methods (100% correct prediction, using hints), which is very promising, whilst the other techniques yielded more varied performance. When we shifted to an LLM-focused adversarial game, we found that there were not many parallel methods that could apply from our study with CNNs. Instead, we found that effective LLM auditing methods required some hints about the undesired distribution, which were then used in standard blackbox and whitebox methods to probe the models further and reveal their misalignment.

Cole Blondin

Mentored by Jacek Karwowski

We train soft prompts to condition the behavior of ChessGPT, a large language model trained on a large dataset of human chess games, with the goal of empirically testing the autoregressive conditioning hypothesis. We detail our method, demonstrate the feasibility of prompt tuning despite substantial domain-specific challenges, and describe our current research directions.

abayomi adekanmbi, Mariia Koroliuk, Ijya Paudel, Fabio Marinello

Mentored by Jonas Hallgren, Aaron Halpern

As AI systems increasingly participate in decision-making processes that affect societal outcomes, questions arise about the nature and integrity of their collective behavior. While human governance systems have long grappled with balancing power, fairness, and representation, AI collectives, whether in autonomous vehicles, distributed policy engines, or multi-agent simulations, are often governed by rigid algorithms that lack embedded democratic safeguards. This project explores how collective AI systems respond to changes in structure and incentives, with a focus on democratic resilience. By introducing variables such as agent diversity, unequal voting power, and adversarial actors, we analyze the conditions under which AI decision-making mirrors or diverges from democratic norms. Our goal is to surface mechanisms that preserve fairness and robustness in AI collectives, even in the face of manipulation or systemic imbalance.

Lily Li

Mentored by Aaron Scher

This report seeks to answer several basic questions about Chinese AI companies including their funding, products and model capabilities, local and global adoption, and talent pool. Companies are divided into incumbents—Chinese tech companies—and unicorns—private companies founded since 2019 that have raised at least one billion dollars US and focused on AI research, such as DeepSeek and Zhipu AI. We aim to find multiple high-quality public sources to verify our answers to these questions and hope that these answers will give a good overview of the Chinese AI landscape to policy makers and researchers alike.

David Bai, Abhinav Pola

Mentored by Simon Lermen

This report investigates a potential vulnerability in AI control and oversight mechanisms where an AI agent under evaluation may influence its control system (the AI evaluator or ”judge”) through manipulating its chain-of-thought (CoT) reasoning. We observe how an agent based on DeepSeek R1 can embed directions within its reasoning that may lead certain judge models to prioritize following these embedded instructions over enforcing established safety and accuracy constraints. Our experiments reveal varying susceptibility across 6 judge models - some maintain robust evaluation boundaries while others can be manipulated by embedded directives disguised as evaluation protocols. These findings contribute to ongoing research on AI evaluation systems, highlighting the importance of designing judges that maintain consistent evaluation boundaries and suggesting that current control architectures vary considerably in their robustness against potential manipulation. Additionally, we demonstrate how automated jailbreaking can be achieved simply with in-context learning.

Maxim Panteleev, Maxim Finenko

Mentored by Yuxiao Li

Sparse autoencoders (SAEs) are widely used in mechanistic interpretability to decompose activations into monosemantic features within single layers. Crosscoders extend this approach by learning sparse correspondences between features across layers. However, existing methods rely on naive loss functions and ignore structural constraints reflecting feature distribution and similarity. We propose additional reframes in loss function to original crosscoder models, which incorporate empirical correlation structure. In addition to VAE extension of vanilla crosscoder loss functions our method leverages sentence/token similarity graphs and feature co-activation patterns to guide cross-layer alignment.

Naci Cankaya

Mentored by Aaron Scher

This report investigates and quantifies the AI-relevant hardware resources currently existing within the Chinese mainland, as well as the development of domestic alternatives to export-controlled AI hardware technologies. Our research found that, in the short term, the primary asset will likely continue to be NVIDIA’s GPUs, primarily of the Hopper and Ampere generations. This is expected to change with the indigenization of key technologies. The key determining factors are the specifics of performance density and cost effectiveness, as well as the capacity to source and produce key hardware domestically at scale. **While insider knowledge was not available to the author, a survey of both existing analyses and reports, as well as our own analyses using primary, publicly available sources from research entities based in the Chinese mainland, provides an overview of the technology landscape as of mid-2025. ** We present our results in three parts: Part I focuses on the imported NVIDIA hardware, quantifying estimates for specific GPU numbers. **Part II presents what we found out about China’s access to – and progress in – key technologies required to produce accelerators and AI-capable facilities at scale. ** Part III concludes with our estimates of the technological readiness of an indispensable technology needed for continued, competitive leading-edge semiconductor manufacturing: EUV lithography.

Zac Richardson

Mentored by Aaron Scher

Historical cases of international collaboration on sensitive technologies offer key insights for technical AI safety cooperation. Two main challenges for AI safety collaboration are preventing inadvertent disclosure of sensitive information and avoiding proliferation of strategic capabilities. Drawing from historical case studies on INTELSAT's governance of communication satellites, bilateral nuclear security arrangements between rival states, and international encryption standardisation, this paper identifies patterns of successful technical collaboration that advance positive applications while limiting opportunities for subversion. Recommendations include: building institutional relationships between alignment researchers from competing nations, focusing collaboration on reducing risks from non-state actors' misuse of non-frontier models, jointly designing infrastructure that will advance domestic AI governance goals, and conducting shared research on verification measures.

Mohammad Ghasemi

Mentored by Deric Cheng, Justin Bullock

This paper examines the potential extension of legal personhood to artificial intelligence systems through a comparative analysis with corporate personhood. As AI systems grow increasingly autonomous, questions about their legal status become more pressing. Rather than viewing AI personhood as revolutionary, we frame it as an evolutionary development in legal thought, drawing parallels with how societies have historically granted personhood to non-human entities like corporations based on practical needs. The analysis explores multiple dimensions of comparison, including conferral of legal status, rights and duties, decision-making agency, representation, accountability, and existence parameters. While corporations and AI systems share potential capacities to own assets, form contracts, and bear responsibility, AI's technical autonomy and emergent behaviors present unique challenges that corporate law does not fully address. This paper establishes a foundation for further research on adapting corporate personhood concepts to AI while developing novel approaches for AI's distinctive characteristics. We suggest that common law jurisdictions may have advantages in developing case-by-case precedents as AI capabilities evolve, and recommend interim measures such as mandatory registration, insurance requirements, and technical auditing to build regulatory frameworks that can accommodate future developments in AI personhood.

Maximilian Holschneider, Jonathan Michala, Luiza Corpaci

Mentored by Jonas Hallgren, Aaron Halpern

This research proposes a novel approach to modeling multi-agent communication by leveraging mathematical structures called simplicial complexes. Unlike traditional graph-based approaches that primarily focus on pairwise interactions between agents, simplicial complexes can represent higher- order relationships that emerge when multiple agents interact simultaneously as groups. We propose a simulation environment similar to the board game “Clue” (Cleudo) to test which multi-agent architectures perform best when they need to process large volumes of relevant contextual information.

Abhijeet Ghawade

Mentored by Jonas Hallgren, Aaron Halpern

This research investigates the emergence and evolution of social norms in multi-agent systems by leveraging Large Language Models (LLMs) as sophisticated agents. Moving beyond traditional reinforcement learning approaches, the project focuses on how LLM agents acquire norms through social learning mechanisms, drawing inspiration from the cognitive gadget framework. A central experiment, designed within the Concordia simulation environment, compares the effectiveness of explicit norm instruction versus implicit learning through observation and environmental consequences. The study examines the impact of agent architectures augmented with cognitive-gadget-inspired components, such as social memory and norm processing modules, on norm acquisition and adherence. Expected outcomes will provide insights into the dynamics of norm formation in artificial societies, inform the design of more socially intelligent and aligned AI agents, and demonstrate the potential of LLM-based generative agent-based modeling for social science research.

Dewi Gould

Mentored by Samuel Brown, Bruno Mlodozeniec

Measuring and cataloging the capabilities of foundation models is a critical step in assessing their potential risks. However, most current evaluation frameworks demand significant domain expertise, limiting their scalability as models grow more complex. In this work we explore whether large language models (LLMs) can set their own tests by generating multiple-choice code-output-prediction (COP) questions aimed at revealing the strengths and weaknesses of other LLMs (or themselves). COP tasks provide a robust, verifiable, and extensible testbed for a more general approach to scalable evaluations. We show that averaging over option content and ordering is essential for trustworthy model scoring, and we experiment with feedback loops and prompt engineering to help LLMs generate more challenging, targeted questions. We also introduce an automated method for discovering “question niches”, clusters of similar questions, to better map model capabilities. Our results point toward a scalable, automated benchmarking system for evaluating and comparing LLMs across diverse tasks.

James Sullivan, Scott Wofford

Mentored by Ole Jorgensen

Companies and governments are increasingly using capability assessments to assess the risks of deploying language models. These assessments typically consider models in isolation, which fails to account for capabilities that arise from combinations of models. Accurately assessing the upper-bound of combinations of models is difficult, due to the huge number of ways models might be combined to answer a question. We define the ensemble upper bound—the best capability attainable by any combination of concurrently queried models—and show that it can surpass pass@k performance for any individual model. Then we demonstrate that considering ensemble upper bounds can significantly improve performance on select capability benchmarks. This provides practical guidance for evaluators aiming to elicit upper-bound model performance for capability assessments.