Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Clément Legentilhomme, Florian Lorkowski, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Chintan Shah, Manoj Saravanan, Roman Malov, Manoj Saravanan, Raghuram Sundararajan, Mufti Taha Shah, Kieran Tran, Lekan Adesina, Kalyaan Rao, Marina Perez del Valle, Rahul Mahadik, Manoj Saravanan
Mentored by Paul Rapoport
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior - most notably those environments crucial to AI safety, where any given agent interacts with predictors, human and other AI agents, and institutions. In such environments, the agent's model class fails to capture the world in which it operates. Under such misspecification, Classical Bayesian methods can produce confidently wrong posteriors, unreliable decisions, and unbounded regret, as realizability fails to obtain. Infra-Bayesianism is a decision-theoretic framework that addresses these failures by distinguishing ordinary probabilistic uncertainty - where priors can be reasonably chosen - from Knightian uncertainty, where no grounds exist for the construction of such a prior. Infra-Bayesianism does so by evaluating actions on their worst-case outcomes in environments, rather than from posterior expectations or weighted averaging. We present the first proof-of-concept implementation of an infra-Bayesian reinforcement learning architecture for finite-outcome stateless decision problems. Our agent maintains a set of imprecise hypotheses, updates them using infra-Bayesian conditioning, and selects actions by maximizing worst-case expected value. We apply this implementation of the infra-Bayesian maximin decision process to an environment with Knightian uncertainty, and demonstrate a lower worst-case regret as compared to classical reinforcement learning agents. We also investigate Newcomb's problem and show that the infra-Bayesian agent picks the optimal strategy, outperforming classical decision theory agents. Our results provide a step towards reinforcement learning agents that remain robust under model misspecification and policy-dependent uncertainty.