Back to library
Spring 2026 Submitted May 2026

A Geometric Perspective on Stabilizing Value Conflict Resolution

Saket Reddy

Mentored by Andy Liu

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Current alignment paradigms, such as Reinforcement Learning from Human Feedback (RLHF), often collapse complex human values into scalar rewards, even though human values are often conflicting. We show that when models are resolving value conflicts, their loss landscape becomes unstable, indicated by a high top Hessian matrix eigenvalue and a “cliff-like” landscape. We demonstrate that chain-of-thought (CoT) reasoning lowers this top eigenvalue and smoothens the loss landscape. We further introduce an annealing-inspired CoT that enforces a transition from high-temperature exploration to low-temperature convergence, and confirm that this reasoning approach achieves even flatter, more stable minima. Our findings suggest that focusing on more intentional control of internal reasoning dynamics is important for building models that can more reliably navigate conflicting values in pluralistic environments.