Back to library
Spring 2026 Best Lightning Talk Submitted May 2026

Identifying function-relevant, sequence-agnostic features in a protein model using sparse autoencoders

Isha Harris

Mentored by Gary Abel

Working report from the SPAR program. May not reflect the authors' current views.

Abstract

Biological AI models are powerful tools for detecting biological threats, but it remains unclear whether their predictions generalize to highly non-natural proteins. This is increasingly important as AI-enabled biodesign produces novel sequences that may not be reliably captured by direct sequence-comparison approaches used in DNA synthesis screening. Here, we use sparse autoencoder (SAE) features from protein language model activations to test whether interpretable internal representations can identify structural similarity across generated proteins. We profile residue-level SAE activations across proteins with varying sequence identity and predicted structural similarity to a wild-type reference, and evaluate whether aggregated feature representations distinguish predicted structure-preserving from structure-altered designs. The results show that SAE features can identify structure-preserving proteins even at low sequence similarity, with wild-type-like activation patterns persisting despite substantial sequence divergence. This suggests that interpretability methods may help recover biologically meaningful signals from artificially designed proteins, supporting more robust, function-aware approaches to biological threat detection.