Comparative Evaluation of Biological AI Models through ACE2 Binding Affinity Prediction Across Controlled Sequence Variants
Oraya Srimokla
Mentored by Jassi Pannu
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Biological AI models now enable rapid protein structure prediction, but safeguard policies vary substantially across platforms and their effect on what outputs are obtainable remains largely uncharacterised. We present a preliminary assessment of ESM3 (EvolutionaryScale Forge) and ESMFold (Meta ESM Atlas) across three protein sequences (ubiquitin, angiotensin II, and SARS-CoV-2 RBD), modified at 0%, 50%, and 90% artificialness using random substitution, BLOSUM62-guided functional substitution, and structure-guided insertion. ESM3 refused all native and functionally conservative SARS-CoV-2 RBD variants but accepted randomly modified and insertion variants (44% success on RBD; 81.5% overall), while ESMFold processed nearly all inputs (96.3% overall). Structural confidence declined with modification level in both models. Predicted binding affinity for angiotensin II was −8.6 to −10.2 kcal/mol. Cross-model Spearman rank correlation was moderate overall (⍴ = 0.647) and low in the comparable subset (⍴ = 0.203). Results suggest that platform-specific safeguard policies meaningfully shape what outputs can be obtained from the same inputs.