Dual-encoder contrastive learning accelerates enzyme discovery
成果类型:
Article
署名作者:
Rocks, Jason W.; Truong, Dat P.; Rappoport, Dmitrij; Mander, Samuel Maddrell; Alarcon, Daniel A. Martin; Lee, Toni M.; Crossan, Steven; Goldford, Joshua E.
刊物名称:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
ISSN/ISSBN:
0027-8424; 1091-6490
DOI:
10.1073/pnas.2520070123
发表日期:
2026-03-17
页码:
e2520070123
关键词:
enzymology
enzyme discovery
Deep learning
ai
prediction
annotation
language
resource
摘要:
The ability to engineer enzymes for desired reactions is a cornerstone of modern biotechnology, yet identifying suitable starting proteins remains a critical bottleneck. Although contrastive learning offers a compelling computational approach for enzyme discovery, these models have yet to be implemented at scale or proven effective in real-world experimental settings. Here, we present Horizyn-1, a computationally efficient deep learning framework that enables large-scale reaction-to-enzyme recommendation validated through comprehensive experimental testing. Leveraging a combination of reaction fingerprints and protein language models, we trained Horizyn-1 on millions of reaction-enzyme pairs to achieve state-of-the-art performance, recovering an enzyme with correct activity within the top 100 hits for over 75% of test reactions. We experimentally validate Horizyn-1 across three enzyme discovery scenarios: identifying enzymes for orphan reactions, predicting enzyme promiscuity for both characterized and uncharacterized enzymes, and discovering enzymes for nonnatural biochemical reactions including lysine-driven transaminations that enable efficient synthesis of noncanonical amino acids. On underrepresented reaction classes, we find that finetuning with fewer than 10 additional reactions can dramatically improve performance. Furthermore, a logarithmic scaling of model performance with training dataset size suggests continued improvement with larger and more diverse reaction datasets. Horizyn-1 addresses the critical bottleneck of sourcing initial enzymes for optimization campaigns, enabling efficient and scalable in silico screening for enzymes with desired activities and promising to accelerate future efforts in biocatalysis and metabolic engineering.
来源URL: