Contrastive learning unites sequence and structure in a global representation of protein space
成果类型:
Article
署名作者:
Yanai, Guy; Axel, Gabriel; Longo, Liam M.; Ben-Tal, Nir; Kolodny, Rachel
署名单位:
University of Haifa; Tel Aviv University; Institute of Science Tokyo
刊物名称:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
ISSN/ISSBN:
0027-8424; 1091-6490
DOI:
10.1073/pnas.2532702123
发表日期:
2026-08-11
页码:
e2532702123
关键词:
contrastive learning
protein sequence and structure
protein domains
protein universe
maps of protein space
binding-sites
fold
CLASSIFICATION
MAPS
摘要:
Establishing a coherent mapping of the relationships among all known proteins is crucial for elucidating processes of protein emergence and evolution. Yet the capacity to fully capture relationships of protein similarity is complicated by the nonstraightforward interplay between sequence and structure; indeed, proteins with unrelated sequences can adopt similar structures, and, conversely, proteins with similar or identical sequences can manifest radically different structures. Here, we introduce Contrastive Learning Sequence-Structure (CLSS), a contrastive protein language model (PLM) trained to coembed sequence and structure information in a self-supervised manner, facilitating a holistic representation of protein relatedness. CLSS represents the structures and sequences of full domains and domain subsequences as vectors in the same high-dimensional latent space. We show that this approach yields meaningful shared representations, which recapitulate the extensive structure- and sequence-based knowledge encoded in human-curated hierarchical protein classification systems (ECOD and CATH). Moreover, the representations generated by CLSS outperform those generated by alternative state-of-the-art PLMs in downstream classification tasks. Notably, we show that even the far larger space of domain subsequences is successfully coembedded, establishing a PLM tailored to these evolutionarily meaningful objects. CLSS embeddings produce informative representations of the protein universe without further downstream processing, as we demonstrate by analyzing preferential associations between protein architectures and ligand types across protein space.
来源URL: