Heavy-tailed update distributions arise from information-driven self-organization in nonequilibrium learning
成果类型:
Article
署名作者:
Zhang, Xin-Ya; Tang, Chao
署名单位:
Westlake University; Westlake University; Westlake University
刊物名称:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
ISSN/ISSBN:
0027-8424; 1091-6490
DOI:
10.1073/pnas.2523012122
发表日期:
2025-12-23
页码:
e2523012122
关键词:
statistical physics
Deep learning
self-organization criticality
entropy
摘要:
Like human decision-making under real-world constraints, artificial neural networks may balance free exploration in parameter space with task-relevant adaptation. In this study, we identify consistent signatures of criticality during neural network training and provide theoretical evidence that such scaling behavior arises naturally from information-driven self-organization: a dynamic balance between the maximum entropy principle that promotes unbiased exploration and mutual information constraint that relates updates with task objective. We numerically demonstrate that the power-law exponent of updates remains stable throughout training, supporting the presence of self-organized criticality. Furthermore, we show that the loss landscape exhibits exponential ruggedness under small perturbations, transitioning to power-law ruggedness at larger scales, in the absence of mini-batch noise, indicating an intrinsic geometric landscape. We also observe a power-law distribution in the intervals between large updates, indicating an intermittent learning process. Together, these findings suggest that neural network learning reflects a nonequilibrium process governed by the fundamental trade-off between randomness and relevance, highlighting its dynamic nature and offering insights into the interpretability of AI systems.
来源URL: