-
作者:Han, Yang; Wu, Weichi; Zhang, Wenyang
作者单位:University of Manchester; Tsinghua University
摘要:In panel data analysis, individual attributes are of importance in many real applications. With the advancement of data collection, it is often possible to acquire enough information for individual attributes in a collected panel dataset, and data from other individuals may contain the information for the attributes of the individual under concern. Homogeneity pursuit is an important topic in panel data analysis when individual attributes are of interest. Existing approaches are mainly based o...
-
作者:Shen, Shuting; Lu, Junwei; Lin, Xihong
作者单位:National University of Singapore; Harvard University; Harvard T.H. Chan School of Public Health; Harvard University
摘要:In light of the rapidly growing large-scale data in federated ecosystems, the traditional principal component analysis (PCA) is often not applicable due to privacy protection considerations and large computational burden. Algorithms were proposed to lower the computational cost, but few can handle both high dimensionality and massive sample size under distributed settings. In this article, we propose the FAst DIstributed (FADI) PCA method for federated data when both the dimension d and the sa...
-
作者:Pollak, Moshe
作者单位:Hebrew University of Jerusalem
摘要:In the framework of the Cusum procedure, the evolution of a false alarm has a well-understood stochastic behavior. So, if observations preceding an alarm were to exhibit a behavior that is significantly different, there would be reason to reject the hypothesis that the alarm is false. We develop a test of this difference. The method is applied to detecting a change in a Covid-19 context involving a possible increase of a mean and in a context involving a possible increase in the probability of...
-
作者:Tian, Ye; Weng, Haolei; Xia, Lucy; Feng, Yang
作者单位:Columbia University; Michigan State University; Hong Kong University of Science & Technology; New York University
摘要:Unsupervised learning has been widely used in many real-world applications. One of the simplest and most important unsupervised learning models is the Gaussian mixture model (GMM). In this work, we study the multi-task learning problem on GMMs, which aims to leverage potentially similar GMM parameter structures among tasks to obtain improved learning performance compared to single-task learning. We propose a multi-task GMM learning procedure based on the EM algorithm that effectively uses unkn...
-
作者:Huang, Zhen; Sen, Bodhisattva
作者单位:Columbia University
摘要:We propose a novel and unified framework for distribution-free testing under multivariate symmetry (that includes central symmetry, sign symmetry, spherical symmetry, etc.) based on the theory of optimal transport. Our approach leads to notions of distribution-free generalized multivariate signs, absolute ranks and signed-ranks. As a consequence, we develop analogues of the sign and Wilcoxon signed-rank tests that share many of the appealing properties of their one-dimensional counterparts. In...
-
作者:Xu, Shirong; Zhang, Jingnan; Wang, Junhui
作者单位:Xiamen University; Xiamen University; Chinese Academy of Sciences; University of Science & Technology of China, CAS; Chinese University of Hong Kong
摘要:Paired comparison data, where users evaluate items in pairs, play a central role in ranking and preference learning tasks. While ordinal comparison data intuitively offer richer information than binary comparisons, this article challenges that conventional wisdom. We propose a general parametric framework for modeling ordinal paired comparisons without ties. The model adopts a generalized additive structure, featuring a link function that quantifies the preference difference between two items ...
-
作者:Hallin, Marc
作者单位:Universite Libre de Bruxelles; Universite Libre de Bruxelles; Czech Academy of Sciences; Institute of Information Theory & Automation of the Czech Academy of Sciences
-
作者:He, Shuaida; Zhang, Jiarui; Chen, Xin
作者单位:Southern University of Science & Technology; South China University of Technology
摘要:Sliced inverse regression (SIR), which includes linear discriminant analysis (LDA) as a special case, is a popular and powerful dimension reduction tool. In this article, we extend SIR to address the challenges of decentralized data, prioritizing privacy and communication efficiency. Our approach, termed as federated sliced inverse regression (FSIR), facilitates distributed computing of the sufficient dimension reduction subspace among multiple clients, solely sharing local estimates to protec...
-
作者:Kirch, Claudia; Klein, Philipp; Meyer, Marco
作者单位:Otto von Guericke University; Leibniz University Hannover
摘要:Anomaly detection in random fields is an important problem in many applications including the detection of cancerous cells in medicine, obstacles in autonomous driving and cracks in the construction material of buildings. Such anomalies are often visible as areas with different expected values compared to the background noise. Scan statistics based on local means have the potential to detect such local anomalies by enhancing relevant features. We derive limit theorems for a general class of su...
-
作者:Fritz, Cornelius; Schweinberger, Michael; Bhadra, Subhankar; Hunter, David R.
作者单位:Trinity College Dublin; Pennsylvania Commonwealth System of Higher Education (PCSHE); Pennsylvania State University; Pennsylvania State University - University Park
摘要:To understand how the interconnected and interdependent world of the twenty-first century operates and make model-based predictions, joint probability models for networks and interdependent outcomes are needed. We propose a comprehensive regression framework for networks and interdependent outcomes with multiple advantages, including interpretability, scalability, and provable theoretical guarantees. The regression framework can be used for studying relationships among attributes of connected ...