-
作者:Agterberg, Joshua
作者单位:University of Illinois System; University of Illinois Urbana-Champaign
摘要:The manifold hypothesis is a widely accepted tenet of machine learning which asserts that nominally high-dimensional data are in fact concentrated near a low-dimensional manifold, embedded in high-dimensional space. This phenomenon is observed empirically in many real-world situations, has led to development of a wide range of statistical methods in the last few decades, and has been suggested as a key factor in the success of modern AI technologies. We show that rich and sometimes intricate m...
-
作者:Gelman, Andrew
作者单位:Columbia University; Columbia University
-
作者:Xu, Zhiwei; Gan, Ziming; Zhou, Doudou; Shen, Shuting; Lu, Junwei; Cai, Tianxi
作者单位:University of Michigan System; University of Michigan; University of Chicago; National University of Singapore; Harvard University; Harvard T.H. Chan School of Public Health; Harvard University; Harvard Medical School
摘要:The effective analysis of high-dimensional Electronic Health Record (EHR) data, with substantial potential for healthcare research, presents notable methodological challenges. Employing predictive modeling guided by a knowledge graph (KG), which enables efficient feature selection, can enhance both statistical efficiency and interpretability. While various methods have emerged for constructing KGs, existing techniques often lack statistical certainty concerning the presence of links between en...
-
作者:Wang, Fan; Li, Wanshan; Madrid Padilla, Oscar Hernan; Yu, Yi; Rinaldo, Alessandro
作者单位:University of Warwick; University of California System; University of California Los Angeles; University of Texas System; University of Texas Austin
摘要:We study the multilayer random dot product graph (MRDPG) model, a generalization of the random dot product graph model to multilayer networks. To estimate the edge probabilities, we deploy a tensor-based methodology and demonstrate its superiority over existing approaches. Moving to dynamic MRDPGs, we formulate and analyse an online change point detection framework, where, at each time point, we observe a realization from an MRDPG. Across layers, we assume fixed shared common node sets and lat...
-
作者:Chen, Shuyan; Feng, Xingdong; Ge, Yeheng; Li, Tao; Wu, Mengyun
作者单位:Chinese Academy of Sciences; University of Science & Technology of China, CAS; Shanghai University of Finance & Economics; Hong Kong Polytechnic University
摘要:In this paper, we study online estimation and inference of regression coefficients in the presence of hidden confounders by leveraging over-parameterized models. Unlike existing offline approaches that rely on factor and sparse models, our closed-form estimator simultaneously removes hidden-confounder bias and is directly applicable to streaming data. Using tools from random matrix theory, we analyse phase transition phenomena in the variance of the coefficient estimator that arise as the samp...
-
作者:He, Junhui; Ma, Guoxuan; Kang, Jian; Yang, Ying
作者单位:Tsinghua University; University of Michigan System; University of Michigan; Tsinghua University
摘要:We establish a scalable manifold learning method and theory, motivated by the problem of estimating functional magnetic resonance imaging activation manifolds in the Human Connectome Project. Our primary contribution is the development of an efficient estimation technique for heat kernel Gaussian processes in the exponential family model. This approach handles large sample sizes n, preserves the intrinsic geometry of data, and significantly reduces computational complexity from O(n3) to O(n) v...
-
作者:Whiteley, Nick; Gray, Annie; Rubin-Delanchy, Patrick
作者单位:University of Bristol; Alan Turing Institute; University of Edinburgh; University of Edinburgh; Heriot Watt University
摘要:The manifold hypothesis is a widely accepted tenet of machine learning which asserts that nominally high-dimensional data are in fact concentrated near a low-dimensional manifold, embedded in high-dimensional space. This phenomenon is observed empirically in many real-world situations, has led to development of a wide range of statistical methods in the last few decades, and has been suggested as a key factor in the success of modern AI technologies. We show that rich and sometimes intricate m...
-
作者:Auddy, Arnab; Cai, T. Tony; Chakraborty, Abhinav
作者单位:University System of Ohio; Ohio State University; University of Pennsylvania; Columbia University
摘要:This paper considers minimax and adaptive transfer learning for nonparametric classification under the posterior drift model with distributed differential privacy constraints. Our study is conducted within a heterogeneous framework, encompassing diverse sample sizes, varying privacy parameters, and data heterogeneity across different servers. We first establish the minimax misclassification rate, precisely characterizing the effects of privacy constraints, source samples, and target samples on...
-
作者:Song, Shanshan; Wang, Tong; Shen, Guohao; Lin, Yuanyuan; Huang, Jian
作者单位:Tongji University; Tongji University; Chinese University of Hong Kong; Hong Kong Polytechnic University; Hong Kong Polytechnic University; Hong Kong Polytechnic University
摘要:In this paper, we propose a new and unified approach for nonparametric regression and conditional distribution learning. Our approach simultaneously estimates a regression function and a conditional generator using a generative learning framework, where a conditional generator is a function that can generate samples from a conditional distribution. The main idea is to estimate a conditional generator satisfying the constraint that it produces a good regression function estimator. We use deep n...
-
作者:Gao, Ming; Tai, Wai Ming; Aragam, Bryon
作者单位:University of Chicago
摘要:We introduce a new method for neighbourhood selection in linear structural equation models that improves over classical methods such as best subset selection (BSS) and the Lasso. Our method, called KL-BSS, takes advantage of the existence of underlying structure in SEM-even when this structure is unknown-and is easily implemented using existing solvers. Under weaker eigenvalue conditions compared to BSS and the Lasso, KL-BSS can provably recover the support of linear models with fewer samples....