-
作者:Pandolfi, Andrea; Papaspiliopoulos, Omiros; Zanella, Giacomo
作者单位:Bocconi University; Bocconi University; Bocconi University
摘要:Generalized linear mixed models (GLMMs) are a widely used tool in statistical analysis. The main bottleneck of many computational approaches lies in the inversion of the high dimensional precision matrices associated with the random effects. Such matrices are typically sparse; however, the sparsity pattern resembles a multi partite random graph, which does not lend itself well to default sparse linear algebra techniques. Notably, we show that, for typical GLMMs, the Cholesky factor is dense ev...
-
作者:Duan, Congyuan; Li, Jingyang; Xia, Dong
作者单位:Hong Kong University of Science & Technology; University of Michigan System; University of Michigan
摘要:Is it possible to make online decisions when personalized covariates are unavailable? We take a collaborative-filtering approach for decision-making based on collective preferences. By assuming low-dimensional latent features, we formulate the covariate-free decision-making problem as a matrix completion bandit. We propose a policy learning procedure that combines an epsilon -greedy policy for decision-making with an online gradient descent algorithm for bandit parameter estimation. Our novel ...
-
作者:Shao, Meijia; Xia, Dong; Zhang, Yuan; Wu, Qiong; Chen, Shuo
作者单位:University System of Ohio; Ohio State University; Hong Kong University of Science & Technology; Pennsylvania Commonwealth System of Higher Education (PCSHE); University of Pittsburgh; University System of Maryland; University of Maryland Baltimore
摘要:Two-sample hypothesis testing for network comparison presents many significant challenges, including: leveraging repeated network observations and known node registration, but without requiring them to operate; relaxing strong structural assumptions; achieving finite-sample higher-order accuracy; handling different network sizes and sparsity levels; fast computation and memory parsimony; controlling false discovery rate (FDR) in multiple testing; and theoretical understandings, particularly re...
-
作者:Xu, Yangjianchen; Zeng, Donglin; Lin, D. Y.
作者单位:University of Waterloo; University of Michigan System; University of Michigan; University of North Carolina; University of North Carolina Chapel Hill
摘要:This article presents a general framework for checking the adequacy of the Cox proportional hazards model with interval-censored data, which arise when the event of interest is known only to occur over a random time interval. Specifically, we construct certain stochastic processes that are informative about various aspects of the model, that is, the functional forms of covariates, the exponential link function and the proportional hazards assumption. We establish their weak convergence to zero...
-
作者:Agnoletto, Davide; Rigon, Tommaso; Dunson, David B.
作者单位:Duke University; University of Milano-Bicocca
摘要:This article is motivated by challenges in conducting Bayesian inferences on unknown discrete distributions, with a particular focus on count data. To avoid the computational disadvantages of traditional mixture models, we develop a novel Bayesian predictive approach. In particular, our Metropolis-adjusted Dirichlet (mad) sequence model characterizes the predictive measure as a mixture of a base measure and Metropolis-Hastings kernels centered on previous data points. The resulting mad sequenc...
-
作者:Liang, Ziyi; Xie, Tianmin; Tong, Xin; Sesia, Matteo
作者单位:University of California System; University of California Irvine; University of Southern California; University of Hong Kong; University of Southern California
摘要:We develop a conformal inference method to construct joint prediction regions for structured groups of missing entries in a sparsely observed matrix, focusing on groups drawn from the same column. The method can be combined with any black-box matrix completion algorithm and makes no distributional assumptions for the underlying data matrix; instead, it obtains rigorous inferences by modeling the missingness mechanism. In the context of recommender systems, for example, it is useful to quantify...
-
作者:Namdari, Jamshid; Manatunga, Amita; Ferrarelli, Fabio; Krafty, Robert T.
作者单位:Emory University; Rollins School Public Health; Pennsylvania Commonwealth System of Higher Education (PCSHE); University of Pittsburgh
摘要:Principal component analysis has been a main tool in multivariate analysis for estimating a low dimensional linear subspace that explains most of the variability in the data. However, in high-dimensional regimes, naive estimates of the principal loadings are not consistent and difficult to interpret. In the context of time series, principal component analysis of spectral density matrices can provide valuable, parsimonious information about the behavior of the underlying process, particularly i...
-
作者:Zhang, Likun; Ma, Xiaoyu; Wikle, Christopher K.; Huser, Raphael
作者单位:University of Missouri System; University of Missouri Columbia; King Abdullah University of Science & Technology
摘要:Many real-world processes have complex tail dependence structures that cannot be adequately characterized using classical Gaussian processes. Alternatively, models motivated by extreme-value theory exhibit appealing extremal dependence properties but are often exceedingly prohibitive to fit and simulate from in high dimensions using classical methods. In this article, we extend the boundaries on computation and modeling of high-dimensional spatial extremes by integrating a new flexible and non...
-
作者:Wang, Zeya; Ye, Chenglong
作者单位:University of Kentucky
摘要:Deep clustering partitions complex high-dimensional data using deep neural networks for clustering. It involves projecting data into lower-dimensional embeddings before partitioning, which embarks unique evaluation challenges. Traditional clustering validation measures, designed for low-dimensional spaces, are problematic for deep clustering for two reasons: (a) the curse of dimensionality when applied to the high-dimensional input data, and (b) unreliable comparison of clustering results when...
-
作者:He, Di; Zou, Hui
作者单位:Nanjing University; University of Minnesota System; University of Minnesota Twin Cities
摘要:Canonical correlation analysis (CCA) is an important statistical technique that explores the linear relationships between two sets of variables. In this article, we propose a new generalization of CCA named sparse Gaussianized CCA (SGCCA) for high-dimensional data analysis. SGCCA has a number of favorable properties. First, it is conceptually easy to comprehend and efficient to implement. Second, it not only yields sparse and nested canonical vectors, but is also invariant against monotone tra...