-
作者:Walther, Guenther; Zhao, Qian
作者单位:Stanford University; University of Massachusetts System; University of Massachusetts Amherst
摘要:Multivariate histograms are difficult to construct due to the curse of dimensionality. Motivated by k-d trees in computer science, we show how to construct an efficient data-adaptive partition of Euclidean space that possesses the following two properties: With high confidence the distribution from which the data are generated is close to uniform on each rectangle of the partition; and despite the data-dependent construction we can give guaranteed finite sample simultaneous confidence interval...
-
作者:He, Shuren; Sang, Huiyan; Zhou, Quan
作者单位:Texas A&M University System; Texas A&M University College Station
摘要:Ensemble decision tree methods such as XGBoost, Random Forest, and Bayesian Additive Regression Trees (BART) have gained enormous popularity in data science for their superior performance in machine learning regression and classification tasks. In this article, we introduce a new Bayesian graph-split additive decision tree method, GS-BART, designed to enhance the performance of axis-parallel split-based BART for dependent data with graph structures. The proposed approach encodes input feature ...
-
作者:Shan, Jiawei; Li, Wei; Ai, Chunrong
作者单位:Renmin University of China; Renmin University of China; Renmin University of China; The Chinese University of Hong Kong, Shenzhen
摘要:Mediation analysis is widely used for exploring treatment mechanisms; however, it faces challenges when nonignorable missing confounders are present. Efficient inference of mediation effects and the efficiency loss due to nonignorable missingness have been rarely studied in the literature because of the difficulties arising from the ill-posed inverse problem. In this article, we propose a general shadow variable framework for identifying mediation effects, allowing shadow variables to be selec...
-
作者:Xia, Junwen; Zhang, Jingxiao; Kong, Dehan
作者单位:Renmin University of China; Renmin University of China; University of Toronto
摘要:Quantile optimal treatment regimes (OTRs) aim to assign treatments that maximize a specified quantile of patients' outcomes. Compared to treatment regimes that target the mean outcomes, quantile OTRs offer fairer regimes when a welower quantile is selected, as it improves outcomes for vulnerable patients. In this article, we propose a novel method for estimating quantile OTRs by reformulating the problem as a successive classification task, solvable via training a sequence of classifiers, each...
-
作者:Zhu, Claire R.; Alt, Ethan M.; Ibrahim, Joseph G.
作者单位:University of North Carolina; University of North Carolina Chapel Hill; University of North Carolina School of Medicine
摘要:An increasingly common approach in clinical trials involves leveraging external data to supplement the sample size of randomized controlled trials. Although designing studies with concurrent data allows direct treatment comparisons, recruiting an adequate number of participants, especially in the context of rare diseases, can be challenging, time-intensive, and financially burdensome. Incorporating external data offers a cost-effective solution to enhance trial data; however, potential non-exc...
-
作者:Fan, Jianqing; Li, Yingying; Xia, Ningning; Zheng, Xinghua
作者单位:Fudan University; Hong Kong University of Science & Technology; Shanghai University of Finance & Economics; Princeton University
摘要:We establish central limit theorems for principal eigenvalues and eigenvectors under divergent spiked covariance models, and develop three two-sample tests for testing (a) the equality of principal eigenvalues, (b) the stability of eigenvalue proportions, and (c) the invariance of principal eigenvectors. One important application of our results is to understand structural breaks in large factor models. That is, when a structural break is suspected or detected, these tests provide unique insigh...
-
作者:Zhu, Junhao; Kong, Dehan; Zhang, Zhaolei; Lin, Zhenhua
作者单位:University of Toronto; National University of Singapore
摘要:In modern interdisciplinary research, manifold time series data have been garnering more attention. A critical question in analyzing such data is stationarity, which reflects the underlying dynamic behavior and is crucial across various fields like cell biology, neuroscience and empirical finance. Yet, there has been an absence of a formal definition of stationarity that is tailored to manifold time series. This work bridges this gap by proposing the first definitions of first-order and second...
-
作者:Cai, T. Tony; Chakraborty, Abhinav; Wang, Yichen
作者单位:University of Pennsylvania; Columbia University
摘要:Data privacy is a central concern in many applications involving ranking from incomplete and noisy pairwise comparisons, such as recommendation systems, educational assessments, and opinion surveys on sensitive topics. In this work, we propose differentially private algorithms for ranking based on pairwise comparisons. Specifically, we develop and analyze ranking methods under two privacy notions: edge differential privacy, which protects the confidentiality of individual comparison outcomes, ...
-
作者:Ye, Shangyuan; Rakshe, Shauna; Liang, Ye
作者单位:State University System of Florida; Florida International University; Oregon Health & Science University; Oklahoma State University System; Oklahoma State University - Stillwater
摘要:Simultaneous variable selection and statistical inference is challenging in high-dimensional data analysis. Most existing post-selection inference methods require explicitly specified regression models, which are often linear, as well as sparsity in the regression model. The performance of such procedures can be poor under either misspecified nonlinear models or a violation of the sparsity assumption. In this article, we propose a sufficient dimension association (SDA) technique that measures ...
-
作者:Yu, Tao; Qin, Jing; Li, Pengfei
作者单位:National University of Singapore; National Institutes of Health (NIH) - USA; NIH National Institute of Allergy & Infectious Diseases (NIAID); University of Waterloo
摘要:Multivariate mixture data analysis presents numerous challenges and constitutes a vital area of interest in the fields of statistics and data science. Research into multivariate mixture structures holds relevance across diverse application domains and plays a pivotal role in the advancement of artificial intelligence and machine learning. In this article, we focus on nonparametric estimation techniques for multivariate mixture data. Specifically, we assume a known number of subpopulations and ...