-
作者:Tian, Xinyu; Shen, Xiaotong
作者单位:University of Minnesota System; University of Minnesota Twin Cities
摘要:Reliable machine learning and statistical analysis rely on diverse, well-distributed training data. However, real-world datasets are often limited in size and exhibit underrepresentation across key subpopulations, leading to biased predictions and reduced performance, particularly in supervised tasks such as classification. To address these challenges, we propose Conditional Data Synthesis Augmentation (CoDSA), a novel framework that leverages generative models, such as diffusion models, to sy...
-
作者:Shi, Jiaxin; Zhu, Xuening; Zhou, Jing; Yu, Baichen; Wang, Hansheng
作者单位:Peking University; Fudan University; Fudan University; Fudan University; Renmin University of China
摘要:We study one particular type of multivariate spatial autoregression (MSAR) model with diverging dimensions in both responses and covariates. This makes the usual MSAR models no longer applicable due to the high computational cost. To address this issue, we propose a factor-augmented spatial autoregression (FSAR) model. FSAR is a special case of MSAR but with a novel factor structure imposed on the high-dimensional random error vector. The latent factors of FSAR are assumed to be of a fixed dim...
-
作者:Magnani, Chiara G.; Sesia, Matteo; Solari, Aldo
作者单位:Bocconi University; University of Southern California; University of Southern California; Universita Ca Foscari Venezia; University of Milano-Bicocca
摘要:This article develops a flexible distribution-free method for collective outlier detection and enumeration, designed for situations in which the presence of outliers can be detected powerfully even though their precise identification may be challenging due to the sparsity, weakness, or elusiveness of their signals. This method builds upon recent developments in conformal inference and integrates classical ideas from other areas, including multiple testing, locally most powerful and adaptive ra...
-
作者:Zhou, Xingcai; Xu, Zinan; Jiang, Bei; Kong, Linglong
作者单位:Nanjing Audit University; University of Alberta
摘要:Glioblastoma multiforme (GBM) is a highly aggressive brain cancer with largely ineffective treatment. It is imperative to explore more effective therapies, such as gene-based treatments. For co-expression QTL studies of GBM, we develop fairness-aware Gaussian graphical regression models (Fair RegGGMs), which can determine how genetic variants modulate subject-level gene networks, and recover both population-level and subject-level gene graphs, while ensuring that the developed learning and inf...
-
作者:Liao, Sijia; Sun, Xiaoxiao; Hao, Ning; Zhang, Hao Helen
作者单位:University of Arizona; University of Arizona; University of Arizona
摘要:The scalar-on-image regression model examines the association between a scalar response and a bivariate function (e.g., images) through the estimation of a bivariate coefficient function. Existing approaches often impose smoothness constraints to control the bias-variance trade-off, and thus prevent overfitting. However, such assumptions can hinder interpretability, especially when only certain regions of an image influence changes in the response. In such a scenario, interpretability can be b...
-
作者:Zhao, Junlong; Zheng, Shengbin; Leng, Chenlei
作者单位:Beijing Normal University; Hong Kong Polytechnic University
摘要:Transfer learning is an emerging paradigm for leveraging multiple sources to improve the statistical inference on a single target. In this article, we propose a novel approach named residual importance weighted transfer learning (RIW-TL) for high-dimensional linear models built on penalized likelihood. Compared to existing methods such as Trans-Lasso that selects sources in an (approximately) all-in-or-all-out manner, RIW-TL includes samples via importance weighting and thus may permit more ef...
-
作者:Wang, Shuqi; Thall, Peter F.; Yuan, Ying; Liu, Suyu
作者单位:University of Texas System; UTMD Anderson Cancer Center
摘要:For first-in-human dose-finding trials, to protect patient safety, regulatory agencies may enforce strict within-cohort staggering rules that require delaying treatment of each patient in the first cohort at an untried dose until dose-limiting toxicities (DLTs) of all previously treated patients have been evaluated. Consequently, many new patients may face therapy delays, which reduces their probability of achieving a response due to disease progression, or be treated off-protocol, which may s...
-
作者:Zhang, Jiazhao; Lin, Chung-Ching; Hung, Ying
作者单位:Rutgers University System; Rutgers University New Brunswick; Microsoft
摘要:Hyperparameter optimization plays a crucial role in the success of neural networks as hyperparameters directly control the behavior and performance of the training algorithms. To obtain efficient tuning, Bayesian optimization based on Gaussian process is widely used. Despite numerous applications in deep learning, the existing methods rely on a convenient but restrictive assumption that the tuning parameters are independent of each other. However, tuning parameters with conditional dependence ...
-
作者:Cai, Junhui; Yang, Dan; Chen, Ran; Shen, Haipeng; Zhao, Linda; Zhu, Wu
作者单位:University of Notre Dame; University of Hong Kong; Washington University (WUSTL); University of Pennsylvania; Tsinghua University
摘要:The centrality in a network is often used to measure nodes' importance and model network effects on a certain outcome. Empirical studies widely adopt a two-stage procedure, which first estimates the centrality from the observed noisy network and then infers the network effect from the estimated centrality, even though it lacks theoretical understanding. We propose a unified modeling framework to study the properties of centrality estimation and inference and the subsequent network regression a...
-
作者:Chang, Xiangyu; Chen, Xi; Lai, Zehua; Li, He; Liu, Zhihong; Zhang, Yichen
作者单位:Xi'an Jiaotong University; New York University; University of Texas System; University of Texas Austin; Purdue University System; Purdue University
摘要:With the fast development of big data, learning the optimal decision rule by recursively updating it and making online decisions has been easier than before. We study the online statistical inference of model parameters in a contextual bandit framework of sequential decision-making. We propose a general framework for an online and adaptive data collection environment that can update decision rules via weighted stochastic gradient descent. We allow different weighting schemes of the stochastic ...