-
作者:Bhaduri, Ritwik; Bhattacharyya, Aabesh; Barber, Rina Foygel; Janson, Lucas
作者单位:Harvard University; University of Chicago
摘要:Tests of goodness of fit are used in nearly every domain where statistics is applied. One powerful and flexible approach is to sample artificial datasets that are exchangeable with the real data under the null hypothesis (but not under the alternative), as this allows the analyst to conduct a valid test using any test statistic they desire. Such sampling is typically done by conditioning on either an exact or an approximate sufficient statistic, but existing methods for doing so have significa...
-
作者:Xu, Congbin; Yu, Yue; Wang, Zhaojun; Zou, Changliang; Ren, Haojie
作者单位:Nankai University; Nankai University; Shanghai Jiao Tong University
摘要:Conformal prediction offers a distribution-free framework for constructing prediction sets with finite-sample coverage. Yet efficiently leveraging multiple nonconformity scores to reduce set sizes remains an open challenge. Instead of selecting a single best score, this work introduces a principled aggregation strategy that intersects multiple conformal prediction sets, with confidence levels allocated optimally to minimize the empirical set size while maintaining asymptotic coverage. Two vari...
-
作者:Chen, Xuyang; Wang, Yinjie; Tang, Weijing
作者单位:University of Pennsylvania; University of Chicago; Carnegie Mellon University
摘要:In many real-world networks, relationships often go beyond simple dyadic presence or absence; they can be positive, such as friendship, alliance and mutualism, or negative, characterized by enmity, disputes and competition. To understand the mechanisms of formation of such signed networks, social balance theory sheds light on the dynamics of positive and negative connections. In particular, it characterizes the proverbs 'a friend of my friend is my friend' and 'an enemy of my enemy is my frien...
-
作者:Heng, Pei; He, Shiyuan; Sun, Yi; Guo, Jianhua
作者单位:Northeast Normal University - China; Beijing Technology & Business University
摘要:Collapsibility provides a principled approach to dimension reduction in contingency tables and graphical models. Madigan & Mosurski (1990) pioneered the study of minimal collapsible sets in decomposable models, but existing algorithms for general graphs remain computationally demanding. We show that a model is collapsible on to a target set precisely when that set contains at least one minimal separator between its nonadjacent vertices. This insight motivates the close minimal separator absorp...
-
作者:Li, Yunchen; Wang, Guanghui; Xu, Shuntuo; Yu, Zhou
作者单位:East China Normal University; Nankai University; Nankai University
摘要:We propose CUSUM-Net, a nonparametric method for changepoint detection based on integral probability metrics and deep neural networks. Our approach learns a critic function by maximizing an aggregate CUSUM objective over candidate changepoints, thereby linking changepoint detection to optimization of two-sample integral probability metrics. The learned critic induces a one-dimensional representation on which changepoints are localized by a classical CUSUM scan. Unlike parametric procedures, CU...
-
作者:Sengupta, Saikat; Khamaru, Koulik; Ghosh, Suvrojit; Dasgupta, Tirthankar
作者单位:Indian Statistical Institute; Indian Statistical Institute Kolkata; Rutgers University System; Rutgers University New Brunswick
摘要:We study the problem of estimating the average treatment effect under sequentially adaptive treatment assignment mechanisms. In contrast to classical completely randomized designs, the setting we consider is one in which the probability of assigning treatment to each experimental unit may depend on prior assignments and observed outcomes. Within the potential outcomes framework (), we propose and analyse two natural estimators for the average treatment effect: the inverse propensity weighted e...
-
作者:Bing, Xin; Kong, Dehan; Li, Bingqing
作者单位:University of Toronto; University of Toronto
摘要:Gaussian mixture models are fundamental statistical tools for modelling heterogeneous data. Due to the nonconcavity of the likelihood function, the expectation-maximization (EM) algorithm is widely used for parameter estimation of each Gaussian component. Existing analyses of the EM algorithm's convergence to the true parameter focus on either the two-component case or multi-component settings with known mixing probabilities and isotropic covariance matrices. In this work, we study the converg...
-
作者:Farina, Rebecca; Tchetgen Tchetgen, Eric; Kuchibhotla, Arun Kumar
作者单位:Carnegie Mellon University; University of Pennsylvania; Carnegie Mellon University
摘要:Our objective is to construct well-calibrated prediction sets for a time-to-event outcome subject to right censoring with guaranteed coverage. Inspired by modern conformal inference, our approach avoids the need for a well-specified parametric or semiparametric survival model. Unlike existing conformal methods for survival data, which assume Type-I censoring with fully observed censoring times, we consider the more common right-censoring setting in which only the censoring time or the event ti...
-
作者:Banerjee, Bilol; Bhattacharya, Bhaswar B.; Ghosh, Anil K.
作者单位:Indian Statistical Institute; Indian Statistical Institute Kolkata; University of Pennsylvania
摘要:In this paper we introduce a new measure of conditional dependence between two random vectors $ {\boldsymbol{X}} $ and $ {\boldsymbol{Y}} $, given another random vector $ \boldsymbol{Z} $ using the ball divergence. Our measure characterizes conditional independence and does not require any moment assumptions. We propose an estimator of the measure using a kernel-averaging technique and derive its asymptotic distribution. Using this estimator, we construct a test for conditional independence ba...
-
作者:Lin, Zeqin; Liu, Yiming; Pan, Guangming; Yao, Chi; Zhou, Jia
作者单位:Nanyang Technological University; Jinan University; Anhui University; Hefei University of Technology
摘要:We consider the problem of identifying the pattern of latent variables in high-dimensional linear latent variable models, which can also be interpreted as determining the source of spiked singular values in the data matrix. Specifically, we test whether the latent variables are continuous or categorical, a distinction that is crucial for data interpretation, but challenging in the high-dimensional regime. To address this inference problem, we analyse the asymptotic behaviour of empirical measu...