-
作者:Kang, Jizhou; Kottas, Athanasios
作者单位:University of California System; University of California Santa Cruz
摘要:We develop a nonparametric Bayesian modeling approach to ordinal regression based on priors placed directly on the discrete distribution of the ordinal responses. The prior probability models are built from a structured mixture of multinomial distributions. We leverage the continuation-ratio logits representation to formulate the mixture kernel, with mixture weights defined through the logit stick-breaking process that incorporates the covariates through a linear function. The implied regressi...
-
作者:Han, Ruijian; Luo, Lan; Luo, Yuanhang; Lin, Yuanyuan; Huang, Jian
作者单位:Hong Kong Polytechnic University; Rutgers University System; Chinese University of Hong Kong; Hong Kong Polytechnic University; Hong Kong Polytechnic University
摘要:Online statistical inference facilitates real-time analysis of sequentially collected data, making it different from traditional methods that rely on static datasets. This article introduces a novel approach to online inference in high-dimensional generalized linear models, where we update regression coefficient estimates and their standard errors upon each new data arrival. In contrast to existing methods that either require full dataset access or large-dimensional summary statistics storage,...
-
作者:Zheng, Lili; Chang, Andersen; Allen, Genevera I.
作者单位:University of Illinois System; University of Illinois Urbana-Champaign; Baylor College of Medicine; Columbia University; University of Illinois System; University of Illinois Urbana-Champaign
摘要:Patchwork learning arises as a new and challenging data collection paradigm where both samples and features are observed in fragmented subsets. Due to technological limitations and measurement expenses, such patchwork data structures are frequently seen in applications like neuroscience, healthcare, and genomics, among others. Instead of analyzing each data patch separately, it is highly desirable to extract comprehensive knowledge from the whole dataset. In this work, we focus on the clusteri...
-
作者:Menicali, Luca; Grace, Andrew P.; Richter, David H.; Castruccio, Stefano
作者单位:University of Notre Dame; University of Notre Dame
摘要:Fluid thermodynamics underpins atmospheric dynamics, climate science, industrial applications, and energy systems. However, direct numerical simulations (DNS) of such systems can be computationally prohibitive. To address this, we present a novel physics-informed spatiotemporal surrogate model for Rayleigh-B & eacute;nard convection (RBC), a canonical example of convective fluid flow. Our approach combines convolutional neural networks, for spatial dimension reduction, with an innovative recur...
-
作者:Grunwald, Peter D.
作者单位:Centrum Wiskunde & Informatica (CWI); Leiden University - Excl LUMC; Leiden University
-
作者:Bodelet, Julien; Blanc, Guillaume; Shan, Jiajun; Muniz Terrera, Graciela; Chen, Oliver Y.
作者单位:University of Lausanne; Centre Hospitalier Universitaire Vaudois (CHUV); University of Lausanne; University of Zurich; University of Geneva; University System of Ohio; Ohio University; University of Edinburgh
摘要:Large and complex datasets, emerging from technological advancements in fields such as genomics and brain imaging, hold ample promise for gaining new scientific insights. Yet, their inherent nonlinearity and high dimensionality present considerable theoretical, methodological, and application challenges to the statistics and machine learning community. This article introduces Statistical Quantile Learning (SQL), a new nonparametric method for estimating large additive latent variable models. T...
-
作者:Gao, Youqian; Dai, Ben
作者单位:Chinese University of Hong Kong
摘要:The technique of word embedding is widely used in natural language processing (NLP) to represent words as numerical vectors in textual datasets. However, the estimation of word embedding may suffer from severe overfitting due to the huge variety of words. To address the issue, this article proposes a novel regularization framework that recognizes and accounts for the word-level distribution discrepancy-a common phenomenon in a range of NLP tasks where word distributions are noticeably disparat...
-
作者:Morey, Richard D.; Davis-Stober, Clintin P.
作者单位:Cardiff University; University of Missouri System; University of Missouri Columbia; University of Missouri System; University of Missouri Columbia
摘要:The P-curve is a widely used suite of meta-analytic tests advertised for detecting problems in sets of studies. They are based on nonparametric combinations of p values (e.g., Marden) across significant (p < .05) studies and are variously claimed to detect evidential value, lack of evidential value, and left skew in p values. We show that these tests do not have the properties ascribed to them. Moreover, they fail basic desiderata for tests, including admissibility and monotonicity. In light o...
-
作者:He, Xianwen; Li, Yao
作者单位:University of North Carolina; University of North Carolina Chapel Hill; University of North Carolina School of Medicine
摘要:With the widespread application of machine learning algorithms in daily life, it is crucial to mitigate the risk of these algorithms producing socially undesirable outcomes that may disproportionately disadvantage certain groups or individuals based on demographic characteristics such as gender, race, or disabilities. In recent years, machine learning fairness has gained increasing attention from both researchers and the public. This article provides a comprehensive overview of fairness-enhanc...
-
作者:Watson, Samuel I.; Smith, Thomas A.
作者单位:University of Birmingham
摘要:In this article, we consider randomized trial design to evaluate interventions with spatially or spatio-temporally heterogeneous effects. A common approach in this setting is the cluster randomized trial. In many cases, clusters are constituted as discrete subdivisions of a contiguous area of interest. However, cluster trials designed in this way may suffer from issues of spillover and may fail to capture the relevant spatial and temporal effects. We define possible randomization schemes and c...