-
作者:Kock, Anders B.; Pedersen, Rasmus S.; Sorensen, Jesper R. -V.
作者单位:University of Oxford; University of Copenhagen; Danish Finance Institute
摘要:Lasso-type estimators are routinely used to estimate high-dimensional time series models. The theoretical guarantees established for these estimators typically require the penalty level to be chosen in a suitable fashion often depending on unknown population quantities. Furthermore, the resulting estimates and the number of variables retained in the model depend crucially on the chosen penalty level. However, there is currently no theoretically founded guidance for this choice in the context o...
-
作者:Grunwald, Peter D.
作者单位:Centrum Wiskunde & Informatica (CWI); Leiden University - Excl LUMC; Leiden University
-
作者:Bodelet, Julien; Blanc, Guillaume; Shan, Jiajun; Muniz Terrera, Graciela; Chen, Oliver Y.
作者单位:University of Lausanne; Centre Hospitalier Universitaire Vaudois (CHUV); University of Lausanne; University of Zurich; University of Geneva; University System of Ohio; Ohio University; University of Edinburgh
摘要:Large and complex datasets, emerging from technological advancements in fields such as genomics and brain imaging, hold ample promise for gaining new scientific insights. Yet, their inherent nonlinearity and high dimensionality present considerable theoretical, methodological, and application challenges to the statistics and machine learning community. This article introduces Statistical Quantile Learning (SQL), a new nonparametric method for estimating large additive latent variable models. T...
-
作者:Gao, Youqian; Dai, Ben
作者单位:Chinese University of Hong Kong
摘要:The technique of word embedding is widely used in natural language processing (NLP) to represent words as numerical vectors in textual datasets. However, the estimation of word embedding may suffer from severe overfitting due to the huge variety of words. To address the issue, this article proposes a novel regularization framework that recognizes and accounts for the word-level distribution discrepancy-a common phenomenon in a range of NLP tasks where word distributions are noticeably disparat...
-
作者:Morey, Richard D.; Davis-Stober, Clintin P.
作者单位:Cardiff University; University of Missouri System; University of Missouri Columbia; University of Missouri System; University of Missouri Columbia
摘要:The P-curve is a widely used suite of meta-analytic tests advertised for detecting problems in sets of studies. They are based on nonparametric combinations of p values (e.g., Marden) across significant (p < .05) studies and are variously claimed to detect evidential value, lack of evidential value, and left skew in p values. We show that these tests do not have the properties ascribed to them. Moreover, they fail basic desiderata for tests, including admissibility and monotonicity. In light o...
-
作者:Watson, Samuel I.; Smith, Thomas A.
作者单位:University of Birmingham
摘要:In this article, we consider randomized trial design to evaluate interventions with spatially or spatio-temporally heterogeneous effects. A common approach in this setting is the cluster randomized trial. In many cases, clusters are constituted as discrete subdivisions of a contiguous area of interest. However, cluster trials designed in this way may suffer from issues of spillover and may fail to capture the relevant spatial and temporal effects. We define possible randomization schemes and c...
-
作者:Betancourt, Brenda
作者单位:George Mason University
-
作者:Ohnishi, Yuki; Li, Fan
作者单位:Yale University; Yale University
摘要:Cluster randomized trials (CRTs) with multiple unstructured mediators present significant methodological challenges for causal inference due to within-cluster correlation, interference among units, and the complexity introduced by multiple mediators. Existing causal mediation methods often fall short in simultaneously addressing these complexities, particularly in disentangling mediator-specific effects under interference that are central to studying complex mechanisms. To address this gap, we...
-
作者:Kim, Rakheon; Zhang, Jingfei
作者单位:Baylor University; Emory University
摘要:While covariance matrices have been widely studied in many scientific fields, relatively limited progress has been made on estimating conditional covariances that permits a large covariance matrix to vary with high-dimensional subject-level covariates. In this article, we present a new sparse covariance regression framework that models the covariance matrix as a function of subject-level covariates. In the context of co-expression quantitative trait locus (QTL) studies, our method can be used ...
-
作者:Montanari, Andrea; Wu, Yuchen
作者单位:Stanford University; Stanford University; University of Pennsylvania
摘要:We consider the problem of sampling from the posterior distribution of a d-dimensional coefficient vector theta, given linear observations y=X theta+epsilon. For sparse Bayesian models, such posteriors are in general multimodal, and therefore challenging to sample from. This observation has prompted the exploration of various heuristics that aim at approximating the posterior distribution. In this article, we study a different approach based on decomposing the posterior distribution into a log...