-
作者:Su, Buxin; Zhang, Jiayao; Collina, Natalie; Yan, Yuling; Li, Didong; Cho, Kyunghyun; Fan, Jianqing; Roth, Aaron; Su, Weijie
作者单位:University of Pennsylvania; University of Pennsylvania; University of Wisconsin System; University of Wisconsin Madison; University of North Carolina; University of North Carolina Chapel Hill; New York University; Princeton University
摘要:We conducted an experiment during the review process of the 2023 International Conference on Machine Learning (ICML), asking authors with multiple submissions to rank their papers based on perceived quality. In total, we received 1342 rankings, each from a different author, covering 2592 submissions. In this article, we present an empirical analysis of how author-provided rankings could be leveraged to improve peer review processes at machine learning conferences. We focus on the Isotonic Mech...
-
作者:Dahl, David B.; Warr, Richard L.; Jensen, Thomas P.
作者单位:Brigham Young University; Berry Consultants, LLC
摘要:Although exchangeable processes from Bayesian nonparametrics have been used as a generating mechanism for random partition models, we deviate from this paradigm to explicitly incorporate clustering information in the formulation of our random partition model. Our shrinkage partition distribution takes any partition distribution and shrinks its probability mass toward a specific anchor partition. We show how this provides a framework to model hierarchically-dependent and temporally-dependent ra...
-
作者:Huang, Xinmeng; Xu, Kan; Lee, Donghwan; Hassani, Hamed; Bastani, Hamsa; Dobriban, Edgar
作者单位:University of Pennsylvania; Arizona State University; Arizona State University-Tempe; University of Pennsylvania; University of Pennsylvania; University of Pennsylvania
摘要:Large and complex datasets are often collected from several, possibly heterogeneous sources. Multitask learning methods improve efficiency by leveraging commonalities across datasets while accounting for possible differences among them. Here, we study multitask linear regression and contextual bandits under sparse heterogeneity, where the source/task-associated parameters are equal to a global parameter plus a sparse task-specific term. We propose a novel two-stage estimator called MOLAR that ...
-
作者:Sarkar, Sanat K.
作者单位:Pennsylvania Commonwealth System of Higher Education (PCSHE); Temple University
-
作者:Imai, Kosuke; Nakamura, Kentaro
作者单位:Harvard University; Harvard University; Harvard University
摘要:In this article, we demonstrate how to enhance the validity of causal inference with unstructured high-dimensional treatments like texts, by leveraging the power of generative Artificial Intelligence (GenAI). Specifically, we propose to use a deep generative model such as large language models (LLMs) to efficiently generate treatments and use their internal representation for subsequent causal effect estimation. We show that the knowledge of this true internal representation helps disentangle ...
-
作者:Kent, Alexander; Berrett, Thomas B.; Yu, Yi
作者单位:University of Warwick
摘要:Most literature on differential privacy considers the item-level case where each user has a single observation, but a growing field is that of user-level privacy where each of the n users holds T observations and wishes to maintain the privacy of their entire collection. We derive a general minimax lower bound, which shows that, for locally private user-level estimation problems, the risk cannot, in general, be made to vanish for a fixed n even for T arbitrarily large. We then derive matching,...
-
作者:Qiao, Xinghao; Wang, Zihan; Yao, Qiwei; Zhang, Bo
作者单位:University of Hong Kong; Tsinghua University; University of London; London School Economics & Political Science; Chinese Academy of Sciences; University of Science & Technology of China, CAS
摘要:The factor modeling for high-dimensional time series is powerful in discovering latent common components for dimension reduction and information extraction. Most available estimation methods can be divided into two categories: the covariance-based under asymptotically-identifiable assumption and the autocovariance-based with white idiosyncratic noise. This article follows the autocovariance-based framework and develops a novel weight-calibrated method to improve the estimation performance. It ...
-
作者:Xiong, Xin; Guo, Zijian; Cai, Tianxi
作者单位:Harvard University; Harvard T.H. Chan School of Public Health; Zhejiang University; Harvard University; Harvard Medical School
摘要:Transfer learning is a critical technique that enables the application of knowledge gained from existing tasks or domains to improve performance on a new one, reducing the need for extensive data and training in each new context. Many existing transfer learning methods rely on leveraging information from source populations closely resembling the target population. However, this approach often overlooks valuable knowledge that may be present in different yet potentially related auxiliary sample...
-
作者:Chen, Ling; Huang, Chengzhu; Gu, Yuqi
作者单位:Columbia University
摘要:This work focuses on the mixed membership models for multivariate categorical data widely used for analyzing survey responses and population genetics data. These grade of membership (GoM) models offer rich modeling power but present significant estimation challenges for high-dimensional polytomous data. Popular existing approaches, such as Bayesian MCMC inference, are not scalable and lack theoretical guarantees in high-dimensional settings. To address this, we first observe that data from thi...
-
作者:Haupt, Andreas; Koyejo, Sanmi
作者单位:Stanford University