-
作者:Weylandt, Michael; Michailidis, George
作者单位:City University of New York (CUNY) System; Baruch College (CUNY); University of California System; University of California Los Angeles
摘要:Network data are commonly collected in a variety of applications, representing either directly measured or statistically inferred connections between subjects or features of interest. In an increasing number of domains, these networks are collected over time, such as repeated interactions between users of a social media platform, or across multiple subjects, such as in multi-subject neuroimaging studies. When analyzing multiple large networks, dimensionality reduction techniques are often used...
-
作者:Cai, Zhongze; Liu, Shang; Wang, Hanzhao; Zhong, Huaiyang; Li, Xiaocheng
作者单位:Imperial College London; University of Sydney; Virginia Polytechnic Institute & State University
摘要:In this article, we study the problem of watermarking large language models (LLMs). We consider the tradeoff between model distortion and detection ability and formulate it as a constrained optimization problem based on the red-green list watermarking algorithm. We show that the optimal solution to the optimization problem enjoys a nice analytical property which provides a better understanding and inspires the algorithm design for the watermarking process. We develop an online dual gradient as...
-
作者:Shen, Xinwei; Buhlmann, Peter; Taeb, Armeen
作者单位:Swiss Federal Institutes of Technology Domain; ETH Zurich; University of Washington; University of Washington Seattle
摘要:Since distribution shifts are common in real-world applications, there is a pressing need to develop prediction models that are robust against such shifts. Existing frameworks, such as empirical risk minimization or distributionally robust optimization, either lack generalizability for unseen distributions or rely on postulated distance measures. Alternatively, causality offers a data-driven and structural perspective to robust predictions. However, the assumptions necessary for causal inferen...
-
作者:Park, Seyoung; Lee, Eun Ryung; Kim, Hyunjin; Zhao, Hongyu
作者单位:Yonsei University; Yonsei University; Sungkyunkwan University (SKKU); Yale University
摘要:In high-dimensional multiple response regression problems, the large dimensionality of the coefficient matrix poses a challenge to parameter estimation. To address this challenge, low-rank matrix estimation methods have been developed to facilitate parameter estimation in the high-dimensional regime, where the number of parameters increases with sample size. Despite these methodological advances, accurately predicting multiple responses with limited target data remains a difficult task. To gai...
-
作者:Dominitz, Jeff; Manski, Charles F.
作者单位:Rice University; Northwestern University; Northwestern University
摘要:The potential impact of non-sampling errors on election polls is well known, but measurement has focused on the margin of sampling error. Statisticians have recommended measurement of total survey error by mean square error (MSE), which jointly measures sampling and non-sampling errors. We suggest use of the square root of maximum MSE to measure the total margin of error (TME). We suggest that measurement of TME should be a standard feature in the reporting of polls. Because the exceedingly lo...
-
作者:Cheng, Gang; Chen, Yen-Chi; Unger, Joseph M.; Till, Cathee; Zhao, Ying-Qi
作者单位:University of Washington; University of Washington Seattle; Fred Hutchinson Cancer Center
摘要:Combining experimental and observational follow-up datasets has received much attention lately. In a survival setting, recent work has used Medicare claims to extend the follow-up period for participants in a prostate cancer clinical trial. This allows the estimation of the long-term effect that cannot be estimated by the trial data alone. In this article, we study the estimation of long-term effect when participants in a clinical trial are linked to an observational follow-up dataset. Such li...
-
作者:Xu, Qi; Qu, Annie
作者单位:Carnegie Mellon University; University of California System; University of California Santa Barbara
摘要:In the era of big data, large-scale, multi-source, multi-modality datasets are increasingly ubiquitous, offering unprecedented opportunities for predictive modeling and scientific discovery. However, these datasets often exhibit complex heterogeneity, such as covariates shift, posterior drift, and blockwise missingness, which worsen predictive performance of existing supervised learning algorithms. To address these challenges simultaneously, we propose a novel Representation Retrieval ( R-2 ) ...
-
作者:Bersson, Elizabeth
作者单位:Massachusetts Institute of Technology (MIT)
-
作者:Gu, Yifan; Yang, Hanfang; Yang, Songshan; Zou, Hui
作者单位:Renmin University of China; Renmin University of China; Renmin University of China; University of Minnesota System; University of Minnesota Twin Cities
摘要:In modern data analysis, an improvement in statistical efficiency is expected via effective collaboration among multiple data holders with non-shared data. In this article, we propose a collaborative score-type test (CST) for testing linear hypotheses, which accommodates potentially high-dimensional nuisance parameters and a diverging number of constraints and target parameters. Through a careful decomposition of the Kiefer-Bahadur representation for the traditional score statistic, we identif...
-
作者:Srivastava, Radhendushka; Sengupta, Debasis
作者单位:Indian Institute of Technology System (IIT System); Indian Institute of Technology (IIT) - Bombay
摘要:In this article, we present a semiparametric model for describing the effect of temperature on Antarctic ice accumulation on a paleoclimatic time scale. The model is motivated by sharp ups and downs in the rate of ice accumulation apparent from ice core data records, which are synchronous with movements of temperature. We prove consistency of the estimators under reasonable conditions. We conduct extensive simulations to assess the performance of the estimators and their bootstrap based standa...