-
作者:Xia, Xintao; Zhang, Linjun; Cai, Zhanrui
作者单位:Iowa State University; Rutgers University System; Rutgers University New Brunswick; University of Hong Kong
摘要:Privacy preservation has become a critical concern in high-dimensional data analysis due to the growing prevalence of data-driven applications. Since its proposal, sliced inverse regression has emerged as a widely used statistical technique to reduce the dimensionality of covariates while maintaining sufficient statistical information. In this used, we propose optimally differentially private algorithms specifically designed to address privacy concerns in the context of sufficient dimension re...
-
作者:Ohnishi, Yuki; Li, Fan
作者单位:Yale University; Yale University
摘要:Cluster randomized trials (CRTs) with multiple unstructured mediators present significant methodological challenges for causal inference due to within-cluster correlation, interference among units, and the complexity introduced by multiple mediators. Existing causal mediation methods often fall short in simultaneously addressing these complexities, particularly in disentangling mediator-specific effects under interference that are central to studying complex mechanisms. To address this gap, we...
-
作者:Kim, Rakheon; Zhang, Jingfei
作者单位:Baylor University; Emory University
摘要:While covariance matrices have been widely studied in many scientific fields, relatively limited progress has been made on estimating conditional covariances that permits a large covariance matrix to vary with high-dimensional subject-level covariates. In this article, we present a new sparse covariance regression framework that models the covariance matrix as a function of subject-level covariates. In the context of co-expression quantitative trait locus (QTL) studies, our method can be used ...
-
作者:Schindl, Kyle; Branson, Zach
作者单位:Iowa State University; Carnegie Mellon University
摘要:When designing a randomized experiment, one way to ensure treatment and control groups exhibit similar covariate distributions is to randomize treatment until some prespecified level of covariate balance is satisfied; this strategy is known as rerandomization. Most rerandomization methods use balance metrics based on a quadratic form v(T)Av , where v is a vector of covariate mean differences and A is a positive semi-definite matrix. In this work, we derive general results for treatment-versus-...
-
作者:Dun, Yuzheng; Chatterjee, Nilanjan; Jin, Jin; Nishimura, Akihiko
作者单位:Johns Hopkins University; Johns Hopkins Bloomberg School of Public Health; Johns Hopkins University; Johns Hopkins Medicine; University of Pennsylvania; Pennsylvania Medicine
摘要:Polygenic risk scores (PRS) developed from genome-wide association studies (GWAS) can be used for risk stratification by quantifying the genetic contribution to disease, and many clinical applications have been proposed. Bayesian methods are popular for building PRS because of their natural ability to regularize models and incorporate external information. In this article, we present new theoretical results, methods, and extensive numerical studies to advance Bayesian methods for PRS applicati...
-
作者:Xu, Yuliang; Johnson, Timothy D.; Nichols, Thomas E.; Kang, Jian
作者单位:University of Chicago; University of Michigan System; University of Michigan; University of Oxford
摘要:Bayesian Image-on-Scalar Regression (ISR) provides flexible, uncertainty-aware neuroimaging analysis. However, applying ISR to large-scale datasets such as the UK Biobank is challenging due to intensive computational demands and the need to handle subject-specific brain masks rather than a common mask. We propose a novel Bayesian ISR model that scales efficiently while accommodating these inconsistent masks. Our method leverages Gaussian process priors with salience area indicators and introdu...
-
作者:Manderson, Andrew A.; Goudie, Robert J. B.
作者单位:MRC Biostatistics Unit; University of Cambridge
摘要:When complex Bayesian models exhibit implausible behavior, one solution is to assemble available information into an informative prior. Challenges arise as prior information is often only available for the observable quantity, or some model-derived marginal quantity, rather than directly pertaining to the (usually latent) parameters in our model. We propose a method for translating available prior information, in the form of an elicited distribution for the observable or model-derived marginal...
-
作者:Gai, Xin; Jiang, Shiyi; Zhang, Anru R.
作者单位:Vanderbilt University; Duke University; Duke University; Duke University
摘要:Electronic Health Records (EHRs) contain extensive patient information that can inform downstream clinical decisions, such as mortality prediction, disease phenotyping, and disease onset prediction. A key challenge in EHR data analysis is the temporal gap between when a condition is first recorded and its actual onset time. Such timeline misalignment can lead to artificially distinct biomarker trends among patients with similar disease progression, undermining the reliability of downstream ana...
-
作者:Montanari, Andrea; Wu, Yuchen
作者单位:Stanford University; Stanford University; University of Pennsylvania
摘要:We consider the problem of sampling from the posterior distribution of a d-dimensional coefficient vector theta, given linear observations y=X theta+epsilon. For sparse Bayesian models, such posteriors are in general multimodal, and therefore challenging to sample from. This observation has prompted the exploration of various heuristics that aim at approximating the posterior distribution. In this article, we study a different approach based on decomposing the posterior distribution into a log...
-
作者:Janvin, Matias; Stensrud, Mats J.
作者单位:University of Oslo; Diakonhjemmet Hospital; Swiss Federal Institutes of Technology Domain; Ecole Polytechnique Federale de Lausanne