Representation Retrieval Learning for Heterogeneous Data Integration

成果类型:
Article; Early Access
署名作者:
Xu, Qi; Qu, Annie
署名单位:
Carnegie Mellon University; University of California System; University of California Santa Barbara
刊物名称:
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION
ISSN/ISSBN:
0162-1459; 1537-274X
DOI:
10.1080/01621459.2026.2696485
发表日期:
2026-09-02
关键词:
Blockwise missing data Excess risk bound Multi-modality data multi-task learning Representation learning inference sparsity joint
摘要:
In the era of big data, large-scale, multi-source, multi-modality datasets are increasingly ubiquitous, offering unprecedented opportunities for predictive modeling and scientific discovery. However, these datasets often exhibit complex heterogeneity, such as covariates shift, posterior drift, and blockwise missingness, which worsen predictive performance of existing supervised learning algorithms. To address these challenges simultaneously, we propose a novel Representation Retrieval ( R-2 ) framework, which integrates a dictionary of representation learning modules (representer dictionary) with data source-specific sparsity-induced machine learning model (learners). Under the R-2 framework, we introduce the notion of integrativeness for each representer, and propose a novel Selective Integration Penalty (SIP) to explicitly encourage more integrative representers to improve predictive performance. Theoretically, we show that the excess risk bound of the R-2 framework is characterized by the integrativeness of representers, and SIP effectively improves the excess risk. Extensive simulation studies validate the superior performance of R-2 framework and the effect of SIP. We further apply our method to two real-world datasets to confirm its empirical success. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
来源URL: