A STRUCTURED ESTIMATOR FOR LARGE COVARIANCE MATRICES IN THE PRESENCE OF PAIRWISE AND SPATIAL COVARIATES
成果类型:
Article
署名作者:
Metodiev, Martin; Perrot-Dockes, Marie; Ouadah, Sarah; Fosdick, Bailey k.; Robin, Stephane; Latouche, Pierre; Raftery, Adrian E.
署名单位:
Universite Clermont Auvergne (UCA); Centre National de la Recherche Scientifique (CNRS); Centre National de la Recherche Scientifique (CNRS); Universite Paris Cite; Centre National de la Recherche Scientifique (CNRS); Sorbonne Universite; Universite Paris Cite; University of Colorado Anschutz; Colorado State University Fort Collins; University of Northern Colorado; Colorado School of Public Health; Institut Universitaire de France; University of Washington; University of Washington Seattle
刊物名称:
ANNALS OF APPLIED STATISTICS
ISSN/ISSBN:
1932-6157; 1941-7330
DOI:
10.1214/26-AOAS2183
发表日期:
2026-06
页码:
1736-1765
关键词:
Large covariance matrix estimation
pairwise covariates
Spatial effects
probabilistic projections
models
car
摘要:
We consider the problem of estimating a high-dimensional covariance matrix from a small number of observations when covariates on pairs of variables are available and the variables can have spatial structure. This is motivated by the problem arising in demography of estimating the covariance matrix of the total fertility rate (TFR) of 195 different countries when only 11 observations are available. We construct an estimator for high-dimensional covariance matrices by exploiting information about pairwise covariates, such as whether pairs of variables belong to the same cluster, or spatial structure of the variables, and interactions between the covariates. We reformulate the problem in terms of a mixed effects model. This requires the estimation of only a small number of parameters, which are easy to interpret and which can be selected using standard procedures. The estimator is consistent under general conditions, and asymptotically normal. It works if the mean and variance structure of the data is already specified or if some of the data are missing. Using simulations, we assess its performance under our model assumptions as well as under model misspecification. We find that it outperforms several popular alternatives. We apply it to the TFR dataset and draw some conclusions.
来源URL: