A Zero-Inflated Logistic Normal Multinomial Model for Extracting Microbial Compositions
成果类型:
Article
署名作者:
Zeng, Yanyan; Pang, Daolin; Zhao, Hongyu; Wang, Tao
署名单位:
Shanghai Jiao Tong University; Yale University; Shanghai Jiao Tong University; Shanghai Jiao Tong University
刊物名称:
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION
ISSN/ISSBN:
0162-1459
DOI:
10.1080/01621459.2022.2044827
发表日期:
2023
页码:
2356-2369
关键词:
rank tensor completion
community detection
regression
decompositions
factorization
SPARSE
摘要:
High throughput sequencing data collected to study the microbiome provide information in the form of relative abundances and should be treated as compositions. Although many approaches including scaling and rarefaction have been proposed for converting raw count data into microbial compositions, most of these methods simply return zero values for zero counts. However, zeros can distort downstream analyses, and they can also pose problems for composition-aware methods. This problem is exacerbated with microbiome abundance data because they are sparse with excessive zeros. In addition to data sparsity, microbial composition estimation depends on other data characteristics such as high dimensionality, over-dispersion, and complex co-occurrence relationships. To address these challenges, we introduce a zero-inflated probabilistic PCA (ZIPPCA) model that accounts for the compositional nature of microbiome data, and propose an empirical Bayes approach to estimate microbial compositions. An efficient iterative algorithm, called classification variational approximation, is developed for carrying out maximum likelihood estimation. Moreover, we study the consistency and asymptotic normality of variational approximation estimator from the perspective of profile M-estimation. Extensive simulations and an application to a dataset from the Human Microbiome Project are presented to compare the performance of the proposed method with that of the existing methods. The method is implemented in R and available at . for this article are available online.