Flexible Data Aggregation for Prediction and Decision Making with Contextual Information: Applications in Retailing

成果类型:
Article
署名作者:
Peng, Zhenkang; Li, Chengzhang; Rong, Ying; Luo, Zichao; Ma, Guangrui; Zhao, Mingyong
署名单位:
Zhejiang University; Shanghai Jiao Tong University
刊物名称:
M&SOM-MANUFACTURING & SERVICE OPERATIONS MANAGEMENT
ISSN/ISSBN:
1523-4614
DOI:
10.1287/msom.2025.0313
发表日期:
2026
关键词:
INVENTORY chains
摘要:
Problem definition: How should online retailers make demand predictions or operational decisions with limited relevant data? Motivated by conventional data aggregation approaches, we develop a flexible data aggregation (FlexDA) framework to adapt to different degrees of heterogeneity across products, thereby striking a better balance between the bias and variance of data-driven predictions and decisions. Methodology/ results: Under the FlexDA framework, we propose aggregating the individual data sets at different levels, potentially based on contextual information, and training submodels with each aggregated data set. A meta-model is trained to integrate the outputs of these submodels with a set of weights. For demand prediction tasks with linear models, we propose a consistent estimate of the optimal weight and theoretically demonstrate the advantage of the FlexDA approach over existing approaches under the small-data large-scale regime. For decision making with nonlinear feature-demand relationships, with a fixed sample size, we show that the optimality gap of the FlexDA approach decays near-linearly in the number of products with high probability. We further validate the FlexDA approach with synthetic data and real data from the Rossmann store. Building on theoretical development and empirical validation, we conducted an internal study with Meituan, focusing on ordering problems in their community group buying business. Our proposed approach achieves an average reduction of more than 10% in both lost sales ratio and inventory ratio for newly launched fresh products, as well as standard products compared with the algorithm implemented by Meituan. Managerial implications: Simply aggregating data from all products and training a shared model reduces the high variance caused by data scarcity but compromises the ability to capture heterogeneity across products. Our study highlights the value of flexible data aggregation for data-driven prediction and decision making, especially for large-scale applications with limited data, with both theoretical and empirical support.