Short-Lived High-Volume Bandits
成果类型:
Article
署名作者:
Jia, Su; Li, Andrew; Ravi, R.; Oli, Nishant; Duff, Paul; Anderson, Ian
署名单位:
Amazon.com; Carnegie Mellon University
刊物名称:
OPERATIONS RESEARCH
ISSN/ISSBN:
0030-364X
DOI:
10.1287/opre.2023.0557
发表日期:
2026
关键词:
Bayesian bandits
A/B testing
field experiment
rates
摘要:
Modern platforms leverage randomized experiments to make informed decisions from a given set of items (treatment arms). As a particularly challenging scenario, these items can (i) arrive in a high volume, with thousands of new items being released per hour, and (ii) have a short lifetime due to their transient nature. We study a Bayesian multiple-play bandit problem that encapsulates the key features of this scenario. In each round, a set of arms arrives. Each arm has a lifetime w and an unknown mean reward. The learner selects a multiset of n arms and receives observable rewards for each play. We aim to minimize the loss due to not knowing the reward rates. We show that if at most n rho arms arrive per round, then our policy has a O(n-min rho,1 of prior distributions for the mean rewards. We complement this by showing that all policies suffer an ohm(n-min rho, 1 a large-scale field experiment on Glance, a content card service platform that faces exactly this challenge. A simple variant of our policy outperformed the current recommender at the time by 4.32% in total duration and 7.48% in total number of click-throughs.
来源URL: