主催: 人工知能学会
会議名: 第96回 人工知能基本問題研究会
回次: 96
開催地: 名古屋工業大学
開催日: 2014/01/13 - 2014/01/14
p. 06-
Stochastic K-armed bandits tries to maximize his cumulative reward in limited number of plays. In this paper, we consider the variant of stochastic K-armed bandits that has action-dependent processing time. For this problem, we propose the policy N-UCB (Normalized UCB), the extension of well-known policy UCB, and shows some fundamental results of its regret analysis.