人工知能学会研究会資料 人工知能基本問題研究会
Online ISSN : 2436-4584
96回 (2014/1)
会議情報

処理時間の長短を考慮した確率的多腕バンディット問題へのUCB戦略の拡張
渡辺 僚中村 篤祥工藤 峰一
著者情報
会議録・要旨集 フリー

p. 06-

詳細
抄録

Stochastic K-armed bandits tries to maximize his cumulative reward in limited number of plays. In this paper, we consider the variant of stochastic K-armed bandits that has action-dependent processing time. For this problem, we propose the policy N-UCB (Normalized UCB), the extension of well-known policy UCB, and shows some fundamental results of its regret analysis.

著者関連情報
© 2015 人工知能学会
前の記事 次の記事
feedback
Top