Abstract
Typical fuzzy reinforcement learning algorithms take value-function based approaches such as fuzzy Q-learning in Markov Decision Processes and use constant or linear functions in the conclusion parts of the fuzzy rules. On the other hand, the policy-gradient approaches design policy functions directly and learn parameters included in the policy functions. Based on one of the policy-gradient approaches, a fuzzy reinforcement learning algorithm is proposed. This algorithm can deal with fuzzy sets even in the conclusion parts and also learn the rule weights of fuzzy rules. This paper's experiments show that the proposed learning algorithm is effective for a decision making problem of a soccer robot that plays in RoboCup Soccer Small Size League. After learning experiments with 30 scenes of a robot holding a ball, the robot control system learns a stochastic policy that agrees with that of human decision making in 25 of the 30 scenes.