IEEJ Transactions on Electronics, Information and Systems
Online ISSN : 1348-8155
Print ISSN : 0385-4221
ISSN-L : 0385-4221
<Intelligence, Robotics>
Hybrid Learning Using Profit Sharing and Genetic Algorithm under the POMDPs
Kohei SuzukiShohei Kato
Author information
JOURNAL FREE ACCESS

2017 Volume 137 Issue 12 Pages 1591-1599

Details
Abstract

Reinforcement learning is generally performed in the Markov decision processes (MDP). However, there is a possibility that the agent can not correctly observe the environment due to the perception ability of the sensor. This is called partially observable Markov decision processes (POMDP). In a POMDP environment, an agent may observe the same information at more than one state. HQ-learning and Episode-based Profit Sharing (EPS) are well known methods for this problem. HQ-learning divides a POMDP environment into subtasks. EPS distributes same reward to state-action pairs in the episode when an agent achieves a goal. However, these methods have disadvantages in learning efficiency and localized solutions. In this paper, we propose a hybrid learning method combining PS and genetic algorithm. We also report the effectiveness of our method by some experiments with partially observable mazes.

Content from these authors
© 2017 by the Institute of Electrical Engineers of Japan
Previous article Next article
feedback
Top