Abstract
Agents select random action about first element, which come up a view area of agent. The proposed learning system has a multi-layer style rule base. If one rule base has no valid rule, other rule is able to output the incomplete actions by slightly valid elements of the rule base. In any cases, agents should take a behavior by the rule system. The proposed Q-learning system continues learning and adds the new element and changes the next layer when the agents has observed new elements. In dynamic cases, agents should work any cases, but old rules as often as dose not work a current state, because new elements are changed any situations for whole agents, dynamically. In this paper, we discuss the reasonable interrelation level about those rules.