Abstract
We proposed a method of Q-learning with dynamic construction facility of the
fuzzy state space with the real number attributes. We initially have no states and gradually
add a new state of fuzzy set for the given attributes. We update Q values with the reward
and the fuzzy sets with TD (Temporal Difference) error and we remove unnecessary states.
When we add a rule in this method, we generate all the fuzzy sets in each attribute, which
may lead to have similar fuzzy sets in a certain attribute. In addition, the states gradually
increase when the success ratio keeps to be high. So, we adjust a parameter to decrease the
occurrences of addition and remove of fuzzy sets when the success ratio is high. Furthermore,
we share fuzzy sets with several rules to prevent similar fuzzy sets. As a result of application
of this method to the pursuit problem in a real number environment, we have suppressed
the increase of the state in the final stage of learning.