2026 Volume 38 Issue 1 Pages 599-606
In reinforcement learning (Q-learning), multiple agents can achieve efficient learning by updating the Q-table cooperatively while simultaneously trying in parallel. In this study, we propose a switching reinforcement learning model by simultaneously analyzing agent clustering and Q-learning for each cluster, assuming that multiple agents are solving a problem in several different environments. But it is also assumed to be unknown in which environment each agent is solving the problem. We calculate fuzzy memberships following the Fuzzy c-Means (FCM) method using the acquired gain based on the policy of each cluster as the clustering criterion, and learn the Q-table for each environment in parallel by updating the Q-values with membership weighting. In addition, by introducing deterministic annealing of the fuzziness degree of the partition, we achieve both robust model estimation and maximization of the acquired gain.