2026 Volume 39 Issue 4 Pages 77-87
In this paper, we propose a data-driven reinforcement learning method for nonlinear systems with process and measurement noise. More specifically, we extend Average Off-Policy Learning, which learns an optimal controller while mitigating the effects of both types of noise, to nonlinear systems using the Koopman operator. Furthermore, we conduct a stability analysis and evaluate the control performance of the proposed method. The effectiveness of the proposed method is validated through a numerical simulation using the Duffing oscillator, demonstrating that it enables more stable learning compared to a conventional method.