In this paper, we propose a data-driven reinforcement learning method for nonlinear systems with process and measurement noise. More specifically, we extend Average Off-Policy Learning, which learns an optimal controller while mitigating the effects of both types of noise, to nonlinear systems using the Koopman operator. Furthermore, we conduct a stability analysis and evaluate the control performance of the proposed method. The effectiveness of the proposed method is validated through a numerical simulation using the Duffing oscillator, demonstrating that it enables more stable learning compared to a conventional method.
We propose a singular value decomposition-based recursive subspace state-space system identification algorithm (R4SID) with fixed input-output data size using the matrix inversion lemma (MIL). The proposed R4SID improves the computational efficiency of recursions because it updates an inverse matrix by the MIL and computes the multiplication of vectors instead of matrices. We clarify the following two effectiveness of the proposed R4SID through numerical experiments. The first effectiveness is that the proposed R4SID can identify a time-varying system disturbed by white Gaussian measurement noise. The second effectiveness is that the accuracy of the proposed R4SID is less sensitive to the choice of data length in the case of the white Gaussian measurement noise whose level is sufficiently smaller than that of input signal.
We propose two singular value decomposition (SVD) based recursive PI-MOESP identification algorithms (PI-MOESP identification algorithms: the multiple-input multiple-output output-error state-space model (MOESP) identification algorithms using an instrumental variable consisting of past input (PI) data) with fixed input-output data size using the matrix inversion lemmas (MILs). It is clarified in the two proposed SVD-based recursive PI-MOESP identification algorithms (RPI- MOESPs) that computation time is reduced by using two MILs for addition of the latest input data and subtraction of the oldest input data separately (separated-type MILs) instead of a single MIL for the addition and the subtraction in an integrated manner (integrated-type MIL). Numerical experiments provide the following two results. The first result is that the computation time of the proposed SVD-based RPI-MOESP using the separated-type MILs is less than that of the proposed SVD-based RPI-MOESP using the integrated-type MIL despite achieving almost the same accuracy in the case of linear state-space model identification. The second result is that the two proposed SVD-based RPI-MOESPs can identify a time-varying system disturbed by colored measurement noise.