2026 Volume 7 Issue 2 Pages 69-80
Teleoperated hydraulic excavators are at high risk of tip-over on sloped terrain due to communication delays and the lack of bodily feedback. In this study, sloped terrain and obstacle environments were constructed in the high-speed simulator Genesis, and an Actor–Critic policy integrating internal states with Bird’s-Eye View (BEV) images generated from LiDAR point clouds was trained using Proximal Policy Optimization (PPO). The evaluation was conducted by averaging the results under a condition in which the slope angle was increased at a constant rate up to 45° and then maintained for 10s. As a result, in the obstacle-free environment, the learned policy maintained stability for an average of 850 steps (approximately 17s), while in the obstacle-containing environment, it avoided tip-over up to nearly the maximum episode length of 1000 steps (20s). Furthermore, both policies were integrated into a single policy through knowledge distillation, and convergence of the loss was confirmed.