2026 年 21 巻 1 号 p. 25-00198
We introduce a deep-learning framework that turns a single moving camera into a reliable 3-D sensor for thermal imaging. Portable or drone-mounted inspections increasingly demand lightweight depth acquisition, and passive light field imaging is attractive because it dispenses with active hardware such as LiDAR. Unfortunately, conventional light field cameras rely on large-scale lens arrays, and simply sweeping a monocular camera in their place introduces subtle inter-view vertical offsets that quickly erode reconstruction accuracy: in our simulation, shifting five of nine views increased the Root-Mean-Square Error (RMSE) by a factor of 2.9. To counter this effect we adopt a two-stage strategy in which a Residual U-Net (ResUNet) first rectifies the displaced images and the resulting aligned sequence is then processed by EPINET to estimate depth. In the initial experiments with a synthetic RGB dataset, the rectifier reduced RMSE by 14.2% on average (up to 33.8% in the best case) and increased Peak Signal-to-Noise Ratio (PSNR) by 0.3 dB. These gains were accompanied by the disappearance of streak-like artifacts and the recovery of clean linear structures in epipolar plane images. Subsequently, in the thermal image simulation, the proposed method consistently improved depth estimation across all relative-depth metrics, while also mitigating systematic errors in isothermal regions and enhancing the clarity of temperature boundaries. Our study is, to the best of our knowledge, the first to combine learned rectification with light field depth inference for calibration-free, single-camera thermal diagnostics. Source code and trained models are available at https://github.com/ryutaroLF/ResUNet_EPI_Stabiliser/ .