2026 年 65 巻 4 号 p. 182-193
This study proposes a metric monocular depth estimation method for indoor RGB images and evaluates its application to 3D point cloud reconstruction. The proposed Swin Transformer-U-Net integrates local feature extraction with global contextual reasoning. Experiments using the NYU Depth v2 dataset show that the proposed method improves the RMSE from 0.702 m for Attention U-Net to 0.352 m. The reconstructed point clouds also show improved planar continuity and spatial consistency. These results indicate that the proposed method is effective for approximate indoor 3D reconstruction from a single RGB image, although it does not replace high-accuracy ranging sensors such as LiDAR or RGB-D cameras.