2026 年 17 巻 3 号 p. 1015-1028
Recent advances in deep generative models have enabled the synthesis of highly photorealistic images, raising concerns about image authenticity and misinformation. While generated image detection has been widely studied for faces and objects, landscape images remain challenging due to their complex semantic structure and strong dependence on global factors such as illumination and perspective. In this study, we investigate the characteristics of AI-generated landscape images and analyze the decision-making basis of convolutional neural networks (CNNs) from an explainable AI perspective. A carefully controlled dataset of real and generated landscape images across ten categories is constructed, and a baseline CNN using only RGB inputs is evaluated. SHAP and Grad-CAM analyses reveal that the model relies primarily on local texture irregularities and smoothness artifacts. Robustness experiments using masking and retraining demonstrate that detection remains stable under blurring and edge removal but degrades significantly when local visual information is completely suppressed. Finally, a multimodal model integrating semantic segmentation masks with RGB inputs is proposed, achieving improved accuracy and interpretability by explicitly leveraging region-level structural information.