2025 年 16 巻 4 号 p. 896-908
Many zero-shot image restoration methods have been proposed by leveraging pre-trained image diffusion models. These methods are capable of performing various image restoration tasks without the need for task-specific training. In general, such methods tend to improve restoration performance by using conditional image diffusion models, such as those based on classes. However, the challenge has been that a separate method is required to determine the appropriate class from degraded images. In this study, we focus on image colorization and propose a method in which the classification of the input grayscale image is performed by applying CLIP, a type of vision-language model, and the resulting class is used as the class condition for a conditional image diffusion models. Through experiments, it was confirmed that the proposed method enables high-precision colorization compared to conventional zero-shot image restoration methods without class conditions, as well as deep neural networks specifically trained for image colorization.