2026 Volume 38 Issue 1 Pages 569-572
In this study, we investigated the applicability of CLIP, a vision-language model, to automate the affective evaluation of product images. For three product categories (chairs, cups, and pens), we calculated impression scores using CLIP using an ensemble of object-conditional prompts and analyzed the correlation with human subjective evaluations (Likert scales). Experimental results confirmed a moderate positive correlation with impression words related to visual atmosphere, such as “cute” and “casual,” demonstrating CLIP’s effectiveness. On the other hand, correlations with physical and cultural attributes, such as “heavy” and “formal,” were weak or even negative. Furthermore, correlation analysis of antonyms revealed that CLIP was unable to preserve semantic oppositional structures (e.g., heavy vs. light). These results suggest that CLIP is capable of expressing impressions of visual surface features, but has limitations in its ability to infer underlying physical properties from visual information.