Journal of Natural Language Processing
Online ISSN : 2185-8314
Print ISSN : 1340-7619
ISSN-L : 1340-7619
General Paper (Peer-Reviewed)
Analyzing the Multilingual Ability of Vision-Language Models to Generate Explanations for Artworks
Shintaro Ozaki, Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe
Author information
JOURNAL FREE ACCESS

2026 Volume 33 Issue 2 Pages 760-808

Details
Abstract

As the performance of Vision-Language Models (VLMs) continues to advance, these models are becoming increasingly capable of generating responses in multiple languages, raising expectations for their use in multilingual explanation generation. However, since the pre-training of Vision Encoders and the joint training of Large Language Models (LLMs) with Vision Encoders are predominantly conducted on English data, it remains uncertain whether VLMs can fully realize their potential when generating explanations in non-English languages. Moreover, existing multilingual QA benchmarks often rely on machine-translated datasets, which can introduce country-specific discrepancies and nuances, limiting their reliability as evaluation tasks. To address these issues, our study constructs an extended multilingual dataset without relying on machine translation. The created dataset incorporates linguistic nuances and country-specific expressions, allowing for a more faithful evaluation of VLMs’ explanation generation capabilities across languages. Additionally, our study analyzes whether instruction tuning performed in resource-rich English can enhance performance in other languages. Our findings on several settings and VLMs demonstrate that models generally underperform in non-English languages compared to English. Furthermore, the results on fine-tuning suggest that VLMs face challenges in effectively transferring and managing knowledge acquired from English data.

Content from these authors
© 2026 The Association for Natural Language Processing
Previous article Next article
feedback
Top