Journal of Natural Language Processing
Online ISSN : 2185-8314
Print ISSN : 1340-7619
ISSN-L : 1340-7619
General Paper (Peer-Reviewed)
Cross-Task Evaluation and Empirical Analysis of Japanese Visual Language Models
Koki Maeda, Issa Sugiura, Yusuke Oda, Shuhei Kurita, Naoaki Okazaki
Author information
JOURNAL FREE ACCESS

2026 Volume 33 Issue 2 Pages 509-536

Details
Abstract

Vision language models (VLMs) are rapidly advancing; however, evaluations in the Japanese language remain fragmented across various tasks and domains, making it difficult to obtain a clear picture of the overall capabilities and to conduct fair comparisons. This study develops a comprehensive evaluation framework to systematically assess the capabilities of Japanese-capable VLMs and presents cross-task empirical results. We propose llm-jp-eval-mm, which consolidates ten Japanese and nine existing English datasets and provides a toolkit that enables consistent evaluation under a unified protocol aligned with predefined capability axes. We describe the capability axes, framework design, dataset selection, and implementation. We identify capability areas in which Japan-developed VLMs are relatively weak and analyze structural patterns in performance correlations between Japanese and English by evaluating 32 publicly available VLMs developed in Japan and elsewhere under identical conditions. The proposed framework enables reliable and comprehensive comparisons of VLMs and offers guidance for future model improvements and benchmark designs.

Content from these authors
© 2026 The Association for Natural Language Processing
Previous article Next article
feedback
Top