2026 Volume 33 Issue 2 Pages 509-536
Vision language models (VLMs) are rapidly advancing; however, evaluations in the Japanese language remain fragmented across various tasks and domains, making it difficult to obtain a clear picture of the overall capabilities and to conduct fair comparisons. This study develops a comprehensive evaluation framework to systematically assess the capabilities of Japanese-capable VLMs and presents cross-task empirical results. We propose llm-jp-eval-mm, which consolidates ten Japanese and nine existing English datasets and provides a toolkit that enables consistent evaluation under a unified protocol aligned with predefined capability axes. We describe the capability axes, framework design, dataset selection, and implementation. We identify capability areas in which Japan-developed VLMs are relatively weak and analyze structural patterns in performance correlations between Japanese and English by evaluating 32 publicly available VLMs developed in Japan and elsewhere under identical conditions. The proposed framework enables reliable and comprehensive comparisons of VLMs and offers guidance for future model improvements and benchmark designs.