Journal of Japan Society for Fuzzy Theory and Intelligent Informatics
Online ISSN : 1881-7203
Print ISSN : 1347-7986
ISSN-L : 1347-7986
Original Papers
Multimodal Translation Method Using Summary-Level Image Utilization
Joji TOYAMAMasanori MISONOMasahiro SUZUKIKeiichi OCHIAIYusuke IWASAWAYutaka MATSUO
Author information
JOURNAL FREE ACCESS

2026 Volume 38 Issue 1 Pages 588-598

Details
Abstract

In this paper, we propose a method for utilizing images at the summary level in multimodal neural machine translation. Conventional models typically extract and use only the image information related to the next token being predicted, but we show that this approach can lead to over-translation. To address this issue, we propose MVNMT, a new model that employs image information to model the features of the entire sentence (summary) and integrates those features into the decoder. MVNMT uses a variational autoencoder to extract a shared latent representation from both text and images. Our experimental results demonstrate that MVNMT outperforms conventional text-only translation models in terms of translation metrics and effectively mitigates the over-translation problem compared to MNMT models that utilize image information at the token level.

Content from these authors
© 2026 Japan Society for Fuzzy Theory and Intelligent Informatics
Previous article Next article
feedback
Top