2026 Volume 38 Issue 1 Pages 588-598
In this paper, we propose a method for utilizing images at the summary level in multimodal neural machine translation. Conventional models typically extract and use only the image information related to the next token being predicted, but we show that this approach can lead to over-translation. To address this issue, we propose MVNMT, a new model that employs image information to model the features of the entire sentence (summary) and integrates those features into the decoder. MVNMT uses a variational autoencoder to extract a shared latent representation from both text and images. Our experimental results demonstrate that MVNMT outperforms conventional text-only translation models in terms of translation metrics and effectively mitigates the over-translation problem compared to MNMT models that utilize image information at the token level.