2026 Volume 33 Issue 2 Pages 591-629
In recent years, there has been active development of general-purpose text embedding models for English and multilingual settings. However, efforts to develop such models for Japanese remain limited, primarily due to the lack of suitable datasets and limited practical know-how for model development. In this paper, we present the development of Ruri, a general-purpose Japanese text embedding model, and describe the process behind it. Specifically, we explain how we construct synthetic datasets using large language models to compensate for the scarcity of training data, train a base model via contrastive pre-training, and subsequently fine-tune it on high-quality data. The resulting text embedding model, Ruri, achieves performance that surpasses existing models on benchmarks for Japanese text embeddings.