2026 年 30 巻 4 号 p. 1093-1101
Pre-trained language models (PLMs) have demonstrated high performance across various tasks and domains. Among these PLMs, Mixture of Experts (MoE) models also exhibit high performance with fewer active parameters. In domain adaptation, generally, continual pre-training is performed with existing models using domain-specific corpora. However, few efforts have been made to transform and train these models into models with MoE architectures. We propose a method to construct domain-adapted MoE models from general pre-trained models that do not initially have MoE architectures. By independently training multiple experts using domain corpora and integrating them into an MoE architecture, we constructed a domain-adapted MoE model. We performed this MoE transformation in the financial domain and verified its effectiveness in financial tasks. The evaluation results indicate that our domain-adapted MoE models perform better than those without MoE architectures. Our domain-adapted MoE models achieved consistent improvements over domain-adaptive pretraining, task-adaptive pretraining, and domain- and task-adaptive pretraining, with an average improvement of approximately 0.08 in F1 across six financial benchmark tasks for encoder-based models.
この記事は最新の被引用情報を取得できません。