2026 年 17 巻 3 号 p. 805-821
In this paper, we aim to show the usefulness of contextual information in distributed representations for the genre classification of modern Japanese literary works. Our previous studies have demonstrated that high-accuracy genre classification can be achieved by using distributed semantic representations generated by the Continuous Bag-of-Words (CBOW) model. Meanwhile, because BERT can acquire distributed semantic representations that vary according to context, it is expected to enable more precise semantic representations and, consequently, more accurate genre classification. Therefore, the purpose of this study is to compare the genre classification performance of BERT and the CBOW model based on their distributed semantic representations, and to clarify the differences in the representations obtained by the two models. The experimental results show that in the “novel vs. poetry” task, where the text forms differ greatly, both models achieved a high accuracy of 99%. On the other hand, in the “novel vs. essay” task, where the text forms are similar, the 500-dimensional BERT model achieved the highest accuracy of 92.25%, outperforming the CBOW model. These results indicate that the contextual representations acquired by BERT capture subtle differences in word usage, thereby contributing to improved classification performance in genre pairs with similar sentence structures.