Transactions of the Japanese Society for Artificial Intelligence
Online ISSN : 1346-8030
Print ISSN : 1346-0714
ISSN-L : 1346-0714
Special Paper: Intelligent Dialogue Systems
Proposal of an Automatic Evaluation Method for Dialogue System Reflecting Individual Tendencies
Keisuke KameyamaKazunori Komatani
Author information
JOURNAL OPEN ACCESS FULL-TEXT HTML

2026 Volume 41 Issue 2 Pages IDS26-C_1-8

Details
Abstract

Recently, many methods have been proposed for automatic evaluation of dialogue systems, which show highcorrelation with human evaluations. Although these methods tend to align well with the average scores of multipleevaluators, the scores may not reflect individual preferences. In this study, we propose an automatic evaluationmethod that incorporates the evaluation tendencies of specific individuals, such as system designers or specific users,in order to realize evaluations that align with individual preferences, rather than average-based evaluations. Wefirst focus on the differences in the aspects that each evaluator emphasizes in dialogue evaluation and computesweights to each sub-metric accordingly. Then, based on the obtained weights, we estimate an overall score foreach dialogue system using the scores for each sub-metric produced by automatic evaluation. Through experimentsinvolving multiple evaluators, we confirmed that our method can produce system evaluations that reflect individualevaluation tendencies. In this process, we utilized a Large Language Model (LLM) for the automatic evaluation andapplied multiple regression analysis to determine the metric weights. The results show that, compared to evaluationby the LLM alone, incorporating individual regression-based weights leads to a reduction in the mean squared errorof the overall score, making it closer to each evaluator’s actual scores.

Content from these authors
© JSAI (The Japanese Society for Artificial Intelligence)
Previous article Next article
feedback
Top