2026 Volume 44 Issue 4 Pages 413-416
In this study, we propose a method for estimating the user's turn-ending intention (turn-end estimation) and turn-taking intention (interruption estimation) in real time, enabling dialogue agents to achieve natural and smooth turn-taking with humans. The proposed method uses linguistic cues (speech content) to extract features using a large language model (LLM), followed by estimation using a classifier (SVM). The implemented estimation models achieved high accuracy (F1 score of approximately 90%) in Japanese dialogue. Furthermore, it was confirmed that the processing time was within tens of milliseconds, demonstrating real-time responsiveness.