Host: The Japanese Society for Artificial Intelligence
Name : The 103rd SIG-SLUD
Number : 103
Location : [in Japanese]
Date : March 20, 2025 - March 22, 2025
Pages 195-200
Our goal is to construct a framework for annotating the process of understanding the meaning of gestures in spoken conversations in a simple and versatile way. As a first step, we conduct sequential annotation of embodied actions on the movement scenes from a multimodal corpus of conversations between science communicators and visitors at the National Museum of Emerging Science and Innovation (Miraikan SC corpus). For a sequence in which a science communicator (SC) prompts a visitor to move by some movement or utterance, and the visitor follows and begins to move, we annotate the SC's first action and the visitor's second action. The first actions of SCs are annotated as walking, pointing, change of orientation, speech, and/or gesture. In this paper, we introduce the purpose and outline of the sequential annotation of embodied actions under development and the results of preliminary analyses based on the annotations.