2025 Volume 37 Issue 4 Pages 725-733
In recent years, databases of acted-emotional speech with labeled acted-emotions have often been used as training data for machine learning tasks in Speech Emotion Recognition and Emotion Speech Synthesis, instead of spontaneous speech databases labeled with the speaker’s own emotions. However, there has been insufficient analysis regarding the extent to which acted emotions correspond with perceived emotions, as well as comparisons of acting methods for naturalistic emotional expression. To address this, we constructed the Hiroshima City University Emotion Speech Database (HCUDB) for comparing and analyzing acted-emotion labels versus perceived emotion labels, and empathetic acting versus technical acting in speech. HCUDB comprises two corpora: one is “Acted-Emotion vs Perceived Emotion Evaluation Corpus,” that contains emotion speeches with both acted-emotion labels and perceived emotion labels annotated (HCUDB1), and the other is “Empathetic vs Technical Acting Method Comparison Corpus,” that contains emotion speeches performed by the same speaker using different acting methods (HCUDB2). This paper describes the methods of voice recording and emotion labeling for these corpora and presents statistical analysis of the acoustic features of the speech data in HCUDB.