抄録
This study investigated the reliability of AI-based transcription for assessing reproduction rates in English shadowing tasks. Although shadowing accuracy has traditionally been scored manually, the process is time-intensive and prone to inter-rater discrepancies. Recent advances in speech-to-text systems, such as Whisper, suggest the potential for automatic scoring; however, validation against human judgment remains limited. In this study, shadowing and parallel reading performances by Japanese EFL learners were transcribed using Whisper and by two independent human raters. After both the Whisper-based and human transcriptions were completed, reproduction-rate scoring was conducted by human coders according to predefined scoring criteria. Reproduction rates and inter-rater reliability were evaluated using percent agreement, Cohen’s κ, Prevalence-Adjusted Bias-Adjusted Kappa (PABAK), Gwet’s AC1, and Intraclass Correlation Coefficient (ICC). Consequently, AI-based ratings closely approximated human judgments, particularly under the parallel condition, although κ values were consistently lower for AI–human comparisons. Overall, AI transcription provides a scalable and reliable tool for evaluating shadowing reproduction rates, reducing instructor workload, and offering immediate feedback to learners. Finally, the limitations and pedagogical implications of this study are discussed.