2022 Volume 24 Issue 2 Pages 37-50
The purpose of this study is to obtain data to improve assessment decision-making by analyzing raters’ factors for situational tasks as part of the validation of a speaking test for placement. In this study, we conducted an experiment on situational tasks with a total of 112 participants: 33 teachers with more than 10 years of teaching experience (type-A teachers), 15 teachers with three to ten years of teaching experience (type-B teachers) and 64 non-teachers. The results showed that although there was no major difference in reliability between teachers and non-teachers, non-teachers had more difficulty in discriminating between intermediate and advanced levels. In addition, the type-A teachers tended to place more importance on audio samples than the non-teachers. Non-teachers did not change their assessments depending on level, whereas teachers, especially type-A, changed their focus depending on the level. For example, “fluency” for elementary-low, “text type” for elementary-high, “expressiveness” for intermediate-low/mid, and “deference” and “fluency” for intermediate-high/advanced). These findings illustrate the need for more meticulous training sessions for non-teachers and the need to emphasize the benchmark criteria for each level in assessment training.