Accurate severity assessment during emergency medical service (EMS) calls is essential for timely dispatch and efficient resource allocation. This study investigates whether callers’ voice acoustics can provide a low-latency, acoustics-only decision-support signal that operates in parallel with protocol-guided triage in operational call centers. Severity labels were defined using post-transport clinical severity at the first medical examination after transport (Japanese shōbyō teido) as recorded in EMS outcome records. This label should be interpreted as a clinically meaningful downstream outcome surrogate rather than a direct observation of urgency at call time. Accordingly, because this outcome-based label may not fully align with call-time dispatch priority, the proposed score is positioned as a protocol-parallel auxiliary alert cue rather than a replacement for protocol triage or a direct model of dispatcher-perceived urgency. Using anonymized Japanese EMS call recordings from the Tokyo Fire Department, we analyzed a balanced subset of 204 calls (102 serious, 102 minor) after excluding moderate cases. We extracted call-level acoustic descriptors (F0 statistics, MFCCs, Mel-band energies, spectral measures, and voice-quality features) and conducted Mann–Whitney U tests with false discovery rate (FDR) correction, revealing systematic differences, particularly in pitch variability and spectral energy patterns. Robustness was evaluated using a waveform-level Wiener-filtering baseline and systematic data augmentation (pitch shifting, gain perturbation, and additive noise) under stratified cross-validation. Logistic regression, RBF-kernel support vector machine, and random forest models were assessed using area under the ROC curve (AUC) and recall-focused operating points selected to satisfy a target recall for serious cases (Recall_serious) of at least 0.90. Within the present internal cross-validation setting, full augmentation yielded the best discrimination, whereas Wiener filtering provided limited and model-dependent gains. The best configuration within the present dataset (random forest with full augmentation) achieved AUC = 0.93 and Recall_serious = 0.933. These results suggest the potential feasibility of integrating an acoustics-only severity score as a complementary alert signal, while highlighting limitations when vocal arousal cues are weak or caller speech is sparse.
抄録全体を表示