Acoustical Science and Technology
Online ISSN : 1347-5177
Print ISSN : 1346-3969
ISSN-L : 0369-4232
ACOUSTICAL LETTERS
Semi-supervised acoustic scene classification based on neural word unigram
Naoki KogaYoshiaki BandoKeisuke ImotoTakao Tsuchiya
Author information
JOURNAL OPEN ACCESS

2026 Volume 47 Issue 5 Pages 474-478

Details
Abstract

This paper presents a semi-supervised method of acoustic scene classification (ASC) using neural word unigram (NWU). Semi-supervised learning is a key technique in ASC to train models using a large amount of low-cost unlabeled data and limited labeled data. When the labeled data is few, the existing semi-supervised methods, such as teacher-student learning, can overfit to the data because there is less constraint on the unlabeled data. To address this issue, we propose NWU that mitigates excessive dependence on labeled data. NWU is a deep generative model of audio clips, parameterized by a unigram and a Gaussian mixture model (GMM), and is utilized for the self-clustering of the unlabeled data. The model is trained as a variational autoencoder by maximizing a variational lower bound. This training is performed jointly with supervised learning for semi-supervised classification of audio clips. We evaluated NWU using 24-hours of real-world data. The experimental results show that NWU achieved classification performance improvements of up to 9.3% points over the supervised baselines and up to 3.6% points over the semi-supervised baselines. They also revealed that NWU maintained classification performance superior to the baselines even when labeled data constituted only 1.6% of the dataset.

Content from these authors
© 2026 by The Acoustical Society of Japan

This article is licensed under a Creative Commons [Attribution-NoDerivatives 4.0 International] license.
https://creativecommons.org/licenses/by-nd/4.0/
Previous article Next article
feedback
Top