Acoustical Science and Technology
Online ISSN : 1347-5177
Print ISSN : 1346-3969
ISSN-L : 0369-4232

この記事には本公開記事があります。本公開記事を参照してください。
引用する場合も本公開記事を引用してください。

Encoder-masking-decoder networks using orthogonal convolutional layer as invertible linear encoder
Ren Uchida, Kohei Yatabe, Tomohiko Nakamura
著者情報
ジャーナル オープンアクセス 早期公開

論文ID: e26.10

この記事には本公開記事があります。
詳細
抄録

End-to-end audio signal processing frameworks based on deep neural networks (DNNs) typically employ trainable 1-D convolutional layers as encoders for feature extraction and decoders for signal reconstruction. This architecture is advantageous because signal representations can be learned directly from training data to maximize performance on the target task. However, standard trained encoder-decoder pairs generally fail to achieve perfect reconstruction; consequently, information is lost even when the encoded signal remains unprocessed. In this paper, to circumvent the information loss caused by such non-invertible encoders, we propose incorporating orthogonal convolutional layers into end-to-end DNNs, combined with a reshaping technique to facilitate practical implementation. The proposed architectures were evaluated on a speech enhancement task using Conv-TasNet and RE-SepFormer. Experimental results demonstrate that the orthogonal layers ensure invertibility and alleviate training difficulties, particularly when the kernel size is large.

著者関連情報
© 2026 by The Acoustical Society of Japan

This article is licensed under a Creative Commons [Attribution-NoDerivatives 4.0 International] license.
https://creativecommons.org/licenses/by-nd/4.0/
feedback
Top