2026 年 47 巻 5 号 p. 442-450
End-to-end audio signal processing frameworks based on deep neural networks (DNNs) typically employ trainable 1-D convolutional layers as encoders for feature extraction and decoders for signal reconstruction. This architecture is advantageous because signal representations can be learned directly from training data to maximize performance on the target task. However, standard trained encoder-decoder pairs generally fail to achieve perfect reconstruction; consequently, information is lost even when the encoded signal remains unprocessed. In this paper, to circumvent the information loss caused by such non-invertible encoders, we propose incorporating orthogonal convolutional layers into end-to-end DNNs, combined with a reshaping technique to facilitate practical implementation. The proposed architectures were evaluated on a speech enhancement task using Conv-TasNet and RE-SepFormer. Experimental results demonstrate that the orthogonal layers ensure invertibility and alleviate training difficulties, particularly when the kernel size is large.