On Time Domain Conformer Models for Monaural Speech Separation in Noisy Reverberant Acoustic Environments

William Ravenscroft, Stefan Goetze, Thomas Hain

2023-10-09Speech Separation

Abstract

Speech separation remains an important topic for multi-speaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech separation. Most recent state-of-the-art (SOTA) separation models have been time-domain audio separation networks (TasNets). A number of successful models have made use of dual-path (DP) networks which sequentially process local and global information. Time domain conformers (TD-Conformers) are an analogue of the DP approach in that they also process local and global context sequentially but have a different time complexity function. It is shown that for realistic shorter signal lengths, conformers are more efficient when controlling for feature dimension. Subsampling layers are proposed to further improve computational efficiency. The best TD-Conformer achieves 14.6 dB and 21.2 dB SISDR improvement on the WHAMR and WSJ0-2Mix benchmarks, respectively.

Results

Task	Dataset	Metric	Value	Model
Speech Separation	WHAMR!	SI-SDRi	14.6	TD-Conformer (XL) + DM
Speech Separation	WHAMR!	SI-SDRi	13.4	TD-Conformer (L) + DM
Speech Separation	WHAMR!	SI-SDRi	12	TD-Confomer (M) + DM
Speech Separation	WHAMR!	SI-SDRi	10.5	TD-Confomer (S)
Speech Separation	WSJ0-2mix	SI-SDRi	21.2	TD-Conformer (XL) + DM

Related Papers

Dynamic Slimmable Networks for Efficient Speech Separation2025-07-08 Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios2025-06-17 SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline2025-05-25 Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers2025-05-22 Single-Channel Target Speech Extraction Utilizing Distance and Room Clues2025-05-20 Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation2025-05-19 SepPrune: Structured Pruning for Efficient Deep Speech Separation2025-05-17 A Survey of Deep Learning for Complex Speech Spectrograms2025-05-13