Voice Separation with an Unknown Number of Multiple Speakers

Eliya Nachmani, Yossi Adi, Lior Wolf

2020-02-29ICML 2020 1Speech Separation

Abstract

We present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing steps, while maintaining the speaker in each output channel fixed. A different model is trained for every number of possible speakers, and the model with the largest number of speakers is employed to select the actual number of speakers in a given sample. Our method greatly outperforms the current state of the art, which, as we show, is not competitive for more than two speakers.

Results

Task	Dataset	Metric	Value	Model
Speech Separation	WHAMR!	SI-SDRi	12.2	VSUNOS
Speech Separation	WSJ0-5mix	SI-SDRi	10.56	Gated DualPathRNN
Speech Separation	WSJ0-2mix	SI-SDRi	20.12	Gated DualPathRNN
Speech Separation	WSJ0-3mix	SI-SDRi	16.85	Gated DualPathRNN
Speech Separation	WSJ0-4mix	SI-SDRi	12.88	Gated DualPathRNN

Related Papers

Dynamic Slimmable Networks for Efficient Speech Separation2025-07-08 Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios2025-06-17 SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline2025-05-25 Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers2025-05-22 Single-Channel Target Speech Extraction Utilizing Distance and Room Clues2025-05-20 Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation2025-05-19 SepPrune: Structured Pruning for Efficient Deep Speech Separation2025-05-17 A Survey of Deep Learning for Complex Speech Spectrograms2025-05-13