TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Utterance Weighted Multi-Dilation Temporal Convolutional N...

Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

William Ravenscroft, Stefan Goetze, Thomas Hain

2022-05-17Speech Dereverberation
PaperPDFCode(official)

Abstract

Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing the number of convolutional blocks. The best performance improvement over the baseline TCN is 0.55 dB scale-invariant signal-to-distortion ratio (SISDR) and the best performing WD-TCN model attains 12.26 dB SISDR on the WHAMR dataset.

Results

TaskDatasetMetricValueModel
Speech EnhancementWHAMR!ESTOI93.5WD-TCN
Speech EnhancementWHAMR!PESQ3.5WD-TCN
Speech EnhancementWHAMR!SI-SDR12.26WD-TCN
Speech EnhancementWHAMR!SRMR8.8WD-TCN

Related Papers

VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification2025-02-11A Hybrid Model for Weakly-Supervised Speech Dereverberation2025-02-06Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising2024-10-30Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models2024-08-14Schrödinger Bridge for Generative Speech Enhancement2024-07-22Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System2024-05-17USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering2024-02-01Distributed Speech Dereverberation Using Weighted Prediction Error2023-12-05