Bunlong Lay

dblp:326/6566 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-0847-7896ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency
Bunlong Lay, Rostilav Makarov, Timo Gerkmann
INTERSPEECH1
2025 Real-Time Diffusion Buffer for Speech Enhancement On A Laptop
Bunlong Lay, Rostilav Makarov, Timo Gerkmann
INTERSPEECH1
2024 Single and Few-Step Diffusion for Generative Speech Enhancement
abstract
Diffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the iterative reverse process. This results in a slow inference process and causes discretization errors that accumulate over the sampling trajectory. In this paper, we address these limitations through a two-stage training approach. In the first stage, we train the diffusion model the usual way using the generative denoising score matching loss. In the second stage, we compute the enhanced signal by solving the reverse process and compare the resulting estimate to the clean speech target using a predictive loss. We show that using this second training stage enables achieving the same performance as the baseline model using only 5 function evaluations instead of 60 function evaluations. While the performance of usual generative diffusion algorithms drops dramatically when lowering the number of function evaluations to obtain single-step diffusion, we show that our proposed method keeps a steady performance and therefore largely outperforms the diffusion baseline in this setting and also generalizes better than its predictive counterpart1.
Bunlong Lay, Jean-Marie Lemercier, Julius Richter, Timo Gerkmann
ICASSP1
2024 EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data
abstract
Speech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on acted-out datasets and parallel data samples, in this work we specifically focus on more challenging in-the-wild scenarios and do not rely on parallel data. To this end, we propose a diffusion-based generative model for speech emotion conversion, the EmoConv-Diff, that is trained to reconstruct an input utterance while also conditioning on its emotion. Subsequently, at inference, a target emotion embedding is employed to convert the emotion of the input utterance to the given target emotion. As opposed to performing emotion conversion on categorical representations, we use a continuous arousal dimension to represent emotions while also achieving intensity control. We validate the proposed methodology on a large in-the-wild dataset, the MSP-Podcast v1.10. Our results show that the proposed diffusion model is indeed capable of synthesizing speech with a controllable target emotion. Crucially, the proposed approach shows improved performance along the extreme values of arousal and thereby addresses a common challenge in the speech emotion conversion literature.
Navin Raj Prabhu, Bunlong Lay, Simon Welker, Nale Lehmann-Willenbrock, Timo Gerkmann
ICASSP2
2024 An Analysis of the Variance of Diffusion-based Speech Enhancement
Bunlong Lay, Timo Gerkmann
INTERSPEECH1
2024 EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker, Bunlong Lay, Shinji Watanabe 0001, Alexander Richard, Timo Gerkmann
INTERSPEECH5
2023 Speech Signal Improvement Using Causal Generative Diffusion Models
abstract
In this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processing, we modify the network architecture of our previous work and replace global normalization with causal adaptive gain control. We generate diverse training data containing a broad range of distortions. This work was performed in the context of an "ICASSP Signal Processing Grand Challenge" and submitted to the non-real-time track of the "Speech Signal Improvement Challenge 2023", where it was ranked fifth.
Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Tal Peer, Timo Gerkmann
ICASSP4
2023 Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement
Bunlong Lay, Simon Welker, Julius Richter, Timo Gerkmann
INTERSPEECH1
2023 Speech Enhancement and Dereverberation With Diffusion-Based Generative Models
abstract
In this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement. We present a detailed overview of the diffusion process that is based on a stochastic differential equation and delve into an extensive theoretical examination of its implications. Opposed to usual conditional generation tasks, we do not start the reverse process from pure Gaussian noise but from a mixture of noisy speech and Gaussian noise. This matches our forward process which moves from clean speech to noisy speech by including a drift term. We show that this procedure enables using only 30 diffusion steps to generate high-quality clean speech estimates. By adapting the network architecture, we are able to significantly improve the speech enhancement performance, indicating that the network, rather than the formalism, was the main limitation of our original approach. In an extensive cross-dataset evaluation, we show that the improved method can compete with recent discriminative models and achieves better generalization when evaluating on a different corpus than used for training. We complement the results with an instrumental evaluation using real-world noisy recordings and a listening experiment, in which our proposed method is rated best. Examining different sampler configurations for solving the reverse process allows us to balance the performance and computational speed of the proposed method. Moreover, we show that the proposed method is also suitable for dereverberation and thus not limited to additive background noise removal. Code and audio examples are available online1https://github.com/sp-uhh/sgmse.
Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Timo Gerkmann
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Speech Enhancement Regularized by a Speaker Verification Model
abstract
In the last years, end-to-end neural networks were employed to enhance single-channel noisy speech. The Conv-Tasnet architecture is such a neural network and it has been successfully trained on the scale invariant signal-to-distortion ratio (SI-SDR) loss. However, we find that in a speech enhancement task Conv-Tasnet trained on the SI-SDR loss introduces distortions to the enhanced signal as the mismatch between training and testing data increases. To mitigate these effects, we investigate on a new loss function, that combines SI-SDR with a pretrained speaker verification (SV) model as a regularizer. As the SV model captures information about important characteristics of clean speech signals, we argue that the proposed regularization ensures that the estimated speech signal is speech-like even if there is an increasing mismatch between training and testing data. Accordingly, we observe considerable improvements in POLQA in mismatched training and testing conditions, e.g. when training on anechoic speech but testing on reverberant data.
Bunlong Lay, Timo Gerkmann
MMSP1