EDBT 2026 Demo / reviewers in the wild / expert
Bunlong Lay
dblp:326/6566
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-0847-7896ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency
Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
INTERSPEECH | 1 |
| 2025 | Real-Time Diffusion Buffer for Speech Enhancement On A Laptop
Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
INTERSPEECH | 1 |
| 2024 | Single and Few-Step Diffusion for Generative Speech EnhancementabstractDiffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the iterative reverse process. This results in a slow inference process and causes discretization errors that accumulate over the sampling trajectory. In this paper, we address these limitations through a two-stage training approach. In the first stage, we train the diffusion model the usual way using the generative denoising score matching loss. In the second stage, we compute the enhanced signal by solving the reverse process and compare the resulting estimate to the clean speech target using a predictive loss. We show that using this second training stage enables achieving the same performance as the baseline model using only 5 function evaluations instead of 60 function evaluations. While the performance of usual generative diffusion algorithms drops dramatically when lowering the number of function evaluations to obtain single-step diffusion, we show that our proposed method keeps a steady performance and therefore largely outperforms the diffusion baseline in this setting and also generalizes better than its predictive counterpart1. Bunlong Lay, Jean-Marie Lemercier, Julius Richter, Timo Gerkmann |
ICASSP | 1 |
| 2024 | EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild DataabstractSpeech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on acted-out datasets and parallel data samples, in this work we specifically focus on more challenging in-the-wild scenarios and do not rely on parallel data. To this end, we propose a diffusion-based generative model for speech emotion conversion, the EmoConv-Diff, that is trained to reconstruct an input utterance while also conditioning on its emotion. Subsequently, at inference, a target emotion embedding is employed to convert the emotion of the input utterance to the given target emotion. As opposed to performing emotion conversion on categorical representations, we use a continuous arousal dimension to represent emotions while also achieving intensity control. We validate the proposed methodology on a large in-the-wild dataset, the MSP-Podcast v1.10. Our results show that the proposed diffusion model is indeed capable of synthesizing speech with a controllable target emotion. Crucially, the proposed approach shows improved performance along the extreme values of arousal and thereby addresses a common challenge in the speech emotion conversion literature. Navin Raj Prabhu, Bunlong Lay, Simon Welker, Nale Lehmann-Willenbrock, Timo Gerkmann |
ICASSP | 2 |
| 2024 | An Analysis of the Variance of Diffusion-based Speech Enhancement
Bunlong Lay, Timo Gerkmann |
INTERSPEECH | 1 |
| 2024 | EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker, Bunlong Lay, Shinji Watanabe 0001, Alexander Richard, Timo Gerkmann |
INTERSPEECH | 5 |
| 2023 | Speech Signal Improvement Using Causal Generative Diffusion ModelsabstractIn this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processing, we modify the network architecture of our previous work and replace global normalization with causal adaptive gain control. We generate diverse training data containing a broad range of distortions. This work was performed in the context of an "ICASSP Signal Processing Grand Challenge" and submitted to the non-real-time track of the "Speech Signal Improvement Challenge 2023", where it was ranked fifth. Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Tal Peer, Timo Gerkmann |
ICASSP | 4 |
| 2023 | Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement
Bunlong Lay, Simon Welker, Julius Richter, Timo Gerkmann |
INTERSPEECH | 1 |
| 2023 | Speech Enhancement and Dereverberation With Diffusion-Based Generative ModelsabstractIn this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement. We present a detailed overview of the diffusion process that is based on a stochastic differential equation and delve into an extensive theoretical examination of its implications. Opposed to usual conditional generation tasks, we do not start the reverse process from pure Gaussian noise but from a mixture of noisy speech and Gaussian noise. This matches our forward process which moves from clean speech to noisy speech by including a drift term. We show that this procedure enables using only 30 diffusion steps to generate high-quality clean speech estimates. By adapting the network architecture, we are able to significantly improve the speech enhancement performance, indicating that the network, rather than the formalism, was the main limitation of our original approach. In an extensive cross-dataset evaluation, we show that the improved method can compete with recent discriminative models and achieves better generalization when evaluating on a different corpus than used for training. We complement the results with an instrumental evaluation using real-world noisy recordings and a listening experiment, in which our proposed method is rated best. Examining different sampler configurations for solving the reverse process allows us to balance the performance and computational speed of the proposed method. Moreover, we show that the proposed method is also suitable for dereverberation and thus not limited to additive background noise removal. Code and audio examples are available online1https://github.com/sp-uhh/sgmse. Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Speech Enhancement Regularized by a Speaker Verification ModelabstractIn the last years, end-to-end neural networks were employed to enhance single-channel noisy speech. The Conv-Tasnet architecture is such a neural network and it has been successfully trained on the scale invariant signal-to-distortion ratio (SI-SDR) loss. However, we find that in a speech enhancement task Conv-Tasnet trained on the SI-SDR loss introduces distortions to the enhanced signal as the mismatch between training and testing data increases. To mitigate these effects, we investigate on a new loss function, that combines SI-SDR with a pretrained speaker verification (SV) model as a regularizer. As the SV model captures information about important characteristics of clean speech signals, we argue that the proposed regularization ensures that the estimated speech signal is speech-like even if there is an increasing mismatch between training and testing data. Accordingly, we observe considerable improvements in POLQA in mismatched training and testing conditions, e.g. when training on anechoic speech but testing on reverberant data. Bunlong Lay, Timo Gerkmann |
MMSP | 1 |