VLDB 2026 Research / reviewers in the wild / expert
Danilo de Oliveira
dblp:323/0116
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0001-3338-7913ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration ModelingabstractSpeech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic properties such as fundamental frequency, spectral envelope and energy, but often lack the ability to control the duration of sounds. To address this, we propose a duration modeling framework using resynthesis-based discrete content representations, enabling modification of speech duration to reflect target emotions and achieve controllable speech rates without using parallel data. Experimental results reveal that the inclusion of the proposed duration modeling framework significantly enhances emotional expressiveness, in the in-the-wild MSP-Podcast dataset. Analyses show that low-arousal emotions correlate with longer durations and slower speech rates, while high-arousal emotions produce shorter, faster speech. Navin Raj Prabhu, Danilo de Oliveira, Nale Lehmann-Willenbrock, Timo Gerkmann |
ASRU | 2 |
| 2025 | Investigating Training Objectives for Generative Speech EnhancementabstractGenerative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrodinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrodinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain1. Julius Richter, Danilo de Oliveira, Timo Gerkmann |
ICASSP | 2 |
| 2025 | Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
Danilo de Oliveira, Julius Richter, Jean-Marie Lemercier, Simon Welker, Timo Gerkmann |
INTERSPEECH | 1 |
| 2024 | Distilling Hubert with LSTMs via Decoupled Knowledge DistillationabstractMuch research effort is being applied to the task of compressing the knowledge of self-supervised models, which are powerful, yet large and memory consuming. Existing works generally distill internal features of self-supervised Transformer models. In this work, aiming at more flexibility in the design of the student model, we apply the method of Knowledge Distillation and its more recently proposed extension, Decoupled Knowledge Distillation, to the task of distilling HuBERT. We achieve this by leveraging the cluster prediction pre-training task of HuBERT, which provides valuable targets for the distillation objective. We thus propose to exploit the acquired flexibility to distill HuBERT’s Transformer layers into an LSTM-based model that reduces the number of parameters even below DistilHuBERT and at the same time shows improved performance in automatic speech recognition. Danilo de Oliveira, Timo Gerkmann |
ICASSP | 1 |
| 2024 | The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
Danilo de Oliveira, Simon Welker, Julius Richter, Timo Gerkmann |
INTERSPEECH | 1 |
| 2023 | Leveraging Semantic Information for Efficient Self-Supervised Emotion Recognition with Audio-Textual Distilled ModelsabstractIn large part due to their implicit semantic modeling, selfsupervised learning (SSL) methods have significantly increased the performance of valence recognition in speech emotion recognition (SER) systems.Yet, their large size may often hinder practical implementations.In this work, we take HuBERT as an example of an SSL model and analyze the relevance of each of its layers for SER.We show that shallow layers are more important for arousal recognition while deeper layers are more important for valence.This observation motivates the importance of additional textual information for accurate valence recognition, as the distilled framework lacks the depth of its large-scale SSL teacher.Thus, we propose an audio-textual distilled SSL framework that, while having only ∼20% of the trainable parameters of a large SSL model, achieves on par performance across the three emotion dimensions (arousal, valence, dominance) on the MSP-Podcast v1.10 dataset. Danilo de Oliveira, Navin Raj Prabhu, Timo Gerkmann |
INTERSPEECH | 1 |
| 2022 | Efficient Transformer-based Speech Enhancement Using Long Frames and STFT MagnitudesabstractThe SepFormer architecture shows very good results in speech separation. Like other learned-encoder models, it uses short frames, as they have been shown to obtain better performance in these cases. This results in a large number of frames at the input, which is problematic; since the SepFormer is transformer-based, its computational complexity drastically increases with longer sequences. In this paper, we employ the SepFormer in a speech enhancement task and show that by replacing the learned-encoder features with a magnitude short-time Fourier transform (STFT) representation, we can use long frames without compromising perceptual enhancement performance. We obtained equivalent quality and intelligibility evaluation scores while reducing the number of operations by a factor of approximately 8 for a 10-second utterance. Danilo de Oliveira, Tal Peer, Timo Gerkmann |
INTERSPEECH | 1 |