EDBT 2026 Demo / reviewers in the wild / expert
Stefan Uhlich
dblp:19/7822
· DBLP profile ↗
17ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0003-3158-4945ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Diffusion Bridges for Unsupervised Musical Audio Timbre TransferabstractMusic timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired monophonic single-instrument audio data. Each diffusion model is trained on a specific instrument with a Gaussian prior. During inference, a model is designated as the source model to map the input audio to its corresponding Gaussian prior, and another model is designated as the target model to reconstruct the target audio from this Gaussian prior, thereby facilitating timbre transfer. We compare our approach against existing unsupervised timbre transfer models such as VAEGAN and Gaussian Flow Bridges (GFB). Experimental results demonstrate that our method achieves both better Fréchet Audio Distance (FAD) and melody preservation, as reflected by lower pitch distances (DPD) compared to VAEGAN and GFB. Additionally, we discover that the noise level from the Gaussian prior, σ, can be adjusted to control the degree of melody preservation and amount of timbre transferred. Michele Mancusi, Yurii Halychanskyi, Kin Wai Cheuk, Eloi Moliner, Chieh-Hsin Lai, Stefan Uhlich, Junghyun Koo, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Giorgio Fabbro, Yuki Mitsufuji |
ICASSP | 6 |
| 2024 | SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
Bac Nguyen, Stefan Uhlich, Fabien Cardinaux, Lukas Mauch, Marzieh Edraki, Aaron C. Courville |
ECCV (69) | 2 |
| 2023 | Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio EffectsabstractWe propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording. All our models are trained in a self-supervised manner from an already-processed wet multitrack dataset with an effective data preprocessing method that alleviates the data scarcity of obtaining unprocessed dry data. We analyze the proposed encoder for the disentanglement capability of audio effects and also validate its performance for mixing style transfer through both objective and subjective evaluations. From the results, we show the proposed system not only converts the mixing style of multitrack audio close to a reference but is also robust with mixture-wise style transfer upon using a music source separation model. Junghyun Koo, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Stefan Uhlich, Kyogu Lee, Yuki Mitsufuji |
ICASSP | 4 |
| 2023 | Autotts: End-to-End Text-to-Speech Synthesis Through Differentiable Duration ModelingabstractParallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they are not jointly trained. In this paper, we propose a differentiable duration method for learning monotonic alignments between input and output sequences. Our method is based on a soft-duration mechanism that optimizes a stochastic process in expectation. Using this differentiable duration method, we introduce AutoTTS, a direct text-to-waveform speech synthesis model. AutoTTS enables high-fidelity speech synthesis through a combination of adversarial training and matching the total ground-truth duration. Experimental results show that our model obtains competitive results while enjoying a much simpler training pipeline. Audio samples are available online1. Bac Nguyen, Fabien Cardinaux, Stefan Uhlich |
ICASSP | 3 |
| 2023 | Improving Self-Supervised Learning for Audio Representations by Feature Diversity and DecorrelationabstractSelf-supervised learning (SSL) has recently shown remarkable results in closing the gap between supervised and unsupervised learning. The idea is to learn robust features that are invariant to distortions of the input data. Despite its success, this idea can suffer from a collapsing issue where the network produces a constant representation. To this end, we introduce SELFIE, a novel Self-supervised Learning approach for audio representation via Feature Diversity and Decorrelation. SELFIE avoids the collapsing issue by ensuring that the representation (i) maintains a high diversity among embeddings and (ii) decorrelates the dependencies between dimensions. SELFIE is pre-trained on the large-scale AudioSet dataset and its embeddings are validated on nine audio downstream tasks, including speech, music, and sound event recognition. Experimental results show that SELFIE outperforms existing SSL methods in several tasks. Bac Nguyen, Stefan Uhlich, Fabien Cardinaux |
ICASSP | 2 |
| 2022 | Music Source Separation With Deep Equilibrium ModelsabstractWhile deep neural network-based music source separation (MSS) is very effective and achieves high performance, its model size is often a problem for practical deployment. Deep implicit architectures such as deep equilibrium models (DEQ) were recently proposed, which can achieve higher performance than their explicit counterparts with limited depth while keeping the number of parameters small. This makes DEQ also attractive for MSS, especially as it was originally applied to sequential modeling tasks in natural language processing and thus should in principle be also suited for MSS. However, an investigation of a good architecture and training scheme for MSS with DEQ is needed as the characteristics of acoustic signals are different from those of natural language data. Hence, in this paper we propose an architecture and training scheme for MSS with DEQ. Starting with the architecture of Open-Unmix (UMX), we replace its sequence model with DEQ. We refer to our proposed method as DEQ-based UMX (DEQ-UMX). Experimental results show that DEQ-UMX performs better than the original UMX while reducing its number of parameters by 30%. Yuichiro Koyama, Naoki Murata, Stefan Uhlich, Giorgio Fabbro, Shusuke Takahashi, Yuki Mitsufuji |
ICASSP | 3 |
| 2022 | TRUNet: Transformer-Recurrent-U Network for Multi-channel Reverberant Sound Source SeparationabstractIn recent years, many deep learning techniques for single-channel sound source separation have been proposed using recurrent, convolutional and transformer networks.When multiple microphones are available, spatial diversity between speakers and background noise in addition to spectro-temporal diversity can be exploited by using multi-channel filters for sound source separation.Aiming at end-to-end multi-channel source separation, in this paper we propose a transformerrecurrent-U network (TRUNet), which directly estimates multi-channel filters from multi-channel input spectra.TRUNet consists of a spatial processing network with an attention mechanism across microphone channels aiming at capturing the spatial diversity, and a spectrotemporal processing network aiming at capturing spectral and temporal diversities.In addition to multi-channel filters, we also consider estimating single-channel filters from multi-channel input spectra using TRUNet.We train the network on a large reverberant dataset using a proposed combined compressed mean-squared error loss function, which further improves the sound separation performance.We evaluate the network on a realistic and challenging reverberant dataset, generated from measured room impulse responses of an actual microphone array.The experimental results on realistic reverberant sound source separation show that the proposed TRUNet outperforms state-of-the-art single-channel and multi-channel source separation methods. Ali Aroudi, Stefan Uhlich, Marc Ferras |
INTERSPEECH | 2 |
| 2021 | All For One And One For All: Improving Music Separation By Bridging NetworksabstractThis paper proposes several improvements for music separation with deep neural networks (DNNs), namely a multi-domain loss (MDL) and two combination schemes. First, by using MDL we take advantage of the frequency and time domain representation of audio signals. Next, we utilize the relationship among instruments by jointly considering them. We do this on the one hand by modifying the network architecture and introducing a CrossNet structure. On the other hand, we consider combinations of instrument estimates by using a new combination loss (CL). MDL and CL can easily be applied to many existing DNN-based separation methods as they are merely loss functions which are only used during training and do not affect the inference step. Experimental results show that the performance of Open-Unmix (UMX), a well-known and state-of-the-art open-source library for music separation, can be improved by utilizing our above schemes. Our modifications of UMX are open-sourced together with this paper. Ryosuke Sawata, Stefan Uhlich, Shusuke Takahashi, Yuki Mitsufuji |
ICASSP | 2 |
| 2020 | Mixed Precision DNNs: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso García, Stephen Tiedemann, Thomas Kemp, Akira Nakamura |
ICLR | 1 |
| 2020 | Multichannel Non-Negative Matrix Factorization Using Banded Spatial Covariance Matrices in Wavenumber DomainabstractBlind source separation exploiting multichannel information has long been a popular topic, and recently proposed methods based on the local Gaussian model have shown promising results despite its high computational cost for the case of many microphone signals. The low updating speed for such a model is mainly due to the inversion of a spatial covariance matrix, for which the complexity increases with the number of microphones, M, and is generally of order O(M3). Several projection-based approaches that attempt to concentrate energy on the diagonal part of the spatial covariance matrix have been introduced to circumvent the matrix inversion, which can reduce the complexity to O(M). In this article, we focus on the fast Fourier transform as a projection method because the energy concentration on the diagonal can be efficiently achieved compared with other projection-based methods. For the case where the diagonalization is imperfect, for example, owing to discontinuities at the edge of a linear array, we also developed a more robust algorithm approximating the tri-diagonal part of the spatial covariance matrix, which requires a complexity of O(M2) for the inversion by applying the Thomas algorithm. To remove the ad-hoc integration of post clustering after the decomposition, we also examine a self-clustering algorithm. Our evaluation shows better results than other previously proposed methods in terms of the separation quality under reverberant conditions as well as higher efficiency than multichannel non-negative matrix factorization. Yuki Mitsufuji, Stefan Uhlich, Norihiro Takamune, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Improving music source separation based on deep neural networks through data augmentation and network blendingabstractThis paper deals with the separation of music into individual instrument tracks which is known to be a challenging problem. We describe two different deep neural network architectures for this task, a feed-forward and a recurrent one, and show that each of them yields themselves state-of-the art results on the SiSEC DSD100 dataset. For the recurrent network, we use data augmentation during training and show that even simple separation networks are prone to overfitting if no data augmentation is used. Furthermore, we propose a blending of both neural network systems where we linearly combine their raw outputs and then perform a multi-channel Wiener filter post-processing. This blending scheme yields the best results that have been reported to-date on the SiSEC DSD100 dataset. Stefan Uhlich, Marcello Porcu, Franck Giron, Michael Enenkl, Thomas Kemp, Naoya Takahashi, Yuki Mitsufuji |
ICASSP | 1 |
| 2015 | NMF-based blind source separation using a linear predictive coding error clustering criterionabstractNon-negative matrix factorization (NMF) based sound source separation involves two phases: First, the signal spectrum is decomposed into components which, in a second step, are clustered in order to obtain estimates of the source signal spectra. The major challenge with this approach is the accuracy of the clustering algorithm in the second step, especially as most previously used clustering algorithms are focusing on the frequency part of NMF only and, hence, are missing the information of the time activation matrix. In this paper, we propose a novel clustering criterion which combines the frequency and time activation part of NMF. It is based on the linear predictive coding compression error and we show that it allows a good clustering of the NMF components while at the same time can be efficiently computed. Our new clustering criterion shows an overall improved performance compared with the current state-of-the-art clustering algorithms as we experiment on the TRIOS dataset. Xin Guo 0010, Stefan Uhlich, Yuki Mitsufuji |
ICASSP | 2 |
| 2015 | Deep neural network based instrument extraction from musicabstractThis paper deals with the extraction of an instrument from music by using a deep neural network. As prior information, we only assume to know the instrument types that are present in the mixture and, using this information, we generate the training data from a database with solo instrument performances. The neural network is built up from rectified linear units where each hidden layer has the same number of nodes as the output layer. This allows a least squares initialization of the layer weights and speeds up the training of the network considerably compared to a traditional random initialization. We give results for two mixtures, each consisting of three instruments, and evaluate the extraction performance using BSS Eval for a varying number of hidden layers. Stefan Uhlich, Franck Giron, Yuki Mitsufuji |
ICASSP | 1 |
| 2014 | Computing Jacobian and Hessian of Estimators and Their Application to Risk ApproximationabstractThis letter gives formulas to compute the Jacobian and Hessian of an estimator that can be written as the maximum of a given scoring function, which includes the important cases of maximum likelihood (ML) and least squares (LS) estimation. We use the knowledge about these derivatives to compute two approximations of the estimator risk and show that the linear risk approximation of an ML estimator coincides with the Cramér–Rao bound for the case of a Gaussian signal model where the underlying loss function that is used for the risk computation is the squared error loss. Stefan Uhlich |
IEEE Signal Process. Lett. | 1 |
| 2011 | Recursive estimation of room impulse responses with energy conservation constraintsabstractThis paper considers the problem of constrained tracking the time-varying room impulse response of a source/microphone pair. The constraint which is used to improve the performance stems from the energy conservation that has to hold for real-world impulse responses. We consider three different recursive estimators and compare their performance with the recursive weighted least squares algorithm which does not take the constraint into account. The simulation results show that exploiting this constraint decreases the mean squared error and is thus interesting for applications, especially in the low SNR regime. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 1 |
| 2009 | MMSE estimation in a linear signal model with ellipsoidal constraintsabstractThe estimation of an unknown parameter vector in a Gaussian linear model is studied in this paper. Two different cases are analyzed: the parameter vector is assumed to lie either in or on a given ellipsoid. The best estimator in terms of the mean squared error is derived. The performance of this estimator is analyzed and compared with the ordinary least squares, the constrained least squares and the linear minimax approach. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 1 |
| 2008 | A generalized optimal correlating transform for multiple description coding and its theoretical analysisabstractThis paper considers a coding scheme for data transmission over erasure channels which is also known as multiple description coding. The LMMSE prefilter method of Romano [1] is reviewed and generalized to allow three different operational modes of the prefilter. They include the possibility to decrease or increase the number of descriptions to be transmitted. We derive explicitly the Hessian matrix for an efficient calculation of the prefilter. We also study the properties of the distortion measure theoretically. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 1 |