Mostafa Sadeghi

dblp:131/6589 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 Diffusion-based Unsupervised Audio-visual Speech Enhancement
abstract
This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on corresponding video data to simulate the speech generative distribution. This pre-trained model is then paired with the NMF-based noise model to estimate clean speech iteratively. Specifically, a diffusion-based posterior sampling approach is implemented within the reverse diffusion process, where after each iteration, a speech estimate is obtained and used to update the noise parameters. Experimental results confirm that the proposed AVSE approach not only outperforms its audio-only counterpart but also generalizes better than a recent supervised-generative AVSE method. Additionally, the new inference algorithm offers a better balance between inference speed and performance compared to the previous diffusion-based method. Code and demo available at: https://jeaneudesayilo.github.io/fast_UdiffSE
Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda
ICASSP2
2025 Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
Can Cui 0007, Paul Magron, Mostafa Sadeghi, Emmanuel Vincent 0001
MMSP3
2025 Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
abstract
Supervised models for speech enhancement are trained using artificially generated mixtures of clean speech and noise signals. However, the synthetic training conditions may not accurately reflect real-world conditions encountered during testing. This discrepancy can result in poor performance when the test domain significantly differs from the synthetic training domain. To tackle this issue, the UDASE task of the 7th CHiME challenge aimed to leverage real-world noisy speech recordings from the test domain for unsupervised domain adaptation of speech enhancement models. Specifically, this test domain corresponds to the CHiME-5 dataset, characterized by real multi-speaker and conversational speech recordings made in noisy and reverberant domestic environments, for which ground-truth clean speech signals are not available. In this paper, we present the objective and subjective evaluations of the systems that were submitted to the CHiME-7 UDASE task, and we provide an analysis of the results. This analysis reveals a limited correlation between subjective ratings and several supervised nonintrusive performance metrics recently proposed for speech enhancement. Conversely, the results suggest that more traditional intrusive objective metrics can be used for in-domain performance evaluation using the reverberant LibriCHiME-5 dataset developed for the challenge. The subjective evaluation indicates that all systems successfully reduced the background noise, but always at the expense of increased distortion. Out of the four speech enhancement methods evaluated subjectively, only one demonstrated an improvement in overall quality compared to the unprocessed noisy speech, highlighting the difficulty of the task. The tools and audio material created for the CHiME-7 UDASE task are shared with the community.
Simon Leglaive, Matthieu Fraticelli, Hend Elghazaly, Léonie Borne, Mostafa Sadeghi, Scott Wisdom, Manuel Pariente, John R. Hershey, Daniel Pressnitzer, Jon Barker
Comput. Speech Lang.5
2025 Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
abstract
We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed likelihood score, combined with the unconditional score via a trade-off hyperparameter. In this work, we propose two alternative algorithms that directly model the conditional reverse transition distribution of diffusion states. The first method integrates the diffusion prior with the observation model in a principled way, removing the need for hyperparameter tuning. The second defines a diffusion process over the noisy speech itself, yielding a fully tractable and exact likelihood score. Experiments on the WSJ0-QUT and VoiceBank-DEMAND datasets demonstrate improved enhancement metrics and greater robustness to domain shifts compared to both supervised and unsupervised baselines.
Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel, Xavier Alameda-Pineda
IEEE Signal Process. Lett.1
2024 Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning Loss
abstract
Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise, usually centered on noisy speech, and subsequently learn a parameterized model to reverse this process, conditionally on noisy speech. Unlike supervised methods, generative-based SE approaches often rely solely on an unsupervised loss, which may result in less efficient incorporation of conditioned noisy speech. To address this issue, we propose augmenting the original diffusion training objective with an ℓ2loss, measuring the discrepancy between ground-truth clean speech and its estimation at each diffusion time-step. Experimental results demonstrate the effectiveness of our proposed methodology.
Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel
ICASSP2
2024 A Weighted-Variance Variational Autoencoder Model for Speech Enhancement
abstract
We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative model, where the speech information is encoded in the variance as a function of a latent variable. In contrast to this commonly used approach, we propose a weighted variance generative model, where the contribution of each spectrogram time-frame in parameter learning is weighted. We impose a Gamma prior distribution on the weights, which would effectively lead to a Student’s t-distribution instead of Gaussian for speech generative modeling. We develop efficient training and speech enhancement algorithms based on the proposed generative model. Our experimental results on spectrogram auto-encoding and speech enhancement demonstrate the effectiveness and robustness of the proposed approach compared to the standard unweighted variance model.
Ali Golmakani, Mostafa Sadeghi, Xavier Alameda-Pineda, Romain Serizel
ICASSP2
2024 Unsupervised Speech Enhancement with Diffusion-Based Generative Models
abstract
Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to unseen conditions. To address this issue, we introduce an alternative approach that operates in an unsupervised manner, leveraging the generative power of diffusion models. Specifically, in a training phase, a clean speech prior distribution is learnt in the short-time Fourier transform (STFT) domain using score-based diffusion models, allowing it to unconditionally generate clean speech from Gaussian noise. Then, we develop a posterior sampling methodology for speech enhancement by combining the learnt clean speech prior with a noise model for speech signal inference. The noise parameters are simultaneously learnt along with clean speech estimation through an iterative expectation-maximisation (EM) approach. To the best of our knowledge, this is the first work exploring diffusion-based generative models for unsupervised speech enhancement, demonstrating promising results compared to a recent variational auto-encoder (VAE)-based unsupervised approach and a state-of-the-art diffusion-based supervised method. It thus opens a new direction for future research in unsupervised speech enhancement.
Berné Nortier, Mostafa Sadeghi, Romain Serizel
ICASSP2
2024 Posterior Sampling Algorithms for Unsupervised Speech Enhancement with Recurrent Variational Autoencoder
abstract
In this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Nevertheless, the involved iterative variational expectation-maximization (VEM) process at test time, which relies on a variational inference method, results in high computational complexity. To tackle this issue, we present efficient sampling techniques based on Langevin dynamics and Metropolis-Hasting algorithms, adapted to the EM-based speech enhancement with RVAE. By directly sampling from the intractable posterior distribution within the EM process, we circumvent the intricacies of variational inference. We conduct a series of experiments, comparing the proposed methods with VEM and a state-of-the-art supervised speech enhancement approach based on diffusion models. The results reveal that our sampling-based algorithms significantly outperform VEM, not only in terms of computational efficiency but also in overall performance. Furthermore, when compared to the supervised baseline, our methods showcase robust generalization performance in mismatched test conditions.
Mostafa Sadeghi, Romain Serizel
ICASSP1
2024 Unsupervised performance analysis of 3D face alignment with a statistically robust confidence test
Mostafa Sadeghi, Xavier Alameda-Pineda, Radu Horaud
Neurocomputing1
2023 End-to-End Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis
abstract
We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame cross-channel attention and a speaker-attributed Transformer-based decoder. To the best of our knowledge, this is the first model that efficiently integrates ASR and speaker identification modules in a multichannel setting. On simulated mixtures of LibriSpeech data, our system reduces the word error rate (WER) by up to 12% and 16% relative compared to previously proposed single-channel and multichannel approaches, respectively. Furthermore, we investigate the impact of different input features, including multichannel magnitude and phase information, on the ASR performance. Finally, our experiments on the AMI corpus confirm the effectiveness of our system for real-world multichannel meeting transcription.
Can Cui 0007, Imran A. Sheikh, Mostafa Sadeghi, Emmanuel Vincent 0001
ASRU3
2023 Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative Model
abstract
Deep latent variable generative models based on variational autoencoder (VAE) have shown promising performance for audio-visual speech enhancement (AVSE). The underlying idea is to learn a VAE-based audio-visual prior distribution for clean speech data, and then combine it with a statistical noise model to recover a speech signal from a noisy audio recording and video (lip images) of the target speaker. Existing generative models developed for AVSE do not take into account the sequential nature of speech data, which prevents them from fully incorporating the power of visual data. In this paper, we present an audio-visual deep Kalman filter (AV-DKF) generative model which assumes a first-order Markov chain model for the latent variables and effectively fuses audio-visual data. Moreover, we develop an efficient inference methodology to estimate speech signals at test time. We conduct a set of experiments to compare different variants of generative models for speech enhancement. The results demonstrate the superiority of the AV-DKF model compared with both its audio-only version and the non-sequential audio-only and audio-visual VAE-based models.
Ali Golmakani, Mostafa Sadeghi, Romain Serizel
ICASSP2
2023 Fast and Efficient Speech Enhancement with Variational Autoencoders
abstract
Unsupervised speech enhancement based on variational autoencoders has shown promising performance compared with the commonly used supervised methods. This approach involves the use of a pre-trained deep speech prior along with a parametric noise model, where the noise parameters are learned from the noisy speech signal with an expectation-maximization (EM)-based method. The E-step involves an intractable latent posterior distribution. Existing algorithms to solve this step are either based on computationally heavy Monte Carlo Markov Chain sampling methods and variational inference, or inefficient optimization-based methods. In this paper, we propose a new approach based on Langevin dynamics that generates multiple sequences of samples and comes with a total variation-based regularization to incorporate temporal correlations of latent vectors. Our experiments demonstrate that the developed framework makes an effective compromise between computational efficiency and enhancement quality, and outperforms existing methods.
Mostafa Sadeghi, Romain Serizel
ICASSP1
2023 Expression-Preserving Face Frontalization Improves Visually Assisted Speech Processing
Zhiqi Kang, Mostafa Sadeghi, Radu Horaud, Xavier Alameda-Pineda
Int. J. Comput. Vis.2
2022 The Impact of Removing Head Movements on Audio-Visual Speech Enhancement
abstract
This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by past and recent studies: they challenge today’s learning-based methods as they often degrade the performance of models that are trained on clean, frontal, and steady face images. To alleviate this problem, we propose to use robust face frontalization (RFF) in combination with an AVSE method based on a variational auto-encoder (VAE) model. We briefly describe the basic ingredients of the proposed pipeline and we perform experiments with a recently released audio-visual dataset. In the light of these experiments, and based on three standard metrics, namely STOI, PESQ and SI-SDR, we conclude that RFF improves the performance of AVSE by a considerable margin.1
Zhiqi Kang, Mostafa Sadeghi, Radu Horaud, Xavier Alameda-Pineda, Jacob Donley, Anurag Kumar 0003
ICASSP2
2022 A Sparsity-promoting Dictionary Model for Variational Autoencoders
abstract
Structuring the latent space in probabilistic deep generative models, e.g., variational autoencoders (VAEs), is important to yield more expressive models and interpretable representations, and to avoid overfitting.One way to achieve this objective is to impose a sparsity constraint on the latent variables, e.g., via a Laplace prior.However, such approaches usually complicate the training phase, and they sacrifice the reconstruction quality to promote sparsity.In this paper, we propose a simple yet effective methodology to structure the latent space via a sparsity-promoting dictionary model, which assumes that each latent code can be written as a sparse linear combination of a dictionary's columns.In particular, we leverage a computationally efficient and tuning-free method, which relies on a zeromean Gaussian latent prior with learnable variances.We derive a variational inference scheme to train the model.Experiments on speech generative modeling demonstrate the advantage of the proposed approach over competing techniques, since it promotes sparsity while not deteriorating the output speech quality.
Mostafa Sadeghi, Paul Magron
INTERSPEECH1
2021 Switching Variational Auto-Encoders for Noise-Agnostic Audio-Visual Speech Enhancement
abstract
Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, which at test time is combined with a noise model, e.g. nonnegative matrix factorization (NMF), whose parameters are learned without supervision. Consequently, the proposed model is agnostic to the noise type. When visual data are clean, audio-visual VAE-based architectures usually outperform the audio-only counterpart. The opposite happens when the visual data are corrupted by clutter, e.g. the speaker not facing the camera. In this paper, we propose to find the optimal combination of these two architectures through time. More precisely, we introduce the use of a latent sequential variable with Markovian dependencies to switch between different VAE architectures through time in an unsupervised manner: leading to switching variational auto-encoder (SwVAE). We propose a variational factorization to approximate the computationally intractable posterior distribution. We also derive the corresponding variational expectation-maximization algorithm to estimate the parameters of the model and enhance the speech signal. Our experiments demonstrate the promising performance of SwVAE.
Mostafa Sadeghi, Xavier Alameda-Pineda
ICASSP1
2020 Low Mutual and Average Coherence Dictionary Learning Using Convex Approximation
abstract
In dictionary learning, a desirable property for the dictionary is to be of low mutual and average coherences. Mutual coherence is defined as the maximum absolute correlation between distinct atoms of the dictionary, whereas the average coherence is a measure of the average correlations. In this paper, we consider a dictionary learning problem regularized with the average coherence and constrained by an upper-bound on the mutual coherence of the dictionary. Our main contribution is then to propose an algorithm for solving the resulting problem based on convexly approximating the cost function over the dictionary. Experimental results demonstrate that the proposed approach has higher convergence rate and lower representation error (with a fixed sparsity parameter) than other methods, while yielding similar mutual and average coherence values.
Javad Parsa, Mostafa Sadeghi, Massoud Babaie-Zadeh, Christian Jutten
ICASSP2
2020 Robust Unsupervised Audio-Visual Speech Enhancement Using a Mixture of Variational Autoencoders
abstract
Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised speech enhancement. When visual data is clean, speech enhancement with audio-visual VAE shows a better performance than with audio-only VAE, which is trained on audio-only data. However, audio-visual VAE is not robust against noisy visual data, e.g., when for some video frames, speaker face is not frontal or lips region is occluded. In this paper, we propose a robust unsupervised audio-visual speech enhancement method based on a per-frame VAE mixture model. This mixture model consists of a trained audio-only VAE and a trained audio-visual VAE. The motivation is to skip noisy visual frames by switching to the audio-only VAE model. We present a variational expectation-maximization method to estimate the parameters of the model. Experiments show the promising performance of the proposed method.
Mostafa Sadeghi, Xavier Alameda-Pineda
ICASSP1
2020 Dictionary learning with low mutual coherence constraint
Mostafa Sadeghi, Massoud Babaie-Zadeh
Neurocomputing1
2020 Audio-Visual Speech Enhancement Using Conditional Variational Auto-Encoders
abstract
Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. One advantage of this generative approach is that it does not require pairs of clean and noisy speech signals at training. In this article, we propose audio-visual variants of VAEs for single-channel and speaker-independent speech enhancement. We develop a conditional VAE (CVAE) where the audio speech generative process is conditioned on visual information of the lip region. At test time, the audio-visual speech generative model is combined with a noise model based on nonnegative matrix factorization, and speech enhancement relies on a Monte Carlo expectation-maximization algorithm. Experiments are conducted with the recently published NTCD-TIMIT dataset as well as the GRID corpus. The results confirm that the proposed audio-visual CVAE effectively fuses audio and visual information, and it improves the speech enhancement performance compared with the audio-only VAE model, especially when the speech signal is highly corrupted by noise. We also show that the proposed unsupervised audio-visual speech enhancement approach outperforms a state-of-the-art supervised deep learning method.
Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Novel efficient full adder and full subtractor designs in quantum cellular automata
Mostafa Sadeghi, Keivan Navi, Mehdi Dolatshahi
J. Supercomput.1
2017 Incoherent Unit-Norm Frame Design via an Alternating Minimization Penalty Method
abstract
This letter is concerned with designing incoherent unit-norm frames, i.e., a set of vectors in a finite dimensional Hilbert space with unit norms and very low absolute pairwise correlations. Due to their widespread use in a variety of applications, including compressed sensing and coding theory, incoherent frame design has received considerable attention, and many algorithms have been proposed to this aim. In this letter, a new algorithm is presented which constructs incoherent frames by minimizing the maximum absolute pairwise correlations (mutual coherence) of the frame vectors. Our strategy is based on an alternating minimization penalty method, which admits efficient solvers using proximal algorithms. Experimental results on designing incoherent frames of various dimensions show that our algorithm outperforms some recent methods in the literature.
Mostafa Sadeghi, Massoud Babaie-Zadeh
IEEE Signal Process. Lett.1
2016 Union of low-rank subspaces detector
abstract
The problem of signal detection using a flexible and general model is considered. Owing to applicability and flexibility of sparse signal representation and approximation, it has attracted a lot of attention in many signal processing areas. In this study, the authors propose a new detection method based on sparse decomposition in a union of subspaces model. Their proposed detector uses a dictionary that can be interpreted as a bank of matched subspaces. This improves the performance of signal detection, as it is a generalisation for detectors. Low‐rank assumption for the desired signals implies that the representations of these signals in terms of some proper bases would be sparse. Their proposed detector exploits sparsity in its decision rule. They demonstrate the high efficiency of their method in the cases of voice activity detection in speech processing.
Mohsen Joneidi, Parvin Ahmadi, Mostafa Sadeghi, Nazanin Rahnavard
IET Signal Process.3
2013 Dictionary Learning for Sparse Representation: A Novel Approach
abstract
A dictionary learning problem is a matrix factorization in which the goal is to factorize a training data matrix, Y, as the product of a dictionary, D, and a sparse coefficient matrix, X, as follows, Y ≃ DX. Current dictionary learning algorithms minimize the representation error subject to a constraint on D (usually having unit column-norms) and sparseness of X. The resulting problem is not convex with respect to the pair (D,X). In this letter, we derive a first order series expansion formula for the factorization, DX. The resulting objective function is jointly convex with respect to D and X. We simply solve the resulting problem using alternating minimization and apply some of the previously suggested algorithms onto our new problem. Simulation results on recovery of a known dictionary and dictionary learning for natural image patches show that our new problem considerably improves performance with a little additional computational load.
Mostafa Sadeghi, Massoud Babaie-Zadeh, Christian Jutten
IEEE Signal Process. Lett.1