VLDB 2026 Research / reviewers in the wild / expert
Ali Aroudi
dblp:119/2540
· DBLP profile ↗
13ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0001-5770-0858ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Advancing Active Speaker Detection for Egocentric VideosabstractThis paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing the lip region in the visual input, and (ii) applying motion blur augmentation. These methods significantly enhance the model’s performance in handling the challenges typical of egocentric videos. We showcase the effectiveness of these techniques on a simple but efficient causal audio-visual model. The proposed model, named EgoASD, demonstrates state-of-the-art performance on the EasyCom dataset, beating the previous SOTA by 1.7% mean Average Precision (mAP) with a model 2.5 times smaller. Our ablations highlight the importance of visual input, motion blur augmentation, the pretraining method and the importance of temporal context. To demonstrate its applicability in the real world, we apply our model to audio-visual speaker diarization, outperforming other baselines on EasyCom. Jaesung Huh, Juan Azcarreta, Anurag Kumar 0003, Ashutosh Pandey 0004, Ali Aroudi, Daniel D. E. Wong, Francesco Nesta, Buye Xu, Jacob Donley |
ICASSP | 5 |
| 2025 | Reexamining the Efficacy of MetricGAN for Speech EnhancementabstractMetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluation exclusively at high SNR. Firstly, we comprehensively assess MetricGAN models using mainstream metrics, surprisingly revealing MetricGAN models produce worse SISDR and STOI than unprocessed noisy speech. Secondly, we demonstrate that training MetricGAN models at low SNR often results in convergence to biased local minima, where PESQ scores are inflated while their SISDR and STOI values deteriorate significantly. In addition, we propose and validate two training tricks to address these issues: SISDR regularization and mixture-of-actor training. We find that these tricks effectively guide MetricGAN models to avoid local minima, thus improving speech quality. Ali Aroudi, Buye Xu, Ashutosh Pandey 0004, Francesco Nesta, Anurag Kumar 0003, Alexander Reich, Ke Tan 0001 |
ICASSP | 2 |
| 2024 | FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, Ke Tan 0001, Ashutosh Pandey 0004, Jung-Suk Lee, Buye Xu, Francesco Nesta |
INTERSPEECH | 2 |
| 2022 | TRUNet: Transformer-Recurrent-U Network for Multi-channel Reverberant Sound Source SeparationabstractIn recent years, many deep learning techniques for single-channel sound source separation have been proposed using recurrent, convolutional and transformer networks.When multiple microphones are available, spatial diversity between speakers and background noise in addition to spectro-temporal diversity can be exploited by using multi-channel filters for sound source separation.Aiming at end-to-end multi-channel source separation, in this paper we propose a transformerrecurrent-U network (TRUNet), which directly estimates multi-channel filters from multi-channel input spectra.TRUNet consists of a spatial processing network with an attention mechanism across microphone channels aiming at capturing the spatial diversity, and a spectrotemporal processing network aiming at capturing spectral and temporal diversities.In addition to multi-channel filters, we also consider estimating single-channel filters from multi-channel input spectra using TRUNet.We train the network on a large reverberant dataset using a proposed combined compressed mean-squared error loss function, which further improves the sound separation performance.We evaluate the network on a realistic and challenging reverberant dataset, generated from measured room impulse responses of an actual microphone array.The experimental results on realistic reverberant sound source separation show that the proposed TRUNet outperforms state-of-the-art single-channel and multi-channel source separation methods. Ali Aroudi, Stefan Uhlich, Marc Ferras |
INTERSPEECH | 1 |
| 2021 | DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source SeparationabstractMany deep learning techniques are available to perform source separation and reduce background noise. However, designing an end-to-end multi-channel source separation method using deep learning and conventional acoustic signal processing techniques still remains challenging. In this paper we propose a direction-of-arrival-driven beamforming network (DBnet) consisting of direction-of-arrival (DOA) estimation and beamforming layers for end-to-end source separation. We propose to train DBnet using loss functions that are solely based on the distances between the separated speech signals and the target speech signals, without a need for the ground-truth DOAs of speakers. To improve the source separation performance, we also propose end-to-end extensions of DBnet which incorporate post masking networks. We evaluate the proposed DBnet and its extensions on a very challenging dataset, targeting realistic far-field sound source separation in reverberant and noisy environments. The experimental results show that the proposed extended DBnet using a convolutional-recurrent post masking network outperforms state-of-the-art source separation methods. Ali Aroudi, Sebastian Braun |
ICASSP | 1 |
| 2020 | Improving Auditory Attention Decoding Performance of Linear and Non-Linear Methods using State-Space ModelabstractIdentifying the target speaker in hearing aid applications is crucial to improve speech understanding. Recent advances in electroencephalography (EEG) have shown that it is possible to identify the target speaker from single-trial EEG recordings using auditory attention decoding (AAD) methods. AAD methods reconstruct the attended speech envelope from EEG recordings, based on a linear least-squares cost function or non-linear neural networks, and then directly compare the reconstructed envelope with the speech envelopes of speakers to identify the attended speaker using Pearson correlation coefficients. Since these correlation coefficients are highly fluctuating, for a reliable decoding a large correlation window is used, which causes a large processing delay. In this paper, we investigate a state-space model using correlation coefficients obtained with a small correlation window to improve the decoding performance of the linear and the non-linear AAD methods. The experimental results show that the state-space model significantly improves the decoding performance. Ali Aroudi, Tobias de Taillez, Simon Doclo |
ICASSP | 1 |
| 2020 | Cognitive-Driven Binaural Beamforming Using EEG-Based Auditory Attention DecodingabstractIdentifying the target speaker in hearing aid applications is an essential ingredient to improve speech intelligibility. Recently, a least-squares-based auditory attention decoding (AAD) method has been proposed to identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers. Aiming at enhancing the target speaker and suppressing the interfering speaker and ambient noise, in this article, we propose a cognitive-driven speech enhancement system, consisting of a binaural beamformer which is steered based on AAD and estimated relative transfer function (RTF) vectors, which require estimates of the direction-of-arrivals (DOAs) of both speakers. For binaural beamforming and to generate reference signals for AAD, we consider either minimum-variance-distortionless-response (MVDR) beamformers or linearly-constrained-minimum-variance (LCMV) beamformers. Contrary to the binaural MVDR beamformer, the binaural LCMV beamformer allows to preserve the spatial impression of the acoustic scene and to control the suppression of the interfering speaker, which is important when intending to switch attention between speakers. The speech enhancement performance of the proposed system is evaluated in terms of the binaural signal-to-interference-plus-noise ratio (SINR) improvement in anechoic and reverberant conditions. Furthermore, we investigate the impact of RTF and DOA estimation errors and AAD errors on the speech enhancement performance. The experimental results show that the proposed system using LCMV beamformers yields a larger decoding performance and binaural SINR improvement compared to using MVDR beamformers. Ali Aroudi, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Cognitive-driven Binaural LCMV Beamformer Using EEG-based Auditory Attention DecodingabstractIdentifying the target speaker in hearing aid applications is an essential ingredient to improve speech intelligibility. To identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method was recently proposed. Aiming at enhancing the target speaker and suppressing the interfering speaker and ambient noise, in this paper we propose a cognitive-driven speech enhancement system, consisting of a direction-of-arrival (DOA) estimator, steerable beamformers and AAD. To preserve the spatial impression of the acoustic scene, which is important when intending to switch attention between speakers, the proposed system only partially suppresses the interfering speaker. The speech enhancement performance of the proposed system is evaluated in terms of the signal-to-interference-plus-noise ratio (SINR) improvement in anechoic and reverberant conditions. The experimental results show that the proposed system can obtain a considerably large SINR improvement (between 3.1 dB and 7.5 dB) in both conditions. Ali Aroudi, Simon Doclo |
ICASSP | 1 |
| 2018 | EEG-Based Auditory Attention Decoding Using Steerable Binaural Superdirective BeamformerabstractDuring the last decades significant progress in multi-microphone speech enhancement algorithms has been made for hearing aids. However, the performance of many algorithms depends on identifying the target speaker to be enhanced. To identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method was recently proposed. This AAD method however requires the clean speech signals of both the attended and the unattended speaker as reference signals for decoding. Since in practice only microphone signals, containing several undesired acoustic components, are available, in this paper we explore the potential of using steerable binaural superdirective beamformer for generating appropriate reference signals for decoding. The experimental results show that using steerable superdirective beamformer output signals improves the decoding performance compared to using the noisy microphone signals as reference signals. Ali Aroudi, Daniel Marquardt, Simon Doclo |
ICASSP | 1 |
| 2017 | EEG-based auditory attention decoding: Impact of reverberation, noise and interference reductionabstractTo identify the attended speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method has recently been proposed. The AAD method requires the clean speech signals of both the attended and the unattended speaker as reference signals for decoding. However, in practice only the binaural signals, containing several undesired acoustic components (reverberation, background noise and interference), and influenced by anechoic head-related transfer functions (HRTFs), are available. To generate appropriate reference signals for decoding from the binaural signals, it is important to understand the impact of these acoustic components on the AAD performance. In this paper, we investigate this impact for decoding several acoustic conditions (anechoic, reverberant, noisy, and reverberant-noisy) by using simulated speech signals in which different acoustic components have been reduced. The experimental results show that for obtaining a good decoding performance the joint suppression of reverberation, background noise and interference as undesired acoustic components is of great importance. Ali Aroudi, Simon Doclo |
SMC | 1 |
| 2016 | Auditory attention decoding with EEG recordings using noisy acoustic reference signalsabstractTo decode auditory attention from electroencephalography (EEG) recordings in a cocktail-party scenario with two competing speakers a least-squares method has recently been proposed, showing a promising decoding accuracy. This method however requires the clean speech signals of both the attended and the unattended speaker to be available as reference signals, which is difficult to achieve from the noisy recorded microphone signals in practice. In addition, optimizing the parameters involved in the spatio-temporal filter design is of crucial importance in order to reach the largest possible decoding performance. In this paper, the influence of noisy acoustic reference signals and the spatio-temporal filter and regularization parameters on the decoding performance is investigated. The results show that to some extent the decoding performance is robust to noisy acoustic reference signals, depending on the noise type. Furthermore, we demonstrate the crucial influence of several parameters on the decoding performance, especially when the acoustic reference signals used for decoding have been corrupted by noise. Ali Aroudi, Bojana Mirkovic, Maarten De Vos, Simon Doclo |
ICASSP | 1 |
| 2015 | Hidden Markov model-based speech enhancement using multivariate Laplace and Gaussian distributionsabstractIn this paper, statistical speech enhancement using hidden Markov model (HMM) is studied and new techniques for applying non‐Gaussian distributions are proposed. The superiority of using non‐Gaussian distributions in online adaptive noise suppression algorithms has been proven; however, in this study, this approach is formulated in an HMM‐based mean‐square error estimator (MMSE) estimator in which a priori models are trained in an off‐line manner. In addition, an analytical study of using different distributions other than autoregressive (AR) Gaussian distribution, such as Laplace, is presented in order to construct an accurate HMM as a priori model for discrete Fourier transform and discrete cosine transform feature vectors of speech signal. In the proposed framework, an HMM‐based MMSE estimator bassed on Gaussian assumption using diagonal covariance matrix is provided rather than AR hypothesis which is employed in the conventional AR‐HMM‐based speech enhancement algorithm. Experimental evaluations of the proposed methods are done in the presence of four different noise types at various signal‐to‐noise ratio levels which demonstrate the superiority of the proposed methods in most conditions in comparison with AR‐HMM. Ali Aroudi, Hadi Veisi, Hossein Sameti |
IET Signal Process. | 1 |
| 2012 | Automatic Noise Recognition Based on Neural Network Using LPC and MFCC Feature Parameters
Reza Haghmaram, Ali Aroudi, Mohammad Hossein Ghezel Aiagh, Hadi Veisi |
FedCSIS | 2 |