VLDB 2026 Research / reviewers in the wild / expert
Mhd Modar Halimeh
dblp:225/9947
· DBLP profile ↗
10ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0001-6113-3067ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Relation Between Speech Quality and Quantized Latent Representations of Neural CodecsabstractNeural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded signals, e.g., speech. In this work, we investigate the relation between the latent representation of the input signal learned by a neural codec and the quality of speech signals. To do so, we introduce Latent-representation-to-Quantization error Ratio (LQR) measures, which quantify the distance from the idealized neural codec’s speech signal model for a given speech signal. We compare the proposed metrics to intrusive measures as well as data-driven supervised methods using two subjective speech quality datasets. This analysis shows that the proposed LQR correlates strongly (up to 0.9 Pearson’s correlation) with the subjective quality of speech. Despite being a non-intrusive metric, this yields a competitive performance with, or even better than, other pre-trained and intrusive measures. These results show that LQR is a promising basis for more sophisticated speech quality measures. Mhd Modar Halimeh, Matteo Torcoli, Philipp Grundhuber, Emanuël A. P. Habets |
ICASSP | 1 |
| 2024 | Odaq: Open Dataset of Audio QualityabstractResearch into the prediction and analysis of perceived audio quality is hampered by the scarcity of openly available datasets of audio signals accompanied by corresponding subjective quality scores. To address this problem, we present the Open Dataset of Audio Quality (ODAQ), a new dataset containing the results of a MUSHRA listening test conducted with expert listeners from 2 international laboratories. ODAQ contains 240 audio samples and corresponding quality scores. Each audio sample is rated by 26 listeners. The audio samples are stereo audio signals sampled at 44.1 or 48 kHz and are processed by a total of 6 method classes, each operating at different quality levels. The processing method classes are designed to generate quality degradations possibly encountered during audio coding and source separation, and the quality levels for each method class span the entire quality range. The diversity of the processing methods, the large span of quality levels, the high sampling frequency, and the pool of international listeners make ODAQ particularly suited for further research into subjective and objective audio quality. The dataset is released with permissive licenses, and the software used to conduct the listening test is also made publicly available. Matteo Torcoli, Chih-Wei Wu, Sascha Dick, Phillip A. Williams, Mhd Modar Halimeh, William Wolcott, Emanuël A. P. Habets |
ICASSP | 5 |
| 2024 | Nonlinear acoustic echo cancellation based on pipelined Hermite filters
Mhd Modar Halimeh, Yi-Fei Pu, Lu Lu 0005, Walter Kellermann |
Signal Process. | 2 |
| 2023 | Exploiting Spatial Information with the Informed Complex-Valued Spatial Autoencoder for Target Speaker ExtractionabstractIn conventional multichannel audio signal enhancement, spatial and spectral filtering are often performed sequentially. In contrast, it has been shown that for neural spatial filtering a joint approach of spectro-spatial filtering is more beneficial. In this contribution, we investigate the spatial filtering performed by such a time-varying spectro-spatial filter. We extend the recently proposed complex-valued spatial autoencoder (COSPA) for the task of target speaker extraction by leveraging its interpretable structure and purposefully informing the network of the target speaker’s position. We show that the resulting informed COSPA (iCOSPA) effectively and flexibly extracts a target speaker from a mixture of speakers. We also find that the proposed architecture is well capable of learning pronounced spatial selectivity patterns and show that the results depend significantly on the training target and the reference signal when computing various evaluation metrics. Annika Briegleb, Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 2 |
| 2022 | Complex-Valued Spatial Autoencoders for Multichannel Speech EnhancementabstractIn this contribution, we present a novel online approach to multichannel speech enhancement. The proposed method estimates the enhanced signal through a filter-and-sum framework. More specifically, complex-valued masks are estimated by a deep complex-valued neural network, termed the complex-valued spatial autoencoder. The proposed network is capable of manipulating both the phase and the amplitude of the microphone signals and hence, the network is able to exploit both spatial and spectral characteristics of the desired source signal resulting in a physically plausible spatial selectivity and superior speech quality. Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 1 |
| 2021 | Combining Adaptive Filtering And Complex-Valued Deep Postfiltering For Acoustic Echo CancellationabstractIn this contribution, we introduce a novel approach to noise-robust acoustic echo cancellation employing a complex-valued Deep Neural Network (DNN) for postfiltering. In a first step, early linear echo components are removed using a double-talk robust adaptive filter. The residual signal is subsequently processed by the proposed post-filter (PF). Due to its complex-valued nature, the PF allows to sup-press unwanted signal components without introducing distortions to the near-end speaker. For training and evaluation, we exclusively use data from the ICASSP 2021 AEC challenge. Exploiting only a moderate amount of training data, we demonstrate the efficacy of the proposed method. Specifically, we show that the PF (i) benefits significantly from a preceding linear adaptive filter and (ii) significantly outperforms a conventional real-valued DNN-based PF. Mhd Modar Halimeh, Thomas Haubner, Annika Briegleb, Alexander Schmidt 0004, Walter Kellermann |
ICASSP | 1 |
| 2020 | Efficient Multichannel Nonlinear Acoustic Echo Cancellation Based on a Cooperative StrategyabstractWhile a common approach to address nonlinear distortions, emitted by multiple loudspeakers and observed by multiple microphones, is to use post-filtering techniques, this paper proposes a cooperative strategy to rather model and then cancel such distortions. In this approach, the overall problem of modeling distortions emitted by a number of loudspeakers is divided into multiple simpler and easier tasks of estimating distortions emitted by subsets of loudspeakers. This approach allows also the exploitation of the physical configuration of the loudspeakers and microphones to select certain microphone signals for estimating the nonlinearity of loudspeakers that contribute the predominant part of the acoustic echo to this microphone signal. The proposed strategy is realized using the elitist resampling particle filter and the Gaussian particle filter. Both variants are evaluated and compared to a linear approach using synthesized and real recordings. Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 1 |
| 2019 | Neural Networks Sequential Training Using Variational Gaussian Particle FilterabstractIn this paper, we propose a sequential training algorithm for feed-forward neural networks based on particle filtering. The proposed algorithm uses variational learning to tailor a proposal density by minimizing the variational energy. This density is then incorporated into the Gaussian particle filter framework. The proposed algorithm and an extension to it using evolutionary resampling are compared to training a neural network using a random walk-based particle filter, an extended Kalman filter, the use of variational learning only, and the backpropagation algorithm, using a synthetic dataset generated by a time-varying random process and a real dataset, where the proposed approach resulted in a moderately lower training and testing errors and a better convergence behavior, rendering the algorithm attractive for uses such as neural networks pre-training. Mhd Modar Halimeh, Andreas Brendel, Walter Kellermann |
ICASSP | 1 |
| 2019 | A Neural Network-Based Nonlinear Acoustic Echo CancellerabstractIn this letter, we introduce a novel approach for nonlinear acoustic echo cancellation. The proposed approach uses the principle of transfer learning to train a neural network that approximates the nonlinear function responsible for the nonlinear distortions and generalizes this network to different acoustic conditions. The topology of the proposed network is inspired by the conventional adaptive filtering approaches for nonlinear acoustic echo cancellation. The network is trained to model the nonlinear distortions using the conventional error backpropagation algorithm. In deployment, and in order to account for any variation or discrepancy between training and deployment conditions, only a subset of the network's parameters is adapted using the significance-aware elitist resampling particle filter. The proposed approach is evaluated and verified using synthesized nonlinear distortions and real nonlinear distortions recorded by a commercial mobile phone. Mhd Modar Halimeh, Christian Huemmer 0001, Walter Kellermann |
IEEE Signal Process. Lett. | 1 |
| 2018 | Nonlinear Acoustic Echo Cancellation Using Elitist Resampling Particle FilterabstractThis paper considers an effective method for nonlinear acoustic echo cancellation (NL-AEC). More specifically, we model the nonlinear echo path by a latent state vector capturing the coefficients of a memoryless processor and a linear finite impulse response filter. To estimate the posterior probability distribution of the state vector, an elitist particle filter based on evolutionary strategies (EPFES) has been proposed, which evaluates realizations of the latent state vector based on long-term fitness measures. This method includes a manually-tuned recursive calculation of the probabilities that the observation has been produced by the state-vector realizations. For avoiding this manual tuning, we introduce a new approach denoted as Elitist Resampling Particle Filtering (ERPF) which can also be shown to combine the advantages of the Sequential Importance Sampling Particle Filter (SIS-PF) and the Sequential Importance Sampling/Resampling Particle Filter (SIR-PF). This new approach allows universal use and leads to superior system identification performance compared to both the original EPFES as well as the SIR-PF, as verified for a simulated scenario and a real smartphone recording. Mhd Modar Halimeh, Christian Huemmer 0001, Walter Kellermann |
ICASSP | 1 |