VLDB 2026 Research / reviewers in the wild / expert
Amin Edraki
dblp:263/4884
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-0843-5522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Intelligibility Prediction for Time-Modified Speech Signals Using Spectro-Temporal Modulation FeaturesabstractReference-based speech intelligibility prediction algorithms (RB-SIPAs) are limited to speech degradations that maintain the time alignment between the clean and degraded speech signals. We address this limitation by augmenting the existing RB-SIPAs framework with time alignment. To achieve robust time alignment at low signal-to-noise ratios, we propose using spectro-temporal modulations (STM) of the speech signals for dynamic time warping (DTW). Our experiments demonstrate that in DTW, the use of specific STM components as features for clean and time-modified degraded speech achieves the baseline time alignment obtained using clean speech and noise-free time-modified speech. Moreover, we propose two methods to incorporate time alignment into existing RB-SIPAs. Using these methods, we compare the output scores of the RB-SIPA with listening test scores and show better correlation results using the STM features as compared to MFCCs. Aymen Bashir, Haolan Wang, Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2025 | Noise-Robust Hearing Aid Voice ControlabstractAdvancing the design of robust hearing aid (HA) voice control is crucial to increase the HA use rate among hard of hearing people as well as to improve HA users' experience. In this work, we contribute towards this goal by, first, presenting a novel HA speech dataset consisting of noisy own voice captured by 2 behind-the-ear (BTE) and 1 in-ear-canal (IEC) microphones. Second, we provide baseline HA voice control results from the evaluation of light, state-of-the-art keyword spotting models utilizing different combinations of HA microphone signals. Experimental results show the benefits of exploiting bandwidth-limited bone-conducted speech (BCS) from the IEC microphone to achieve noise-robust HA voice control. Furthermore, results also demonstrate that voice control performance can be boosted by assisting BCS by the broader-bandwidth BTE microphone signals. Aiming at setting a baseline upon which the scientific community can continue to progress, the HA noisy speech dataset has been made publicly available. Iván López-Espejo, Eros Roselló, Amin Edraki, Naomi Harte, Jesper Jensen 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Speaker Adaptation For Enhancement Of Bone-Conducted SpeechabstractDeep neural network (DNN)-based speech enhancement models often face challenges in maintaining their performance for speakers not encountered during training. This challenge is exacerbated in applications such as enhancement and bandwidth extension of bone-conducted speech, where the distortion exhibits a close correlation with speaker-specific characteristics. We address this issue by introducing a bottleneck module aimed at disentangling speaker-specific characteristics from speech content in speech enhancement DNNs. A DNN model is trained for enhancement of bone-conducted speech and modified with the proposed bottleneck module. We evaluate the DNN’s adaptability to unseen speakers through fine-tuning the network with a limited amount of adaptation data. The results show that the proposed bottleneck module can enhance adaptation performance to new unseen speakers, especially when limited amount of speaker-specific adaptation data is available. Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty |
ICASSP | 1 |
| 2024 | No-Reference Speech Intelligibility Prediction Leveraging a Noisy-Speech ASR Pre-Trained ModelabstractRecent advances in deep learning have improved the capabilities of data-driven speech intelligibility prediction (SIP) algorithms. Nevertheless, the scarcity of speech intelligibility datasets limits the development of data-driven algorithms. This study introduces a set of no-reference SIP algorithms leveraging a pre-trained wav2vec 2.0 backbone. We adapt wav2vec 2.0 for automatic speech recognition under additive noise conditions with a parameter-efficient methodology, low-rank adaptation. We demonstrate no-reference SIP algorithms designed with this approach using a moderate amount of training data. The best designs perform on par or even better than a state-of-the-art reference-based SIP algorithm across a variety of datasets comprising different degradation types. Haolan Wang, Amin Edraki, Wai-Yip Chan, Iván López-Espejo, Jesper Jensen 0001 |
INTERSPEECH | 2 |
| 2023 | On the deficiency of intelligibility metrics as proxies for subjective intelligibilityabstractA recent trend in deep neural network (DNN)-based speech enhancement consists of using intelligibility and quality metrics as loss functions for model training with the aim of achieving high subjective speech intelligibility and perceptual quality in real-life conditions. In this study, we analyze a variety of loss functions, including some based on state-of-the-art intelligibility and quality metrics, to train an end-to-end speech enhancement system based on a fully convolutional neural network. The loss functions include perceptual metric for speech quality evaluation (PMSQE), scale-invariant signal-to-distortion ratio (SI-SDR), SI-SDR integrating speech pre-emphasis, short-time objective intelligibility (STOI), extended STOI (ESTOI), spectro-temporal glimpsing index (STGI), and a composite loss function combining STGI and SI-SDR. While DNNs trained with these loss functions produce notable speech intelligibility (and quality) gains according to pertinent objective metrics, we conduct a subjective intelligibility test that contradicts this result, showing no intelligibility improvement. From the results of this study, our conclusion is twofold: (1) subjective intelligibility evaluation is currently not replaceable by objective intelligibility evaluation, and (2) both the development of meaningful intelligibility metrics and DNN-based speech enhancement systems that can consistently improve the intelligibility of noisy speech for human listening remain open problems. Iván López-Espejo, Amin Edraki, Wai-Yip Chan, Zheng-Hua Tan, Jesper Jensen 0001 |
Speech Commun. | 2 |
| 2021 | A Spectro-Temporal Glimpsing Index (STGI) for Speech Intelligibility PredictionabstractWe propose a monaural intrusive speech intelligibility prediction (SIP) algorithm called STGI based on detecting glimpses in short-time segments in a spectro-temporal modulation decomposition of the input speech signals. Unlike existing glimpse-based SIP methods, the application of STGI is not limited to additive uncorrelated noise; STGI can be employed in a broad range of degradation conditions. Our results show that STGI performs consistently well across 15 datasets covering degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, checkerboard noise, and gated noise. Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty |
Interspeech | 1 |
| 2021 | Speech Intelligibility Prediction Using Spectro-Temporal Modulation AnalysisabstractSpectro-temporal modulations are believed to mediate the analysis of speech sounds in the human primary auditory cortex. Inspired by humans' robustness in comprehending speech in challenging acoustic environments, we propose an intrusive speech intelligibility prediction (SIP) algorithm, wSTMI, for normal-hearing listeners based on spectro-temporal modulation analysis (STMA) of the clean and degraded speech signals. In the STMA, each of 55 modulation frequency channels contributes an intermediate intelligibility measure. A sparse linear model with parameters optimized using Lasso regression results in combining the intermediate measures of 8 of the most salient channels for SIP. In comparison with a suite of 10 SIP algorithms, wSTMI performs consistently well across 13 datasets, which together cover degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, and speech interruption. We show that the optimized parameters of wSTMI may be interpreted in terms of modulation transfer functions of the human auditory system. Thus, the proposed approach offers evidence affirming previous studies of the perceptual characteristics underlying speech signal intelligibility. Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Improvement and Assessment of Spectro-Temporal Modulation Analysis for Speech Intelligibility EstimationabstractSeveral recent high-performing intelligibility estimators of acoustically degraded speech signals employ temporal modulation analysis. In this paper, we investigate the utility of using both spectro- and temporal-modulation for estimating speech intelligibility. We modified a pre-existing speech intelligibility estimation scheme (STMI) that was inspired by human auditory spectro-temporal modulation analysis. We produced several variants of the modified STMI and assessed their intelligibility prediction accuracy, in comparison with several high-performing estimators. Among the estimators tested, one of the STMI variants and eSTOI performed consistently well on both noisy and reverberated speech. These results suggest that spectro-temporal modulation analysis is useful for certain degradation conditions such as modulated noise and reverberation. Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty |
INTERSPEECH | 1 |