Iván López-Espejo

dblp:153/4531 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
7since 2021 · last 2027
0000-0001-8634-7897ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2027 Advancing listening effort-based evaluation of speech enhancement systems
abstract
This study investigates response time as a behavioral indicator related to listening effort (LE) for evaluating speech enhancement (SE) systems. English and Norwegian intelligibility matrix tests were conducted within a single-task paradigm that incorporated click-time recording (logging the precise time of all participant clicks), enabling simultaneous estimation of speech intelligibility and LE-related temporal behavior. Three temporal proxy measures for LE were examined—time per stimulus, reaction time, and word click time—across a broad range of input signal-to-noise ratios (SNRs) and for both discriminative and generative enhancement approaches. Time per stimulus showed an inverted-U pattern across SNRs, whereas reaction time and word click time exhibited monotonic behavior, providing more directly interpretable metrics for comparative evaluation. Analyses of pairwise SNR comparisons revealed that increases in our LE-related temporal measures at higher SNRs precede measurable intelligibility declines, suggesting that these temporal metrics can be more sensitive than intelligibility in this regime. Overall, the proposed framework—where LE-related measurements remain unknown to participants—offers a comprehensive and nuanced behavioral tool for SE evaluation, complementing intelligibility particularly under realistic, moderate-to-high SNR conditions. • A unified behavioral framework that jointly captures intelligibility and listening-effort-related information within a single listening task, enabling more expressive evaluation of speech enhancement systems. • Three temporal proxy measures for listening effort (time per stimulus, reaction time, and word click time) that capture complementary aspects of cognitive effort during system evaluation. • Evidence that temporal effort-related measures increase at higher signal-to-noise ratios (SNRs) before observable intelligibility declines, providing complementary insight under moderate-to-high SNR conditions.
Iván López-Espejo, Femke B. Gelderblom, Tron V. Tronstad, Christoffer Blomberg Skiaker, Naomi Harte
Comput. Speech Lang.1
2025 Noise-Robust Hearing Aid Voice Control
abstract
Advancing the design of robust hearing aid (HA) voice control is crucial to increase the HA use rate among hard of hearing people as well as to improve HA users' experience. In this work, we contribute towards this goal by, first, presenting a novel HA speech dataset consisting of noisy own voice captured by 2 behind-the-ear (BTE) and 1 in-ear-canal (IEC) microphones. Second, we provide baseline HA voice control results from the evaluation of light, state-of-the-art keyword spotting models utilizing different combinations of HA microphone signals. Experimental results show the benefits of exploiting bandwidth-limited bone-conducted speech (BCS) from the IEC microphone to achieve noise-robust HA voice control. Furthermore, results also demonstrate that voice control performance can be boosted by assisting BCS by the broader-bandwidth BTE microphone signals. Aiming at setting a baseline upon which the scientific community can continue to progress, the HA noisy speech dataset has been made publicly available.
Iván López-Espejo, Eros Roselló, Amin Edraki, Naomi Harte, Jesper Jensen 0001
IEEE Signal Process. Lett.1
2024 Anti-spoofing Ensembling Model: Dynamic Weight Allocation in Ensemble Models for Improved Voice Biometrics Security
abstract
This paper proposes an ensembling model as spoofed speech countermeasure, with a particular focus on synthetic voice. Despite the recent advances in speaker verification based on deep neural networks, this technology is still susceptible to various malicious attacks, so that some kind of countermeasures are needed. While an increasing number of anti-spoofing techniques can be found in the literature, the combination of multiple models, or ensemble models, still proves to be one of the best approaches. However, current iterations often rely on fixed weight assignments, potentially neglecting the unique strengths of each individual model. In response, we propose a novel ensembling model, an adaptive neural network-based approach that dynamically adjusts weights based on input utterances. Our experimental findings show that this approach outperforms traditional weighted score averaging techniques, showcasing its ability to adapt to diverse audio characteristics effectively.
Eros Roselló, Ángel M. Gómez, Iván López-Espejo, Antonio M. Peinado, Juan M. Martín-Doñas
INTERSPEECH3
2024 No-Reference Speech Intelligibility Prediction Leveraging a Noisy-Speech ASR Pre-Trained Model
abstract
Recent advances in deep learning have improved the capabilities of data-driven speech intelligibility prediction (SIP) algorithms. Nevertheless, the scarcity of speech intelligibility datasets limits the development of data-driven algorithms. This study introduces a set of no-reference SIP algorithms leveraging a pre-trained wav2vec 2.0 backbone. We adapt wav2vec 2.0 for automatic speech recognition under additive noise conditions with a parameter-efficient methodology, low-rank adaptation. We demonstrate no-reference SIP algorithms designed with this approach using a moderate amount of training data. The best designs perform on par or even better than a state-of-the-art reference-based SIP algorithm across a variety of datasets comprising different degradation types.
Haolan Wang, Amin Edraki, Wai-Yip Chan, Iván López-Espejo, Jesper Jensen 0001
INTERSPEECH4
2023 Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting
abstract
In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels is severely decreased. Reducing the number of channels might yield certain KWS performance drop, but also a substantial energy consumption reduction, which is key when deploying common always-on KWS on low-resource devices. Experimental results on a noisy version of the Google Speech Commands Dataset show that filterbank learning adapts to noise characteristics to provide a higher degree of robustness to noise, especially when dropout is integrated. Thus, switching from typically used 40-channel log-Mel features to 8channel learned features leads to a relative KWS accuracy loss of only 3.5% while simultaneously achieving a 6.3× energy consumption reduction.
Iván López-Espejo, Ram C. M. C. Shekar, Zheng-Hua Tan, Jesper Jensen 0001, John H. L. Hansen
ICASSP1
2023 On the deficiency of intelligibility metrics as proxies for subjective intelligibility
abstract
A recent trend in deep neural network (DNN)-based speech enhancement consists of using intelligibility and quality metrics as loss functions for model training with the aim of achieving high subjective speech intelligibility and perceptual quality in real-life conditions. In this study, we analyze a variety of loss functions, including some based on state-of-the-art intelligibility and quality metrics, to train an end-to-end speech enhancement system based on a fully convolutional neural network. The loss functions include perceptual metric for speech quality evaluation (PMSQE), scale-invariant signal-to-distortion ratio (SI-SDR), SI-SDR integrating speech pre-emphasis, short-time objective intelligibility (STOI), extended STOI (ESTOI), spectro-temporal glimpsing index (STGI), and a composite loss function combining STGI and SI-SDR. While DNNs trained with these loss functions produce notable speech intelligibility (and quality) gains according to pertinent objective metrics, we conduct a subjective intelligibility test that contradicts this result, showing no intelligibility improvement. From the results of this study, our conclusion is twofold: (1) subjective intelligibility evaluation is currently not replaceable by objective intelligibility evaluation, and (2) both the development of meaningful intelligibility metrics and DNN-based speech enhancement systems that can consistently improve the intelligibility of noisy speech for human listening remain open problems.
Iván López-Espejo, Amin Edraki, Wai-Yip Chan, Zheng-Hua Tan, Jesper Jensen 0001
Speech Commun.1
2021 A Novel Loss Function and Training Strategy for Noise-Robust Keyword Spotting
abstract
The development of keyword spotting (KWS) systems that are accurate in noisy conditions remains a challenge. Towards this goal, in this paper we propose a novel training strategy relying on multi-condition training for noise-robust KWS. By this strategy, we think of the state-of-the-art KWS models as the composition of a keyword embedding extractor and a linear classifier that are successively trained. To train the keyword embedding extractor, we also propose a new (CN,2+1)-pair loss function extending the concept behind related loss functions like triplet and N-pair losses to reach larger inter-class and smaller intra-class variation. Experimental results on a noisy version of the Google Speech Commands Dataset show that our proposal achieves around 12% KWS accuracy relative improvement with respect to standard end-to-end multi-condition training when speech is distorted by unseen noises. This performance improvement is achieved without increasing the computational complexity of the KWS model.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions
abstract
The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks. In this paper, we present a study on different methods for linear compensation of embeddings making use of Gaussian mixture models to cluster shouted and normal speech domains. These compensation techniques are borrowed from the area of robustness for automatic speech recognition and, in this work, we apply them to compensate the mismatch between shouted and normal conditions in speaker verification. Before compensation, shouted condition is automatically detected by means of logistic regression. The process is computationally light and it is performed in the back-end of an x-vector system. Experimental results show that applying the proposed approach in the presence of vocal effort mismatch yields up to 13.8% equal error rate relative improvement with respect to a system that applies neither shouted speech detection nor compensation.
Santi Prieto, Alfonso Ortega Giménez, Iván López-Espejo, Eduardo Lleida
INTERSPEECH3
2020 Improved External Speaker-Robust Keyword Spotting for Hearing Assistive Devices
abstract
For certain applications, keyword spotting (KWS) requires some degree of personalization. This is the case for KWS for hearing assistive devices, e.g., hearing aids, where only the device user should be allowed to trigger the KWS system. In this paper, we first develop a new realistic hearing aid experimental framework. Next, using this framework we show that the performance of a state-of-the-art multi-task deep learning architecture exploiting cepstral features for joint KWS and users' own-voice/external speaker detection drops significantly. To overcome this problem, we use phase difference information through GCC-PHAT (Generalized Cross-Correlation with PHAse Transform)-based coefficients along with log-spectral magnitude features. In addition, we demonstrate that working in the perceptually-motivated constant-Q transform (CQT) domain instead of in the short-time Fourier transform (STFT) domain allows for the generation of compact and coherent features which provide superior KWS performance. Our experimental results show that our CQT-based proposal achieves a relative KWS accuracy improvement of around 18% compared to using cepstral features while dramatically decreasing the number of multiplications in the multi-task architecture, which is key in the context of low-resource devices like hearing assistive devices.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Keyword Spotting for Hearing Assistive Devices Robust to External Speakers
abstract
Keyword spotting (KWS) is experiencing an upswing due to the pervasiveness of small electronic devices that allow interaction with them via speech. Often, KWS systems are speaker-independent, which means that any person --user or not-- might trigger them. For applications like KWS for hearing assistive devices this is unacceptable, as only the user must be allowed to handle them. In this paper we propose KWS for hearing assistive devices that is robust to external speakers. A state-of-the-art deep residual network for small-footprint KWS is regarded as a basis to build upon. By following a multi-task learning scheme, this system is extended to jointly perform KWS and users' own-voice/external speaker detection with a negligible increase in the number of parameters. For experiments, we generate from the Google Speech Commands Dataset a speech corpus emulating hearing aids as a capturing device. Our results show that this multi-task deep residual network is able to achieve a KWS accuracy relative improvement of around 32% with respect to a system that does not deal with external speakers.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH1
2017 Dual-channel DNN-based speech enhancement for smartphones
abstract
Speech communications in real-world scenarios need high performance enhancement algorithms to address the distortions that can degrade the intelligibility and quality of the speech signal. Current portable devices usually integrate multiple microphones that can conveniently be exploited to improve the signal quality. In this paper we present a dual-microphone speech enhancement approach suitable for smartphones with primary (front) and reference (back) microphones. Our proposal is based on the use of deep neural networks which are able to obtain a non-linear mapping function between noisy and clean speech signals. We explore two different architectures: a feedforward deep neural network (DNN) with temporal context and a gated recurrent unit (GRU) recurrent neural network (RNN). The proposed system is evaluated under different acoustic conditions in close- and far-talk device positions. A comparison with other single- and dual-channel approaches shows that our proposal obtains the best performance in terms of perceptual quality.
Juan M. Martín-Doñas, Ángel M. Gómez, Iván López-Espejo, Antonio M. Peinado
MMSP3
2017 Dual-channel VTS feature compensation for noise-robust speech recognition on mobile devices
abstract
One way to improve automatic speech recognition (ASR) performance on the latest mobile devices, which can be employed on a variety of noisy environments, consists of taking advantage of the small microphone arrays embedded in them. Since the performance of the classic beamforming techniques with small microphone arrays is rather limited, specific techniques are being developed to efficiently exploit this novel feature for noise‐robust ASR purposes. In this study, a novel dual‐channel minimum mean square error‐based feature compensation method relying on a vector Taylor series (VTS) expansion of a dual‐channel speech distortion model is proposed. In contrast to the single‐channel VTS approach (which can be considered as the state‐of‐the‐art for feature compensation), the authors’ technique particularly benefits from the spatial properties of speech and noise. Their proposal is assessed on a dual‐microphone smartphone (a particular case of interest) by means of the AURORA2‐2C synthetic corpus. Word recognition results, also validated with real noisy speech data, demonstrate the higher accuracy of their method by clearly outperforming minimum variance distortionless response beamforming and a single‐channel VTS feature compensation approach, especially at low signal‐to‐noise ratios.
Iván López-Espejo, Antonio M. Peinado, Ángel M. Gómez, José A. González 0001
IET Signal Process.1