EDBT 2026 Demo / reviewers in the wild / expert
Ignacio López-Moreno
dblp:333/2058
· DBLP profile ↗
32ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalizing Keyword Spotting with Speaker InformationabstractKeyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows models to learn from different data inputs and features. We explore both Text-Dependent and Text-Independent speaker recognition systems to extract speaker information, and we experiment on extracting this information from both the input audio and pre-enrolled user audio. Evaluating our systems on a diverse dataset, our primary approach yields a notable 2.6% relative improvement on Equal Error Rate overall, particularly improving performance by 5.9% for children under 12 years old and up to 24% for underrepresented speaker groups. Moreover, our proposed approach only requires a small 1% increase in the number of parameters, with a minimum impact on latency and computational cost, which makes it a practical solution for real-world applications. Beltran Labrador, Pai Zhu, Guanlong Zhao, Angelo Scorza Scarpati, Alicia Lozano-Diez, Ignacio López-Moreno |
ICASSP | 7 |
| 2024 | FedAQT: Accurate Quantized Training with Federated LearningabstractFederated learning (FL) has been widely used to train neural networks with the decentralized training procedure where data is only accessed on clients’ devices for privacy preservation. However, the limited computation resources on clients’ devices prevent FL of large models. To overcome the constraint, one possible method is to reduce the computation memory usage with quantized neural networks such as quantization aware training on a centralized server. However, directly applying the quantization aware methods does not reduce the memory consumption on the clients’ devices of FL because the full-precision model is still used in the forward propagation of the model computation. To enable FL of the Conformer based ASR models, we propose FedAQT, an accurate quantized training framework under FL by training with quantized variables directly on clients’ devices. We empirically show that our method can achieve comparable WER with only 60% memory of the full-precision model. Renkun Ni, Yonghui Xiao, Phoenix Meadowlark, Oleg Rybakov, Tom Goldstein, Ananda Theertha Suresh, Ignacio López-Moreno, Mingqing Chen, Rajiv Mathews |
ICASSP | 7 |
| 2024 | Version control of speaker recognition systems
Ignacio López-Moreno |
J. Syst. Softw. | 2 |
| 2023 | Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword SpottingabstractIn this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the keyword in the text transcription with a special tokenand training the system to detect thetoken in an audio stream. At inference time, we create a decision function inspired by conventional KWS approaches, to make our approach more suitable for the KWS task. Furthermore, we introduce a specific keyword spotting loss by adapting the sequence-discriminative Minimum Bayes-Risk training technique. We find that our approach significantly outperforms ASR based KWS systems. When compared with a conventional keyword spotting system, our proposal has similar performance while bringing the advantages and flexibility of sequence-to-sequence training. Additionally, when combined with the conventional KWS system, our approach can improve the performance at any operation point. Beltran Labrador, Guanlong Zhao, Ignacio López-Moreno, Angelo Scorza Scarpati, Liam Fowl |
ICASSP | 3 |
| 2023 | Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training LossabstractIn this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to the sparsity of the speaker changes in the training data, the conventional T-T based SCD model loss leads to sub-optimal detection accuracy. To mitigate this issue, we use a customized edit-distance algorithm to estimate the token-level SCD false accept (FA) and false reject (FR) rates during training and optimize model parameters to minimize a weighted combination of the FA and FR, focusing the model on accurately predicting speaker changes. We also propose a set of evaluation metrics that align better with commercial use cases. Experiments on a group of challenging real-world datasets show that the proposed training method can significantly improve the overall performance of the SCD model with the same number of parameters. Guanlong Zhao, Han Lu 0003, Ignacio López-Moreno |
ICASSP | 5 |
| 2023 | Locale Encoding for Scalable Multilingual Keyword Spotting ModelsabstractA Multilingual Keyword Spotting (KWS) system detects spoken keywords over multiple locales. Conventional monolingual KWS approaches do not scale well to multilingual scenarios because of high development/maintenance costs and lack of resource sharing. To overcome this limit, we propose two locale-conditioned universal models with locale feature concatenation and feature-wise linear modulation (FiLM). We compare these models with two baseline methods: locale-specific monolingual KWS, and a single universal model trained over all data. Experiments over 10 localized language datasets show that locale-conditioned models substantially improve accuracy over baseline methods across all locales in different noise conditions. FiLM performed the best, improving on average FRR by 61% (relative) compared to monolingual KWS models of similar sizes. Pai Zhu, Hyun Jin Park, Alex Park 0001, Angelo Scorza Scarpati, Ignacio López-Moreno |
ICASSP | 5 |
| 2022 | Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn DetectionabstractIn this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, represent each speaker turn by a speaker embedding, then cluster these embeddings with constraints from the detected speaker turns. Compared with conventional clustering-based diarization systems, our system largely reduces the computational cost of clustering due to the sparsity of speaker turns. Unlike other supervised speaker diarization systems which require annotations of time-stamped speaker labels for training, our system only requires including speaker turn tokens during the transcribing process, which largely reduces the human efforts involved in data collection. Han Lu 0003, Anshuman Tripathi, Ignacio López-Moreno, Hasim Sak |
ICASSP | 6 |
| 2022 | Production federated keyword spotting via distillation, filtering, and joint federated-centralized training
Andrew Hard, Kurt Partridge, Neng Chen, Sean Augenstein, Aishanee Shah, Hyun Jin Park, Alex Park 0001, Sara Ng, Jessica Nguyen, Ignacio López-Moreno, Rajiv Mathews, Françoise Beaufays |
INTERSPEECH | 10 |
| 2021 | SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification SystemabstractIn this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of different languages together for multilingual generalization and reducing development cycles; (2) A novel triage mechanism between text-dependent and text-independent models to reduce runtime cost and expected latency. To the best of our knowledge, this is the first study of speaker verification systems at the scale of 46 languages. The problem is framed from the perspective of using a smart speaker device with interactions consisting of a wake-up keyword (text-dependent) followed by a speech query (text-independent). Experimental evidence suggests that training on multiple languages can generalize to unseen varieties while maintaining performance on seen varieties. We also found that it can reduce computational requirements for training models by an order of magnitude. Furthermore, during model inference on English data, we observe that leveraging a triage framework can reduce the number of calls to the more computationally expensive text-independent system by 73% (and reduce latency by 59%) while maintaining an EER no worse than the text-independent setup. Roza Chojnacka, Jason Pelecanos, Ignacio López-Moreno |
Interspeech | 4 |
| 2021 | Noisy Student-Teacher Training for Robust Keyword SpottingabstractWe propose self-training with noisy student-teacher approach for streaming keyword spotting, that can utilize large-scale unlabeled data and aggressive data augmentation. The proposed method applies aggressive data augmentation (spectral augmentation) on the input of both student and teacher and utilize unlabeled data at scale, which significantly boosts the accuracy of student against challenging conditions. Such aggressive augmentation usually degrades model performance when used with supervised training with hard-labeled data. Experiments show that aggressive spec augmentation on baseline supervised training method degrades accuracy, while the proposed self-training with noisy student-teacher training improves accuracy of some difficult-conditioned test sets by as much as 60%. Hyun-Jin Park, Pai Zhu, Ignacio López-Moreno, Niranjan Subrahmanya |
Interspeech | 3 |
| 2021 | Dr-Vectors: Decision Residual Networks and an Improved Loss for Speaker RecognitionabstractMany neural network speaker recognition systems model each speaker using a fixed-dimensional embedding vector. These embeddings are generally compared using either linear or 2nd-order scoring and, until recently, do not handle utterance-specific uncertainty. In this work we propose scoring these representations in a way that can capture uncertainty, enroll/test asymmetry and additional non-linear information. This is achieved by incorporating a 2nd-stage neural network (known as a decision network) as part of an end-to-end training regimen. In particular, we propose the concept of decision residual networks which involves the use of a compact decision network to leverage cosine scores and to model the residual signal that's needed. Additionally, we present a modification to the generalized end-to-end softmax loss function to target the separation of same/different speaker scores. We observed significant performance gains for the two techniques. Jason Pelecanos, Ignacio López-Moreno |
Interspeech | 3 |
| 2020 | Training Keyword Spotting Models on Non-IID Data with Federated LearningabstractWe demonstrate that a production-quality keyword-spotting model can be trained on-device using federated learning and achieve comparable false accept and false reject rates to a centrally-trained model. To overcome the algorithmic constraints associated with fitting on-device data (which are inherently non-independent and identically distributed), we conduct thorough empirical studies of optimization algorithms and hyperparameter configurations using large-scale federated simulations. To overcome resource constraints, we replace memory intensive MTR data augmentation with SpecAugment, which reduces the false reject rate by 56%. Finally, to label examples (given the zero visibility into on-device data), we explore teacher-student training. Andrew Hard, Kurt Partridge, Cameron Nguyen, Niranjan Subrahmanya, Aishanee Shah, Pai Zhu, Ignacio López-Moreno, Rajiv Mathews |
INTERSPEECH | 7 |
| 2020 | VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech RecognitionabstractWe introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system.Delivering such a model presents numerous challenges: It should improve the performance when the input signal consists of overlapped speech, and must not hurt the speech recognition performance under all other acoustic conditions.Besides, this model must be tiny, fast, and perform inference in a streaming fashion, in order to have minimal impact on CPU, memory, battery and latency.We propose novel techniques to meet these multi-faceted requirements, including using a new asymmetric loss, and adopting adaptive runtime suppression strength.We also show that such a model can be quantized as a 8-bit integer model and run in realtime. Ignacio López-Moreno, Mert Saglam, Kevin W. Wilson, Alan Chiao, Yanzhang He, Wei Li 0133, Jason Pelecanos, Marily Nika, Alexander Gruenstein |
INTERSPEECH | 2 |
| 2019 | Tuplemax Loss for Language IdentificationabstractIn many scenarios of a language identification task, the user will specify a small set of languages which he/she can speak instead of a large set of all possible languages. We want to model such prior knowledge into the way we train our neural networks, by replacing the commonly used softmax loss function with a novel loss function named tuplemax loss. As a matter of fact, a typical language identification system launched in North America has about 95% users who could speak no more than two languages. Using the tuplemax loss, our system achieved a 2.33% error rate, which is a relative 39.4% improvement over the 3.85% error rate of standard softmax loss method. Li Wan 0004, Prashant Sridhar, Ignacio López-Moreno |
ICASSP | 5 |
| 2019 | Improving Keyword Spotting and Language Identification via Neural Architecture Search at Scale
Hanna Mazzawi, Xavi Gonzalvo, Aleks Kracun, Prashant Sridhar, Niranjan Subrahmanya, Ignacio López-Moreno, Hyun-Jin Park, Patrick Violette |
INTERSPEECH | 6 |
| 2019 | VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram MaskingabstractIn this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings;(2) A spectrogram masking network that takes both noisy spectrogram and speaker embedding as input, and produces a mask.Our system significantly reduces the speech recognition WER on multi-speaker signals, with minimal WER degradation on single-speaker signals. Hannah Muckenhirn, Kevin W. Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, Ignacio López-Moreno |
INTERSPEECH | 10 |
| 2018 | Attention-Based Models for Text-Dependent Speaker VerificationabstractAttention-based models have recently shown great performance on a range of tasks, such as speech recognition, machine translation, and image captioning due to their ability to summarize relevant information that expands through the entire length of an input sequence. In this paper, we analyze the usage of attention mechanisms to the problem of sequence summarization in our end-to-end text-dependent speaker recognition system. We explore different topologies and their variants of the attention layer, and compare different pooling methods on the attention weights. Ultimately, we show that attention-based models can improves the Equal Error Rate (EER) of our speaker verification system by relatively 14% compared to our non-attention LSTM baseline model. F. A. Rezaur Rahman Chowdhury, Ignacio López-Moreno, Li Wan 0004 |
ICASSP | 3 |
| 2018 | Generalized End-to-End Loss for Speaker VerificationabstractIn this paper, we propose a new loss function called generalized end-to-end (GE2E) loss, which makes the training of speaker verification models more efficient than our previous tuple-based end-to-end (TE2E) loss function. Unlike TE2E, the GE2E loss function updates the network in a way that emphasizes examples that are difficult to verify at each step of the training process. Additionally, the GE2E loss does not require an initial stage of example selection. With these properties, our model with the new loss function decreases speaker verification EER by more than 10%, while reducing the training time by 60% at the same time. We also introduce the MultiReader technique, which allows us to do domain adaptation - training a more accurate model that supports multiple keywords (i.e., “OK Google” and “Hey Google”) as well as multiple dialects. Li Wan 0004, Alan Papir, Ignacio López-Moreno |
ICASSP | 4 |
| 2018 | Speaker Diarization with LSTMabstractFor many years, i-vector based audio embedding techniques were the dominant approach for speaker verification and speaker diarization applications. However, mirroring the rise of deep learning in various domains, neural network based audio embeddings, also known asd-vectors, have consistently demonstrated superior speaker verification performance. In this paper, we build on the success of d-vector based speaker verification systems to develop a new d-vector based approach to speaker diarization. Specifically, we combine LSTM-based d-vector audio embeddings with recent work in non-parametric clustering to obtain a state-of-the-art speaker diarization system. Our system is evaluated on three standard public datasets, suggesting that d-vector based diarization systems offer significant advantages over traditional i-vector based systems. We achieved a 12.0% diarization error rate on NIST SRE 2000 CALLHOME, while our model is trained with out-of-domain data from voice search logs. Carlton Downey, Li Wan 0004, Philip Andrew Mansfield, Ignacio López-Moreno |
ICASSP | 5 |
| 2018 | Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech SynthesisabstractWe describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently trained components: (1) a speaker encoder network, trained on a speaker verification task using an independent dataset of noisy speech from thousands of speakers without transcripts, to generate a fixed-dimensional embedding vector from seconds of reference speech from a target speaker; (2) a sequence-to-sequence synthesis network based on Tacotron 2, which generates a mel spectrogram from text, conditioned on the speaker embedding; (3) an auto-regressive WaveNet-based vocoder that converts the mel spectrogram into a sequence of time domain waveform samples. We demonstrate that the proposed model is able to transfer the knowledge of speaker variability learned by the discriminatively-trained speaker encoder to the new task, and is able to synthesize natural speech from speakers that were not seen during training. We quantify the importance of training the speaker encoder on a large and diverse speaker set in order to obtain the best generalization performance. Finally, we show that randomly sampled speaker embeddings can be used to synthesize speech in the voice of novel speakers dissimilar from those used in training, indicating that the model has learned a high quality speaker representation. Ye Jia, Yu Zhang 0033, Ron J. Weiss, Jonathan Shen, Patrick Nguyen, Ruoming Pang, Ignacio López-Moreno |
NeurIPS | 10 |
| 2016 | On the use of deep feedforward neural networks for automatic language identificationabstractIn this work, we present a comprehensive study on the use of deep neural networks (DNNs) for automatic language identification (LID). Motivated by the recent success of using DNNs in acoustic modeling for speech recognition, we adapt DNNs to the problem of identifying the language in a given utterance from its short-term acoustic features. We propose two different DNN-based approaches. In the first one, the DNN acts as an end-to-end LID classifier, receiving as input the speech features and providing as output the estimated probabilities of the target languages. In the second approach, the DNN is used to extract bottleneck features that are then used as inputs for a state-of-the-art i-vector system. Experiments are conducted in two different scenarios: the complete NIST Language Recognition Evaluation dataset 2009 (LRE'09) and a subset of the Voice of America (VOA) data from LRE'09, in which all languages have the same amount of training data. Results for both datasets demonstrate that the DNN-based systems significantly outperform a state-of-art i-vector system when dealing with short-duration utterances. Furthermore, the combination of the DNN-based and the classical i-vector system leads to additional performance improvements (up to 45% of relative improvement in both EER and Cavg on 3s and 10s conditions, respectively). Ignacio López-Moreno, Javier Gonzalez-Dominguez, David Martinez, Oldrich Plchot, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
Comput. Speech Lang. | 1 |
| 2015 | Locally-connected and convolutional neural networks for small footprint speaker recognition
Yu-hsin Chen, Ignacio López-Moreno, Tara N. Sainath, Mirkó Visontai, Raziel Alvarez, Carolina Parada |
INTERSPEECH | 2 |
| 2015 | Frame-by-frame language identification in short utterances using deep neural networks
Javier Gonzalez-Dominguez, Ignacio López-Moreno, Pedro J. Moreno 0001, Joaquín González-Rodríguez |
Neural Networks | 2 |
| 2014 | Automatic language identification using deep neural networksabstractThis work studies the use of deep neural networks (DNNs) to address automatic language identification (LID). Motivated by their recent success in acoustic modelling, we adapt DNNs to the problem of identifying the language of a given spoken utterance from short-term acoustic features. The proposed approach is compared to state-of-the-art i-vector based acoustic systems on two different datasets: Google 5M LID corpus and NIST LRE 2009. Results show how LID can largely benefit from using DNNs, especially when a large amount of training data is available. We found relative improvements up to 70%, in Cavg, over the baseline system. Ignacio López-Moreno, Javier Gonzalez-Dominguez, Oldrich Plchot, David Martinez, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
ICASSP | 1 |
| 2014 | Large-scale speaker identificationabstractSpeaker identification is one of the main tasks in speech processing. In addition to identification accuracy, large-scale applications of speaker identification give rise to another challenge: fast search in the database of speakers. In this paper, we propose a system based on i-vectors, a current approach for speaker identification, and locality sensitive hashing, an algorithm for fast nearest neighbor search in high dimensions. The connection between the two techniques is the cosine distance: on the one hand, we use the cosine distance to compare i-vectors, on the other hand, locality sensitive hashing allows us to quickly approximate the cosine distance in our retrieval procedure. We evaluate our approach on a realistic data set from YouTube with about 1,000 speakers. The results show that our algorithm is approximately one to two orders of magnitude faster than a linear search while maintaining the identification accuracy of an i-vector-based system. Ludwig Schmidt, Matthew Sharifi, Ignacio López-Moreno |
ICASSP | 3 |
| 2014 | Improving DNN speaker independence with I-vector inputsabstractWe propose providing additional utterance-level features as inputs to a deep neural network (DNN) to facilitate speaker, channel and background normalization. Modifications of the basic algorithm are developed which result in significant reductions in word error rates (WERs). The algorithms are shown to combine well with speaker adaptation by backpropagation, resulting in a 9% relative WER reduction. We address implementation of the algorithm for a streaming task. Andrew W. Senior, Ignacio López-Moreno |
ICASSP | 2 |
| 2014 | Deep neural networks for small footprint text-dependent speaker verificationabstractIn this paper we investigate the use of deep neural networks (DNNs) for a small footprint text-dependent speaker verification task. At development stage, a DNN is trained to classify speakers at the framelevel. During speaker enrollment, the trained DNN is used to extract speaker specific features from the last hidden layer. The average of these speaker features, or d-vector, is taken as the speaker model. At evaluation stage, a d-vector is extracted for each utterance and compared to the enrolled speaker model to make a verification decision. Experimental results show the DNN based speaker verification system achieves good performance compared to a popular i-vector system on a small footprint text-dependent speaker verification task. In addition, the DNN based system is more robust to additive noise and outperforms the i-vector system at low False Rejection operating points. Finally the combined system outperforms the i-vector system by 14% and 25% relative in equal error rate (EER) for clean and noisy conditions respectively. Ehsan Variani, Erik McDermott, Ignacio López-Moreno, Javier Gonzalez-Dominguez |
ICASSP | 4 |
| 2014 | Automatic language identification using long short-term memory recurrent neural networksabstractThis work explores the use of Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) for automatic lan-guage identification (LID). The use of RNNs is motivated by their better ability in modeling sequences with respect to feed forward networks used in previous works. We show that LSTM RNNs can effectively exploit temporal dependencies in acoustic data, learning relevant features for language discrimination pur-poses. The proposed approach is compared to baseline i-vector and feed forward Deep Neural Network (DNN) systems in the NIST Language Recognition Evaluation 2009 dataset. We show LSTM RNNs achieve better performance than our best DNN system with an order of magnitude fewer parameters. Further, the combination of the different systems leads to significant per-formance improvements (up to 28%). 1. Javier Gonzalez-Dominguez, Ignacio López-Moreno, Hasim Sak, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2011 | Von Mises-Fisher Models in the Total Variability Subspace for Language RecognitionabstractThis letter proposes a new modeling approach for the Total Variability subspace within a Language Recognition task. Motivated by previous works in directional statistics, von Mises-Fisher distributions are used for assigning language-conditioned probabilities to language data, assumed to be spherically distributed in this subspace. The two proposed methods use Kernel Density Functions or Finite Mixture Models of such distributions. Experiments conducted on NIST LRE 2009 show that the proposed techniques significantly outperform the baseline cosine distance approach in most of the considered experimental conditions, including different speech conditions, durations and the presence of unseen languages. Ignacio López-Moreno, Daniel Ramos-Castro, Javier Gonzalez-Dominguez, Joaquín González-Rodríguez |
IEEE Signal Process. Lett. | 1 |
| 2009 | Speaker dependent emotion recognition using prosodic supervectorsabstractProceedings of Interspeech 2009, Brighton (United Kingdom) Ignacio López-Moreno, Carlos Ortego-Resa, Joaquín González-Rodríguez, Daniel Ramos-Castro |
INTERSPEECH | 1 |
| 2008 | Anchor-model fusion for language recognitionabstractProceedings of Interspeech 2008, Brisbane (Australia) Ignacio López-Moreno, Daniel Ramos-Castro, Joaquín González-Rodríguez, Doroteo T. Toledano |
INTERSPEECH | 1 |
| 2007 | Support vector regression for speaker verificationabstractProceedings of Interspeech 2007, Antwerp (Belgium) Ignacio López-Moreno, Ismael Mateos-Garcia, Daniel Ramos-Castro, Joaquín González-Rodríguez |
INTERSPEECH | 1 |