EDBT 2026 Demo / reviewers in the wild / expert
Daniel Garcia-Romero
dblp:59/2493
· DBLP profile ↗
60ranked-venue papers
17as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 16 first-author · 7 since 2021Artificial intelligence and machine learning · 35 · 8 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyper-adapter for Parameter-Efficient Multilingual ASR AdaptationabstractThis work proposes a new parameter-efficient adaptation approach for multilingual ASR based on the hyper-network. Existing multilingual ASR adaptation methods apply either one residual adapter for all the languages, or language dependent adapters for each individual language. The residual adapter cannot compete the full finetuning in terms of WER, because it is agnostic to language information. Whereas the language dependent adapters introduce high parameter overhead without a parameter sharing strategy. In contrast, we leverage a hyper-network to generate the weights for the adapters across different languages. To achieve the best parameter sharing strategy that scales with a large number of languages, we propose multi-level conditioning vector fusion, orthogonal regularization to improve the hyper-network output diversity, and language loss weighting during the model training. The proposed approach demonstrates comparable or better WER and better parameter efficiency compared to previous multilingual ASR adaptation approaches on commonly used multilingual ASR benchmarks. Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han |
ICASSP | 2 |
| 2025 | Zero-resource Speech Translation and Recognition with LLMsabstractDespite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language. Karel Mundnich, Xing Niu 0001, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff |
ICASSP | 11 |
| 2025 | Knowledge Distillation From Ensemble for Spoken Language IdentificationabstractSpoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework to address this challenge. By distilling an ensemble of large LID models into a single, more efficient student, we achieve comparable or even superior performance while reducing computational cost by 67%. Our approach yields a student model with less than 10% the size of a 200M+ parameter teacher ensemble, yet outperforming a 140M parameter teacher by 13% relative. Additionally, combining our distillation technique with decoupled knowledge distillation leads to substantial gains (50% relative), especially for confusable and low-resource languages in the FLEURS dataset. Raghuveer Peri, Seyed Omid Sadjadi, Daniel Garcia-Romero, Srikanth Vishnubhotla, Kyu J. Han |
ICASSP | 3 |
| 2025 | Contextual ASR with Retrieval Augmented Large Language ModelabstractAutomatic speech recognition (ASR) systems can benefit from incorporating contextual information to improve recognition accuracy, especially for uncommon words or phrases. Current approaches like custom vocabularies or prompting with previous transcript segments provide limited contextual control. Compared to existing context biasing methods, RAG promises more flexible and scalable contextual control by leveraging LLMs’ broad knowledge. To this end, we propose leveraging large language models (LLMs) and retrieval-augmented generation (RAG) to enhance the contextual capabilities of ASR systems. Specifically, we propose systems based on text and audio LLMs to perform contextual error correction with context retrieved by querying a text-based retriever using the ASR module’s firstpass ASR hypotheses and a frequency-based custom vocabulary (CV) list. Our experiments reveal that the fine-tuned system has effectively learned to extract the relevant context to perform error correction while maintaining robustness against noise. Cihan Xiao, Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han |
ICASSP | 3 |
| 2024 | Revisiting Convolution-free Transformer for Speech Recognition
Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff |
INTERSPEECH | 4 |
| 2024 | The VoxCeleb Speaker Recognition Challenge: A RetrospectiveabstractThe VoxCeleb Speaker Recognition Challenges (VoxSRC) were a series of challenges and workshops that ran annually from 2019 to 2023. The challenges primarily evaluated the tasks of speaker recognition and diarisation under various settings including: closed and open training data; as well as supervised, self-supervised, and semi-supervised training for domain adaptation. The challenges also provided publicly available training and evaluation datasets for each task and setting, with new test sets released each year. In this paper, we provide a review of these challenges that covers: what they explored; the methods developed by the challenge participants and how these evolved; and also the current state of the field for speaker verification and diarisation. We chart the progress in performance over the five installments of the challenge on a common evaluation dataset and provide a detailed analysis of how each year's special focus affected participants' performance. This paper is aimed both at researchers who want an overview of the speaker recognition and diarisation field, and also at challenge organisers who want to benefit from the successes and avoid the mistakes of the VoxSRC challenges. We end with a discussion of the current strengths of the field and open challenges. Jaesung Huh, Joon Son Chung, Arsha Nagrani, Andrew Brown 0006, Jee-Weon Jung, Daniel Garcia-Romero, Andrew Zisserman |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | Directed speech separation for automatic speech recognition of long form conversational speechabstractMany of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap.Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads to difficulty in accurately stitching homogenous speaker chunks for downstream tasks like Automatic Speech Recognition (ASR).Also, most of these models are trained with synthetic mixtures and do not generalize to real conversational data.In this paper, we propose a speaker conditioned separator trained on speaker embeddings extracted directly from the mixed signal using an over-clustering based approach.This model naturally regulates the order of the separated chunks without the need for an additional stitching step.We also introduce a data sampling strategy with real and synthetic mixtures which generalizes well to real conversation speech.With this model and data sampling technique, we show significant improvements in speaker-attributed word error rate (SA-WER) on Hub5 data. Rohit Paturi, Sundararajan Srinivasan, Katrin Kirchhoff, Daniel Garcia-Romero |
INTERSPEECH | 4 |
| 2021 | Recent Developments on Espnet Toolkit Boosted By ConformerabstractIn this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, such as automatic speech recognition (ASR), speech translations (ST), speech separation (SS) and text-to-speech (TTS). Our experiments reveal various training tips and significant performance benefits obtained with the Conformer on different tasks. These results are competitive or even outperform the current state-of-art Transformer models. We are preparing to release all-in-one recipes using open source and publicly available corpora for all the above tasks with pre-trained models. Our aim for this work is to contribute to our research community by reducing the burden of preparing state-of-the-art research environments usually requiring high resources. Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi 0003, Shinji Watanabe 0001, Wangyou Zhang, Yuekai Zhang |
ICASSP | 9 |
| 2020 | Jhu-HLTCOE System for the Voxsrc Speaker Recognition ChallengeabstractThe VoxSRC speaker recognition challenge comprises data obtained from YouTube videos of celebrity interviews in a wide range of recording environments. The challenge provides FIXED and OPEN training conditions to allow cross-system comparisons and to characterize the effects of additional amounts of training data on system performance. This paper describes our submission to this challenge where we have explored x-vector extractor topologies, classification head alternatives, data augmentation, and angular margin penalty. Our final entry to the FIXED condition (which achieved 2nd place) is the score average of 4 diverse systems. We find that this system outperforms a large single DNN with similar number of parameters. Daniel Garcia-Romero, Alan McCree, David Snyder, Gregory Sell |
ICASSP | 1 |
| 2020 | State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, L. Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, Najim Dehak |
Comput. Speech Lang. | 4 |
| 2019 | Speaker Recognition for Multi-speaker Conversations Using X-vectorsabstractRecently, deep neural networks that map utterances to fixed-dimensional embeddings have emerged as the state-of-the-art in speaker recognition. Our prior work introduced x-vectors, an embedding that is very effective for both speaker recognition and diarization. This paper combines our previous work and applies it to the problem of speaker recognition on multi-speaker conversations. We measure performance on Speakers in the Wild and report what we believe are the best published error rates on this dataset. Moreover, we find that diarization substantially reduces error rate when there are multiple speakers, while maintaining excellent performance on single-speaker recordings. Finally, we introduce an easily implemented method to remove the domain-sensitive threshold typically used in the clustering stage of a diarization system. The proposed method is more robust to domain shifts, and achieves similar results to those obtained using a well-tuned threshold. David Snyder, Daniel Garcia-Romero, Gregory Sell, Alan McCree, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 2 |
| 2019 | Script Identification using Across- and Within-Image Distribution EstimationabstractIn this paper, we apply several modifications to script identification, several of which inspired by techniques from the similar audio task of spoken language recognition. Specifically, we alter the architecture of a convolutional network with global average pooling to include variance pooling as well, we utilize score calibration of the output scores of the network, and we utilize prior distribution estimation to condition the calibrated scores. We show that these methods are effective in script identification, with the use of priors showing especially promising improvements. Furthermore, in the domain of script identification, several additional extensions of distribution estimation are available which consider the distribution within each image, and we demonstrate much larger improvements when employing these extensions. Finally, we also show that an embedding-plus-classifier approach performs similarly to the full network, and so its potential for increased flexibility may be beneficial for future consideration. With all modifications, overall accuracy on the ICDAR 2017 validation dataset increases from 89.7% to 93.6%. Gregory Sell, David Etter, Daniel Garcia-Romero, Alan McCree |
ICDAR | 3 |
| 2019 | x-Vector DNN Refinement with Full-Length Recordings for Speaker Recognition
Daniel Garcia-Romero, David Snyder, Gregory Sell, Alan McCree, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2019 | Speaker Recognition Benchmark Using the CHiME-5 Corpus
Daniel Garcia-Romero, David Snyder, Shinji Watanabe 0001, Gregory Sell, Alan McCree, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2019 | Speaker Diarization Using Leave-One-Out Gaussian PLDA Clustering of DNN Embeddings
Alan McCree, Gregory Sell, Daniel Garcia-Romero |
INTERSPEECH | 3 |
| 2019 | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 4 |
| 2018 | Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus ProgramabstractCurrently, datasets that support audio-visual recognition of people in videos are scarce and limited. In this paper, we introduce an expansion of video data from the IARPA Janus program to support this research area. We refer to the expanded set, which adds labels for voice to the already-existing face labels, as the Janus Multimedia dataset. We first describe the speaker labeling process, which involved a combination of automatic and manual criteria. We then discuss two evaluation settings for this data. In the core condition, the voice and face of the labeled individual are present in every video. In the full condition, no such guarantee is made. The power of audiovisual fusion is then shown using these publicly-available videos and labels, showing significant improvement over only recognizing voice or face alone. In addition to this work, several other possible paths for future research with this dataset are discussed. Gregory Sell, Kevin Duh, David Snyder, Dave Etter, Daniel Garcia-Romero |
ICASSP | 5 |
| 2018 | X-Vectors: Robust DNN Embeddings for Speaker RecognitionabstractIn this paper, we use data augmentation to improve performance of deep neural network (DNN) embeddings for speaker recognition. The DNN, which is trained to discriminate between speakers, maps variable-length utterances to fixed-dimensional embeddings that we call x-vectors. Prior studies have found that embeddings leverage large-scale training datasets better than i-vectors. However, it can be challenging to collect substantial quantities of labeled data for training. We use data augmentation, consisting of added noise and reverberation, as an inexpensive method to multiply the amount of training data and improve robustness. The x-vectors are compared with i-vector baselines on Speakers in the Wild and NIST SRE 2016 Cantonese. We find that while augmentation is beneficial in the PLDA classifier, it is not helpful in the i-vector extractor. However, the x-vector DNN effectively exploits data augmentation, due to its supervised training. As a result, the x-vectors achieve superior performance on the evaluation datasets. David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 2 |
| 2018 | Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba 0001, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe 0001, Sanjeev Khudanpur |
INTERSPEECH | 4 |
| 2018 | Fast Variational Bayes for Heavy-tailed PLDA Applied to i-vectors and x-vectorsabstractThe standard state-of-the-art backend for text-independent speaker recognizers that use i-vectors or x-vectors, is Gaussian PLDA (G-PLDA), assisted by a Gaussianization step involving length normalization. G-PLDA can be trained with both generative or discriminative methods. It has long been known that heavy-tailed PLDA (HT-PLDA), applied without length normalization, gives similar accuracy, but at considerable extra computational cost. We have recently introduced a fast scoring algorithm for a discriminatively trained HT-PLDA backend. This paper extends that work by introducing a fast, variational Bayes, generative training algorithm. We compare old and new backends, with and without length-normalization, with i-vectors and x-vectors, on SRE'10, SRE'16 and SITW. Anna Silnova, Niko Brümmer, Daniel Garcia-Romero, David Snyder, Lukás Burget |
INTERSPEECH | 3 |
| 2017 | Speaker diarization using deep neural network embeddingsabstractSpeaker diarization is an important front-end for many speech technologies in the presence of multiple speakers, but current methods that employ i-vector clustering for short segments of speech are potentially too cumbersome and costly for the front-end role. In this work, we propose an alternative approach for learning representations via deep neural networks to remove the i-vector extraction process from the pipeline entirely. The proposed architecture simultaneously learns a fixed-dimensional embedding for acoustic segments of variable length and a scoring function for measuring the likelihood that the segments originated from the same or different speakers. Through tests on the CALLHOME conversational telephone speech corpus, we demonstrate that, in addition to streamlining the diarization architecture, the proposed system matches or exceeds the performance of state-of-the-art baselines. We also show that, though this approach does not respond as well to unsupervised calibration strategies as previous systems, the incorporation of well-founded speaker priors sufficiently mitigates this shortcoming. Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, Alan McCree |
ICASSP | 1 |
| 2017 | Extended Variability Modeling and Unsupervised Adaptation for PLDA Speaker Recognition
Alan McCree, Gregory Sell, Daniel Garcia-Romero |
INTERSPEECH | 3 |
| 2017 | Deep Neural Network Embeddings for Text-Independent Speaker Verification
David Snyder, Daniel Garcia-Romero, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2016 | Stacked Long-Term TDNN for Spoken Language Recognition
Daniel Garcia-Romero, Alan McCree |
INTERSPEECH | 1 |
| 2016 | Priors for Speaker Counting and Diarization with AHC
Gregory Sell, Alan McCree, Daniel Garcia-Romero |
INTERSPEECH | 3 |
| 2016 | Deep neural network-based speaker embeddings for end-to-end speaker verificationabstractIn this study, we investigate an end-to-end text-independent speaker verification system. The architecture consists of a deep neural network that takes a variable length speech segment and maps it to a speaker embedding. The objective function separates same-speaker and different-speaker pairs, and is reused during verification. Similar systems have recently shown promise for text-dependent verification, but we believe that this is unexplored for the text-independent task. We show that given a large number of training speakers, the proposed system outperforms an i-vector baseline in equal error-rate (EER) and at low miss rates. Relative to the baseline, the end-to-end system reduces EER by 13% average and 29% pooled across test conditions. The fused system achieves a reduction of 32% average and 38% pooled. David Snyder, Pegah Ghahremani, Daniel Povey, Daniel Garcia-Romero, Yishay Carmiel, Sanjeev Khudanpur |
SLT | 4 |
| 2015 | Time delay deep neural network-based universal background models for speaker recognitionabstractRecently, deep neural networks (DNN) have been incorporated into i-vector-based speaker recognition systems, where they have significantly improved state-of-the-art performance. In these systems, a DNN is used to collect sufficient statistics for i-vector extraction. In this study, the DNN is a recently developed time delay deep neural network (TDNN) that has achieved promising results in LVCSR tasks. We believe that the TDNN-based system achieves the best reported results on SRE10 and it obtains a 50% relative improvement over our GMM baseline in terms of equal error rate (EER). For some applications, the computational cost of a DNN is high. Therefore, we also investigate a lightweight alternative in which a supervised GMM is derived from the TDNN posteriors. This method maintains the speed of the traditional unsupervised-GMM, but achieves a 20% relative improvement in EER. David Snyder, Daniel Garcia-Romero, Daniel Povey |
ASRU | 2 |
| 2015 | Topic Identification and Discovery on Text and SpeechabstractWe compare the multinomial i-vector framework from the speech community with LDA, SAGE, and LSA as feature learners for topic ID on multinomial speech and text data.We also compare the learned representations in their ability to discover topics, quantified by distributional similarity to gold-standard topics and by human interpretability.We find that topic ID and topic discovery are competing objectives.We argue that LSA and i-vectors should be more widely considered by the text processing community as pre-processing steps for downstream tasks, and also speculate about speech processing tasks that could benefit from more interpretable representations like SAGE. Chandler May, Francis Ferraro, Alan McCree, Jonathan Wintrode, Daniel Garcia-Romero, Benjamin Van Durme |
EMNLP | 5 |
| 2015 | Diarization resegmentation in the factor analysis subspaceabstractResegmentation is an important post-processing step to refine the rough boundaries of diarization systems that rely on segment clustering of an initial uniform segmentation. Past work has primarily used a Viterbi resegmentation with MFCC features for this purpose. In this paper, we examine an algorithm for resegmentation that operates instead in factor analysis subspace. By combining this system with a speaker clustering front-end, we yield a diarization error rate of 11.5% on the CALLHOME conversational telephone speech corpus. Gregory Sell, Daniel Garcia-Romero |
ICASSP | 2 |
| 2015 | Content-based recommender systems for spoken documentsabstractContent-based recommender systems use preference ratings and features that characterize media to model users' interests or information needs for making future recommendations. While previously developed in the music and text domains, we present an initial exploration of content-based recommendation for spoken documents using a corpus of public domain internet audio. Unlike familiar speech technologies of topic identification and spoken document retrieval, our recommendation task requires a more comprehensive notion of document relevance than bags-of-words would supply. Inspired by music recommender systems, we automatically extract a wide variety of content-based features to characterize non-linguistic aspects of the audio such as speaker, language, gender, and environment. To combine these heterogeneous information sources into a single relevance judgement, we evaluate feature, score, and hybrid fusion techniques. Our study provides an essential first exploration of the task and clearly demonstrates the value of a multisource approach over a bag-of-words baseline. Jonathan Wintrode, Gregory Sell, Aren Jansen, Michelle Fox, Daniel Garcia-Romero, Alan McCree |
ICASSP | 5 |
| 2015 | Analysis of the second phase of the 2013-2014 i-vector machine learning challenge
Désiré Bansé, George R. Doddington, Daniel Garcia-Romero, John J. Godfrey, Craig S. Greenberg, Jaime Hernandez-Cordero, John M. Howard, Alvin F. Martin, Lisa P. Mason, Alan McCree, Douglas A. Reynolds |
INTERSPEECH | 3 |
| 2015 | Insights into deep neural networks for speaker recognition
Daniel Garcia-Romero, Alan McCree |
INTERSPEECH | 1 |
| 2015 | DNN senone MAP multinomial i-vectors for phonotactic language recognition
Alan McCree, Daniel Garcia-Romero |
INTERSPEECH | 2 |
| 2015 | Speaker diarization with i-vectors from DNN senone posteriors
Gregory Sell, Daniel Garcia-Romero, Alan McCree |
INTERSPEECH | 2 |
| 2014 | Generative modelling for unsupervised score calibrationabstractScore calibration enables automatic speaker recognizers to make cost-effective accept / reject decisions. Traditional calibration requires supervised data, which is an expensive resource. We propose a 2-component GMM for unsupervised calibration and demonstrate good performance relative to a supervised baseline on NIST SRE'10 and SRE'12. A Bayesian analysis demonstrates that the uncertainty associated with the unsupervised calibration parameter estimates is surprisingly small. Niko Brümmer, Daniel Garcia-Romero |
ICASSP | 2 |
| 2014 | Supervised domain adaptation for I-vector based speaker recognitionabstractIn this paper, we present a comprehensive study on supervised domain adaptation of PLDA based i-vector speaker recognition systems. After describing the system parameters subject to adaptation, we study the impact of their adaptation on recognition performance. Using the recently designed domain adaptation challenge, we observe that the adaptation of the PLDA parameters (i.e. across-class and within-class co variances) produces the largest gains. Nonetheless, length-normalization is also important; whereas using an indomani UBM and T matrix is not crucial. For the PLDA adaptation, we compare four approaches. Three of them are proposed in this work, and a fourth one was previously published. Overall, the four techniques are successful at leveraging varying amounts of labeled in-domain data and their performance is quite similar. However, our approaches are less involved, and two of them are applicable to a larger class of models (low-rank across-class). Daniel Garcia-Romero, Alan McCree |
ICASSP | 1 |
| 2014 | Unsupervised idiolect discovery for speaker recognitionabstractShort-time spectral characterizations of the human voice have proven to be the most dependable features available to modern speaker recognition systems. However, it is well-known that highlevel linguistic information such as word usage and pronunciation patterns can provide complementary discriminative power. In an automatic setting, the availability of these idiolectal cues is dependent on access to a word or phonetic tokenizer, ideally in the given language and domain. In this paper, we propose a novel approach to speaker recognition that leverages recently developed zero-resource term discovery algorithms to identify speaker-characteristic lexical and phrasal acoustic patterns without the need for any supervised speech recognition tools. We use the enrollment audio itself to score each trial and perform no model training (supervised or unsupervised) at any stage of the processing, allowing immediate application to any language or domain. We evaluate our approach on the extended 8-conversation core condition of the 2010 NIST SRE and demonstrate a 16% relative (0.06 absolute) reduction in minDCF when combined with a state-of-the-art unsupervised i-vector cosine system. Aren Jansen, Daniel Garcia-Romero, Pascal Clark, Jaime Hernandez-Cordero |
ICASSP | 2 |
| 2014 | Summary and initial results of the 2013-2014 speaker recognition i-vector machine learning challengeabstractDuring late-2013 through early-2014 NIST coordinated a special i-vector challenge based on data used in previous NIST Speaker Recognition Evaluations (SREs). Unlike evaluations in the SRE series, the i-vector challenge was run entirely online and used fixed-length feature vectors projected into a low-dimensional space (i-vectors) rather than audio recordings. These changes made the challenge more readily accessible, especially to participants from outside the audio processing field. Compared to the 2012 SRE, the i-vector challenge saw an increase in the number of participants by nearly a factor of two, and a two orders of magnitude increase in the number of systems submitted for evaluation. Initial results indicate the leading system achieved an approximate 37% improvement relative to the baseline system. Désiré Bansé, George R. Doddington, Daniel Garcia-Romero, John J. Godfrey, Craig S. Greenberg, Alvin F. Martin, Alan McCree, Mark A. Przybocki, Douglas A. Reynolds |
INTERSPEECH | 3 |
| 2014 | Improving speaker recognition performance in the domain adaptation challenge using deep neural networksabstractTraditional i-vector speaker recognition systems use a Gaussian mixture model (GMM) to collect sufficient statistics (SS). Recently, replacing this GMM with a deep neural network (DNN) has shown promising results. In this paper, we explore the use of DNNs to collect SS for the unsupervised domain adaptation task of the Domain Adaptation Challenge (DAC).We show that collecting SS with a DNN trained on out-of-domain data boosts the speaker recognition performance of an out-of-domain system by more than 25%. Moreover, we integrate the DNN in an unsupervised adaptation framework, that uses agglomerative hierarchical clustering with a stopping criterion based on unsupervised calibration, and show that the initial gains of the out-of-domain system carry over to the final adapted system. Despite the fact that the DNN is trained on the out-of-domain data, the final adapted system produces a relative improvement of more than 30% with respect to the best published results on this task. Daniel Garcia-Romero, Xiaohui Zhang 0007, Alan McCree, Daniel Povey |
SLT | 1 |
| 2014 | Speaker diarization with plda i-vector scoring and unsupervised calibrationabstractSpeaker diarization via unsupervised i-vector clustering has gained popularity in recent years. In this approach, i-vectors are extracted from short clips of speech segmented from a larger multi-speaker conversation and organized into speaker clusters, typically according to their cosine score. In this paper, we propose a system that incorporates probabilistic linear discriminant analysis (PLDA) for i-vector scoring, a method already frequently utilized in speaker recognition tasks, and uses unsupervised calibration of the PLDA scores to determine the clustering stopping criterion. We also demonstrate that denser sampling in the i-vector space with overlapping temporal segments provides a gain in the diarization task. We test our system on the CALLHOME conversational telephone speech corpus, which includes multiple languages and a varying number of speakers, and we show that PLDA scoring outperforms the same system with cosine scoring, and that overlapping segments reduce diarization error rate (DER) as well. Gregory Sell, Daniel Garcia-Romero |
SLT | 2 |
| 2013 | Subspace-constrained supervector PLDA for speaker verification
Daniel Garcia-Romero, Alan McCree |
INTERSPEECH | 1 |
| 2013 | A Symmetric Kernel Partial Least Squares Framework for Speaker RecognitionabstractI-vectors are concise representations of speaker characteristics. Recent progress in i-vectors related research has utilized their ability to capture speaker and channel variability to develop efficient automatic speaker verification (ASV) systems. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker verification requires a good modeling of these non-linearities and can be cast as a machine learning problem. Kernel partial least squares (KPLS) can be used for discriminative training in the i-vector space. However, this framework suffers from training data imbalance and asymmetric scoring. We use “one shot similarity scoring” (OSS) to address this. The resulting ASV system (OSS-KPLS) is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA), Probabilistic Linear Discriminant Analysis (PLDA), and Cosine Distance Scoring (CDS) classifiers. Improvements are shown. Balaji Vasan Srinivasan, Yuancheng Luo, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Multicondition training of Gaussian PLDA models in i-vector space for noise and reverberation robust speaker recognitionabstractWe present a multicondition training strategy for Gaussian Probabilistic Linear Discriminant Analysis (PLDA) modeling of i-vector representations of speech utterances. The proposed approach uses a multicondition set to train a collection of individual subsystems that are tuned to specific conditions. A final verification score is obtained by combining the individual scores according to the posterior probability of each condition given the trial at hand. The performance of our approach is demonstrated on a subset of the interview data of NIST SRE 2010. Significant robustness to the adverse noise and reverberation conditions included in the multicondition training set are obtained. The system is also shown to generalize to unseen conditions. Daniel Garcia-Romero, Xinhui Zhou, Carol Y. Espy-Wilson |
ICASSP | 1 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 1 |
| 2012 | Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulationsabstractOral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Depending on the tumor size, location and staging, patients are treated by radical surgery, radiology, chemotherapy or a combination of those treatments. As a result, their anatomical structures for speech are impaired and this leads to some negative impact on their speech intelligibility. As a part of the INTERSPEECH 2012 speaker trait Pathology sub-challenge, this study explored the use of auditory-inspired spectro-temporal modulation features for automatic speech intelligibility assessment of those pathologic speech. The averaged spectro-temporal modulations of speech considered as either intelligible or non-intelligible in the challenge database were analyzed and it was found that the non-intelligible speech tends to have its modulation amplitude peaks shift towards a smaller rate and scale. Based on SVM and GMM, variants of spectro-temporal modulation features were tested on the speaker trait challenge problem and the resulting performances on both the development and the test datasets are comparable to the baseline performance. Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen Stone 0001, Carol Y. Espy-Wilson, Shihab A. Shamma |
INTERSPEECH | 2 |
| 2011 | Linear versus mel frequency cepstral coefficients for speaker recognitionabstractMel-frequency cepstral coefficients (MFCC) have been dominantly used in speaker recognition as well as in speech recognition. However, based on theories in speech production, some speaker characteristics associated with the structure of the vocal tract, particularly the vocal tract length, are reflected more in the high frequency range of speech. This insight suggests that a linear scale in frequency may provide some advantages in speaker recognition over the mel scale. Based on two state-of-the-art speaker recognition back-end systems (one Joint Factor Analysis system and one Probabilistic Linear Discriminant Analysis system), this study compares the performances between MFCC and LFCC (Linear frequency cepstral coefficients) in the NIST SRE (Speaker Recognition Evaluation) 2010 extended-core task. Our results in SRE10 show that, while they are complementary to each other, LFCC consistently outperforms MFCC, mainly due to its better performance in the female trials. This can be explained by the relatively shorter vocal tract in females and the resulting higher formant frequencies in speech. LFCC benefits more in female speech by better capturing the spectral characteristics in the high frequency region. In addition, our results show some advantage of LFCC over MFCC in reverberant speech. LFCC is as robust as MFCC in the babble noise, but not in the white noise. It is concluded that LFCC should be more widely used, at least for the female trials, by the mainstream of the speaker recognition community. Xinhui Zhou, Daniel Garcia-Romero, Ramani Duraiswami, Carol Y. Espy-Wilson, Shihab A. Shamma |
ASRU | 2 |
| 2011 | Analysis of i-vector Length Normalization in Speaker Recognition SystemsabstractWe present a method to boost the performance of probabilistic generative models that work with i-vector representations. The proposed approach deals with the nonGaussian behavior of i-vectors by performing a simple length normalization. This non-linear transformation allows the use of probabilistic models with Gaussian assumptions that yield equivalent performance to that of more complicated systems based on Heavy-Tailed assumptions. Significant performance improvements are demonstrated on the telephone portion of NIST SRE 2010. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2011 | Kernel Partial Least Squares for Speaker RecognitionabstractI-vectors are a concise representation of speaker characteristics. Recent advances in speaker recognition have utilized their ability to capture speaker and channel variability to develop efficient recognition engines. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker recognition requires a good modeling of these non-linearities and can be cast as a machine learning problem. In this paper, we propose a kernel partial least squares (kernel PLS, or KPLS) framework for modeling speakers in the i-vectors space. The resulting recognition system is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA), Balaji Vasan Srinivasan, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami |
INTERSPEECH | 2 |
| 2011 | Automatic Speech Codec Identification with Applications to Tampering Detection of Speech RecordingsabstractIn this work many versions of CELP codecs are explored, and an observation is made that different codebooks are used to encode noisy part of residual. Taking advantage of noise patterns they generated, an algorithm was proposed to detect GSM-AMR,EFR,HR and SILK codecs. Another partly knowledge-based and partly data driven algorithm is also proposed to improve the performance for SILK. Then it's extended to identify subframe offset to do tampering detection of cellphone speech recordings. Jingting Zhou, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2010 | Automatic acquisition device identification from speech recordingsabstractIn this paper we present a study on the automatic identification of acquisition devices when only access to the output speech recordings is possible. A statistical characterization of the frequency response of the device contextualized by the speech content is proposed. In particular, the intrinsic characteristics of the device are captured by a template, constructed by appending together the means of a Gaussian mixture trained on the device speech recordings. This study focuses on two classes of acquisition devices, namely, landline telephone handsets and microphones. Three publicly available databases are used to assess the performance of linear- and mel-scaled cepstral coefficients. A Support Vector Machine classifier was used to perform closed-set identification experiments. The results show classification accuracies higher than 90 percent among the eight telephone handsets and eight microphones tested. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
ICASSP | 1 |
| 2008 | Language detection in audio content analysisabstractExperiments have shown that Language Identification systems for telephonic speech using shifted delta cepstra as the feature set and Gaussian mixture models as the backend, offers superior performance than other competing techniques. This paper aims to address the task of Language Identification for audio signals. The abundance of digital music from the Internet calls for a reliable real-time system for analyzing and properly categorizing them. Previous research has mainly focused on categorizing audio files into appropriate genres; however genre types vary with language. This paper proposes a systematic audio content analysis strategy by initially detecting whether an audio file has any vocals present in it and, if present, then detecting the language of the song. Given the language of the song, genre detection becomes a closed set classification problem. Vikramjit Mitra, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2008 | Intersession variability in speaker recognition: a behind the scene analysisabstractThe representation of a speaker’s identity by means of Gaussian supervectors (GSV) is at the heart of most of the state-of-the-art recognition systems. In this paper we present a novel procedure for the visualization of GSV by which qualitative insight about the information being captured can be obtained. Based on this visualization approach, the Switchboard-I database (SWB-I) is used to study the relationship between a data-driven partition of the acoustic space and a knowledge based partition (i.e., broad phonetic classes). Moreover, the structure of an intersession variability subspace (IVS), computed from the SWB-I database, is analyzed by displaying the projection of a speaker’s GSV into the set of eigenvectors with highest eigenvalues. This analysis reveals a strong presence of linguistic information in the IVS components with highest energy. Finally, after projecting away the information contained in the IVS from the speaker’s GSV, a visualization of the resulting GSV provides information about the characteristic patterns of spectral allocation of energy of a speaker. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2008 | Language and genre detection in audio content analysisabstractThis paper presents an audio genre detection framework that can be used for a multi-language audio corpus. Cepstral coefficients are considered and analyzed as the feature set for both a language dependent and language independent genre identification (GID) task. Language information is found to increase the overall detection accuracy on an average by at least 2.6% from its language independent counterpart. Melfrequency cepstral coefficients have been widely used for Music Information Retrieval (MIR), however, the present study shows that Linear-frequency cepstral coefficients (LFCC) with a higher number of frequency bands can improve the detection accuracy. Two other GID architectures have also been considered, but the results show that the logenergy amplitudes from triangular linearly spaced filter banks and their deltas can offer average detection accuracy as high as 98.2%, when language information is taken into account. Vikramjit Mitra, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2006 | Using quality measures for multilevel speaker recognition
Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia |
Comput. Speech Lang. | 1 |
| 2005 | Bayesian adaptation for user-dependent multimodal biometric authentication
Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez |
Pattern Recognit. | 2 |
| 2005 | Adapted user-dependent multimodal biometric authentication exploiting general information
Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez |
Pattern Recognit. Lett. | 2 |
| 2004 | Exploiting general knowledge in user-dependent fusion strategies for multimodal biometric verificationabstractA novel strategy for combining general and user-dependent knowledge in a multimodal biometric verification system is presented. It is based on SVM classifiers and trade-off coefficients introduced in the standard SVM training problem. Experiments are reported on a bimodal biometric system based on fingerprint and on-line signature traits. A comparison between three fusion strategies, namely user-independent, user-dependent and the proposed adapted user-dependent, is carried out. As a result, the suggested approach outperforms the former ones. In particular, a highly remarkable relative improvement of 68% in the EER with respect to the user-independent approach is achieved. The severe and very common problem of training data scarcity in the user-dependent strategy is also relaxed by the proposed scheme, resulting in a relative improvement of 40% in the EER compared to the raw user-dependent strategy. Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez |
ICASSP (5) | 2 |
| 2003 | Support vector machine fusion of idiolectal and acoustic speaker information in Spanish conversational speechabstractThis paper proposes a support vector machine (SVM) based combining scheme that incorporates ideolectal and acoustic characteristics for speaker recognition. Two statistical model paradigms, namely GMM for acoustic modeling and bigrams for language modeling, provide multilevel speaker information that affords a better classification performance when SVM-based fusion is accomplished. This combining approach is useful for all speaker recognition tasks where a considerable amount of data is available. Motivated by the absence of Spanish databases that made feasible our research experiments, more than nine hours of Spanish conversational speech was collected and manually transcribed from broadcasted radio talk shows. Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia |
ICASSP (2) | 1 |
| 2003 | Support vector machine fusion of idiolectal and acoustic speaker information in Spanish conversational speechabstractThis paper proposed a support vector machine (SVM) based combining scheme that incorporates idiolectal and acoustic characteristics for speaker recognition. Two statistical model paradigms, namely GMM for acoustic modeling and bigrams for language modeling, provide multilevel speaker information that affords a better classification performance when SVM-based fusion is accomplished. This combining approach is useful for all speaker recognition tasks where a considerable amount of data is available. Motivated by the absence of Spanish databases that made feasible our research experiments, more than nine hours of Spanish conversational speech was collected and manually transcribed from broadcasted radio talk shows. Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia |
ICME | 1 |
| 2003 | Robust likelihood ratio estimation in Bayesian forensic speaker recognition
Joaquín González-Rodríguez, Daniel Garcia-Romero, Marta Garcia-Gomar, Daniel Ramos-Castro, Javier Ortega-Garcia |
INTERSPEECH | 2 |