EDBT 2026 Demo / reviewers in the wild / expert
L. Paola García-Perera
dblp:23/4175 · also Leibny Paola García, Leibny Paola García-Perera, Paola García 0001
· DBLP profile ↗
69ranked-venue papers
6as first author
49since 2021 · last 2026
0000-0002-7449-5726ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 5 first-author · 42 since 2021Artificial intelligence and machine learning · 41 · 3 first-author · 29 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceabstractAudio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro. Sonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavácek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, L. Paola García-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David F. Harwath, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami |
AAAI | 24 |
| 2026 | Recent trends in distant conversational speech recognition: A review of CHiME-7 and 8 DASR challenges
Samuele Cornell, Christoph Böddeker, Taejin Park, He Huang 0012, Desh Raj, Matthew Wiesner, Yoshiki Masuyama, Xuankai Chang, Zhongqiu Wang 0001, Stefano Squartini, L. Paola García-Perera, Shinji Watanabe 0001 |
Comput. Speech Lang. | 11 |
| 2025 | GenVC: Self-Supervised Zero-Shot Voice ConversionabstractMost current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.11Audio samples, code, and model checkpoints are available at https://caizexin.github.io/GenVC/index.html Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews |
ASRU | 4 |
| 2025 | WST: Weakly Supervised Transducer for Automatic Speech RecognitionabstractThe Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available. Dongji Gao, Chenda Liao, Changliang Liu, Matthew Wiesner, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
ASRU | 5 |
| 2025 | Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution ShiftsabstractWe address the challenge of detecting synthesized speech under distribution shifts—arising from unseen synthesis methods, speakers, languages, or audio conditions—relative to the training data. Fewshot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust fewshot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed fewshot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as $\mathbf{1 0}$ in-distribution samples—achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset. Ashi Garg, Zexin Cai, Henry Li Xinyuan, L. Paola García-Perera, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews |
ASRU | 4 |
| 2025 | The JHU-MIT System for NIST SRE24: Post-Evaluation AnalysisabstractWe present the JHU-MIT submission for NIST SRE24, along with post-evaluation analysis and key insights. In the audio fixed condition, our system used Res2Net50 and ResNet100 embeddings; the open condition additionally included an ECAPA-TDNN with a multilingual Wav2Vec2 front-end, which emerged as the best single system. The audio back-ends consisted of either PLDA adapted to SRE24 Dev or a mixture of PLDA models tuned to different subconditions. To avoid overfitting, we optimized back-end hyperparameters via twofold cross-validation. For the visual condition, we leveraged pretrained ResNet100-Subcenter-ArcFace embeddings. Agglomerative clustering was used to diarize speaker and face identities in multi-speaker videos. The primary audio fixed system achieved Act. Cp=0.574, while the open condition reached Cp=0.366 on SRE24 Eval. The visual system yielded Cp=0.169, and audiovisual fusion further improved performance, achieving Cp=0.101 (fixed) and Cp=0.087 (open). Jesús Villalba 0001, Jonas Borgstrom, Prabhav Singh, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak |
ASRU | 4 |
| 2025 | CASPER: A Large Scale Spontaneous Speech DatasetabstractThe success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community. Cihan Xiao, Ruixing Liang, Xiangyu Zhang 0005, Mehmet Emre Tiryaki, Veronica Bae, Lavanya Shankar, Ethan Poon, Emmanuel Dupoux, Sanjeev Khudanpur, L. Paola García-Perera |
ASRU | 11 |
| 2025 | Scalable Controllable Accented TTSabstractWe propose a method to scale accented TTS training to large, accent-diverse datasets that often lack consistent, high-quality accent labels. Our approach relies on a speech geolocation model to infer accent labels directly from audio. To improve speaker generalization and encourage disentangling speaker from accent we explore timbre augmentation through kNN voice conversion. We validate our approach on CommonVoice by fine-tuning XTTS-v2 with accent labels inferred or improved via geolocation. According to various automated metrics based on embeddings extracted from an accent identification model, the resulting accented TTS model produces speech with better accent fidelity compared to XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, or other existing accented TTS models. According to human evaluation, it was clear that the geolocation model based data discovery and enhancement improved the naturalness and accent fidelity of generated speech. However, the effect of different data augmentation strategies was less clear. Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
ASRU | 5 |
| 2025 | Constructing Datasets From Public Police Body Camera FootageabstractThe enormous potential of body-worn cameras to improve accountability in policing remains largely unrealized due to large volumes of unreviewed footage. Transcription and diarization tools could aid in reviewing footage, but lack of public data hinders their development. We develop a pipeline to construct public datasets, making use of the small number of videos publicly released by police departments, with capacity to update the data as footage gets released or removed. Our pipeline produces two datasets, a large one with transcriptions automatically extracted from department-generated captions, and a smaller test set where we manually validated transcripts and alignment. We benchmark ASR models, including models fine-tuned on our data, on our test set, to show applications of our datasets and continued challenges of this domain. Our work presents a new vision for leveraging public body-worn camera footage—even when it can’t be rereleased—to help address this critical social issue. Jamie Rosas-Smith, Martijn Bartelds, Ruizhe Huang, L. Paola García-Perera, Karen Livescu, Daniel Jurafsky, Anjalie Field |
ICASSP | 4 |
| 2025 | HLTCOE Submission to the VoicePrivacy Attacker ChallengeabstractWe describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By simultaneously employing one or more of these methods, we were able to achieve a significant reduction in EER against all of the submitted anonymization systems in the VoicePrivacy Challenge. Henry Li Xinyuan, Ashi Garg, Zexin Cai, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
ICASSP | 5 |
| 2025 | The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso |
INTERSPEECH | 8 |
| 2024 | ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech RecognitionabstractKnowing the particular context associated with a conversation can help improving the performance of an automatic speech recognition (ASR) system. For example, if we are provided with a list of in-context words or phrases — such as the speaker’s contacts or recent song playlists — during inference, we can bias the recognition process towards this list. There are many works addressing contextual ASR; however, there is few publicly available real benchmark for evaluation, making it difficult to compare different solutions. To this end, we provide a corpus (“ConEC”) and baselines to evaluate contextual ASR approaches, grounded on real-world applications. The ConEC corpus is based on public-domain earnings calls (ECs) and associated supplementary materials, such as presentation slides, earnings news release as well as a list of meeting participants’ names and affiliations. We demonstrate that such real contexts are noisier than artificially synthesized contexts that contain the ground truth, yet they still make great room for future improvement of contextual ASR technology Ruizhe Huang, Mahsa Yarmohammadi, Jan Trmal, Desh Raj, L. Paola García-Perera, Alexei V. Ivanov, Patrick Ehlen, Mingzhi Yu, Daniel Povey, Sanjeev Khudanpur |
LREC/COLING | 6 |
| 2024 | Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelabstractXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola Garcia Perera, EngSiong Chng, Lina Yao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xiangyu Zhang 0005, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, L. Paola García-Perera, Chng Eng Siong |
EMNLP | 6 |
| 2024 | Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual CortexabstractWhile significant advancements in artificial intelligence (AI) have catalyzed progress across various domains, its full potential in understanding visual perception remains underexplored. We propose an artificial neural network dubbed VISION, an acronym for "Visual Interface System for Imaging Output of Neural activity," to mimic the human brain and show how it can foster neuroscientific inquiries. Using visual and contextual inputs, this multimodal model predicts the brain's functional magnetic resonance imaging (fMRI) scan response to natural images. VISION successfully predicts human hemodynamic responses as fMRI voxel values to visual inputs with an accuracy exceeding state-of-the-art performance by 45%. We further probe the trained networks to reveal representational biases in different visual areas, generate experimentally testable hypotheses, and formulate an interpretable metric to associate these hypotheses with cortical functions. With both a model and evaluation metric, the cost and time burdens associated with designing and implementing functional analysis on the visual cortex could be reduced. Our work suggests that the evolution of computational models may shed light on our fundamental understanding of the visual cortex and provide a viable approach toward reliable brain-machine interfaces. Ruixing Liang, Xiangyu Zhang 0005, Hexin Liu, Avisha Kumar, Kelley M. Kempski Leadingham, Joshua Punnoose, L. Paola García-Perera, Amir Manbachi |
ICASSP | 9 |
| 2024 | Enhancing Code-Switching Speech Recognition With Interactive Language BiasesabstractLanguages usually switch within a multilingual speech signal, especially in a bilingual society. This phenomenon is referred to as code-switching (CS), making automatic speech recognition (ASR) challenging under a multilingual scenario. We propose to improve CS-ASR by biasing the hybrid CTC/attention ASR model with multi-level language information comprising frame-and token-level language posteriors. The interaction between various resolutions of language biases is subsequently explored in this work. We conducted experiments on datasets from the ASRU 2019 code-switching challenge. Compared to the baseline, the proposed interactive language biases (ILB) method achieves higher performance and ablation studies highlight the effects of different language biases and their interactions. In addition, the results presented indicate that language bias implicitly enhances internal language modeling, leading to performance degradation after employing an external language model. Hexin Liu, L. Paola García-Perera, Xiangyu Zhang 0005, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 2 |
| 2024 | Enhancing Neural Transducer for Multilingual ASR with Synchronized Language Diarization
Amir Hussein, Desh Raj, Matthew Wiesner, Daniel Povey, L. Paola García-Perera, Sanjeev Khudanpur |
INTERSPEECH | 5 |
| 2024 | Bridging Child-Centered Speech Language Identification and Language Diarization via Phonetics
Hexin Liu, L. Paola García-Perera |
INTERSPEECH | 3 |
| 2024 | Where are you from? Geolocating Speech and Applications to Language IdentificationabstractPatrick Foley, Matthew Wiesner, Bismarck Odoom, Leibny Paola Garcia Perera, Kenton Murray, Philipp Koehn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Patrick Foley, Matthew Wiesner, Bismarck Bamfo Odoom, L. Paola García-Perera, Kenton Murray, Philipp Koehn |
NAACL-HLT | 4 |
| 2024 | Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker AnonymizationabstractAdvances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its utility, including linguistic and paralinguistic aspects. However, anonymizing speech while maintaining emotional state remains challenging. We explore this problem in the context of the VoicePrivacy 2024 challenge. Specifically, we developed various speaker anonymization pipelines and find that approaches either excel at anonymization or preserving emotion state, but not both simultaneously. Achieving both would require an in-domain emotion recognizer. Additionally, we found that it is feasible to train a semi-effective speaker verification system using only emotion representations, demonstrating the challenge of separating these two modalities. Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
SLT | 4 |
| 2023 | Learning From Flawed Data: Weakly Supervised Automatic Speech RecognitionabstractTraining automatic speech recognition (ASR) systems requires large amounts of well-curated paired data. However, human annotators usually perform “non-verbatim” transcription, which can result in poorly trained models. In this paper, we propose Omni-temporal Classification (OTC), a novel training criterion that explicitly incorporates label uncertainties originating from such weak supervision. This allows the model to effectively learn speech-text alignments while accommodating errors present in the training transcripts. OTC extends the conventional CTC objective for imperfect transcripts by leveraging weighted finite state transducers. Through experiments conducted on the LibriSpeech and LibriVox datasets, we demonstrate that training ASR models with OTC avoids performance degradation even with transcripts containing up to 70% errors, a scenario where CTC models fail completely. Our implementation is available at https://github.com/k2-fsa/icefall. Dongji Gao, Hainan Xu, Desh Raj, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
ASRU | 4 |
| 2023 | Euro: Espnet Unsupervised ASR Open-Source ToolkitabstractThis paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, which leverages self-supervised speech representations and adversarial training. In addition to wav2vec2, EURO extends the functionality and promotes reproducibility for UASR tasks by integrating S3PRL and k2, resulting in flexible frontends from 27 self-supervised models and various graph-based decoding strategies. EURO is implemented in ESPnet and follows its unified pipeline to provide UASR recipes with a complete setup. This improves the pipeline’s efficiency and allows EURO to be easily applied to existing datasets in ESPnet. Extensive experiments on three mainstream self-supervised models demonstrate the toolkit’s effectiveness and achieve state-of-the-art UASR performance on TIMIT and LibriSpeech datasets. EURO will be publicly available at https://github.com/espnet/espnet, aiming to promote this exciting and emerging research area based on UASR through open-source activity. Dongji Gao, Jiatong Shi, Shun-Po Chuang, L. Paola García-Perera, Hung-yi Lee, Shinji Watanabe 0001, Sanjeev Khudanpur |
ICASSP | 4 |
| 2023 | Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker EmbeddingsabstractSelf-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to the domain mismatch — which severely limits their use for such applications. In this paper, we investigate the adaptation of upstream SSL models to the multi-talker automatic speech recognition (ASR) task under two conditions. First, when segmented utterances are given, we show that adding a target speaker extraction (TSE) module based on enrollment embeddings is complementary to mixture-aware pre-training. Second, for unsegmented mixtures, we propose a novel joint speaker modeling (JSM) approach, which aggregates information from all speakers in the mixture through their embeddings. With controlled experiments on Libri2Mix, we show that using speaker embeddings provides relative WER improvements of 9.1% and 42.1% over strong baselines for the segmented and unsegmented cases, respectively. We also demonstrate the effectiveness of our models for real conversational mixtures through experiments on the AMI dataset. Our code and models are open-sourced on https://github.com/HuangZiliAndy/SSL_for_multitalker. Zili Huang, Desh Raj, L. Paola García-Perera, Sanjeev Khudanpur |
ICASSP | 3 |
| 2023 | Building Keyword Search System from End-To-End Asr SystemsabstractKeyword search (KWS) systems are commonly built on top of existing automatic speech recognition (ASR) systems. However, end-to-end (E2E) ASR models are not naturally equipped with word-level timing information or confidence. Existing methods for re-purposing E2E ASR systems for KWS are largely heuristic or model-specific. In this paper, we describe a general KWS pipeline, applicable to any ASR model that generates N-best lists. We extract timing information using either external word-aligners, or time-preserving weighted finite-state transducer-based decoders. We show that our light-weight, ASR-agnostic approach for confidence estimation based on N-best lists outperforms other commonly used heuristics, such as using the decoder’s softmax probability, and even a more complicated dedicated confidence estimation model (CEM). Finally, we compare our performance to hybrid ASR models, extensively evaluating the impact of word-level timing, confidence, and recall on KWS performance. Our KWS pipeline is available online1, suitable for evaluating the aforementioned ASR components as downstream tasks. Ruizhe Huang, Matthew Wiesner, L. Paola García-Perera, Daniel Povey, Jan Trmal, Sanjeev Khudanpur |
ICASSP | 3 |
| 2023 | PQLM - Multilingual Decentralized Portable Quantum Language ModelabstractWith careful manipulation, malicious agents can reverse engineer private information encoded in pre-trained language models. Security concerns motivate the development of quantum pre-training. In this work, we propose a highly portable quantum language model (PQLM) that can easily transmit information to downstream tasks on classical machines. The framework consists of a cloud PQLM built with random Variational Quantum Classifiers (VQC) and local models for downstream applications. We demonstrate the ad hoc portability of the quantum model by extracting only the word embeddings and effectively applying them to downstream tasks on classical machines. Our PQLM exhibits comparable performance to its classical counterpart on both intrinsic evaluation (loss, perplexity) and extrinsic evaluation (multilingual sentiment analysis accuracy) metrics. We also perform ablation studies on the factors affecting PQLM performance to analyze model stability. Our work establishes a theoretical foundation for a portable quantum pre-trained language model that could be trained on private data and made available for public use with privacy protection guarantees. Shuyue Stella Li, Xiangyu Zhang 0005, Hongchao Shu, Ruixing Liang, Hexin Liu, L. Paola García-Perera |
ICASSP | 7 |
| 2023 | Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language DiarizationabstractCode-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporate language information within the CS-ASR model by dynamically biasing the model with token-level language posteriors corresponding to outputs of a sequence-to-sequence auxiliary language diarization (LD) module. In contrast, the disentangling process reduces the difference between languages via adversarial training so as to normalize two languages. We conduct experiments on the SEAME dataset. Compared to the baseline model, both the joint optimization with LD and the language posterior bias achieve performance improvement. Comparison of the proposed methods indicates that incorporating language information is more effective than disentangling for reducing language confusion in CS speech. Hexin Liu, Haihua Xu 0001, L. Paola García-Perera, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 3 |
| 2023 | Bridging Speech and Textual Pre-Trained Models With Unsupervised ASRabstractSpoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised models and textual pre-trained models, which have shown reasonable improvements to various SLU tasks. However, because of the mismatched modalities between speech signals and text tokens, previous methods usually need complex designs of the frameworks. This work proposes a simple yet efficient unsupervised paradigm that connects speech and textual pre-trained models, resulting in an unsupervised speech-to-semantic pre-trained model for various tasks in SLU. To be specific, we propose to use unsupervised automatic speech recognition (ASR) as a connector that bridges different modalities used in speech and textual pre-trained models. Our experiments show that unsupervised ASR itself can improve the representations from speech self-supervised models. More importantly, it is shown as an efficient connector between speech and textual pre-trained models, improving the performances of five different SLU tasks. Notably, on spoken question answering, we reach the state-of-the-art result over the challenging NMSQA benchmark. Jiatong Shi, Chan-Jan Hsu, Ho-Lam Chung, Dongji Gao, L. Paola García-Perera, Shinji Watanabe 0001, Ann Lee 0001, Hung-yi Lee |
ICASSP | 5 |
| 2023 | A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive FiltersabstractThe detection of abnormal fetal heartbeats during pregnancy is important for monitoring the health conditions of the fetus. While adult ECG has made several advances in modern medicine, noninvasive fetal electrocardiography (FECG) remains a great challenge. In this paper, we introduce a new method based on affine combinations of adaptive filters to extract FECG signals. The affine combination of multiple filters is able to precisely fit the reference signal, and thus obtain more accurate FECGs. We proposed a method to combine the Least Mean Square (LMS) and Recursive Least Squares (RLS) filters. Our approach found that the Combined Recursive Least Squares (CRLS) filter achieves the best performance among all proposed combinations. In addition, we found that CRLS is more advantageous in extracting FECG from abdominal electrocardiograms (AECG) with a small signal-to-noise ratio (SNR). Compared with the state-of-the-art Multiple Sub-Filter Adaptive Noise Canceller (MSF-ANC) method, CRLS shows improved performance. The sensitivity, accuracy and F1 score are improved by 3.58%, 2.39% and 1.36%, respectively. Yu Xuan, Xiangyu Zhang 0005, Shuyue Stella Li, Zihan Shen, L. Paola García-Perera, Roberto Togneri |
ICASSP | 6 |
| 2023 | Advances in Language Recognition in Low Resource African Languages: The JHU-MIT Submission for NIST LRE22
Jesús Villalba 0001, Jonas Borgstrom, Maliha Jahan, Saurabh Kataria 0001, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak |
INTERSPEECH | 5 |
| 2023 | MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarizationabstractTo enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on datasets that feature diverse language registers and speech patterns. We present the MERLIon CCS challenge, featuring a first-of-its-kind Zoom video call dataset of parent-child shared book reading, of over 30 hours with over 300 recordings, annotated by multilingual transcribers using a high-fidelity linguistic transcription protocol. The audio corpus features spontaneous and in-the-wild English-Mandarin code-switching, child-directed speech in non-standard accents with diverse language-mixing patterns recorded in a variety of home environments. This report describes the corpus, as well as LID and LD results for our baseline and several systems submitted to the MERLIon CCS challenge using the corpus. Yi Han Victoria Chua, Hexin Liu, L. Paola García-Perera, Fei Ting Woon, Jinyi Wong, Xiangyu Zhang 0005, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels, Suzy J. Styles |
INTERSPEECH | 3 |
| 2023 | Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts
Dongji Gao, Matthew Wiesner, Hainan Xu, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 4 |
| 2023 | Investigating model performance in language identification: beyond simple error statisticsabstractLanguage development experts need tools that can automatically identify languages from fluent, conversational speech and provide reliable estimates of usage rates at the level of an individual recording. However, LID systems are typically evaluated on metrics such as equal error rate and balanced accuracy, applied at the level of an entire speech corpus. These overview metrics do not provide information about model performance at the level of individual speakers, recordings, or units of speech with different linguistic characteristics. Overview statistics may mask systematic errors in model performance for some subsets of the data, and consequently, have worse performance on data derived from some subsets of human speakers, creating a kind of algorithmic bias. Here, we investigate how well a number of LID systems perform on individual recordings and speech units with different linguistic properties in the MERLIon CCS Challenge featuring accented code-switched child-directed speech. Suzy J. Styles, Yi Han Victoria Chua, Fei Ting Woon, Hexin Liu, L. Paola García-Perera, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels |
INTERSPEECH | 5 |
| 2023 | Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local AttractorsabstractA method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label classification problem. It has also been extended for a flexible number of speakers by introducing speaker-wise attractors. However, the output number of speakers of attractor-based EEND is empirically capped; it cannot deal with cases where the number of speakers appearing during inference is higher than that during training because its speaker counting is trained in a fully supervised manner. Our method, EEND-GLA, solves this problem by introducing unsupervised clustering into attractor-based EEND. In the method, the input audio is first divided into short blocks, then attractor-based diarization is performed for each block, and finally, the results of each block are clustered on the basis of the similarity between locally-calculated attractors. While the number of output speakers is limited within each block, the total number of speakers estimated for the entire input can be higher than the limitation. To use EEND-GLA in an online manner, our method also extends the speaker-tracing buffer, which was originally proposed to enable online inference of conventional EEND. We introduce a block-wise buffer update to make the speaker-tracing buffer compatible with EEND-GLA. Finally, to improve online diarization, our method improves the buffer update method and revisits the variable chunk-size training of EEND. The experimental results demonstrate that EEND-GLA can perform speaker diarization of an unseen number of speakers in both offline and online inferences. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yuki Takashima, Yohei Kawaguchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Multi-Channel End-To-End Neural Diarization with Distributed MicrophonesabstractRecent progress on end-to-end neural diarization (EEND) has en-abled overlap-aware speaker diarization with a single neural net-work. This paper proposes to enhance EEND by using multi-channel signals from distributed microphones. We replace Transformer en-coders in EEND with two types of encoders that process a multi-channel input: spatio-temporal and co-attention encoders. Both are independent of the number and geometry of microphones and suitable for distributed microphone settings. We also propose a model adaptation method using only single-channel recordings. With simulated and real-recorded datasets, we demonstrated that the proposed method outperformed conventional EEND when a multi-channel in-put was given while maintaining comparable performance with a single-channel input. We also showed that the proposed method performed well even when spatial information is inoperative given multi-channel inputs, such as in hybrid meetings in which the utterances of multiple remote participants are played back from the same loudspeaker. Shota Horiguchi, Yuki Takashima, L. Paola García-Perera, Shinji Watanabe 0001, Yohei Kawaguchi |
ICASSP | 3 |
| 2022 | Investigating Self-Supervised Learning for Speech Enhancement and SeparationabstractSpeech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a great number of supervised learning-based enhancement and separation methods having been proposed and achieving good performance, studies on applying self-supervised learning (SSL) to enhancement and separation are limited. In this paper, we evaluate 13 SSL upstream methods on speech enhancement and separation downstream tasks. Our experimental results on Voicebank-DEMAND and Libri2Mix show that some SSL representations consistently outperform baseline features including the short-time Fourier transform (STFT) magnitude and log Mel filterbank (FBANK). Furthermore, we analyze the factors that make existing SSL frameworks difficult to apply to speech enhancement and separation and discuss the representation properties desired for both tasks. Our study is included as the official speech enhancement and separation downstreams for SUPERB. Zili Huang, Shinji Watanabe 0001, Shu-Wen Yang, L. Paola García-Perera, Sanjeev Khudanpur |
ICASSP | 4 |
| 2022 | PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language IdentificationabstractWe propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training.In this model, named PHO-LID, a self-supervised phoneme segmentation task and a LID task share a convolutional neural network (CNN) module, which encodes both language identity and sequential phonemic information in the input speech to generate an intermediate sequence of "phonotactic" embeddings.These embeddings are then fed into transformer encoder layers for utterance-level LID.We call this architecture CNN-Trans.We evaluate it on AP17-OLR data and the MLS14 set of NIST LRE 2017, and show that the PHO-LID model with multitask optimization exhibits the highest LID performance among all models, achieving over 40% relative improvement in terms of average cost on AP17-OLR data compared to a CNN-Trans model optimized only for LID.The visualized confusion matrices imply that our proposed method achieves higher performance on languages of the same cluster in NIST LRE 2017 data than the CNN-Trans model.A comparison between predicted phoneme boundaries and corresponding audio spectrograms illustrates the leveraging of phoneme information for LID. Hexin Liu, L. Paola García-Perera, Andy W. H. Khong, Suzy J. Styles, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2022 | Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End ASR ModelsabstractIn this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model.Conventional approaches require extra parameters of the same size as the model for optimization, and it is difficult to apply these approaches to end-to-end ASR models because they have a huge amount of parameters.To solve this problem, we first investigate which parts of end-to-end ASR models contribute to high accuracy in the target domain while preventing catastrophic forgetting.We conduct experiments on incremental domain adaptation from the LibriSpeech dataset to the AMI meeting corpus with two popular end-to-end ASR models and found that adapting only the linear layers of their encoders can prevent catastrophic forgetting.Then, on the basis of this finding, we develop an element-wise parameter selection focused on specific layers to further reduce the number of fine-tuning parameters.Experimental results show that our approach consistently prevents catastrophic forgetting compared to parameter selection from the whole model. Yuki Takashima, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yohei Kawaguchi |
INTERSPEECH | 4 |
| 2022 | Mutual Learning of Single- and Multi-Channel End-to-End Neural DiarizationabstractDue to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is also known that single-channel speech data can benefit multi-channel models by mixing it with multi-channel speech data during training or by using it for model pretraining. This paper focuses on speaker diarization and proposes to conduct the above bi-directional knowledge transfer alternately. We first introduce an end-to-end neural diarization model that can handle both single- and multi-channel inputs. Using this model, we alternately conduct i) knowledge distillation from a multi-channel model to a single-channel model and ii) finetuning from the distilled single-channel model to a multi-channel model. Experimental results on two-speaker data show that the proposed method mutually improved single- and multi-channel speaker diarization performances. Shota Horiguchi, Yuki Takashima, Shinji Watanabe 0001, L. Paola García-Perera |
SLT | 4 |
| 2022 | On Compressing Sequences for Self-Supervised Speech ModelsabstractCompressing self-supervised models has become increasingly necessary, as self-supervised models become larger. While previous approaches have primarily focused on compressing the model size, shortening sequences is also effective in reducing the computational cost. In this work, we study fixed-length and variable-length subsampling along the time axis in self-supervised learning. We explore how individual downstream tasks are sensitive to input frame rates. Subsampling while training self-supervised models not only improves the overall performance on downstream tasks under certain frame rates, but also brings significant speed-up in inference. Variable-length subsampling performs particularly well under low frame rates. In addition, if we have access to phonetic boundaries, we find no degradation in performance for an average frame rate as low as 10 Hz. Yen Meng, Hsuan-Jui Chen, Jiatong Shi, Shinji Watanabe 0001, L. Paola García-Perera, Hung-yi Lee, Hao Tang 0002 |
SLT | 5 |
| 2022 | Joint speaker diarization and speech recognition based on region proposal networks
Zili Huang, Marc Delcroix, L. Paola García-Perera, Shinji Watanabe 0001, Desh Raj, Sanjeev Khudanpur |
Comput. Speech Lang. | 3 |
| 2022 | Encoder-Decoder Based Attractors for End-to-End Neural DiarizationabstractThis paper investigates an end-to-end neural diarization (EEND) method for an unknown number of speakers. In contrast to the conventional cascaded approach to speaker diarization, EEND methods are better in terms of speaker overlap handling. However, EEND still has a disadvantage in that it cannot deal with a flexible number of speakers. To remedy this problem, we introduce encoder-decoder-based attractor calculation module (EDA) to EEND. Once frame-wise embeddings are obtained, EDA sequentially generates speaker-wise attractors on the basis of a sequence-to-sequence method using an LSTM encoder-decoder. The attractor generation continues until a stopping condition is satisfied; thus, the number of attractors can be flexible. Diarization results are then estimated as dot products of the attractors and embeddings. The embeddings from speaker overlaps result in larger dot product values with multiple attractors; thus, this method can deal with speaker overlaps. Because the maximum number of output speakers is still limited by the training set, we also propose an iterative inference method to remove this restriction. Further, we propose a method that aligns the estimated diarization results with the results of an external speech activity detector, which enables fair comparison against cascaded approaches. Extensive evaluations on simulated and real datasets show that EEND-EDA outperforms the conventional cascaded approach. Shota Horiguchi, Yusuke Fujita, Shinji Watanabe 0001, Yawen Xue, L. Paola García-Perera |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local AttractorsabstractAttractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the case where the number of speakers is larger than the one observed during training. This is because its speaker counting relies on supervised learning. In this work, we introduce an unsupervised clustering process embedded in the attractor-based end-to-end diarization. We first split a sequence of frame-wise embeddings into short subsequences and then perform attractor-based diarization for each subsequence. Given subsequence-wise diarization results, inter-subsequence speaker correspondence is obtained by unsupervised clustering of the vectors computed from the attractors from all the subsequences. This makes it possible to produce diarization results of a large number of speakers for the whole recording even if the number of output speakers for each subsequence is limited. Experimental results showed that our method could produce accurate diarization results of an unseen number of speakers. Our method achieved 11.84 %, 28.33 %, and 19.49 % on the CALLHOME, DI-HARD II, and DIHARD III datasets, respectively, each of which is better than the conventional end-to-end diarization methods. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yawen Xue, Yuki Takashima, Yohei Kawaguchi |
ASRU | 3 |
| 2021 | End-To-End Speaker Diarization as Post-ProcessingabstractThis paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech because each frame is assigned to one speaker. On the other hand, some end-to-end diarization methods can handle overlapping speech by treating the problem as multi-label classification. Although some methods can treat a flexible number of speakers, they do not perform well when the number of speakers is large. To compensate for each other’s weakness, we propose to use a two-speaker end-to-end diarization method as post-processing of the results obtained by a clustering-based method. We iteratively select two speakers from the results and update the results of the two speakers to improve the overlapped region. Experimental results show that the proposed algorithm consistently improved the performance of the state-of-the-art methods across CALLHOME, AMI, and DIHARD II datasets. Shota Horiguchi, L. Paola García-Perera, Yusuke Fujita, Shinji Watanabe 0001, Kenji Nagamatsu |
ICASSP | 2 |
| 2021 | End-to-End Language Diarization for Bilingual Code-Switching SpeechabstractWe propose two end-to-end neural configurations for language diarization on bilingual code-switching speech. The first, a BLSTM-E2E architecture, includes a set of stacked bidirectional LSTMs to compute embeddings and incorporates the deep clustering loss to enforce grouping of languages belonging to the same class. The second, an XSA-E2E architecture, is based on an x-vector model followed by a self-attention encoder. The former encodes frame-level features into segmentlevel embeddings while the latter considers all those embeddings to generate a sequence of segment-level language labels. We evaluated the proposed methods on the dataset obtained from the shared task B in WSTCSMC 2020 and our handcrafted simulated data from the SEAME dataset. Experimental results show that our proposed XSA-E2E architecture achieved a relative improvement of 12.1% in equal error rate and a 7.4% relative improvement on accuracy compared with the baseline algorithm in the WSTCSMC 2020 dataset. Our proposed XSA-E2E architecture achieved an accuracy of 89.84% with a baseline of 85.60% on the simulated data derived from the SEAME dataset. Hexin Liu, L. Paola García-Perera, Justin Dauwels, Andy W. H. Khong, Sanjeev Khudanpur, Suzy J. Styles |
Interspeech | 2 |
| 2021 | Semi-Supervised Training with Pseudo-Labeling for End-To-End Neural DiarizationabstractIn this paper, we present a semi-supervised training technique using pseudo-labeling for end-to-end neural diarization (EEND).The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech.However, to get a welltuned model, EEND requires labeled data for all the joint speech activities of every speaker at each time frame in a recording.In this paper, we explore a pseudo-labeling approach that employs unlabeled data.First, we propose an iterative pseudolabel method for EEND, which trains the model using unlabeled data of a target condition.Then, we also propose a committeebased training method to improve the performance of EEND.To evaluate our proposed method, we conduct the experiments of model adaptation using labeled and unlabeled data.Experimental results on the CALLHOME dataset show that our proposed pseudo-label achieved a 37.4% relative diarization error rate reduction compared to a seed model.Moreover, we analyzed the results of semi-supervised adaptation with pseudo-labeling.We also show the effectiveness of our approach on the third DI-HARD dataset. Yuki Takashima, Yusuke Fujita, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 5 |
| 2021 | Training Hybrid Models on Noisy Transliterated Transcripts for Code-Switched Speech Recognition
Matthew Wiesner, Mousmita Sarma, Ashish Arora, Desh Raj, Dongji Gao, Ruizhe Huang, Supreet Preet, Moris Johnson, Zikra Iqbal, Nagendra K. Goel, Jan Trmal, L. Paola García-Perera, Sanjeev Khudanpur |
Interspeech | 12 |
| 2021 | Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers
Yawen Xue, Shota Horiguchi, Yusuke Fujita, Yuki Takashima, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 6 |
| 2021 | DOVER-Lap: A Method for Combining Overlap-Aware Diarization OutputsabstractSeveral advances have been made recently towards handling overlapping speech for speaker diarization. Since speech and natural language tasks often benefit from ensemble techniques, we propose an algorithm for combining outputs from such diarization systems through majority voting. Our method, DOVER-Lap, is inspired from the recently proposed DOVER algorithm, but is designed to handle overlapping segments in diarization outputs. We also modify the pair-wise incremental label mapping strategy used in DOVER, and propose an approximation algorithm based on weighted k-partite graph matching, which performs this mapping using a global cost tensor. We demonstrate the strength of our method by combining outputs from diverse systems - clustering-based, region proposal networks, and target-speaker voice activity detection - on AMI and LibriCSS datasets, where it consistently outperforms the single best system. Additionally, we show that DOVER-Lap can be used for late fusion in multichannel diarization, and compares favorably with early fusion methods like beamforming. Desh Raj, L. Paola García-Perera, Zili Huang, Shinji Watanabe 0001, Daniel Povey, Andreas Stolcke, Sanjeev Khudanpur |
SLT | 2 |
| 2021 | End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap DetectionabstractIn this paper, we present a conditional multitask learning method for end-to-end neural speaker diarization (EEND). The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech. In this paper, to further improve the performance of the EEND system, we propose a novel multitask learning framework that solves speaker diarization and a desired subtask while explicitly considering the task dependency. We optimize speaker diarization conditioned on speech activity and overlap detection that are subtasks of speaker diarization, based on the probabilistic chain rule. Experimental results show that our proposed method can leverage a subtask to effectively model speaker diarization, and outperforms conventional EEND systems in terms of diarization error rate. Yuki Takashima, Yusuke Fujita, Shinji Watanabe 0001, Shota Horiguchi, L. Paola García-Perera, Kenji Nagamatsu |
SLT | 5 |
| 2021 | Online End-To-End Neural Diarization with Speaker-Tracing BufferabstractThis paper proposes a novel online speaker diarization algorithm based on a fully supervised self-attention mechanism (SA-EEND). Online diarization inherently presents a speaker's permutation problem due to the possibility to assign speaker regions incorrectly across the recording. To circumvent this inconsistency, we proposed a speaker-tracing buffer mechanism that selects several input frames representing the speaker permutation information from previous chunks and stores them in a buffer. These buffered frames are stacked with the input frames in the current chunk and fed into a self-attention network. Our method ensures consistent diarization outputs across the buffer and the current chunk by checking the correlation between their corresponding outputs. Additionally, we trained SA-EEND with variable chunk-sizes to mitigate the mismatch between training and inference introduced by the speaker-tracing buffer mechanism. Experimental results, including online SA-EEND and variable chunk-size, achieved DERs of 12.54% for CALLHOME and 20.77% for CSJ with 1.4 s actual latency. Yawen Xue, Shota Horiguchi, Yusuke Fujita, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
SLT | 5 |
| 2020 | Overlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech DetectionabstractWe address the problem of effectively handling overlapping speech in a diarization system. First, we detail a neural Long Short-Term Memory- based architecture for overlap detection. Secondly, detected overlap regions are exploited in conjunction with a frame-level speaker posterior matrix to make two-speaker assignments for overlapped frames in the resegmentation step. The overlap detection module achieves state-of-the-art performance on the AMI, DIHARD, and ETAPE corpora. We apply overlap-aware resegmentation on AMI, resulting in a 20% relative DER reduction over the baseline system. While this approach is by no means an end-all solution to overlap-aware diarization, it reveals promising directions for handling overlap. Latané Bullock, Hervé Bredin, L. Paola García-Perera |
ICASSP | 3 |
| 2020 | Speaker Diarization with Region Proposal NetworkabstractSpeaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline. Zili Huang, Shinji Watanabe 0001, Yusuke Fujita, L. Paola García-Perera, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 4 |
| 2020 | Feature Enhancement with Deep Feature Losses for Speaker VerificationabstractSpeaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which optimizes the enhancement network in the hidden activation space of a pre-trained auxiliary speaker embedding network. We experimentally verify the approach on simulated and real data. A simulated testing setup is created using various noise types at different SNR levels. For evaluation on real data, we choose BabyTrain corpus which consists of children recordings in uncontrolled environments. We observe consistent gains in every condition over the state-of-the-art augmented Factorized-TDNN x-vector system. On BabyTrain corpus, we observe relative gains of 10.38% and 12.40% in minDCF and EER respectively. Saurabh Kataria 0001, Phani S. Nidadavolu, Jesús Villalba 0001, Nanxin Chen, L. Paola García-Perera, Najim Dehak |
ICASSP | 5 |
| 2020 | Unsupervised Feature Enhancement for Speaker VerificationabstractThe task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with using both real speech recorded in adverse environments and degraded speech obtained by simulation to train the enhancement systems. The effectiveness of this approach was shown by testing on several real, simulated noisy, and reverberant test sets. The approach yielded significant improvements on both real and simulated sets when data augmentation was not used in speaker verification pipeline. We also experimented with training the x-vector and PLDA systems with enhanced augmented features instead of augmented features and observed better performance on real test conditions (4.2% relative improvement in minDCF on SRI). Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, L. Paola García-Perera, Najim Dehak |
ICASSP | 4 |
| 2020 | End-to-End Domain-Adversarial Voice Activity DetectionabstractInternational audience Marvin Lavechin, Marie-Philippe Gill, Ruben Bousbib, Hervé Bredin, L. Paola García-Perera |
INTERSPEECH | 5 |
| 2020 | State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, L. Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, Najim Dehak |
Comput. Speech Lang. | 8 |
| 2019 | Using ASR Methods for OCRabstractHybrid deep neural network hidden Markov models (DNN-HMM) have achieved impressive results on large vocabulary continuous speech recognition (LVCSR) tasks. However, the recent approaches using DNN-HMM models are not explored much for text recognition. Inspired by the current work in automatic speech recognition (ASR) and machine translation, we present an open vocabulary sub-word text recognition system. The sub-word lexicon and sub-word language model (LM) helps in overcoming the challenge of recognizing out of vocabulary (OOV) words, and a time delay neural network (TDNN) and convolution neural network (CNN) based DNN-HMM optical model (OM) efficiently models the sequence dependency in the line image. We present results on 12 datasets with training data varying from 6k lines to 600k lines. The system is built for 8 languages, i.e., English, French, Arabic, Chinese, Farsi, Tamil, Russian, and Korean. We report competitive results on several commonly used handwritten and printed text datasets. Ashish Arora, L. Paola García-Perera, Shinji Watanabe 0001, Vimal Manohar, Yiwen Shao, Sanjeev Khudanpur, Chun-Chieh Chang, Babak Rekabdar, Bagher BabaAli, Daniel Povey, David Etter, Desh Raj, Hossein Hadian, Jan Trmal |
ICDAR | 2 |
| 2019 | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 12 |
| 2019 | Advances in Automatic Speech Recognition for Child Speech Using Factored Time Delay Neural Network
L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2019 | Multi-PLDA Diarization on Children's Speech
Jiamin Xie, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2017 | DNN Bottleneck Features for Speaker Clustering
Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera |
INTERSPEECH | 2 |
| 2017 | Analysis and Description of ABC Submission to NIST SRE 2016
Oldrich Plchot, Pavel Matejka, Anna Silnova, Ondrej Novotný, Mireia Díez, Johan Rohdin, Ondrej Glembek, Niko Brümmer, Albert Swart, Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera, Patrick Kenny, Jahangir Alam 0001, Gautam Bhattacharya |
INTERSPEECH | 11 |
| 2013 | Optimization of the DET curve in speaker verification under noisy conditionsabstractThe increasing need for secure authentication systems has motivated recent interest in effective algorithms for Speaker Verification (SV). In particular, there is increasing need for noise robust algorithms for SV, which will allow SV systems to operate successfully in real conditions, which are typically noisy. L. Paola García-Perera, Bhiksha Raj, Juan A. Nolazco-Flores |
ICASSP | 1 |
| 2013 | Ensemble approach in speaker verification
L. Paola García-Perera, Bhiksha Raj, Juan A. Nolazco-Flores |
INTERSPEECH | 1 |
| 2012 | Optimization of the DET curve in speaker verificationabstractSpeaker verification systems are, in essence, statistical pattern detectors which can trade off false rejections for false acceptances. Any operating point characterized by a specific tradeoff between false rejections and false acceptances may be chosen. Training paradigms in speaker verification systems however either learn the parameters of the classifier employed without actually considering this tradeoff, or optimize the parameters for a particular operating point exemplified by the ratio of positive and negative training instances supplied. In this paper we investigate the optimization of training paradigms to explicitly consider the tradeoff between false rejections and false acceptances, by minimizing the area under the curve of the detection error tradeoff curve. To optimize the parameters, we explicitly minimize a mathematical characterization of the area under the detection error tradeoff curve, through generalized probabilistic descent. Experiments on the NIST 2008 database show that for clean signals the proposed optimization approach is at least as effective as conventional learning. On noisy data, verification performance obtained with the proposed approach is considerably better than that obtained with conventional learning methods. L. Paola García-Perera, Juan A. Nolazco-Flores, Bhiksha Raj, Richard M. Stern |
SLT | 1 |
| 2010 | Speech Magnitude-Spectrum Information-Entropy (MSIE) for Automatic Speech Recognition in Noisy EnvironmentsabstractThe Magnitude-Spectrum Information-Entropy (MSIE) of the speech signal is presented as an alternative representation of the speech that can be used to mitigate the mismatch between training and testing conditions. The speech-magnitude spectrum is considered as a random variable from which entropy coefficients can be calculated for each frame. By concatenating these entropic coefficients to its corresponding MFCC vector, then calculating the dynamic coefficients, Δ and ΔΔ, the results show an improvement compared to a baseline. The MSIE effectiveness was tested under the Aurora 2 database audio files. When trained in clean speech, the experimental results obtained by the MSIE concatenated to the MFCC outperform the results obtained with the MFCC baseline system for selected types of noises at different SNRs. For this selected group of noises the overall improvement performance in the range 0 dB to 20 dB for the Aurora 2 database is of 15.06%. Juan A. Nolazco-Flores, Roberto A. Aceves L., L. Paola García-Perera |
ICPR | 3 |
| 2008 | Enhancing acoustic models for robust speaker verificationabstractAcoustic model enhancement (AME) refers to adapting the acoustic models to compensate for the distortion induced by a speech enhancement technique. This work extends the AME technique for speaker verification recently presented by incorporating the corresponding adaptation of the model variances, and by exploring the trade off between noise over-estimation and flooring distortion in the verification error. By using spectral subtraction (SS) as the speech enhancement technique, the extended AME highly outperformed SS alone particularly at moderately low SNRs (0 dB - 15 dB), where the adaptation of the variance was found to considerably improve the equal error rate (EER). Juan A. Nolazco-Flores, L. Paola García-Perera |
ICASSP | 2 |
| 2005 | Multi-speaker voice cryptographic key generationabstractSummary form only given. In this work we present a procedure for generating binary vectors, which can be used as keys for cryptographic purposes. This research is based on automatic speech recognition technology and support vector machines. Keys bits are produced by making a distinction among the phoneme features of the users employing hyperplanes. The implementation of the method is described and the statistics are computed for several numbers of users. The results obtained show that the proposed method is sufficiently robust to reliably regenerate the key. L. Paola García-Perera, J. Carlos Mex-Perera, Juan A. Nolazco-Flores |
AICCSA | 1 |
| 2005 | Phoneme Spotting for Speech-Based Crypto-key Generation
L. Paola García-Perera, Juan A. Nolazco-Flores, J. Carlos Mex-Perera |
CIARP | 1 |
| 2004 | SVM Applied to the Generation of Biometric Speech Key
L. Paola García-Perera, J. Carlos Mex-Perera, Juan A. Nolazco-Flores |
CIARP | 1 |