EDBT 2026 Demo / reviewers in the wild / expert
Petr Motlícek
dblp:15/3204
· DBLP profile ↗
117ranked-venue papers
16as first author
40since 2021 · last 2026
0000-0001-6467-1119ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 16 first-author · 29 since 2021Artificial intelligence and machine learning · 68 · 10 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OSCAIL-OpenScience Communication through AI in EU LanguagesabstractThe Anglocentric nature of scholarly communication has many implications, such as limiting publication, discoverability and access from other language communities (even for major languages); putting minoritized languages at risk in the academic domain; and excluding many from peer review. The OSCAIL project addresses these challenges by exploring how machine translation (MT) enhanced by large language model (LLM)–based technologies can support access to scientific knowledge. Outputs will include evaluation datasets, protocols and best practices for MT in scholarly communication, and a prototype integration of MT tools into Open Journal Systems, the world’s most widely used open-source scholarly publishing platform. Sheila Castilho, Susanna Fiorini, Lynne Bowker, Petr Motlícek, Joss Moorkens, Lieve Macken, Dairazalia Sanchez-Cortes, Janne Pölönen, Sami Syrjämäki, Mikael Laakso, Mark Fishel, Anastasia Stasenko |
EAMT (2) | 4 |
| 2026 | When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
Hasindri Watawana, Sergio Burdisso, Diego Aarón Moreno-Galván, Fernando Sánchez-Vega, Adrián Pastor López-Monroy, Petr Motlícek, Esaú Villatoro-Tello |
LREC | 6 |
| 2025 | TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task ActivationabstractToken-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance. Shashi Kumar, Srikanth R. Madikeri, Esaú Villatoro-Tello, Sergio Burdisso, Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Petr Motlícek, D. S. Karthik Pandia, Shankar Venkatesan, Kadri Hacioglu, Andreas Stolcke |
ASRU | 7 |
| 2025 | XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained ModelsabstractSelf-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are trained with full attention context. In this paper, we introduce XLSR-Transducer, where the XLSR-53 model is used as encoder in transducer setup. Our experiments on the AMI dataset reveal that the XLSR-Transducer achieves 4% absolute WER improvement over Whisper large-v2 and 8% over a Zipformer transducer model trained from scratch. To enable streaming capabilities, we investigate different attention masking patterns in the self-attention computation of transformer layers within the XLSR-53 model. We validate XLSR-Transducer on AMI and 5 languages from CommonVoice under low-resource scenarios. Finally, with the introduction of attention sinks, we reduce the left context by half while achieving a relative 12% improvement in WER. Shashi Kumar, Srikanth R. Madikeri, Juan Zuluaga-Gomez, Esaú Villatoro-Tello, Iuliia Thorbecke, Petr Motlícek, Manjunath K. E, Aravind Ganapathiraju |
ICASSP | 6 |
| 2025 | Speech Data Selection for Efficient ASR Fine-Tuning using Domain Classifier and Pseudo-Label FilteringabstractIn real-world speech data processing, the scarcity of annotated data and the abundance of unlabelled speech data present a significant challenge. To address this, we propose an efficient data selection pipeline for fine-tuning ASR models by generating pseudo-labels using WhisperX pipeline and selecting efficient labels for fine-tuning. In our work, we propose a domain classifier system developed with a computationally inexpensive TFIDF and classical machine learning algorithm. Later, we filter data from the classifier output using a novel metric that assesses word ratio and perplexity distribution. The filtered pseudo labels are then used for fine-tuning standard encoder-decoder Whisper models and Zipformer. Our proposed data selection pipeline reduces the dataset size by approximately 1/100thwhile maintaining performance comparable to the full dataset, outperforming random domain-independent selection strategies. Pradeep Rangappa, Juan Zuluaga-Gomez, Srikanth R. Madikeri, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Sergio Burdisso, Shashi Kumar, Esaú Villatoro-Tello, Iuliia Nigmatulina, Petr Motlícek, D. S. Karthik Pandia, Aravind Ganapathiraju |
ICASSP | 10 |
| 2025 | Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data FilteringabstractFine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost. Roberto Andrés Vasco Carofilis, Pradeep Rangappa, Srikanth R. Madikeri, Shashi Kumar, Sergio Burdisso, Jeena J. Prakash, Esaú Villatoro-Tello, Petr Motlícek, Bidisha Sharma, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke |
INTERSPEECH | 8 |
| 2025 | Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage FilteringabstractFine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English. Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Shashi Kumar, Sergio Burdisso, Srikanth R. Madikeri, Esaú Villatoro-Tello, Bidisha Sharma, Petr Motlícek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke |
INTERSPEECH | 9 |
| 2025 | Latent Space Factorization in LoRAabstractLow-rank adaptation (LoRA) is a widely used method for parameter-efficient finetuning.
However, existing LoRA variants lack mechanisms to explicitly disambiguate task-relevant information within the learned low-rank subspace, potentially limiting downstream performance.
We propose Factorized Variational Autoencoder LoRA (FVAE-LoRA), which leverages a VAE to learn two distinct latent spaces.
Our novel Evidence Lower Bound formulation explicitly promotes factorization between the latent spaces, dedicating one latent space to task-salient features and the other to residual information.
Extensive experiments on text, audio, and image tasks demonstrate that FVAE-LoRA consistently outperforms standard LoRA.
Moreover, spurious correlation evaluations confirm that FVAE-LoRA better isolates task-relevant signals, leading to improved robustness under distribution shifts.
Our code is publicly available at: https://github.com/idiap/FVAE-LoRA Shashi Kumar, Yacouba Kaloga, John Mitros, Petr Motlícek, Ina Kodrasi |
NeurIPS | 4 |
| 2024 | Entity Matching Across Small Networks Using Node AttributesabstractEntity matching, also known as user identity linkage, is a critical task in data integration. While established techniques primarily focus on large-scale networks, there are several applications where small networks pose challenges due to limited training data and sparsity. This study addresses entity matching in the field of criminology, where small networks are common and the number of known matching nodes is restricted. To support this research, we exploit a multimodal dataset, collected as part of a security-related project, consisting of an intercepted telephone calls network (i.e., ROXSD data) and a network of social forum interactions (i.e., ROXHOOD data) collected in a simulated environment, although following real investigation scenario. To improve accuracy and efficiency, we propose a novel approach for entity matching across these two small networks using node attributes. Existing techniques often merely focus on topology consistency between two networks and overlook valuable information, such as network node attributes, making them vulnerable to structural changes. Inspired by the remarkable success of deep learning, we present UGC-DeepLink, an end-to-end semi-supervised learning framework that leverages user-generated content. UGC-DeepLink encodes network nodes into vector representations, capturing both local and global network structures to align anchor nodes using deep neural networks. A dual learning paradigm and the policy gradient method transfer knowledge and update the linkage. Additionally, node attributes, such as call contents and forum exchanged texts, enhance the ranking of matching nodes. Experimental results on ROXSD and ROXHOOD demonstrate that UGC-DeepLink surpasses baselines and state-of-the-art methods in terms of identity-match ranking. The code and dataset are available at https://github.com/erichoang/UGC-DeepLink. Zahra Ahmadi, Sergio Burdisso, Srikanth R. Madikeri, Petr Motlícek, Erinç Dikici, Gerhard Backfried, Marek Kovác, Kvetoslav Malý, Daniel Kudenko |
ECAI | 6 |
| 2024 | Dialog2Flow: Pre-training Soft-Contrastive Action-Driven Sentence Embeddings for Automatic Dialog Flow ExtractionabstractEfficiently deriving structured workflows from unannotated dialogs remains an underexplored and formidable challenge in computational linguistics. Automating this process could significantly accelerate the manual design of workflows in new domains and enable the grounding of large language models in domain-specific flowcharts, enhancing transparency and controllability.In this paper, we introduce Dialog2Flow (D2F) embeddings, which differ from conventional sentence embeddings by mapping utterances to a latent space where they are grouped according to their communicative and informative functions (i.e., the actions they represent). D2F allows for modeling dialogs as continuous trajectories in a latent space with distinct action-related regions. By clustering D2F embeddings, the latent space is quantized, and dialogs can be converted into sequences of region/action IDs, facilitating the extraction of the underlying workflow.To pre-train D2F, we build a comprehensive dataset by unifying twenty task-oriented dialog datasets with normalized per-turn action annotations. We also introduce a novel soft contrastive loss that leverages the semantic information of these actions to guide the representation learning process, showing superior performance compared to standard supervised contrastive loss.Evaluation against various sentence embeddings, including dialog-specific ones, demonstrates that D2F yields superior qualitative and quantitative results across diverse domains. Sergio Burdisso, Srikanth R. Madikeri, Petr Motlícek |
EMNLP | 3 |
| 2024 | TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASRabstractShashi Kumar, Srikanth Madikeri, Juan Pablo Zuluaga Gomez, Iuliia Thorbecke, Esaú Villatoro-tello, Sergio Burdisso, Petr Motlicek, Karthik Pandia D S, Aravind Ganapathiraju. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shashi Kumar, Srikanth R. Madikeri, Juan Zuluaga-Gomez, Iuliia Thorbecke, Esaú Villatoro-Tello, Sergio Burdisso, Petr Motlícek, Karthik S, Aravind Ganapathiraju |
EMNLP | 7 |
| 2024 | Contextual Biasing Methods for Improving Rare Word Detection in Automatic Speech RecognitionabstractIn specialized domains like Air Traffic Control (ATC), a notable challenge in porting a deployed Automatic Speech Recognition (ASR) system from one airport to another is the alteration in the set of crucial words that must be accurately detected in the new environment. Typically, such words have limited occurrences in training data, making it impractical to retrain the ASR system. This paper explores innovative word-boosting techniques to improve the detection rate of such rare words in the ASR hypotheses for the ATC domain. Two acoustic models are investigated: a hybrid CNN-TDNNF model trained from scratch and a pre-trained wav2vec2-based XLSR model fine-tuned on a common ATC dataset. The word boosting is done in three ways. First, an out-of-vocabulary word addition method is explored. Second, G-boosting is explored, which amends the language model before building the decoding graph. Third, the boosting is performed on the fly during decoding using lattice re-scoring. The results indicate that the G-boosting method performs best and provides an approximately 30-43% relative improvement in recall of the boosted words. Moreover, a relative improvement of up to 48% is obtained upon combining G-boosting and lattice-rescoring. Mrinmoy Bhattacharjee, Iuliia Nigmatulina, Amrutha Prasad, Pradeep Rangappa, Srikanth R. Madikeri, Petr Motlícek, Hartmut Helmke, Matthias Kleinert |
ICASSP | 6 |
| 2024 | Multitask Speech Recognition and Speaker Change Detection for Unknown Number of SpeakersabstractTraditionally, automatic speech recognition (ASR) and speaker change detection (SCD) systems have been independently trained to generate comprehensive transcripts accompanied by speaker turns. Recently, joint training of ASR and SCD systems, by inserting speaker turn tokens in the ASR training text, has been shown to be successful. In this work, we present a multitask alternative to the joint training approach. Results obtained on the mix-headset audios of AMI corpus show that the proposed multitask training yields an absolute improvement of 1.8% in coverage and purity based F1 score on SCD task without ASR degradation. We also examine the trade-offs between the ASR and SCD performance when trained using multitask criteria. Additionally, we validate the speaker change information in the embedding spaces obtained after different transformer layers of a self-supervised pre-trained model, such as XLSR-53, by integrating an SCD classifier at the output of specific transformer layers. Results reveal that the use of different embedding spaces from XLSR-53 model for multitask ASR and SCD is advantageous.1 Shashi Kumar, Srikanth R. Madikeri, Iuliia Nigmatulina, Esaú Villatoro-Tello, Petr Motlícek, D. S. Karthik Pandia, S. Pavankumar Dubagunta, Aravind Ganapathiraju |
ICASSP | 5 |
| 2024 | Fine-Tuning Self-Supervised Models for Language Identification Using Orthonormal ConstraintabstractSelf-supervised models trained with high linguistic diversity, such as the XLS-R model, can be effectively fine-tuned for the language recognition task. Typically, a back-end classifier followed by statistics pooling layer are added during training. Commonly used back-end classifiers require a large number of parameters to be trained, which is not ideal in limited data conditions. In this work, we explore smaller parameter back-ends using factorized Time Delay Neural Network (TDNN-F). The TDNN-F architecture is also integrated into Emphasized Channel Attention, Propagation and Aggregation- TDNN (ECAPA-TDNN) models, termed ECAPA-TDNN-F, reducing the number of parameters by 30 to 50% absolute, with competitive accuracies and no change in minimum cost. The results show that the ECAPA-TDNN-F can be extended to tasks where ECAPA-TDNN is suitable. We also test the effectiveness of a linear classifier and a variant, the Orthonormal linear classifier, previously used in x-vector type systems. The models are trained with NIST LRE17 data and evaluated on NIST LRE17, LRE22 and the ATCO2 LID datasets. Both linear classifiers outperform conventional back-ends with improvements in accuracy between 0.9% and 9.1%. Amrutha Prasad, Roberto Andrés Vasco Carofilis, Geoffroy Vanderreydt, Driss Khalil, Srikanth R. Madikeri, Petr Motlícek, Christof Schüpbach |
ICASSP | 6 |
| 2024 | Probability-Aware Word-Confusion-Network-To-Text Alignment Approach for Intent ClassificationabstractSpoken Language Understanding (SLU) technologies have greatly improved due to the effective pretraining of speech representations. A common requirement of industry-based solutions is the portability to deploy SLU models in voice-assistant devices. Thus, distilling knowledge from large text-based language models has become an attractive solution for achieving good performance and guaranteeing portability. In this paper, we introduce a novel architecture that uses a cross-modal attention mechanism to extract bin-level contextual embeddings from a word-confusion network (WNC) encoding such that these can be directly compared and aligned with traditional text-based contextual embeddings. This alignment is achieved using a recently proposed tokenwise constrastive loss function. We validate our architecture’s effectiveness by fine-tuning our WCN-based pretrained model to do intent classification (IC) on the well-known SLURP dataset. Obtained accuracy on the IC task (81%), depicts a 9.4% relative improvement compared to a recent/equivalent E2E method. Esaú Villatoro-Tello, Srikanth R. Madikeri, Bidisha Sharma, Driss Khalil, Shashi Kumar, Iuliia Nigmatulina, Petr Motlícek, Aravind Ganapathiraju |
ICASSP | 7 |
| 2024 | Detecting Criminal Networks via Non-content Communication Data Analysis Techniques from the TRACY Project
Pradeep Rangappa, Amanda Muscat, Alejandra Sanchez Lara, Petr Motlícek, Michaela Antonopoulou, Ioannis Fourfouris, Antonios Skarlatos, Nikos Avgerinos, Manolis Tsangaris, Kasia Kostka |
ICDF2C (1) | 4 |
| 2024 | Speech and Language Recognition with Low-rank Adaptation of Pretrained Models
Amrutha Prasad, Srikanth R. Madikeri, Driss Khalil, Petr Motlícek, Christof Schüpbach |
INTERSPEECH | 4 |
| 2024 | Reliability Estimation of News Media Sources: Birds of a Feather Flock TogetherabstractSergio Burdisso, Dairazalia Sanchez-cortes, Esaú Villatoro-tello, Petr Motlicek. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sergio Burdisso, Dairazalia Sanchez-Cortes, Esaú Villatoro-Tello, Petr Motlícek |
NAACL-HLT | 4 |
| 2023 | Parameter-Efficient Tuning with Adaptive Bottlenecks for Automatic Speech RecognitionabstractTransfer learning from large multilingual pretrained models, like XLSR, has become the new paradigm for Automatic Speech Recognition (ASR). Considering their ever-increasing size, fine-tuning all the weights has become impractical when the computing budget is limited. Adapters are lightweight trainable modules inserted between layers while the pre-trained part is kept frozen. They form a parameter-efficient fine-tuning method, but they still require a large bottleneck size to match standard fine-tuning performance. In this paper, we propose ABSADAPTER, a method to further reduce the parameter budget for equal task performance. Specifically, ABSADAPTER uses an Adaptive Bottleneck Scheduler to redistribute the adapter’s weights to the layers that need adaptation the most. By training only 8% of the XLSR model, ABSADAPTER achieves close to standard fine-tuning performance on a domain-shifted Air-Traffic Communication (ATC) ASR task. Geoffroy Vanderreydt, Amrutha Prasad, Driss Khalil, Srikanth R. Madikeri, Kris Demuynck, Petr Motlícek |
ASRU | 6 |
| 2023 | Effectiveness of Text, Acoustic, and Lattice-Based Representations in Spoken Language Understanding TasksabstractIn this paper, we perform an exhaustive evaluation of different representations to address the intent classification problem in a Spoken Language Understanding (SLU) setup. We benchmark three types of systems to perform the SLU intent detection task: 1) text-based, 2) lattice-based, and a novel 3) multimodal approach. Our work provides a comprehensive analysis of what could be the achievable performance of different state-of-the-art SLU systems under different circumstances, e.g., automatically- vs. manually-generated transcripts. We evaluate the systems on the publicly available SLURP spoken language resource corpus. Our results indicate that using richer forms of Automatic Speech Recognition (ASR) outputs, namely word-consensus-networks, allows the SLU system to improve in comparison to the 1-best setup (5.5% relative improvement). However, crossmodal approaches, i.e., learning from acoustic and text embeddings, obtains performance similar to the oracle setup, a relative improvement of 17.8% over the 1-best configuration, being a recommended alternative to overcome the limitations of working with automatically generated transcripts. Esaú Villatoro-Tello, Srikanth R. Madikeri, Juan Zuluaga-Gomez, Bidisha Sharma, Seyyed Saeed Sarfjoo, Iuliia Nigmatulina, Petr Motlícek, Alexei V. Ivanov, Aravind Ganapathiraju |
ICASSP | 7 |
| 2023 | Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical InterviewsabstractWe propose a simple approach for weighting selfconnecting edges in a Graph Convolutional Network (GCN) and show its impact on depression detection from transcribed clinical interviews.To this end, we use a GCN for modeling non-consecutive and long-distance semantics to classify the transcriptions into depressed or control subjects.The proposed method aims to mitigate the limiting assumptions of locality and the equal importance of self-connections vs. edges to neighboring nodes in GCNs, while preserving attractive features such as low computational cost, data agnostic, and interpretability capabilities.We perform an exhaustive evaluation in two benchmark datasets.Results show that our approach consistently outperforms the vanilla GCN model as well as previously reported results, achieving an F1=0.84 on both datasets.Finally, a qualitative analysis illustrates the interpretability capabilities of the proposed approach and its alignment with previous findings in psychology. Sergio Burdisso, Esaú Villatoro-Tello, Srikanth R. Madikeri, Petr Motlícek |
INTERSPEECH | 4 |
| 2023 | HyperConformer: Multi-head HyperMixer for Efficient Speech Recognition
Florian Mai, Juan Zuluaga-Gomez, Titouan Parcollet, Petr Motlícek |
INTERSPEECH | 4 |
| 2023 | Implementing Contextual Biasing in GPU Decoder for Online ASR
Iuliia Nigmatulina, Srikanth R. Madikeri, Esaú Villatoro-Tello, Petr Motlícek, Juan Zuluaga-Gomez, D. S. Karthik Pandia, Aravind Ganapathiraju |
INTERSPEECH | 4 |
| 2022 | A Two-Step Approach to Leverage Contextual Data: Speech Recognition in Air-Traffic CommunicationsabstractAutomatic Speech Recognition (ASR), as the assistance of speech communication between pilots and air-traffic controllers, can significantly reduce the complexity of the task and increase the reliability of transmitted information. ASR application can lead to a lower number of incidents caused by misunderstanding and improve air traffic management (ATM) efficiency. Evidently, high accuracy predictions, especially, of key information, i.e., callsigns and commands, are required to minimize the risk of errors. We prove that combining the benefits of ASR and Natural Language Processing (NLP) methods to make use of surveillance data (i.e. additional modality) helps to considerably improve the recognition of callsigns (named entity). In this paper, we investigate a two-step callsign boosting approach: (1) at the 1ststep (ASR), weights of probable callsign n-grams are reduced in G.fst and/or in the decoding FST (lattices), (2) at the 2ndstep (NLP), callsigns extracted from the improved recognition outputs with Named Entity Recognition (NER) are correlated with the surveillance data to select the most suitable one. Boosting callsign n-grams with the combination of ASR and NLP methods eventually leads up to 53.7% of an absolute, or 60.4% of a relative, improvement in callsign recognition. Iuliia Nigmatulina, Juan Zuluaga-Gomez, Amrutha Prasad, Seyyed Saeed Sarfjoo, Petr Motlícek |
ICASSP | 5 |
| 2022 | An End-to-End Multilingual System for Automatic Minuting of Multi-Party Dialogues
Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh, Petr Motlícek |
PACLIC | 4 |
| 2022 | Bio-Medical Multi-label Scientific Literature Classification using LWAN and Dual-attention module
Deepanshu Khanna, Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh, Petr Motlícek |
PACLIC | 5 |
| 2022 | Expanded Lattice Embeddings for Spoken Document Retrieval on Informal MeetingsabstractIn this paper, we evaluate different alternatives to process richer forms of Automatic Speech Recognition (ASR) output based on lattice expansion algorithms for Spoken Document Retrieval (SDR). Typically, SDR systems employ ASR transcripts to index and retrieve relevant documents. However, ASR errors negatively affect the retrieval performance. Multiple alternative hypotheses can also be used to augment the input to document retrieval to compensate for the erroneous one-best hypothesis. In Weighted Finite State Transducer-based ASR systems, using the n-best output (i.e. the top "n'' scoring hypotheses) for the retrieval task is common, since they can easily be fed to a traditional Information Retrieval (IR) pipeline. However, the n-best hypotheses are terribly redundant, and do not sufficiently encapsulate the richness of the ASR output, which is represented as an acyclic directed graph called the lattice. In particular, we utilize the lattice's constrained minimum path cover to generate a minimum set of hypotheses that serve as input to the reranking phase of IR. The novelty of our proposed approach is the incorporation of the lattice as an input for neural reranking by considering a set of hypotheses that represents every arc in the lattice. The obtained hypotheses are encoded through sentence embeddings using BERT-based models, namely SBERT and RoBERTa, and the final ranking of the retrieved segments is obtained with a max-pooling operation over the computed scores among the input query and the hypotheses set. We present our evaluation on the publicly available AMI meeting corpus. Our results indicate that the proposed use of hypotheses from the expanded lattice improves the SDR performance significantly over the n-best ASR output. Esaú Villatoro-Tello, Srikanth R. Madikeri, Petr Motlícek, Aravind Ganapathiraju, Alexei V. Ivanov |
SIGIR | 3 |
| 2022 | How Does Pre-Trained Wav2Vec 2.0 Perform on Domain-Shifted Asr? an Extensive Benchmark on Air Traffic Control CommunicationsabstractRecent work on self-supervised pre-training focus on leveraging large-scale unlabeled speech data to build robust end-to-end (E2E) acoustic models (AM) that can be later fine-tuned on downstream tasks e.g., automatic speech recognition (ASR). Yet, few works investigated the impact on performance when the data properties substantially differ between the pre-training and fine-tuning phases, termed domain shift. We target this scenario by analyzing the robustness of Wav2Vec 2.0 and XLS-R models on downstream ASR for a completely unseen domain, air traffic control (ATC) communications. We benchmark these two models on several open-source and challenging ATC databases with signal-to-noise ratio between 5 to 20 dB. Relative word error rate (WER) reductions between 20% to 40% are obtained in comparison to hybrid-based ASR baselines by only fine-tuning E2E acoustic models with a smaller fraction of labeled data. We analyze WERs on the low-resource scenario and gender bias carried by one ATC dataset. Juan Zuluaga-Gomez, Amrutha Prasad, Iuliia Nigmatulina, Seyyed Saeed Sarfjoo, Petr Motlícek, Matthias Kleinert, Hartmut Helmke, Oliver Ohneiser, Qingran Zhan |
SLT | 5 |
| 2022 | Bertraffic: Bert-Based Joint Speaker Role and Speaker Change Detection for Air Traffic Control CommunicationsabstractAutomatic speech recognition (ASR) allows transcribing the communications between air traffic controllers (ATCOs) and aircraft pilots. The transcriptions are used later to extract ATC named entities, e.g., aircraft callsigns. One common challenge is speech activity detection (SAD) and speaker diarization (SD). In the failure condition, two or more segments remain in the same recording, jeopardizing the overall performance. We propose a system that combines SAD and a BERT model to perform speaker change detection and speaker role detection (SRD) by chunking ASR transcripts, i.e., SD with a defined number of speakers together with SRD. The proposed model is evaluated on real-life public ATC databases. Our BERT SD model baseline reaches up to 10% and 20% token-based Jaccard error rate (JER) in public and private ATC databases. We also achieved relative improvements of 32% and 7.7% in JERs and SD error rate (DER), respectively, compared to VBx, a well-known SD system.11Our code is stored in the following public GitHub repository: https://github.com/idiap/bert-text-diarization-atc Juan Zuluaga-Gomez, Seyyed Saeed Sarfjoo, Amrutha Prasad, Iuliia Nigmatulina, Petr Motlícek, Karel Ondrej, Oliver Ohneiser, Hartmut Helmke |
SLT | 5 |
| 2021 | A Comparison of Methods for OOV-Word Recognition on a New Public DatasetabstractA common problem for automatic speech recognition systems is how to recognize words that they did not see during training. Currently there is no established method of evaluating different techniques for tackling this problem.We propose using the CommonVoice dataset to create test sets for multiple languages which have a high out-of-vocabulary (OOV) ratio relative to a training set and release a new tool for calculating relevant performance metrics. We then evaluate, within the context of a hybrid ASR system, how much better subword models are at recognizing OOVs, and how much benefit one can get from incorporating OOV-word information into an existing system by modifying WFSTs. Additionally, we propose a new method for modifying a subword-based language model so as to better recognize OOV-words. We showcase very large improvements in OOV-word recognition and make both the data and code available. Rudolf A. Braun, Srikanth R. Madikeri, Petr Motlícek |
ICASSP | 3 |
| 2021 | ROXANNE Research Platform: Automate Criminal Investigations
Maël Fabien, Shantipriya Parida, Petr Motlícek, Aravind Krishnan |
Interspeech | 3 |
| 2021 | Multi-Task Neural Network for Robust Multiple Speaker Embedding ExtractionabstractThis paper introduces a novel approach for extracting speaker embeddings from audio mixtures of multiple overlapping voices. This approach is based on a multi-task neural network. The network first extracts a latent feature for each direction. This feature is used for detecting sound sources as well as identifying speakers. In contrast to traditional approaches, the proposed method does not rely on explicit sound source separation. The neural network model learns from data to extract the most suitable features of the sounds at different directions. The experiments using audio recordings of overlapping sound sources show that the proposed approach outperforms a beamforming-based traditional method. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
Interspeech | 2 |
| 2021 | Boosting of Contextual Information in ASR for Air-Traffic Call-Sign RecognitionabstractContextual adaptation of ASR can be very beneficial for multi-accent and often noisy Air-Traffic Control (ATC) speech. Our focus is call-sign recognition, which can be used to track conversations of ATC operators with individual airplanes. We developed a two-stage boosting strategy, consisting of HCLG boosting and Lattice boosting. Both are implemented as WFST compositions and the contextual information is specific to each utterance. In HCLG boosting we give score discounts to individual words, while in Lattice boosting the score discounts are given to word sequences. The context data have origin in surveillance database of OpenSky Network. From this, we obtain lists of call-signs that are made more likely to appear in the best hypothesis of ASR. This also improves the accuracy of the NLU module that recognizes the call-signs from the best hypothesis of ASR. Martin Kocour, Karel Veselý, Alexander Blatt, Juan Zuluaga-Gomez, Igor Szöke, Jan Cernocký, Dietrich Klakow, Petr Motlícek |
Interspeech | 8 |
| 2021 | Multitask Adaptation with Lattice-Free MMI for Multi-Genre Speech Recognition of Low Resource LanguagesabstractIn this paper, we develop Automatic Speech Recognition (ASR) systems for multi-genre speech recognition of low-resource languages where training data is predominantly conversational speech but test data can be in one of the following genres: news broadcast, topical broadcast and conversational speech. ASR for low-resource languages is often developed by adapting a pre-trained model to a target language. When training data is predominantly from one genre and limited, the system's performance for other genres suffer. To handle such out-of-domain scenarios, we employ multitask adaptation by using auxiliary conversational speech data from other languages in addition to the target-language data. We aim to (1) improve adaptation through implicit data augmentation by adding other languages as auxiliary tasks, and (2) prevent the acoustic model from overfitting to the dominant genre in the training set. Pre-trained parameters are obtained from a multilingual model trained with data from 18 languages using the Lattice-Free Maximum Mutual Information (LF-MMI) criterion. The adaptation is performed with the LF-MMI criterion. We present results on MATERIAL datasets for three languages: Kazakh and Farsi and Pashto. Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard |
Interspeech | 2 |
| 2021 | Robust Command Recognition for Lithuanian Air Traffic Control Tower UtterancesabstractThe maturity of automatic speech recognition (ASR) systems at controller working positions is currently a highly relevant technological topic in air traffic control (ATC). However, ATC service providers are less interested in pure word error rate (WER). They want to see benefits of ASR applications for ATC. Such applications transform recognized word sequences into semantic meanings, i.e., a number of related concepts such as callsign, type, value, unit, etc., which are combined to form commands. Digitized concepts or recognized commands can enter ATC systems based on an ontology for utterance annotation agreed between European ATC stakeholders. Command recognition (CR) has already been performed in approach control. However, spoken utterances of tower controllers are longer, include more free speech, and contain other command types than in approach. An automatic CR rate of 95.8% is achievable on perfect word recognition, i.e., manually transcribed audio recordings (gold transcriptions), taken from Lithuanian controllers in a multiple remote tower environment. This paper presents CR results for various speech-to-text models with different WERs on tower utterances. Although WERs were around 9%, we achieve CR rates of 85%. CR rates only slightly decrease with higher WERs, which enables to bring ASR applications closer to operational ATC environment. Oliver Ohneiser, Seyyed Saeed Sarfjoo, Hartmut Helmke, Shruthi Shetty, Petr Motlícek, Matthias Kleinert, Heiko Ehr, Sarunas Murauskas |
Interspeech | 5 |
| 2021 | Speech Activity Detection Based on Multilingual Speech Recognition SystemabstractTo better model the contextual information and increase the generalization ability of Speech Activity Detection (SAD) system, this paper leverages a multi-lingual Automatic Speech Recognition (ASR) system to perform SAD. Sequence discriminative training of Acoustic Model (AM) using Lattice-Free Maximum Mutual Information (LF-MMI) loss function, effectively extracts the contextual information of the input acoustic frame. Multi-lingual AM training, causes the robustness to noise and language variabilities. The index of maximum output posterior is considered as a frame-level speech/non-speech decision function. Majority voting and logistic regression are applied to fuse the language-dependent decisions. The multi-lingual ASR is trained on 18 languages of BABEL datasets and the built SAD is evaluated on 3 different languages. On out-of-domain datasets, the proposed SAD model shows significantly better performance with respect to baseline models. On the Ester2 dataset, without using any in-domain data, this model outperforms the WebRTC, phoneme recognizer based VAD (Phn Rec), and Pyannote baselines (respectively by 7.1, 1.7, and 2.7% absolute) in Detection Error Rate (DetER) metrics. Similarly, on the LiveATC dataset, this model outperforms the WebRTC, Phn Rec, and Pyannote baselines (respectively by 6.4, 10.0, and 3.7% absolutely) in DetER metrics. Seyyed Saeed Sarfjoo, Srikanth R. Madikeri, Petr Motlícek |
Interspeech | 3 |
| 2021 | Late Fusion of the Available Lexicon and Raw Waveform-Based Acoustic Modeling for Depression and Dementia RecognitionabstractMental disorders, e.g. depression and dementia, are categorized as priority conditions according to the World Health Organization (WHO). When diagnosing, psychologists employ structured questionnaires/interviews, and different cognitive tests. Although accurate, there is an increasing necessity of developing digital mental health support technologies to alleviate the burden faced by professionals. In this paper, we propose a multi-modal approach for modeling the communication process employed by patients being part of a clinical interview or a cognitive test. The language-based modality, inspired by the Lexical Availability (LA) theory from psycho-linguistics, identifies the most accessible vocabulary of the interviewed subject and use it as features in a classification process. The acoustic-based modality is processed by a Convolutional Neural Network (CNN) trained on signals of speech that predominantly contained voice source characteristics. In the end, a late fusion technique, based on majority voting, assigns the final classification. Results show the complementarity of both modalities, reaching an overall Macro-F1 of 84% and 90% for Depression and Alzheimer's dementia respectively. Esaú Villatoro-Tello, S. Pavankumar Dubagunta, Julian Fritsch, Gabriela Ramírez-de-la-Rosa, Petr Motlícek, Mathew Magimai-Doss |
Interspeech | 5 |
| 2021 | Contextual Semi-Supervised Learning: An Approach to Leverage Air-Surveillance and Untranscribed ATC Data in ASR SystemsabstractAir traffic management and specifically air-traffic control (ATC) rely mostly on voice communications between Air Traffic Controllers (ATCos) and pilots. In most cases, these voice communications follow a well-defined grammar that could be leveraged in Automatic Speech Recognition (ASR) technologies. The callsign used to address an airplane is an essential part of all ATCo-pilot communications. We propose a two-step approach to add contextual knowledge during semi-supervised training to reduce the ASR system error rates at recognizing the part of the utterance that contains the callsign. Initially, we represent in a WEST the contextual knowledge (i.e. air-surveillance data) of an ATCo-pilot communication. Then, during Semi-Supervised Learning (SSL) the contextual knowledge is added by second-pass decoding (i.e. lattice re-scoring). Results show that 'unseen domains' (e.g. data from airports not present in the supervised training data) are further aided by contextual SSL when compared to standalone SSL. For this task, we introduce the Callsign Word Error Rate (CA-WER) as an evaluation metric, which only assesses ASR performance of the spoken callsign in an utterance. We obtained a 32.1% CA-WER relative improvement applying SSL with an additional 17.5% CA-WER improvement by adding contextual knowledge during SSL on a challenging ATC-based test set gathered from LiveATC. Juan Zuluaga-Gomez, Iuliia Nigmatulina, Amrutha Prasad, Petr Motlícek, Karel Veselý, Martin Kocour, Igor Szöke |
Interspeech | 4 |
| 2021 | IEEE SLT 2021 Alpha-Mini Speech Challenge: Open Datasets, Tracks, Rules and BaselinesabstractThe IEEE Spoken Language Technology Workshop (SLT) 2021 Alpha-mini Speech Challenge (ASC) is intended to improve research on keyword spotting (KWS) and sound source location (SSL) on humanoid robots. Many publications report significant improvements in deep learning based KWS and SSL on open source datasets in recent years. For deep learning model training, it is necessary to expand the data coverage to improve the model robustness. Thus, simulating multi-channel noisy and reverberant data from single-channel speech, noise, echo and room impulsive response (RIR) is widely adopted. However, this approach may generate mismatch between simulated data and recorded data in real application scenarios, especially echo data. In this challenge, we open source a sizable speech, keyword, echo and noise corpus for promoting data-driven methods, particularly deep-learning approaches on KWS and SSL. We also choose Alpha-mini, a humanoid robot produced by UBTECH equipped with a built-in four-microphone array on its head, to record development and evaluation sets under the actual Alpha-mini robot application scenario, including environ-mental noise as well as echo and mechanical noise generated by the robot itself for model evaluation. Furthermore, we illustrate the rules, evaluation methods and baselines for re-searchers to quickly assess their achievements and optimize their models. Yihui Fu, Zhuoyuan Yao, Weipeng He, Jian Wu 0027, Zhanheng Yang, Lei Xie 0001, Dong-Yan Huang, Hui Bu, Petr Motlícek, Jean-Marc Odobez |
SLT | 11 |
| 2021 | Neural Network Adaptation and Data Augmentation for Multi-Speaker Direction-of-Arrival EstimationabstractDeep neural networks have been successfully applied to sound direction-of-arrival estimation under challenging conditions. However, such a learning-based approach requires a large amount of labeled training data, which is difficult to acquire. To address this problem, we propose a novel approach for multi-speaker direction-of-arrival estimation with data augmentation and weakly-supervised domain adaptation. We generate source domain data with simulation, and collect real data annotated with the number of sound sources as the weak labels. The real data are further augmented by mixing single-source segments. Then, weakly-supervised domain adaptation is applied to models pre-trained on the simulated data. We define a loss function for the adaptation process which exploits the weak labels and the mixture component information in the augmented data. Experiments with real robot audio data show that our proposed approach achieves similar performance as if the fully-labeled real data are used. This paper suggests an effective development procedure for DOA estimation models applied to new types of microphone arrays with minimal data collection efforts. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Incremental Semi-Supervised Learning for Multi-Genre Speech RecognitionabstractIn this work, we explore a data scheduling strategy for semi-supervised learning (SSL) for acoustic modeling in automatic speech recognition. The conventional approach uses a seed model trained with supervised data to automatically recognize the entire set of unlabeled (auxiliary) data to generate new labels for subsequent acoustic model training. In this paper, we propose an approach in which the unlabelled set is divided into multiple equal-sized subsets. These subsets are processed in an incremental fashion: for each iteration a new subset is added to the data used for SSL, starting from only one subset in the first iteration. The acoustic model from the previous iteration becomes the seed model for the next one. This scheduling strategy is compared to the approach employing all unlabeled data in one-shot for training. Experiments using lattice-free maximum mutual information based acoustic model training on Fisher English gives 80% word error recovery rate. On the multi-genre evaluation sets on Lithuanian and Bulgarian relative improvements of up to 17.2% in word error rate are observed. Banriskhem K. Khonglah, Srikanth R. Madikeri, Subhadeep Dey, Hervé Bourlard, Petr Motlícek, Jayadev Billa |
ICASSP | 5 |
| 2020 | Lattice-Free Maximum Mutual Information Training of Multilingual Speech Recognition SystemsabstractMultilingual acoustic model training combines data from multiple languages to train an automatic speech recognition system.Such a system is beneficial when training data for a target language is limited.Lattice-Free Maximum Mutual Information (LF-MMI) training performs sequence discrimination by introducing competing hypotheses through a denominator graph in the cost function.The standard approach to train a multilingual model with LF-MMI is to combine the acoustic units from all languages and use a common denominator graph.The resulting model is either used as a feature extractor to train an acoustic model for the target language or directly fine-tuned.In this work, we propose a scalable approach to train the multilingual acoustic model using a typical multitask network for the LF-MMI framework.A set of language-dependent denominator graphs is used to compute the cost function.The proposed approach is evaluated under typical multilingual ASR tasks using GlobalPhone and BABEL datasets.Relative improvements up to 13.2% in WER are obtained when compared to the corresponding monolingual LF-MMI baselines.The implementation is made available as a part of the Kaldi speech recognition toolkit. Srikanth R. Madikeri, Banriskhem K. Khonglah, Sibo Tong, Petr Motlícek, Hervé Bourlard, Daniel Povey |
INTERSPEECH | 4 |
| 2020 | Supervised Domain Adaptation for Text-Independent Speaker Verification Using Limited DataabstractTo adapt the speaker verification (SV) system to a target domain with limited data, this paper investigates the transfer learning of the model pre-trained on the source domain data.To that end, layer-by-layer adaptation with transfer learning from the initial and final layers of the pre-trained model is investigated.We show that the model adapted from the initial layers outperforms the model adapted from the final layers.Based on this evidence, and inspired by the works in image recognition field, we hypothesize that low-level convolutional neural network (CNN) layers characterize domain-specific component while high-level CNN layers are domain-independent and have more discriminative power.For adapting these domain-specific components, angular margin softmax (AMSoftmax) applied on the CNN-based implementation of the x-vector architecture.In addition, to reduce the problem of over-fitting on the limited target data, transfer learning on the batch norm layers is investigated.Mean shift and covariance estimation of batch norm allows to map the represented components of the target domain to the source domain.Using TDNN and E-TDNN versions of the x-vectors as baseline models, the adapted models on the development set of NIST SRE 2018 outperformed the baselines with relative improvements of 11.0 and 13.8 %, respectively. Seyyed Saeed Sarfjoo, Srikanth R. Madikeri, Petr Motlícek, Sébastien Marcel |
INTERSPEECH | 3 |
| 2020 | Automatic Speech Recognition Benchmark for Air-Traffic CommunicationsabstractAdvances in Automatic Speech Recognition (ASR) over the last decade opened new areas of speech-based automation such as in Air-Traffic Control (ATC) environments. Currently, voice communication and Controller Pilot Data Link Communications are the only way of contact between pilots and Air-Traffic Controllers (ATCo), where the former is the most widely used and the latter is a non-speech method mandatory for oceanic messages and limited for some domestically issues. ASR systems on ATCo environments inherit increasing complexity due to accents from non-English speakers, cockpit noise, speaker-dependent biases and small in-domain ATC databases for training. In this paper, we review the last advances related to ASR on ATCo communication. Then, we introduce CleanSky EC H2020 ATCO2, a project that aims to develop a platform to collect, organize and automatically pre-process ATCo data from air space. We apply transfer learning from out-of-domain corpus coupled with adaptation on seven command-related corpora. The acoustic modelling is based on conventional TDNN-HMMs trained using lattice-free MMI objective function. The developed ASR achieves relative improvement in word error rates of 29% when using transfer learning and an additional 36% when adapting the model with seven command-related databases, these results obtained from EC H2020 SESAR project MALORCA Vienna database. Juan Zuluaga-Gomez, Petr Motlícek, Qingran Zhan, Karel Veselý, Rudolf A. Braun |
INTERSPEECH | 2 |
| 2020 | The MuMMER Data Set for Robot Perception in Multi-party HRI ScenariosabstractThis paper presents the MuMMER data set, a data set for human-robot interaction scenarios that is available for research purposes1. It comprises 1h 29 min of multimodal recordings of people interacting with the social robot Pepper in entertainment scenarios, such as quiz, chat, and route guidance. In the 33 clips (of 1 to 4 min long) recorded from the robot point of view, the participants are interacting with the robot in an unconstrained manner.The data set exhibits interesting features and difficulties, such as people leaving the field of view, robot moving (head rotation with embedded camera in the head), different illumination conditions. The data set contains color and depth videos from a Kinect v2, an Intel D435, and the video from Pepper.All the visual faces and the identities in the data set were manually annotated, making the identities consistent across time and clips. The goal of the data set is to evaluate perception algorithms in multi-party human/robot interaction, in particular the re-identification part when a track is lost, as this ability is crucial for keeping the dialog history. The data set can easily be extended with other types of annotations.We also present a benchmark on this data set that should serve as a baseline for future comparison. The baseline system, IHPER2(Idiap Human Perception system) is available for research and is evaluated on the MuMMER data set. We show that an identity precision and recall of ~80% and a MOTA score above 80% are obtained. Olivier Canévet, Weipeng He, Petr Motlícek, Jean-Marc Odobez |
RO-MAN | 3 |
| 2019 | Abstract Text Summarization: A Low Resource ChallengeabstractShantipriya Parida, Petr Motlicek. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shantipriya Parida, Petr Motlícek |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Adaptation of Multiple Sound Source Localization Neural Networks with Weak Supervision and Domain-adversarial TrainingabstractDespite the recent success of deep neural network-based approaches in sound source localization, these approaches suffer the limitations that the required annotation process is costly, and the mismatch between the training and test conditions undermines the performance. This paper addresses the question of how models trained with simulation can be exploited for multiple sound source localization in real scenarios by domain adaptation. In particular, two domain adaptation methods are investigated: weak supervision and domain-adversarial training. Our experiments show that the weak supervision with the knowledge of the number of sources can significantly improve the performance of an unadapted model. However, the domain-adversarial training does not yield significant improvement for this particular problem. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
ICASSP | 2 |
| 2019 | A Bayesian Approach to Inter-task Fusion for Speaker RecognitionabstractIn i-vector based speaker recognition systems, back-end classifiers are trained to factor out nuisance information and retain only the speaker identity. As a result, variabilities arising due to gender, language and accent (among many others) are suppressed. Inter-task fusion, in which such metadata information obtained from automatic systems is used, has been shown to improve speaker recognition performance. In this paper, we explore a Bayesian approach towards inter-task fusion. Speaker similarity score for a test recording is obtained by marginalizing the posterior probability of a speaker. Gender and language probabilities for the test audio are combined with speaker posteriors to obtain a final speaker score. The proposed approach is demonstrated for speaker verification and speaker identification tasks on the NIST SRE 2008 dataset. Relative improvements of up to 10% and 8% are obtained when fusing gender and language information, respectively. Srikanth R. Madikeri, Petr Motlícek, Subhadeep Dey |
ICASSP | 2 |
| 2019 | Exploiting Semi-Supervised Training Through a Dropout Regularization in End-to-End Speech RecognitionabstractIn this paper, we explore various approaches for semi supervised learning in an end to end automatic speech recognition (ASR) framework. The first step in our approach involves training a seed model on the limited amount of labelled data. Additional unlabelled speech data is employed through a data selection mechanism to obtain the best hypothesized output, further used to retrain the seed model. However, uncertainties of the model may not be well captured with a single hypothesis. As opposed to this technique, we apply a dropout mechanism to capture the uncertainty by obtaining multiple hypothesized text transcripts of an speech recording. We assume that the diversity of automatically generated transcripts for an utterance will implicitly increase the reliability of the model. Finally, the data selection process is also applied on these hypothesized transcripts to reduce the uncertainty. Experiments on freely available TEDLIUM corpus and proprietary Adobe's internal dataset show that the proposed approach significantly reduces ASR errors, compared to the baseline model. Subhadeep Dey, Petr Motlícek, Trung Bui, Franck Dernoncourt |
INTERSPEECH | 2 |
| 2019 | End-to-End Accented Speech Recognition
Thibault Viglino, Petr Motlícek, Milos Cernak |
INTERSPEECH | 2 |
| 2018 | DNN Based Speaker Embedding Using Content Information for Text-Dependent Speaker VerificationabstractIn this paper, we are interested in exploring Deep Neural Network (DNN) based speaker embedding for Random-digit task using content information. To this end, a technique is applied to automatically select common phonetic units between the enrollment and test data to produce speaker verification scores. Furthermore, a novel approach is proposed to incorporate content information in the DNN directly. It is hypothesized that features extracted using this DNN will be helpful for the task. Experiments on the RSR dataset show that the proposed method outperforms the baseline i-vector system by 43% relative equal error rate. Subhadeep Dey, Takafumi Koshinaka, Petr Motlícek, Srikanth R. Madikeri |
ICASSP | 3 |
| 2018 | Deep Neural Networks for Multiple Speaker Detection and LocalizationabstractWe propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization methods require fewer strong assumptions about the environment. Previous neural network-based methods have been focusing on localizing a single sound source, which do not extend to multiple sources in terms of detection and localization. In this paper, we thus propose a likelihood-based encoding of the network output, which naturally allows the detection of an arbitrary number of sources. In addition, we investigate the use of sub-band cross-correlation information as features for better localization in sound mixtures, as well as three different network architectures based on different motivations. Experiments on real data recorded from a robot show that our proposed methods significantly outperform the popular spatial spectrum-based approaches. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
ICRA | 2 |
| 2018 | End-to-end Text-dependent Speaker Verification Using Novel Distance MeasuresabstractThis paper explores novel ideas in building end-to-end deep neural network (DNN) based text-dependent speaker verification (SV) system. The baseline approach consists of mapping a variable length speech segment to a fixed dimensional speaker vector by estimating the mean of hidden representations in DNN structure. The distance between two utterances is obtained by computing L2 norm between the vectors. This approach performs worse than the conventional Gaussian Mixture Model-Universal Background Model (GMM-UBM) based SV on a publicly available corpora. We believe that a degraded performance is due to the employed averaging operation, which may not capture the phonetic information of an utterance. Recent studies indicate that techniques exploiting phonetic information in addition to speaker is beneficial for this task. This paper therefore proposes to incorporate content information of the speech signal by computing distance function with linguistic units co-occuring between enrollment and test data. The whole network is optimized by employing a triplet-loss objective in an end-to-end fashion to estimate SV scores. Experiments on the RSR2015 dataset indicate that the proposed approach outperforms GMM-UBM system by 48% and 36% relative equal error rate for fixed-phrase and random-digit conditions respectively. Subhadeep Dey, Srikanth R. Madikeri, Petr Motlícek |
INTERSPEECH | 3 |
| 2018 | Joint Localization and Classification of Multiple Sound Sources Using a Multi-task Neural NetworkabstractWe propose a novel multi-task neural network-based approach for joint sound source localization and speech/non-speech classification in noisy environments. The network takes raw short time Fourier transform as input and outputs the likelihood values for the two tasks, which are used for the simultaneous detection, localization and classification of an unknown number of overlapping sound sources, Tested with real recorded data, our method achieves significantly better performance in terms of speech/non-speech classification and localization of speech sources, compared to method that performs localization and classification separately. In addition, we demonstrate that incorporating the temporal context can further improve the performance. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
INTERSPEECH | 2 |
| 2018 | Analysis of Language Dependent Front-End for Speaker RecognitionabstractIn Deep Neural Network (DNN) i-vector based speaker recognition systems, acoustic models trained for Automatic Speech Recognition are employed to estimate sufficient statistics for i-vector modeling. The DNN based acoustic model is typically trained on a wellresourced language like English. In evaluation conditions where enrollment and test data are not in English, as in the NIST SRE 2016 dataset, a DNN acoustic model generalizes poorly. In such conditions, a conventional Universal Background Model/Gaussian Mixture Model (UBM/GMM) based i-vector extractor performs better than the DNN based i-vector system. In this paper, we address the scenario in which one can develop a Automatic Speech Recognizer with limited resources for a language present in the evaluation condition, thus enabling the use of a DNN acoustic model instead of UBM/GMM. Experiments are performed on the Tagalog subset of the NIST SRE 2016 dataset assuming an open training condition. With a DNN i-vector system trained for Tagalog, a relative improvement of 12.1% is obtained over a baseline system trained for English. Srikanth R. Madikeri, Subhadeep Dey, Petr Motlícek |
INTERSPEECH | 3 |
| 2018 | Iterative Learning of Speech Recognition Models for Air Traffic ControlabstractAutomatic Speech Recognition (ASR) has recently proved to \nbe a useful tool to reduce the workload of air traffic controllers \nleading to significant gains in operational efficiency. Air Traffic Control (ATC) systems in operation rooms around the world \ngenerate large amounts of untranscribed speech and radar data \neach day, which can be utilized to build and improve ASR models. In this paper, we propose an iterative approach that utilizes \nincreasing amounts of untranscribed data to incrementally build \nthe necessary ASR models for an ATC operational area. Our approach uses a semi-supervised learning framework to combine \nspeech and radar data to iteratively update the acoustic model, \nlanguage model and command prediction model (i.e. prediction \nof possible commands from radar data for a given air traffic \nsituation) of an ASR system. Starting with seed models built \nwith a limited amount of manually transcribed data, we simulate an operational scenario to adapt and improve the models \nthrough semi-supervised learning. Experiments on two independent ATC areas (Vienna and Prague) demonstrate the utility \nof our proposed methodology that can scale to operational environments with minimal manual effort for learning and adaptation. Ajay Srinivasamurthy, Petr Motlícek, Mittul Singh, Youssef Oualil, Matthias Kleinert, Heiko Ehr, Hartmut Helmke |
INTERSPEECH | 2 |
| 2017 | A context-aware speech recognition and understanding system for air traffic control domainabstractAutomatic Speech Recognition and Understanding (ASRU) systems can generally use temporal and situational context information to improve their performance for a given task. This is typically done by rescoring the ASR hypotheses or by dynamically adapting the ASR models. For some domains, such as Air Traffic Control (ATC), this context information can be, however, small in size, partial and available only as abstract concepts (e.g. airline codes), which are difficult to map into full possible spoken sentences to perform rescoring or adaptation. This paper presents a multi-modal ASRU system, which dynamically integrates partial temporal and situational ATC context information to improve its performance. This is done either by 1) extracting word sequences which carry relevant ATC information from ASR N-best Lists and then perform a context-based rescoring on the extracted ATC segments or 2) by a partial adaptation of the language model. Experiments conducted on 4 hours of test data from Prague and Vienna approach (arrivals) showed a relative reduction of the ATC command error rate metric by 30% to 50%. Youssef Oualil, Dietrich Klakow, György Szaszák, Ajay Srinivasamurthy, Hartmut Helmke, Petr Motlícek |
ASRU | 6 |
| 2017 | Exploiting sequence information for text-dependent Speaker VerificationabstractModel-based approaches to Speaker Verification (SV), such as Joint Factor Analysis (JFA), i-vector and relevance Maximum-a-Posteriori (MAP), have shown to provide state-of-the-art performance for text-dependent systems with fixed phrases. The performance of i-vector and JFA models has been further enhanced by estimating posteriors from Deep Neural Network (DNN) instead of Gaussian Mixture Model (GMM). While both DNNs and GMMs aim at incorporating phonetic information of the phrase with these posteriors, model-based SV approaches ignore the sequence information of the phonetic units of the phrase. In this paper, we tackle this issue by applying dynamic time warping using speaker-informative features. We propose to use i-vectors computed from short segments of each speech utterance, also called online i-vectors, as feature vectors. The proposed approach is evaluated on the RedDots database and provides an improvement of 75% relative equal error rate over the best model-based SV baseline system in a content-mismatch condition. Subhadeep Dey, Petr Motlícek, Srikanth R. Madikeri, Marc Ferras |
ICASSP | 2 |
| 2017 | Intra-class covariance adaptation in PLDA back-ends for speaker verificationabstractMulti-session training conditions are becoming increasingly common in recent benchmark datasets for both text-independent and text-dependent speaker verification. In the state-of-the-art i-vector framework for speaker verification, such conditions are addressed by simple techniques such as averaging the individual i-vectors, averaging scores, or modifying the Probabilistic Linear Discriminant Analysis (PLDA) scoring hypothesis for multi-session enrollment. The aforementioned techniques fail to exploit the speaker variabilities observed in the enrollment data for target speakers. In this paper, we propose to exploit the multi-session training data by estimating a speaker-dependent covariance matrix and updating the intra-speaker covariance during PLDA scoring for each target speaker. The proposed method is further extended by combining covariance adaptation and score averaging. In this method, the individual examples of the target speaker are compared against the test data as opposed to an averaged i-vector, and the scores obtained are then averaged. The proposed methods are evaluated on the NIST SRE 2012 dataset. Relative improvements of up to 29% in equal error rate are obtained. Srikanth R. Madikeri, Marc Ferras, Petr Motlícek, Subhadeep Dey |
ICASSP | 3 |
| 2017 | Content Normalization for Text-Dependent Speaker VerificationabstractSubspace based techniques, such as i-vector and Joint Factor Analysis (JFA) have shown to provide state-of-the-art performance for fixed phrase based text-dependent speaker verification.However, the error rates of such systems on the random digit task of RSR dataset are higher than that of Gaussian Mixture Model-Universal Background Model (GMM-UBM).In this paper, we aim at improving i-vector system by normalizing the content of the enrollment data to match the test data.We estimate i-vectors for each frames of a speech utterance (also called online i-vectors).The largest similarity scores across frames between enrollment and test are taken using these online i-vectors to obtain speaker verification scores.Experiments on Part3 of RSR corpora show that the proposed approach achieves 12% relative improvement in equal error rate over a GMM-UBM based baseline system. Subhadeep Dey, Srikanth R. Madikeri, Petr Motlícek, Marc Ferras |
INTERSPEECH | 3 |
| 2017 | Semi-Supervised Learning with Semantic Knowledge Extraction for Improved Speech Recognition in Air Traffic ControlabstractAutomatic Speech Recognition (ASR) can introduce higher levels \nof automation into Air Traffic Control (ATC), where spoken \nlanguage is still the predominant form of communication. \nWhile ATC uses standard phraseology and a limited vocabulary, \nwe need to adapt the speech recognition systems to local \nacoustic conditions and vocabularies at each airport to reach \noptimal performance. Due to continuous operation of ATC systems, \na large and increasing amount of untranscribed speech \ndata is available, allowing for semi-supervised learning methods \nto build and adapt ASR models. In this paper, we first identify \nthe challenges in building ASR systems for specific ATC \nareas and propose to utilize out-of-domain data to build baseline \nASR models. Then we explore different methods of data \nselection for adapting baseline models by exploiting the continuously \nincreasing untranscribed data. We develop a basic approach \ncapable of exploiting semantic representations of ATC \ncommands. We achieve relative improvement in both word error \nrate (23.5%) and concept error rates (7%) when adapting \nASR models to different ATC conditions in a semi-supervised \nmanner. Ajay Srinivasamurthy, Petr Motlícek, Ivan Himawan, György Szaszák, Youssef Oualil, Hartmut Helmke |
INTERSPEECH | 2 |
| 2017 | Template-matching for text-dependent speaker verification
Subhadeep Dey, Petr Motlícek, Srikanth R. Madikeri, Marc Ferras |
Speech Commun. | 2 |
| 2016 | Deep neural network based posteriors for text-dependent speaker verificationabstractThe i-vector and Joint Factor Analysis (JFA) systems for text-dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statistics average the acoustic information over the utterance thereby losing all the sequence information. In this paper, we study explicit content matching using Dynamic Time Warping (DTW) and present the best achievable error rates for speaker-dependent and speaker-independent content matching. For this purpose, a Deep Neural Network/Hidden Markov Model Automatic Speech Recognition (DNN/HMM ASR) system is used to extract content-related posterior probabilities. This approach outperforms systems using Gaussian mixture model posteriors by at least 50% Equal Error Rate (EER) on the RSR2015 in content mismatch trials. DNN posteriors are also used in i-vector and JFA systems, obtaining EERs as low as 0.02%. Subhadeep Dey, Srikanth R. Madikeri, Marc Ferras, Petr Motlícek |
ICASSP | 4 |
| 2016 | Information theoretic clustering for unsupervised domain-adaptationabstractThe aim of the domain-adaptation task for speaker verification is to exploit unlabelled target domain data by using the labelled source domain data effectively. The i-vector based Probabilistic Linear Discriminant Analysis (PLDA) framework approaches this task by clustering the target domain data and using each cluster as a unique speaker to estimate PLDA model parameters. These parameters are then combined with the PLDA parameters from the source domain. Typically, agglomerative clustering with cosine distance measure is used. In tasks such as speaker diarization that also require unsupervised clustering of speakers, information-theoretic clustering measures have been shown to be effective. In this paper, we employ the Information Bottleneck (IB) clustering technique to find speaker clusters in the target domain data. This is achieved by optimizing the IB criterion that minimizes the information loss during the clustering process. The greedy optimization of the IB criterion involves agglomerative clustering using the Jensen-Shannon divergence as the distance metric. Our experiments in the domain-adaptation task indicate that the proposed system outperforms the baseline by about 14% relative in terms of equal error rate. Subhadeep Dey, Srikanth R. Madikeri, Petr Motlícek |
ICASSP | 3 |
| 2016 | System fusion and speaker linking for longitudinal diarization of TV showsabstractPerforming speaker diarization while uniquely identifying the speakers in a collection of audio recordings is a challenging task. Based on our previous work on speaker diarization and linking, we developed a system for diarizing longitudinal TV show data sets based on the fusion of speaker diarization system outputs and speaker linking. Agreement between multiple diarization outputs is found prior to speaker linking, largely reducing the diarization error rate at the expense of keeping some speech data unlabelled. To deal with noisy clusters, a linear prediction based technique was used to label speakers after linking. Considerable gains for both fusion and labelling are reported. Despite the challenges of the longitudinal diarization task, this system obtained similar performance for linked and non-linked tasks under moderate session variability, highlighting the viability of a linking approach to longitudinal diarization of speech in the presence of noise, music and special audio effects. Marc Ferras, Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard |
ICASSP | 3 |
| 2016 | Inter-Task System Fusion for Speaker Recognition
Marc Ferras, Srikanth R. Madikeri, Subhadeep Dey, Petr Motlícek, Hervé Bourlard |
INTERSPEECH | 4 |
| 2016 | Idlak Tangle: An Open Source Kaldi Based Parametric Speech Synthesiser Based on DNN
Blaise Potard, Matthew P. Aylett, David A. Baude, Petr Motlícek |
INTERSPEECH | 4 |
| 2016 | Feature mapping using far-field microphones for distant speech recognition
Ivan Himawan, Petr Motlícek, David Imseng, Sridha Sridharan |
Speech Commun. | 2 |
| 2016 | A Large-Scale Open-Source Acoustic Simulator for Speaker RecognitionabstractThe state-of-the-art speaker-recognition systems suffer from significant performance loss on degraded speech conditions and acoustic mismatch between enrolment and test phases. Past international evaluation campaigns, such as the NIST speaker recognition evaluation (SRE), have partly addressed these challenges in some evaluation conditions. This work aims at further assessing and compensating for the effect of a wide variety of speech-degradation processes on speaker-recognition performance. We present an open-source simulator generating degraded telephone, VoIP, and interview-speech recordings using a comprehensive list of narrow-band, wide-band, and audio codecs, together with a database of over 60 h of environmental noise recordings and over 100 impulse responses collected from publicly available data. We provide speaker-verification results obtained with an i-vector-based system using either a clean or degraded PLDA back-end on a NIST SRE subset of data corrupted by the proposed simulator. While error rates increase considerably under degraded speech conditions, large relative equal error rate (EER) reductions were observed when using a PLDA model trained with a large number of degraded sessions per speaker. Marc Ferras, Srikanth R. Madikeri, Petr Motlícek, Subhadeep Dey, Hervé Bourlard |
IEEE Signal Process. Lett. | 3 |
| 2015 | Towards utterance-based neural network adaptation in acoustic modelingabstractDespite the superior classification ability of deep neural networks (DNN), the performance of DNN suffers when there is a mismatch between training and testing conditions. Many speaker adaptation techniques have been proposed for DNN acoustic modeling but in case of environmental robustness the progress is still limited. It is also possible to use techniques developed for adapting speakers to handle the impact of environments at the same time, or to combine both approaches. Directly adapting the large number of DNN parameters is challenging when the adaptation set is small. The learning hidden unit contributions (LHUC) technique for unsupervised speaker adaptation of DNN introduces speaker dependent parameters to the existing speaker independent network to increase the automatic speech recognition (ASR) performance of the target speaker using small amounts of adaptation data. This paper investigates the LHUC to adapt the speech recognizer to target speakers and environments where the impacts of speakers and noise differences are quantified separately. Our finding shows that the LHUC is capable of adapting to both speaker and noise conditions at the same time. Compared to the speaker independent model, about 9% to 13% relative word error rate (WER) improvement are observed for all test conditions using AMI meeting corpus. Ivan Himawan, Petr Motlícek, Marc Ferras, Srikanth R. Madikeri |
ASRU | 2 |
| 2015 | Learning feature mapping using deep neural network bottleneck features for distant large vocabulary speech recognitionabstractAutomatic speech recognition from distant microphones is a difficult task because recordings are affected by reverberation and background noise. First, the application of the deep neural network (DNN)/hidden Markov model (HMM) hybrid acoustic models for distant speech recognition task using AMI meeting corpus is investigated. This paper then proposes a feature transformation for removing reverberation and background noise artefacts from bottleneck features using DNN trained to learn the mapping between distant-talking speech features and close-talking speech bottleneck features. Experimental results on AMI meeting corpus reveal that the mismatch between close-talking and distant-talking conditions is largely reduced, with about 16% relative improvement over conventional bottleneck system (trained on close-talking speech). If the feature mapping is applied to close-talking speech, a minor degradation of 4% relative is observed. Ivan Himawan, Petr Motlícek, David Imseng, Blaise Potard, Namhoon Kim |
ICASSP | 2 |
| 2015 | Combining SGMM speaker vectors and KL-HMM approach for speaker diarizationabstractIn this paper, a method to use SGMM speaker vectors for speaker diarization is introduced. The architecture of the Information Bottleneck (IB) based speaker diarization is utilized for this purpose. The audio for speaker diarization is split into short uniform segments. Speaker vectors are obtained from a Subspace Gaussian Mixture Model (SGMM) system trained on meeting data. The speaker vectors are clustered using the K-means algorithm. Two types of distance measures are explored in the clustering step: cosine distance of the speaker vectors and that of the vectors in a space projected by Probabilistic Linear Discriminant Analysis (PLDA). The clustering output is used as an initialization step for the Kullback Leibler-Hidden Markov Model (KL-HMM) based speech segmentation approach commonly used in the IB system for diarization. The proposed method is compared to clustering the segments using the IB based approach. A relative improvement of approximately 14% is obtained on the diarization performance for the proposed approach using SGMM speaker vectors with PLDA on the NIST RT 09 dataset. Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard |
ICASSP | 2 |
| 2015 | Employment of Subspace Gaussian Mixture Models in speaker recognitionabstractThis paper presents Subspace Gaussian Mixture Model (SGMM) approach employed as a probabilistic generative model to estimate speaker vector representations to be subsequently used in the speaker verification task. SGMMs have already been shown to significantly outperform traditional HMM/GMMs in Automatic Speech Recognition (ASR) applications. An extension to the basic SGMM framework allows to robustly estimate low-dimensional speaker vectors and exploit them for speaker adaptation. We propose a speaker verification framework based on low-dimensional speaker vectors estimated using SGMMs, trained in ASR manner using manual transcriptions. To test the robustness of the system, we evaluate the proposed approach with respect to the state-of-the-art i-vector extractor on the NIST SRE 2010 evaluation set and on four different length-utterance conditions: 3sec-10sec, 10 sec-30 sec, 30 sec-60 sec and full (untruncated) utterances. Experimental results reveal that while i-vector system performs better on truncated 3sec to 10sec and 10 sec to 30 sec utterances, noticeable improvements are observed with SGMMs especially on full length-utterance durations. Eventually, the proposed SGMM approach exhibits complementary properties and can thus be efficiently fused with i-vector based speaker verification system. Petr Motlícek, Subhadeep Dey, Srikanth R. Madikeri, Lukás Burget |
ICASSP | 1 |
| 2015 | Channel selection in the short-time modulation domain for distant speech recognitionabstractAutomatic speech recognition from multiple distant microphones poses significant challenges because of noise and reverberations.The quality of speech acquisition may vary between microphones because of movements of speakers and channel distortions.This paper proposes a channel selection approach for selecting reliable channels based on selection criterion operating in the short-term modulation spectrum domain.The proposed approach quantifies the relative strength of speech from each microphone and speech obtained from beamforming modulations.The new technique is compared experimentally in the real reverb conditions in terms of perceptual evaluation of speech quality (PESQ) measures and word error rate (WER).Overall improvement in recognition rate is observed using delay-sum and superdirective beamformers compared to the case when the channel is selected randomly using circular microphone arrays. Ivan Himawan, Petr Motlícek, Sridha Sridharan, David Dean, Dian Tjondronegoro |
INTERSPEECH | 2 |
| 2015 | Integrating online i-vector extractor with information bottleneck based speaker diarization systemabstractConventional approaches to speaker diarization use short-term features such as Mel Frequency Cepstral Co-efficients (MFCC).Features such as i-vectors have been used on longer segments (minimum 2.5 seconds of speech).Using i-vectors for speaker diarization has been shown to be beneficial as it models speaker information explicitly.In this paper, the i-vector modelling technique is adapted to be used as short term features for diarization by estimating i-vectors over a short window of MFCCs.The Information Bottleneck (IB) approach provides a convenient platform to integrate multiple features together for fast and accurate diarization of speech.Speaker models are estimated over a window of 10 frames of speech and used as features in the IB system.Experiments on the NIST RT datasets show absolute improvements of 3.9% in the best case when ivectors are used as auxiliary features to MFCC.Further, discriminative training algorithms such as LDA and PLDA are applied on the i-vectors.A best case performance improvement of 5% in absolute terms is obtained on the RT datasets. Srikanth R. Madikeri, Ivan Himawan, Petr Motlícek, Marc Ferras |
INTERSPEECH | 3 |
| 2015 | Incremental Syllable-Context Phonetic VocodingabstractCurrent very low bit rate speech coders are, due to complexity limitations, designed to work off-line. This paper investigates incremental speech coding that operates real-time and incrementally (i.e., encoded speech depends only on already-uttered speech without the need of future speech information). Since human speech communication is asynchronous (i.e., different information flows being simultaneously processed), we hypothesized that such an incremental speech coder should also operate asynchronously. To accomplish this task, we describe speech coding that reflects the human cortical temporal sampling that packages information into units of different temporal granularity, such as phonemes and syllables, in parallel. More specifically, a phonetic vocoder-cascaded speech recognition and synthesis systems-extended with syllable-based information transmission mechanisms is investigated. There are two main aspects evaluated in this work, the synchronous and asynchronous coding. Synchronous coding refers to the case when the phonetic vocoder and speech generation process depend on the syllable boundaries during encoding and decoding respectively. On the other hand, asynchronous coding refers to the case when the phonetic encoding and speech generation processes are done independently of the syllable boundaries. Our experiments confirmed that the asynchronous incremental speech coding performs better, in terms of intelligibility and overall speech quality, mainly due to better alignment of the segmental and prosodic information. The proposed vocoding operates at an uncompressed bit rate of 213 bits/sec and achieves an average communication delay of 243 ms. Milos Cernak, Philip N. Garner, Alexandros Lazaridis, Petr Motlícek, Xingyu Na |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Exploiting un-transcribed foreign data for speech recognition in well-resourced languagesabstractManual transcription of audio databases for automatic speech recognition (ASR) training is a costly and time-consuming process. State-of-the-art hybrid ASR systems that are based on deep neural networks (DNN) can exploit un-transcribed foreign data during unsupervised DNN pre-training or semi-supervised DNN training. We investigate the relevance of foreign data characteristics, in particular domain and language. Using three different datasets of the MediaParl and Ester databases, our experiments suggest that domain and language are equally important. Foreign data recorded under matched conditions (language and domain) yields the most improvement. The resulting ASR system yields about 5% relative improvement compared to the baseline system only trained on transcribed data. Our studies also reveal that the amount of foreign data used for semi-supervised training can be significantly reduced without degrading the ASR performance if confidence measure based data selection is employed. David Imseng, Blaise Potard, Petr Motlícek, Alexandre Nanchen, Hervé Bourlard |
ICASSP | 3 |
| 2014 | Multilingual deep neural network based acoustic modeling for rapid language adaptationabstractThis paper presents a study on multilingual deep neural network (DNN) based acoustic modeling and its application to new languages. We investigate the effect of phone merging on multilingual DNN in context of rapid language adaptation. Moreover, the combination of multilingual DNNs with Kullback-Leibler divergence based acoustic modeling (KL-HMM) is explored. Using ten different languages from the Globalphone database, our studies reveal that crosslingual acoustic model transfer through multilingual DNNs is superior to unsupervised RBM pre-training and greedy layer-wise supervised training. We also found that KL-HMM based decoding consistently outperforms conventional hybrid decoding, especially in low-resource scenarios. Furthermore, the experiments indicate that multilingual DNN training equally benefits from simple phoneset concatenation and manually derived universal phonesets. Ngoc Thang Vu, David Imseng, Daniel Povey, Petr Motlícek, Tanja Schultz, Hervé Bourlard |
ICASSP | 4 |
| 2014 | Stress and accent transmission in HMM-based syllable-context very low bit rate speech codingabstractLIDIAP Milos Cernak, Alexandros Lazaridis, Philip N. Garner, Petr Motlícek |
INTERSPEECH | 4 |
| 2014 | Development of bilingual ASR system for MediaParl corpusabstractThe development of an Automatic Speech Recognition (ASR) system for the bilingual MediaParl corpus is challenging for several reasons: (1) reverberant recordings, (2) accented speech, and (3) no prior information about the language. In that context, we employ frequency domain linear prediction-based (FDLP) features to reduce the effect of reverberation, exploit bilingual deep neural networks applied in Tandem and hybrid acoustic modeling approaches to significantly improve ASR for accented speech and develop a fully bilingual ASR system using entropy-based decoding-graph selection. Our experiments indicate that the proposed bilingual ASR system performs similar to a language-specific ASR system if approximately five seconds of speech are available. Petr Motlícek, David Imseng, Milos Cernak, Namhoon Kim |
INTERSPEECH | 1 |
| 2014 | Phoneme background model for information bottleneck based speaker diarizationabstractAcoustic variability of speakers arises due to differences in their vocal tract characteristics.These individual speaker characteristics are reflected in a speech signal when speakers pronounce a given phoneme.The current work hypothesizes that clusters within a phoneme spoken by multiple speakers roughly correspond to different speakers.Based on this hypothesis, a Gaussian mixture model (GMM) based phoneme background model (PBM) is estimated.The components of such a PBM are used as a set of relevance variables in information bottleneck based speaker diarization system.Experiments are done using phone transcripts obtained from ground-truth and automatic speech recognition (ASR) system to estimate the PBM.The diarization experiments done on meeting recordings from AMI and NIST-RT corpora show that the proposed method achieves significant improvements over the system using a background model which ignores phoneme information. Sree Harsha Yella, Petr Motlícek, Hervé Bourlard |
INTERSPEECH | 2 |
| 2014 | The DBOX Corpus Collection of Spoken Human-Human and Human-Machine Dialogues
Volha Petukhova, Martin Gropp, Dietrich Klakow, Gregor Eigner, Mario Topf, Stefan Srb, Petr Motlícek, Blaise Potard, John Dines, Olivier Deroo, Ronny Egeler, Uwe Meinz, Steffen Liersch, Anna Schmidt |
LREC | 7 |
| 2014 | Using out-of-language data to improve an under-resourced speech recognizer
David Imseng, Petr Motlícek, Hervé Bourlard, Philip N. Garner |
Speech Commun. | 2 |
| 2013 | Impact of deep MLP architecture on different acoustic modeling techniques for under-resourced speech recognitionabstractPosterior based acoustic modeling techniques such as Kullback-Leibler divergence based HMM (KL-HMM) and Tandem are able to exploit out-of-language data through posterior features, estimated by a Multi-Layer Perceptron (MLP). In this paper, we investigate the performance of posterior based approaches in the context of under-resourced speech recognition when a standard three-layer MLP is replaced by a deeper five-layer MLP. The deeper MLP architecture yields similar gains of about 15% (relative) for Tandem, KL-HMM as well as for a hybrid HMM/MLP system that directly uses the posterior estimates as emission probabilities. The best performing system, a bilingual KL-HMM based on a deep MLP, jointly trained on Afrikaans and Dutch data, performs 13% better than a hybrid system using the same bilingual MLP and 26% better than a subspace Gaussian mixture system only trained on Afrikaans data. David Imseng, Petr Motlícek, Philip N. Garner, Hervé Bourlard |
ASRU | 2 |
| 2013 | On the (UN)importance of the contextual factors in HMM-based speech synthesis and codingabstractThis paper presents an evaluation of the contextual factors of HMMbased speech synthesis and coding systems.Two experimental setups are proposed that are based on successive context addition from phonetic to full-context.The aim was to investigate the impact of the individual contextual factors on the speech quality.In that sense important and unimportant (i.e., not having significant impact on speech quality, also called weak) contextual factors were identified.The results imply that in speech coding the improvement in quality can be achieved just with reconstruction of syllable contexts.The sentence and utterance contexts are unimportant on the decoder side, and it is not necessary to deal with them.Although in speech coding the wider context was not necessary, in speech synthesis current syllable and utterance contexts are more important over others (previous and next word/phrase contexts). Milos Cernak, Petr Motlícek, Philip N. Garner |
ICASSP | 2 |
| 2013 | Accent adaptation using Subspace Gaussian Mixture ModelsabstractThis paper investigates employment of Subspace Gaussian Mixture Models (SGMMs) for acoustic model adaptation towards different accents for English speech recognition. The SGMMs comprise globally-shared and state-specific parameters which can efficiently be employed for various kinds of acoustic parameter tying. Research results indicate that well-defined sharing of acoustic model parameters in SGMMs can significantly outperform adapted systems based on conventional HMM/GMMs. Furthermore, SGMMs rapidly achieve target acoustic models with small amounts of data. Experiments performed with US and UK English versions of the Wall Street Journal (WSJ) corpora indicate that SGMMs lead to approximately 20% and 8% relative improvements with respect to speaker-independent and speaker-adapted acoustic models respectively over conventional HMM/GMMs. Finally, we demonstrate that SGMMs adapted only with 1.5 hours can reach performance of HMM/GMMs trained with 18 hours. Petr Motlícek, Philip N. Garner, Namhoon Kim, Jeongmi Cho |
ICASSP | 1 |
| 2013 | Feature and score level combination of subspace Gaussinas in LVCSR taskabstractIn this paper, we investigate employment of discriminatively trained acoustic features modeled by Subspace Gaussian Mixture Models (SGMMs) for Rich Transcription meeting recognition. More specifically, first, we focus on exploiting various types of complex features estimated using neural network combined with conventional cepstral features and modeled by standard HMM/GMMs and SGMMs. Then, outputs (word sequences) from individual recognizers trained using different features are also combined on a score-level using ROVER for the both acoustic modeling techniques. Experimental results indicate three important findings: (1) SGMMs consistently outperform HMM/GMMs (relative improvement on average by about 6% in terms of WER) when both techniques are exploited on single features; (2) SGMMs benefit much less from feature-level combination (1% relative improvement) as opposed to HMM/GMMs (4% relative improvement) which can eventually match the performance of SGMMs; (3) SGMMs can be significantly improved when individual systems are combined on a score-level. This suggests that the SGMM systems provide complementary recognition outputs. Overall relative improvements of the combined SGMMand HMM/GMM systems are 21% and 17% respectively compared to a standard ASR baseline. Petr Motlícek, Daniel Povey, Martin Karafiát |
ICASSP | 1 |
| 2013 | Crosslingual tandem-SGMM: exploiting out-of-language data for acoustic model and feature level adaptationabstractRecent studies have shown that speech recognizers may benefit from data in languages other than the target language through efficient acoustic model- or feature-level adaptation. Crosslingual Tandem-Subspace Gaussian Mixture Models (SGMM) are successfully able to combine acoustic model- and feature-level adaptation techniques. More specifically, we focus on under-resourced languages (Afrikaans in our case) and perform feature-level adaptation through the estimation of phone class posterior features with a Multilayer Perceptron that was trained on data from a similar language with large amounts of available speech data (Dutch in our case). The same Dutch data can also be exploited on an acoustic model-level by training globally-shared SGMM parameters in a crosslingual way. The two adaptation techniques are indeed complementary and result in a crosslingual Tandem-SGMM system that yields relative improvement of about 22% compared to a standard speech recognizer on an Afrikaans phoneme recognition task. Interestingly, eventual score-level combination of the individual SGMM systems yields additional 3% relative improvement. Petr Motlícek, David Imseng, Philip N. Garner |
INTERSPEECH | 1 |
| 2013 | A Simple Continuous Pitch Estimation AlgorithmabstractRecent work in text to speech synthesis has pointed to the benefit of using a continuous pitch estimate; that is, one that records pitch even when voicing is not present. Such an approach typically requires interpolation. The purpose of this letter is to show that a continuous pitch estimation is available from a combination of otherwise well known techniques. Further, in the case of an autocorrelation based estimate, the continuous requirement negates the need for other heuristics to correct for common errors. An algorithm is suggested, illustrated, and demonstrated using a parametric vocoder. Philip N. Garner, Milos Cernak, Petr Motlícek |
IEEE Signal Process. Lett. | 3 |
| 2012 | Improving acoustic based keyword spotting using LVCSR latticesabstractThis paper investigates detection of English keywords in a conversational scenario using a combination of acoustic and LVCSR based keyword spotting systems. Acoustic KWS systems search predefined words in parameterized spoken data. Corresponding confidences are represented by likelihood ratios given the keyword models and a background model. First, due to the especially high number of false-alarms, the acoustic KWS system is augmented with confidence measures estimated from corresponding LVCSR lattices. Then, various strategies to combine scores estimated by the acoustic and several LVCSR based KWS systems are explored. We show that a linear regression based combination significantly outperforms other (model-based) techniques. Due to that, the relative number of false-alarms of the combined KWS system decreased by more than 50% compared to the acoustic KWS system. Finally, an attention is also paid to the complexities of the KWS systems enabling them to potentially be exploited in real-detection tasks. Petr Motlícek, Fabio Valente, Igor Szöke |
ICASSP | 1 |
| 2012 | Generating exact lattices in the WFST frameworkabstractWe describe a lattice generation method that is exact, i.e. it satisfies all the natural properties we would want from a lattice of alternative transcriptions of an utterance. This method does not introduce substantial overhead above one-best decoding. Our method is most directly applicable when using WFST decoders where the WFST is “fully expanded”, i.e. where the arcs correspond to HMM transitions. It outputs lattices that include HMM-state-level alignments as well as word labels. The general idea is to create a state-level lattice during decoding, and to do a special form of determinization that retains only the best-scoring path for each word sequence. This special determinization algorithm is a solution to the following problem: Given a WFST A, compute a WFST B that, for each input-symbol-sequence of A, contains just the lowest-cost path through A. Daniel Povey, Mirko Hannemann, Gilles Boulianne, Lukás Burget, Arnab Ghoshal, Milos Janda, Martin Karafiát, Stefan Kombrink, Petr Motlícek, Yanmin Qian, Korbinian Riedhammer, Karel Veselý, Ngoc Thang Vu |
ICASSP | 9 |
| 2012 | Bi-modal authentication in mobile environments using session variability modelling
Petr Motlícek, Laurent El Shafey, Roy Wallace, Chris McCool, Sébastien Marcel |
ICPR | 1 |
| 2012 | Comparing different acoustic modeling techniques for multilingual boostingabstractIn this paper, we explore how different acoustic modeling techniques can benefit from data in languages other than the target language. We propose an algorithm to perform decision tree state clustering for the recently proposed Kullback-Leibler divergence based hidden Markov models (KL-HMM) and compare it to subspace Gaussian mixture modeling (SGMM). KL-HMM can exploit multilingual information in the form of universal phoneme posterior features and SGMM benefits from a universal background model that can be trained on multilingual data. Taking the Greek SpeechDat(II) data as an example, we show that KL-HMM performs best for small amounts of target language data. David Imseng, John Dines, Petr Motlícek, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 3 |
| 2012 | Supervised and unsupervised Web-based language model domain adaptationabstractDomain language model adaptation consists in re-estimating probabilities of a baseline LM in order to better match the specifics of a given broad topic of interest. To do so, a common strategy is to retrieve adaptation texts from the Web based on a given domain-representative seed text. In this paper, we study how the selection of this seed text influences the adaptation process and the performances of resulting adapted language models in automatic speech recognition. More precisely, the goal of this original study is to analyze the differences of our Web-based adaptation approach between the supervised case, in which the seed text is manually generated, and the unsupervised case, where the seed text is given by an automatic transcript. Experiments were carried out on data sourced from a real-world use case, more specifically, videos produced for a university YouTube channel. Results show that our approach is quite robust since the unsupervised adaptation provides similar performance to the supervised case in terms of the overall perplexity and word error rate. Gwénolé Lecorvé, John Dines, Thomas Hain, Petr Motlícek |
INTERSPEECH | 4 |
| 2012 | Conversion of Recurrent Neural Network Language Models to Weighted Finite State Transducers for Automatic Speech RecognitionabstractInternational audience Gwénolé Lecorvé, Petr Motlícek |
INTERSPEECH | 2 |
| 2012 | Annotation and Recognition of Personality Traits in Spoken Conversations from the AMI Meetings CorpusabstractLIDIAP Fabio Valente, Samuel Kim, Petr Motlícek |
INTERSPEECH | 3 |
| 2012 | Multimodal Cue Detection Engine for Orchestrated Entertainment
Danil Korchagin, Stefan Duffner, Petr Motlícek, Carl Scheffler |
MMM | 3 |
| 2012 | Assessing the impact of language style on emergent leadership perception from ubiquitous audioabstractLeaders stand out for what they say and how they say it. This work describes the impact of the language style of emergent leaders in small group discussions based on 7 hours of audio from English spoken discussions recorded with a ubiquitous platform. For the language style analysis, word categories are extracted from manual transcriptions of the discussions as well as from automatically detected keywords. The most relevant word categories are then used to predict the emergent leader in each group. Our findings reveal that non-privacy sensitive word categories like amount of words, conjunctions and assent are good predictors of emergent leadership. The emergent leader can be correctly inferred in a fully automatic approach with up to 82% accuracy using categories derived from keywords, and up to 86% using categories derived from full manual transcriptions. Dairazalia Sanchez-Cortes, Petr Motlícek, Daniel Gatica-Perez |
MUM | 2 |
| 2011 | Speaker diarization of meetings based on speaker role n-gram modelsabstractSpeaker diarization of meeting recordings is generally based on acoustic information ignoring that meetings are instances of conversations. Several recent works have shown that the sequence of speakers in a conversation and their roles are related and statistically predictable. This paper proposes the use of speaker roles n-gram model to capture the conversation patterns probability and investigates its use as prior information into a state-of-the-art diarization system. Experiments are run on the AMI corpus annotated in terms of roles. The proposed technique reduces the diarization speaker error by 19% when the roles are known and by 17% when they are estimated. Furthermore the paper investigates how the n-gram models generalize to different settings like those from the Rich Transcription campaigns. Experiments on 17 meetings reveal that the speaker error can be reduced by 12% also in this case thus the n-gram can generalize across corpora. Fabio Valente, Deepu Vijayasenan, Petr Motlícek |
ICASSP | 3 |
| 2011 | Multistream speaker diarization through Information Bottleneck system outputs combinationabstractSpeaker diarization of meetings recorded with Multiple Distant Microphones makes extensive use of multiple feature streams like MFCC and Time Delay of Arrivals (TDOA). Typically the combination happens using separate models for each feature stream. This work investigates if the combination of multiple feature streams can happen through the combination of multiple diarization systems performed using those features. The paper extends the previously proposed Information Bottleneck method to handle the combination of several probabilistic diarization outputs. In contrast to the conventional model-based feature combination, this technique is referred as system-based combination. Furthermore the paper introduces an hybrid model-system combination. Experiments are run on data from the Rich Transcription campaigns and show that the system based combination largely outperforms the model based combination by 37% relative. The hybrid approaches improve by 10-20%. The analysis of errors shows that the improvements come from the recordings where the individual MFCC and TDOA systems provide very different performances. Deepu Vijayasenan, Fabio Valente, Petr Motlícek |
ICASSP | 3 |
| 2011 | Just-in-time multimodal association and fusion from home entertainmentabstractIn this paper, we describe a real-time multimodal analysis system with just-in-time multimodal association and fusion for a living room environment, where multiple people may enter, interact and leave the observable world with no constraints. It comprises detection and tracking of up to 4 faces, detection and localisation of verbal and paralinguistic events, their association and fusion. The system is designed to be used in open, unconstrained environments like in next generation video conferencing systems that automatically "orchestrate" the transmitted video streams to improve the overall experience of interaction between spatially separated families and friends. Performance levels achieved to date on hand-labelled dataset have shown sufficient reliability at the same time as fulfilling real-time processing requirements. Danil Korchagin, Petr Motlícek, Stefan Duffner, Hervé Bourlard |
ICME | 2 |
| 2010 | Application of out-of-language detection to spoken term detectionabstractThis paper investigates the detection of English spoken terms in a conversational multi-language scenario. The speech is processed using a large vocabulary continuous speech recognition system. The recognition output is represented in the form of word recognition lattices which are then used to search required terms. Due to the potential multi-lingual speech segments at the input, the spoken term detection system is combined with a module performing out-of-language detection to adjust its confidence scores. First, experimental results of spoken term detection are provided on the conversational telephone speech database distributed by NIST in 2006. Then, the system is evaluated on a multi-lingual database with and without employment of the out-of-language detection module, where we are only interested in detecting English terms (stored in the index database). Several strategies to combine these two systems in an efficient way are proposed and evaluated. Around 7% relative improvement over a stand-alone STD is achieved. Petr Motlícek, Fabio Valente |
ICASSP | 1 |
| 2010 | Variational Bayesian speaker diarization of meeting recordingsabstractThis paper investigates the use of the Variational Bayesian (VB) framework for speaker diarization of meetings data extending previous related works on Broadcast News audio. VB learning aims at maximizing a bound, known as Free Energy, on the model marginal likelihood and allows joint model learning and model selection according to the same objective function. While the BIC is valid only in the asymptotic limit, the Free Energy is always a valid bound. The paper proposes the use of Free Energy as objective function in speaker diarization. It can be used to select dynamically without any supervision or tuning, elements that typically affect the diarization performance i.e. the inferred number of speakers, the size of the GMM and the initialization. The proposed approach is compared with a conventional state-of-the-art system on the RT06 evaluation data for meeting recordings diarization and shows an improvement of 8.4% relative in terms of speaker error. Fabio Valente, Petr Motlícek, Deepu Vijayasenan |
ICASSP | 2 |
| 2010 | Hands free audio analysis from home entertainmentabstractIn this paper, we describe a system developed for hands free audio analysis for a living room environment. It comprises detection and localisation of the verbal and paralinguistic events, which can augment the behaviour of virtual director and improve the overall experience of interactions between spatially separated families and friends. The results show good performance in reverberant environments and fulfil real-time requirements. Index Terms: real-time audio processing, direction of arrival, speech meta-data Danil Korchagin, Philip N. Garner, Petr Motlícek |
INTERSPEECH | 3 |
| 2010 | English spoken term detection in multilingual recordingsabstractThis paper investigates the automatic detection of English spoken terms in a multi-language scenario over real lecture recordings. Spoken Term Detection (STD) is based on an LVCSR where the output is represented in the form of word lattices. The lattices are then used to search the required terms. Processed lectures are mainly composed of English, French and Italian recordings where the language can also change within one recording. Therefore, the English STD system uses an Out-Of-Language (OOL) detection module to filter out non-English input segments. OOL detection is evaluated w.r.t. various confidence measures estimated from word lattices. Experimental studies of OOL detection followed by English STD are performed on several hours of multilingual recordings. Significant improvement of OOL+STD over a stand-alone STD system is achieved (relatively more than 50% in EER). Finally, an additional modality (text slides in the form of PowerPoint presentations) is exploited to improve STD. Petr Motlícek, Fabio Valente, Philip N. Garner |
INTERSPEECH | 1 |
| 2010 | Autoregressive Models of Amplitude Modulations in Audio CompressionabstractWe present a scalable medium bit-rate wide-band audio coding technique based on frequency-domain linear prediction (FDLP). FDLP is an efficient method for representing the long-term amplitude modulations of speech/audio signals using autoregressive models. For the proposed audio codec, relatively long temporal segments (1000 ms) of the input audio signal are decomposed into a set of critically sampled sub-bands using a quadrature mirror filter (QMF) bank. The technique of FDLP is applied on each sub-band to model the sub-band temporal envelopes. The residual of the linear prediction, which represents the frequency modulations in the sub-band signal, are encoded and transmitted along with the envelope parameters. These steps are reversed at the decoder to reconstruct the signal. The proposed codec utilizes a simple signal independent nonadaptive compression mechanism for a wide class of speech and audio signals. The subjective and objective quality evaluations show that the reconstruction signal quality for the proposed FDLP codec compares well with the state-of-the-art audio codecs in the 32-64 kbps range. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Automatic out-of-language detection based on confidence measures derived from LVCSR word and phone latticesabstractConfidence Measures (CMs) estimated from Large Vocabulary Continuous Speech Recognition (LVCSR) outputs are commonly used metrics to detect incorrectly recognized words. In this paper, we propose to exploit CMs derived from frame-based word and phone posteriors to detect speech segments containing pronunciations from non-target (alien) languages. The LVCSR system used is built for English, which is the target language, with medium-size recognition vocabulary (5k words). The efficiency of detection is tested on a set comprising speech from three different languages (English, German, Czech). Results achieved indicate that employment of specific temporal context (integrated in the word or phone level) significantly increases the detection accuracies. Furthermore, we show that combination of several CMs can also improve the efficiency of detection. Petr Motlícek |
INTERSPEECH | 1 |
| 2009 | Arithmetic coding of sub-band residuals in FDLP speech/audio codecabstractA speech/audio codec based on Frequency Domain Linear Prediction (FDLP) exploits auto-regressive modeling to approximate instantaneous energy in critical frequency sub-bands of relatively long input segments. The current version of the FDLP codec operating at 66 kbps has been shown to provide comparable subjective listening quality results to state-of-the-art codecs on similar bit-rates even without employing standard blocks such as entropy coding or simultaneous masking. This paper describes an experimental work to increase compression efficiency of the FDLP codec by employing entropy coding. Unlike conventional Huffman coding employed in current speech/audio coding systems, we describe an efficient way to exploit arithmetic coding to entropy compress quantized spectral magnitudes of the sub-band FDLP residuals. Such an approach provides 11% (∼ 3 kbps) bit-rate reduction compared to the Huffman coding algorithm (∼ 1 kbps). Petr Motlícek, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 1 |
| 2008 | Temporal masking for bit-rate reduction in audio codec based on Frequency Domain Linear PredictionabstractAudio coding based on frequency domain linear prediction (FDLP) uses auto-regressive model to approximate Hilbert envelopes in frequency sub-bands for relatively long temporal segments. Although the basic technique achieves good quality of the reconstructed signal, there is a need for improving the coding efficiency. In this paper, we present a novel method for the application of temporal masking to reduce the bit-rate in a FDLP based codec. Temporal masking refers to the hearing phenomenon, where the exposure to a sound reduces response to following sounds for a certain period of time (up to 200 ms). In the proposed version of the codec, a first order forward masking model of the human ear is implemented and informal listening experiments using additive white noise are performed to obtain the exact noise masking thresholds. Subsequently, this masking model is employed in encoding the sub- band FDLP carrier signal. Application of the temporal masking in the FDLP codec results in a bit-rate reduction of about 10% without degrading the quality. Performance evaluation is done with perceptual evaluation of audio quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
ICASSP | 2 |
| 2008 | The DIRAC AWEAR audio-visual platform for detection of unexpected and incongruent eventsabstractIt is of prime importance in everyday human life to cope with and respond appropriately to events that are not foreseen by prior experience. Machines to a large extent lack the ability to respond appropriately to such inputs. An important class of unexpected events is defined by incongruent combinations of inputs from different modalities and therefore multimodal information provides a crucial cue for the identification of such events, e.g., the sound of a voice is being heard while the person in the field-of-view does not move her lips. In the project DIRAC ("Detection and Identification of Rare Audio-visual Cues") we have been developing algorithmic approaches to the detection of such events, as well as an experimental hardware platform to test it. An audio-visual platform ("AWEAR" - audio-visual wearable device) has been constructed with the goal to help users with disabilities or a high cognitive load to deal with unexpected events. Key hardware components include stereo panoramic vision sensors and 6-channel worn-behind-the-ear (hearing aid) microphone arrays. Data have been recorded to study audio-visual tracking, a/v scene/object classification and a/v detection of incongruencies. Jörn Anemüller, Jörg-Hendrik Bach, Barbara Caputo, Michal Havlena, Jie Luo 0018, Hendrik Kayser, Bastian Leibe, Petr Motlícek, Tomás Pajdla, Misha Pavel, Akihiko Torii, Luc Van Gool, Alon Zweig, Hynek Hermansky |
ICMI | 8 |
| 2008 | Spectral noise shaping: improvements in speech/audio codec based on linear prediction in spectral domainabstractAudio coding based on Frequency Domain Linear Prediction (FDLP) uses auto-regressive models to approximate Hilbert envelopes in frequency sub-bands. Although the basic technique achieves good coding efficiency, there is a need to improve the reconstructed signal quality for tonal signals with impulsive spectral content. For such signals, the quantization noise in the FDLP codec appears as frequency components not present in the input signal. In this paper, we propose a technique of Spectral Noise Shaping (SNS) for improving the quality of tonal signals by applying a Time Domain Linear Prediction (TDLP) filter prior to the FDLP processing. The inverse TDLP filter at the decoder shapes the quantization noise to reduce the artifacts. Application of the SNS technique to the FDLP codec improves the quality of the tonal signals without affecting the bit-rate. Performance evaluation is done with Perceptual Evaluation of Audio Quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
INTERSPEECH | 2 |
| 2007 | Unsupervised Speech/Non-Speech Detection for Automatic Speech Recognition in Meeting RoomsabstractThe goal of this work is to provide robust and accurate speech detection for automatic speech recognition (ASR) in meeting room settings. The solution is based on computing long-term modulation spectrum, and examining specific frequency range for dominant speech components to classify speech and non-speech signals for a given audio signal. Manually segmented speech segments, short-term energy, short-term energy and zero-crossing based segmentation techniques, and a recently proposed multi layer perceptron (MLP) classifier system are tested for comparison purposes. Speech recognition evaluations of the segmentation methods are performed on a standard database and tested in conditions where the signal-to-noise ratio (SNR) varies considerably, as in the cases of close-talking headset, lapel, distant microphone array output, and distant microphone. The results reveal that the proposed method is more reliable and less sensitive to mode of signal acquisition and unforeseen conditions. Hari Krishna Maganti, Petr Motlícek, Daniel Gatica-Perez |
ICASSP (4) | 2 |
| 2007 | Wide-Band Perceptual Audio Coding Based on Frequency-Domain Linear PredictionabstractIn this paper we propose an extension of the very low bit-rate speech coding technique, exploiting predictability of the temporal evolution of spectral envelopes, for wide-band audio coding applications. Temporal envelopes in critically band-sized sub-bands are estimated using frequency domain linear prediction applied on relatively long time segments. The sub-band residual signals, which play an important role in acquiring high quality reconstruction, are processed using a heterodyning-based signal analysis technique. For reconstruction, their optimal parameters are estimated using a closed-loop analysis-by-synthesis technique driven by a perceptual model emulating simultaneous masking properties of the human auditory system. We discuss the advantages of the approach and show some properties on challenging audio recordings. The proposed technique is capable of encoding high quality, variable rate audio signals on bit-rates below 1 bit/sample. Petr Motlícek, Vijay Ullal, Hynek Hermansky |
ICASSP (1) | 1 |
| 2005 | Non-parametric speaker turn segmentation of meeting dataabstractAn extension of conventional speaker segmentation framework is presented for a scenario in which a number of microphones record the activity of speakers present at a meeting (one microphone per speaker). Although each microphone can receive speech from both the participant wearing the microphone (local speech) and other participants (cross-talk), the recorded audio can be broadly classified in three ways: local speech, cross-talk, and silence. This paper proposes a technique which takes into account cross-correlations, values of its maxima, and energy differences as features to identify and segment speaker turns. In particular, we have used classical cross-correlation functions, time smoothing and in part temporal constraints to sharpen and disambiguate timing differences between microphone channels that may be dominated by noise and reverberation. Experimental results show that proposed technique can be successively used for speaker segmentation of data collected from a number of different setups. 1. Petr Motlícek, Lukás Burget, Jan Cernocký |
INTERSPEECH | 1 |
| 2003 | Time-domain based temporal processing with application of orthogonal transformations
Petr Motlícek, Jan Cernocký |
INTERSPEECH | 1 |
| 2003 | Autoregressive modeling based feature extraction for Aurora3 DSR task
Petr Motlícek, Jan Cernocký |
INTERSPEECH | 1 |
| 2002 | Noise estimation for efficient speech enhancement and robust speech recognition
Petr Motlícek, Lukás Burget |
INTERSPEECH | 1 |