VLDB 2026 Research / reviewers in the wild / expert
Lukás Burget
dblp:76/686
· DBLP profile ↗
196ranked-venue papers
8as first author
49since 2021 · last 2026
0000-0002-4951-5908ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 177 · 7 first-author · 44 since 2021Artificial intelligence and machine learning · 112 · 4 first-author · 26 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trainable multi-channel front-ends for joint beamforming and speaker embedding extractionabstractMulti-channel speaker verification (SV), employing numerous microphones for capturing enrollment and/or test recordings, gained attention for its benefits in far-field scenarios. While some studies approach the problem by designing multi-channel embedding extractors, we focus on building and thoroughly analyzing a framework integrating beamforming pre-processing paired with single-channel embedding extraction. This strategy benefits from accommodating both multi-channel and single-channel inputs. Furthermore, it provides human-interpretable intermediate output — enhanced speech — that can be independently evaluated and related to SV performance. We first focus on the front-end, taking advantage of deep-learning source separation for direct or indirect mask estimation required by the beamformer. We alternate single-channel network architectures, subsequently extended to multi-channel ones by reference channel attention (RCA). We also analyze the impact of beamformer and network output fusion. Finally, we show improvements brought by end-to-end fine-tuning the entire architecture facilitated by our newly designed multi-channel corpus, MultiSV2, extending our previous MultiSV dataset. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký, Meng Yu 0003 |
Comput. Speech Lang. | 3 |
| 2026 | DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition
Alexander Polok, Dominik Klement, Martin Kocour, Jiangyu Han, Federico Landini, Bolaji Yusuf, Matthew Wiesner, Sanjeev Khudanpur, Jan Cernocký, Lukás Burget |
Comput. Speech Lang. | 10 |
| 2025 | State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb DataabstractIn this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments be-longing to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation. Sara Barahona, Ladislav Mosner, Themos Stafylakis, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký |
ASRU | 6 |
| 2025 | DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech RecognitionabstractThis paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and generalization in both in- and out-of-domain settings. The proposed method, Decoder-Centric Regularization in EncoderDecoder (DeCRED), adds auxiliary classifiers to the decoder, enabling next token prediction via intermediate logits. Empirically, DeCRED reduces the mean internal LM BPE perplexity by 36.6% relative to 11 test sets. Furthermore, this translates into actual WER improvements over the baseline in 5 of 7 in-domain and 3 of 4 out-of-domain test sets, reducing macro WER from 6.4% to 6.3% and 18.2 % to 16.2 %, respectively. On TEDLIUM3, DeCRED achieves 7.0 % WER, surpassing the baseline and encoder-centric InterCTC regularization by 0.6 % and 0.5%, respectively. Finally, we compare DeCRED with OWSM v3.1 and Whisper-medium, showing competitive WERs despite training on much less data with fewer parameters. Alexander Polok, Santosh Kesiraju, Karel Benes, Bolaji Yusuf, Lukás Burget, Jan Cernocký |
ASRU | 5 |
| 2025 | Leveraging Self-Supervised Learning for Speaker DiarizationabstractEnd-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several downstream tasks, but their application on speaker diarization is somehow limited. In this work, we explore using WavLM to alleviate the problem of data scarcity for neural diarization training. We use the same pipeline as Pyannote and improve the local end-to-end neural diarization with WavLM and Conformer. Experiments on far-field AMI, AISHELL-4, and AliMeeting datasets show that our method substantially outperforms the Pyannote baseline and achieves new state-of-the-art results on AMI and AISHELL4, respectively. In addition, by analyzing the system performance under different data quantity scenarios, we show that WavLM representations are much more robust against data scarcity than filterbank features, enabling less data hungry training strategies. Furthermore, we found that simulated data, usually used to train end-to-end diarization models, does not help when using WavLM in our experiments. Additionally, we also evaluate our model on the recent CHiME8 NOTSOFAR-1 task where it achieves better performance than the Pyannote baseline. Our source code is publicly available at https://github.com/BUTSpeechFIT/DiariZen. Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Lukás Burget |
ICASSP | 6 |
| 2025 | CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker VerificationabstractSelf-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42%, 0.48%, and 0.96% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.1 Junyi Peng, Ladislav Mosner, Lin Zhang 0054, Oldrich Plchot, Themos Stafylakis, Lukás Burget, Jan Cernocký |
ICASSP | 6 |
| 2025 | Target Speaker ASR with WhisperabstractWe propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single-speaker ASR models into target-speaker ASR models. Our approach also supports speaker-attributed ASR by sequentially generating transcripts for each speaker in a diarization output. This simplified method outperforms baseline speech separation and diarization cascade by 12.9 % absolute ORC-WER on the NOTSOFAR-1 dataset. Alexander Polok, Dominik Klement, Matthew Wiesner, Sanjeev Khudanpur, Jan Cernocký, Lukás Burget |
ICASSP | 6 |
| 2025 | Text-dependent Speaker Verification Challenge 2024: Exploring Shared and User-defined PassphrasesabstractIn contrast to text-independent speaker verification, which has received significant attention from researchers and has many competitions dedicated to it, text-dependent speaker verification (TdSV) has been less explored recently. The TdSV Challenge 2024 was organized to analyze and explore novel methods for this type of speaker verification and aims to motivate participants to develop new approaches to TdSV, conduct comprehensive analyses, and investigate advanced techniques such as self-supervised learning. This challenge builds on the achievements of the short-duration speaker verification (SdSV) Challenges held in 2020 and 2021 and focuses specifically on TdSV in two distinct scenarios. The first scenario involves conventional TdSV, while the second focuses on speaker enrollment using user-defined passphrases. This paper provides a detailed description of both tasks, introduces the evaluation rules, and presents a comprehensive analysis of the results obtained from this challenge. Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
ICASSP | 4 |
| 2025 | Analysis of ABC Frontend Audio Systems for the NIST-SRE24abstractSection: Speaker Recognition Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek |
INTERSPEECH | 11 |
| 2025 | Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Jan Cernocký, Lukás Burget |
INTERSPEECH | 7 |
| 2024 | Hystoc: Obtaining Word Confidences for Fusion of End-To-End ASR SystemsabstractEnd-to-end (e2e) systems have recently gained wide popularity in automatic speech recognition. However, these systems do generally not provide well-calibrated word-level confidences. In this paper, we propose Hystoc, a simple method for obtaining word-level confidences from hypothesis-level scores. Hystoc is an iterative alignment procedure which turns hypotheses from an n-best output of the ASR system into a confusion network. Eventually, word-level confidences are obtained as posterior probabilities in the individual bins of the confusion network. We show that Hystoc provides confidences that correlate well with the accuracy of the ASR hypothesis. Furthermore, we show that utilizing Hystoc in fusion of multiple e2e ASR systems increases the gains from the fusion by up to 1 % WER absolute on Spanish RTVE2020 dataset. Finally, we experiment with using Hystoc for direct fusion of n-best outputs from multiple systems, but we only achieve minor gains when fusing very similar systems. Karel Benes, Martin Kocour, Lukás Burget |
ICASSP | 3 |
| 2024 | Diacorrect: Error Correction Back-End for Speaker DiarizationabstractIn this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and a transformer-based decoder. By exploiting the interactions between the input recording and the initial system’s outputs, DiaCorrect can automatically correct the initial speaker activities to minimize the diarization errors. Experiments on 2-speaker telephony data show that the proposed DiaCorrect can effectively improve the initial model’s results. Our source code is publicly available at https://github.com/BUTSpeechFIT/diacorrect. Jiangyu Han, Federico Landini, Johan Rohdin, Mireia Díez, Lukás Burget, Yuhang Cao, Jan Cernocký |
ICASSP | 5 |
| 2024 | Discriminative Training of VBx DiarizationabstractBayesian HMM clustering of x-vector sequences (VBx) has become a widely adopted diarization baseline model in publications and challenges. It uses an HMM to model speaker turns, a generatively trained probabilistic linear discriminant analysis (PLDA) for speaker distribution modeling, and Bayesian inference to estimate the assignment of x-vectors to speakers. This paper presents a new framework for updating the VBx parameters using discriminative training, which directly optimizes a predefined loss. We also propose a new loss that better correlates with the diarization error rate compared to binary cross-entropy — the default choice for diarization end-to-end systems. Proof-of-concept results across three datasets (AMI, CALLHOME, and DIHARD II) demonstrate the method’s capability of automatically finding hyperparameters, achieving comparable performance to those found by extensive grid search, which typically requires additional hyperparameter behavior knowledge. Moreover, we show that discriminative fine-tuning of PLDA can further improve the model’s performance. We release the source code with this publication. Dominik Klement, Mireia Díez, Federico Landini, Lukás Burget, Anna Silnova, Marc Delcroix, Naohiro Tawara |
ICASSP | 4 |
| 2024 | Multi-Channel Extension of Pre-trained Models for Speaker VerificationabstractInternational audience Ladislav Mosner, Romain Serizel, Lukás Burget, Oldrich Plchot, Emmanuel Vincent 0001, Junyi Peng, Jan Cernocký |
INTERSPEECH | 3 |
| 2024 | Challenging margin-based speaker embedding extractors by using the variational information bottleneck
Themos Stafylakis, Anna Silnova, Johan Rohdin, Oldrich Plchot, Lukás Burget |
INTERSPEECH | 5 |
| 2024 | DiaPer: End-to-End Neural Diarization With Perceiver-Based AttractorsabstractUntil recently, the field of speaker diarization was dominated by cascaded systems. Due to their limitations, mainly regarding overlapped speech and cumbersome pipelines, end-to-end models have gained great popularity lately. One of the most successful models is end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA). In this work, we replace the EDA module with a Perceiver-based one and show its advantages over EEND-EDA; namely obtaining better performance on the largely studied Callhome dataset, finding the quantity of speakers in a conversation more accurately, and faster inference time. Furthermore, when exhaustively compared with other methods, our model, DiaPer, reaches remarkable performance with a very lightweight design. Besides, we perform comparisons with other works and a cascaded baseline across more than ten public wide-band datasets. Together with this publication, we release the code of DiaPer as well as models trained on public and free data. Federico Landini, Mireia Díez, Themos Stafylakis, Lukás Burget |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label SmoothingabstractWhen recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels. Self-supervised pre-trained representations can robustly capture information from speech enabling state-of-the-art results in many downstream tasks including emotion recognition. However, better ways of aggregating the information across time need to be considered as the relevant emotion information is likely to appear piecewise and not uniformly across the signal. For the labels, we need to take into account that there is a substantial degree of noise that comes from the subjective human annotations. In this paper, we propose a novel approach to attentive pooling based on correlations between the representations’ coefficients combined with label smoothing, a method aiming to reduce the confidence of the classifier on the training labels. We evaluate our proposed approach on the benchmark dataset IEMOCAP, and demonstrate high performance surpassing that in the literature. The code to reproduce the results is available at github.com/skakouros/s3prl_attentive_correlation. Sofoklis Kakouros, Themos Stafylakis, Ladislav Mosner, Lukás Burget |
ICASSP | 4 |
| 2023 | Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural DiarizationabstractEnd-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once. Many flavors of end-to-end models have been proposed but all of them require (so far non-existing) large amounts of annotated data for training. The compromise solution consists in generating synthetic data and the recently proposed simulated conversations (SC) have shown remarkable improvements over the original simulated mixtures (SM). In this work, we create SC with multiple speakers per conversation and show that they allow for substantially better performance than SM, also reducing the dependence on a fine-tuning stage. We also create SC with wide-band public audio sources and present an analysis on several evaluation sets. Together with this publication, we release the recipes for generating such data and models trained on public sets as well as the implementation to efficiently handle multiple speakers per conversation and an auxiliary voice activity detection loss. Federico Landini, Mireia Díez, Alicia Lozano-Diez, Lukás Burget |
ICASSP | 4 |
| 2023 | Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using AdaptersabstractRecently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model size grows and sometimes results in over-fitting on small datasets. In this paper, we conduct a comprehensive analysis of applying parameter-efficient transfer learning (PETL) methods to reduce the required learnable parameters for adapting to speaker verification tasks. Specifically, during the fine-tuning process, the pre-trained models are frozen, and only lightweight modules inserted in each Transformer block are trainable (a method known as adapters). Moreover, to boost the performance in a cross-language low-resource scenario, the Transformer model is further tuned on a large intermediate dataset before directly fine-tuning it on a small dataset. With updating fewer than 4% of parameters, (our proposed) PETL-based methods achieve comparable performances with full fine-tuning methods (Vox1-O: 0.55%, Vox1-E: 0.82%, Vox1-H:1.73%). Junyi Peng, Themos Stafylakis, Rongzhi Gu, Oldrich Plchot, Ladislav Mosner, Lukás Burget, Jan Cernocký |
ICASSP | 6 |
| 2023 | Toroidal Probabilistic Spherical Discriminant AnalysisabstractIn speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring back-ends are commonly used, namely cosine scoring and PLDA. We have recently proposed PSDA, an analog to PLDA that uses Von Mises-Fisher distributions instead of Gaussians. In this paper, we present toroidal PSDA (T-PSDA). It extends PSDA with the ability to model within and between-speaker variabilities in toroidal submanifolds of the hypersphere. Like PLDA and PSDA, the model allows closed-form scoring and closed-form EM updates for training. On VoxCeleb, we find T-PSDA accu-racy on par with cosine scoring, while PLDA accuracy is inferior. On NIST SRE’21 we find that T-PSDA gives large accuracy gains compared to both cosine scoring and PLDA.1 Anna Silnova, Niko Brümmer, Albert Swart, Lukás Burget |
ICASSP | 4 |
| 2023 | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki |
INTERSPEECH | 8 |
| 2023 | Description and Analysis of ABC Submission to NIST LRE 2022
Pavel Matejka, Anna Silnova, Josef Slavícek, Ladislav Mosner, Oldrich Plchot, Michal Klco, Junyi Peng, Themos Stafylakis, Lukás Burget |
INTERSPEECH | 9 |
| 2023 | Multi-Channel Speech Separation with Cross-Attention and Beamforming
Ladislav Mosner, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký |
INTERSPEECH | 4 |
| 2023 | Improving Speaker Verification with Self-Pretrained Transformer Models
Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký |
INTERSPEECH | 5 |
| 2022 | DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and ExtractionabstractIn recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks. In this paper, from the time-frequency domain perspective, we propose a densely-connected pyramid complex convolutional network, termed DPCCN, to improve the robustness of speech separation under complicated conditions. Furthermore, we generalize the DPCCN to target speech extraction (TSE) by integrating a new specially designed speaker encoder. Moreover, we also investigate the robustness of DPCCN to unsupervised cross-domain TSE tasks. A Mixture-Remix approach is proposed to adapt the target domain acoustic characteristics for fine-tuning the source model. We evaluate the proposed methods not only under noisy and reverberant in-domain condition, but also in clean but cross-domain conditions. Results show that for both speech separation and extraction, the DPCCN-based systems achieve significantly better performance and robustness than the currently dominating time-domain methods, especially for the cross-domain tasks. Particularly, we find that the Mixture-Remix fine-tuning with DPCCN significantly outperforms the TD-SpeakerBeam for unsupervised cross-domain TSE, with around 3.5 dB SISNR improvement on target domain test set, without any source domain performance degradation. Jiangyu Han, Yanhua Long, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2022 | Multisv: Dataset for Far-Field Multi-Channel Speaker VerificationabstractMotivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experiments with dereverberation, denoising, and speech enhancement. We tackled the ever-present problem of the lack of multi-channel training data by utilizing data simulation on top of clean parts of the Voxceleb corpus. The development and evaluation trials are based on a retransmitted Voices Obscured in Complex Environmental Settings (VOiCES) corpus, which we modified to provide multi-channel trials. We publish full recipes that create the dataset from public sources as the MultiSV dataset, and we provide results with two of our multi-channel speaker verification systems with neural network-based beamforming based either on predicting ideal binary masks or the more recent Conv-TasNet. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2022 | Multi-Channel Speaker Verification with Conv-Tasnet Based BeamformerabstractWe focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, and we adapt it to perform speech enhancement by forcing it to separate speech and additive noise. We experiment with using the STFT of Conv-TasNet outputs to obtain SCMs of speech and noise, and finally, we fine-tune this multi-channel frontend w.r.t. speaker verification objective. We successfully tackle the problem of the lack of a realistic multichannel training set by using simulated data of MultiSV corpus. The analysis is performed on its retransmitted and simulated test parts. We achieve consistent improvements with a 2.7 times smaller model than the baseline based on a scheme with mask estimating NN. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2022 | GPU-Accelerated Forward-Backward Algorithm with Application to Lattice-Free MMIabstractWe propose to express the forward-backward algorithm in terms of operations between sparse matrices in a specific semiring. This new perspective naturally leads to a GPU-friendly algorithm which is easy to implement in Julia or any programming languages with native support of semiring algebra. We use this new implementation to train a TDNN with the LF-MMI objective function and we compare the training time of our system with PyChain—a recently introduced C++/CUDA implementation of the LF-MMI loss. Our implementation is about two times faster while not having to use any approximation such as the "leaky-HMM". Lucas Ondel Yang, Léa-Marie Lam-Yee-Mui, Martin Kocour, Caio F. Corro, Lukás Burget |
ICASSP | 5 |
| 2022 | Speaker adaptation for Wav2vec2 based dysarthric ASR
Murali Karthick Baskar, Tim Herzig, Diana Nguyen, Mireia Díez, Tim Polzehl, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2022 | Probabilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddingsabstractIn speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA.Both have advantages and disadvantages, depending on the context.Cosine scoring follows naturally from the spherical geometry, but for PLDA the blessing is mixed-length normalization Gaussianizes the between-speaker distribution, but violates the assumption of a speaker-independent within-speaker distribution.We propose PSDA, an analogue to PLDA that uses Von Mises-Fisher distributions on the hypersphere for both within and between-class distributions.We show how the self-conjugacy of this distribution gives closed-form likelihood-ratio scores, making it a drop-in replacement for PLDA at scoring time.All kinds of trials can be scored, including single-enroll and multienroll verification, as well as more complex likelihood-ratios that could be used in clustering and diarization.Learning is done via an EM-algorithm with closed-form updates.We explain the model and present some first experiments. Niko Brümmer, Albert Swart, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Themos Stafylakis, Lukás Burget |
INTERSPEECH | 7 |
| 2022 | Revisiting joint decoding based multi-talker speech recognition with DNN acoustic modelabstractIn typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker decoder which is applied on each speaker-specific output stream separately. In this work, we argue that such a scheme is sub-optimal and propose a principled solution that decodes all speakers jointly. We modify the acoustic model to predict joint state posteriors for all speakers, enabling the network to express uncertainty about the attribution of parts of the speech signal to the speakers. We employ a joint decoder that can make use of this uncertainty together with higher-level language information. For this, we revisit decoding algorithms used in factorial generative models in early multi-talker speech recognition systems. In contrast with these early works, we replace the GMM acoustic model with DNN, which provides greater modeling power and simplifies part of the inference. We demonstrate the advantage of joint decoding in proof of concept experiments on a mixed-TIDIGITS dataset. Martin Kocour, Katerina Zmolíková, Lucas Ondel Yang, Jan Svec, Marc Delcroix, Tsubasa Ochiai, Lukás Burget, Jan Cernocký |
INTERSPEECH | 7 |
| 2022 | From Simulated Mixtures to Simulated Conversations as Training Data for End-to-End Neural DiarizationabstractEnd-to-end neural diarization (EEND) is nowadays one of the most prominent research topics in speaker diarization. EEND presents an attractive alternative to standard cascaded diarization systems since a single system is trained at once to deal with the whole diarization problem. Several EEND variants and approaches are being proposed, however, all these models require large amounts of annotated data for training but available annotated data are scarce. Thus, EEND works have used mostly simulated mixtures for training. However, simulated mixtures do not resemble real conversations in many aspects. In this work we present an alternative method for creating synthetic conversations that resemble real ones by using statistics about distributions of pauses and overlaps estimated on genuine conversations. Furthermore, we analyze the effect of the source of the statistics, different augmentations and amounts of data. We demonstrate that our approach performs substantially better than the original one, while reducing the dependence on the fine-tuning stage. Experiments are carried out on 2-speaker telephone conversations of Callhome and DIHARD 3. Together with this publication, we release our implementations of EEND and the method for creating simulated conversations. Index Terms: speaker diarization, end-to-end neural diarization, simulated conversations Federico Landini, Alicia Lozano-Diez, Mireia Díez, Lukás Burget |
INTERSPEECH | 4 |
| 2022 | Learnable Sparse Filterbank for Speaker Verification
Junyi Peng, Rongzhi Gu, Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
INTERSPEECH | 5 |
| 2022 | Training speaker embedding extractors using multi-speaker audio with unknown speaker boundariesabstractIn this paper, we demonstrate a method for training speaker embedding extractors using weak annotation.More specifically, we are using the full VoxCeleb recordings and the name of the celebrities appearing on each video without knowledge of the time intervals the celebrities appear in the video.We show that by combining a baseline speaker diarization algorithm that requires no training or parameter tuning, a modified loss with aggregation over segments, and a two-stage training approach, we are able to train a competitive ResNet-based embedding extractor.Finally, we experiment with two different aggregation functions and analyze their behaviour in terms of their gradients. Themos Stafylakis, Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Anna Silnova, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2022 | An Attention-Based Backend Allowing Efficient Fine-Tuning of Transformer Models for Speaker VerificationabstractIn recent years, self-supervised learning paradigm has received extensive attention due to its great success in various down-stream tasks. However, the fine-tuning strategies for adapting those pre-trained models to speaker verification task have yet to be fully explored. In this paper, we analyze several feature extraction approaches built on top of a pre-trained model, as well as regularization and a learning rate scheduler to stabilize the fine-tuning process and further boost performance: multi-head factorized attentive pooling is proposed to factorize the comparison of speaker representations into multiple phonetic clusters. We regularize towards the parameters of the pre-trained model and we set different learning rates for each layer of the pre-trained model during fine-tuning. The experimental results show our method can significantly shorten the training time to 4 hours and achieve SOTA performance: 0.59%, 0.79% and 1.77% EER on Vox1-O, Vox1-E and Vox1-H, respectively.11Code is available at https://github.com/JunyiPeng00/IEEE-SLT22-Pretrained-Model-for-SV. Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký |
SLT | 5 |
| 2022 | Extracting Speaker and Emotion Information from Self-Supervised Speech Models via Channel-Wise CorrelationsabstractSelf-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first - and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation. Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot, Lukás Burget, Jan Cernocký |
SLT | 5 |
| 2022 | Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: Theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Díez, Lukás Burget |
Comput. Speech Lang. | 4 |
| 2022 | Spelling-Aware Word-Based End-to-End ASRabstractWe propose a new end-to-end architecture for automatic speech recognition that expands the “listen, attend and spell” (LAS) paradigm. While the main word-predicting network is trained to predict words, the secondary, speller network, is optimized to predict word spellings from inner representations of the main network (e.g. word embeddings or context vectors from the attention module). We show that this joint training improves the word error rate of a word-based system and enables solving additional tasks, such as out-of-vocabulary word detection and recovery. The tests are conducted on LibriSpeech dataset consisting of 1000 h of read speech. Ekaterina Egorova, Hari Krishna Vydana, Lukás Burget, Jan Cernocký |
IEEE Signal Process. Lett. | 3 |
| 2022 | Non-Parametric Bayesian Subspace Models for Acoustic Unit DiscoveryabstractThis work investigates subspace non-parametric models for the task of learning a set of acoustic units from unlabeled speech recordings. We constrain the base-measure of a Dirichlet-Process mixture with a phonetic subspace—estimated from other source languages—to build aneducated prior, thereby forcing the learned acoustic units to resemble phones of known source languages. Two types of models are proposed: (i) the Subspace HMM (SHMM) which assumes that the phonetic subspace is the same for every language, (ii) the Hierarchical-Subspace HMM (H-SHMM) which relaxes this assumption and allows to have a language-specific subspace estimated on the unlabeled target data. These models are applied on 3 languages: English, Yoruba and Mboshi and they are compared with various competitive acoustic units discovery baselines. Experimental results show that both subspace models outperform other systems in terms of clustering quality and segmentation accuracy. Moreover, we observe that the H-SHMM provides results superior to the SHMM supporting the idea that language-specific priors are preferable to language-agnostic priors for acoustic unit discovery. Lucas Ondel Yang, Bolaji Yusuf, Lukás Burget, Murat Saraclar |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Eat: Enhanced ASR-TTS for Self-Supervised Speech RecognitionabstractSelf-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the ASR hypotheses before forwarding it to TTS. 2) In the TTS→ASR direction, a hyper-parameter is introduced to scale the attention context from synthesized speech before sending it to ASR to handle out-of-domain data. Training strategies and the effectiveness of the EAT model are explored under out-of-domain data conditions. The results show that EAT reduces the performance gap between supervised and self-supervised training significantly by absolute 2.6% and 2.7% on Librispeech and BABEL respectively. Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Jan Cernocký |
ICASSP | 2 |
| 2021 | Analysis of the but Diarization System for Voxconverse ChallengeabstractThis paper describes the system developed by the BUT team for the fourth track of the VoxCeleb Speaker Recognition Challenge, focusing on diarization on the VoxConverse dataset. The system consists of signal pre-processing, voice activity detection, speaker embedding extraction, an initial agglomerative hierarchical clustering followed by diarization using a Bayesian hidden Markov model, a reclustering step based on per-speaker global embeddings and overlapped speech detection and handling. We provide comparisons for each of the steps and share the implementation of the most relevant modules of our system. Our system scored second in the challenge in terms of the primary metric (diarization error rate) and first according to the secondary metric (Jaccard error rate). Federico Landini, Ondrej Glembek, Pavel Matejka, Johan Rohdin, Lukás Burget, Mireia Díez, Anna Silnova |
ICASSP | 5 |
| 2021 | Jointly Trained Transformers Models for Spoken Language TranslationabstractEnd-to-End and cascade (ASR-MT) spoken language translation (SLT) systems are reaching comparable performances, however, a large degradation is observed when translating the ASR hypothesis in comparison to using oracle input text. In this work, degradation in performance is reduced by creating an End-to-End differentiable pipeline between the ASR and MT systems. In this work, we train SLT systems with ASR objective as an auxiliary loss and both the networks are connected through the neural hidden representations. This training has an End-to-End differentiable path with respect to the final objective function and utilizes the ASR objective for better optimization. This architecture has improved the BLEU score from 41.21 to 44.69. Ensembling the proposed architecture with independently trained ASR and MT systems further improved the BLEU score from 44.69 to 46.9. All the experiments are reported on English-Portuguese speech translation task using the How2 corpus. The final BLEU score is on-par with the best speech translation system on How2 dataset without using any additional training data and language model and using fewer parameters. Hari Krishna Vydana, Martin Karafiát, Katerina Zmolíková, Lukás Burget, Jan Cernocký |
ICASSP | 4 |
| 2021 | A Hierarchical Subspace Model for Language-Attuned Acoustic Unit DiscoveryabstractIn this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach, we frame the task as one of learning embeddings on a low-dimensional phonetic subspace, and simultaneously specify the subspace itself as an embedding on a hyper-subspace. We train the hyper-subspace on a set of transcribed languages and transfer it to the target language. In the target language, we infer both the language and unit embeddings in an unsupervised manner, and in so doing, we simultaneously learn a subspace of units specific to that language and the units that dwell on it. We conduct experiments on TIMIT and two low-resource languages: Mboshi and Yoruba. Results show that our model outperforms major acoustic unit discovery techniques, both in terms of clustering quality and segmentation accuracy. Bolaji Yusuf, Lucas Ondel Yang, Lukás Burget, Jan Cernocký, Murat Saraclar |
ICASSP | 3 |
| 2021 | Text Augmentation for Language Models in High Error Recognition ScenarioabstractWe examine the effect of data augmentation for training of language models for speech recognition. We compare augmentation based on global error statistics with one based on per-word unigram statistics of ASR errors and observe that it is better to only pay attention the global substitution, deletion and insertion rates. This simple scheme also performs consistently better than label smoothing and its sampled variants. Additionally, we investigate into the behavior of perplexity estimated on augmented data, but conclude that it gives no better prediction of the final error rate. Our best augmentation scheme increases the absolute WER improvement from second-pass rescoring from 1.1 % to 1.9 % absolute on the CHiMe-6 challenge. Karel Benes, Lukás Burget |
Interspeech | 2 |
| 2021 | Out-of-Vocabulary Words Detection with Attention and CTC Alignments in an End-to-End ASR System
Ekaterina Egorova, Hari Krishna Vydana, Lukás Burget, Jan Cernocký |
Interspeech | 3 |
| 2021 | Effective Phase Encoding for End-To-End Speaker Verification
Junyi Peng, Xiaoyang Qu, Rongzhi Gu, Jianzong Wang, Jing Xiao 0006, Lukás Burget, Jan Cernocký |
Interspeech | 6 |
| 2021 | ICSpk: Interpretable Complex Speaker Embedding Extractor from Raw Waveform
Junyi Peng, Xiaoyang Qu, Jianzong Wang, Rongzhi Gu, Jing Xiao 0006, Lukás Burget, Jan Cernocký |
Interspeech | 6 |
| 2021 | Speaker Embeddings by Modeling Channel-Wise CorrelationsabstractSpeaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling along the time axis. In this paper we examine an alternative pooling method, where pairwise correlations between channels for given frequencies are used as statistics. The method is inspired by style-transfer methods in computer vision, where the style of an image, modeled by the matrix of channel-wise correlations, is transferred to another image, in order to produce a new image having the style of the first and the content of the second. By drawing analogies between image style and speaker characteristics, and between image content and phonetic sequence, we explore the use of such channel-wise correlations features to train a ResNet architecture in an end-to-end fashion. Our experiments on VoxCeleb demonstrate the effectiveness of the proposed pooling method in speaker recognition. Themos Stafylakis, Johan Rohdin, Lukás Burget |
Interspeech | 3 |
| 2021 | Integration of Variational Autoencoder and Spatial Clustering for Adaptive Multi-Channel Neural Speech SeparationabstractIn this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in several works. As the spectral model, previous works used either factorial generative models of the mixed speech or discriminative neural networks. In our work, we combine the strengths of both approaches, by building a factorial model based on a generative neural network, a variational autoencoder. By doing so, we can exploit the modeling power of neural networks, but at the same time, keep a structured model. Such a model can be advantageous when adapting to new noise conditions as only the noise part of the model needs to be modified. We show experimentally, that our model significantly outperforms previous factorial model based on Gaussian mixture model (DOLPHIN), performs comparably to integration of permutation invariant training with spatial clustering, and enables us to easily adapt to new noise conditions. Katerina Zmolíková, Marc Delcroix, Lukás Burget, Tomohiro Nakatani, Jan Cernocký |
SLT | 3 |
| 2020 | Investigation of Specaugment for Deep Speaker Embedding LearningabstractSpecAugment is a newly proposed data augmentation method for speech recognition. By randomly masking bands in the log Mel spectogram this method leads to impressive performance improvements. In this paper, we investigate the usage of SpecAugment for speaker verification tasks. Two different models, namely 1-D convolutional TDNN and 2-D convolutional ResNet34, trained with either Softmax or AAM-Softmax loss, are used to analyze SpecAugment's effectiveness. Experiments are carried out on the Voxceleb and NIST SRE 2016 dataset. By applying SpecAugment to the original clean data in an on-the-fly manner without complex off-line data augmentation methods, we obtained 3.72% and 11.49% EER for NIST SRE 2016 Cantonese and Tagalog, respectively. For Voxceleb1 evaluation set, we obtained 1.47% EER. Shuai Wang 0016, Johan Rohdin, Oldrich Plchot, Lukás Burget, Kai Yu 0004, Jan Cernocký |
ICASSP | 4 |
| 2020 | Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization ChallengeabstractThis paper presents an analysis of our diarization system winning the second DIHARD speech diarization challenge, track 1. This system is based on clustering x-vector speaker embeddings extracted every 0.25s from short segments of the input recording. In this paper, we focus on the two x-vector clustering methods employed, namely Agglomerative Hierarchical Clustering followed by a clustering based on Bayesian Hidden Markov Model (BHMM). Even though the system submitted to the challenge had further post-processing steps, we will show that using this BHMM solely is enough to achieve the best performance in the challenge. The analysis will show improvements achieved by optimizing individual processing steps, including a simple procedure to effectively perform "domain adaptation" by Probabilistic Linear Discriminant Analysis model interpolation. All experiments are performed in the DIHARD II evaluation framework. Mireia Díez, Lukás Burget, Federico Landini, Shuai Wang 0016, Jan Cernocký |
ICASSP | 2 |
| 2020 | But System for the Second Dihard Speech Diarization ChallengeabstractThis paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering based on Bayes hidden Markov model and variational Bayes inference. We provide a comparison of the improvement given by each step and share the implementation of the core of the system. For tracks 3 and 4 with recordings from the Fifth CHiME Challenge, we explored different approaches for doing multi-channel diarization and our best performance was obtained when applying AHC on the fusion of per channel probabilistic linear discriminant analysis scores. Federico Landini, Shuai Wang 0016, Mireia Díez, Lukás Burget, Pavel Matejka, Katerina Zmolíková, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Ondrej Novotný, Hossein Zeinali, Johan Rohdin |
ICASSP | 4 |
| 2020 | BUT Text-Dependent Speaker Verification System for SdSV Challenge 2020abstractIn this paper, we present the winning BUT submission for the text-dependent task of the SdSV challenge 2020. Given the large amount of training data available in this challenge, we explore successful techniques from text-independent systems in the text-dependent scenario. In particular, we trained x-vector extractors on both in-domain and out-of-domain datasets and combine them with i-vectors trained on concatenated MFCCs and bottleneck features, which have proven effective for the text-dependent scenario. Moreover, we proposed the use of phrase-dependent PLDA backend for scoring and its combination with a simple phrase recognizer, which brings up to 63% relative improvement on our development set with respect to using standard PLDA. Finally, we combine our different i-vector and x-vector based systems using a simple linear logistic regression score level fusion, which provides 28% relative improvement on the evaluation set with respect to our best single system Alicia Lozano-Diez, Anna Silnova, Bhargav Pulugundla, Johan Rohdin, Karel Veselý, Lukás Burget, Oldrich Plchot, Ondrej Glembek, Ondrej Novotný, Pavel Matejka |
INTERSPEECH | 6 |
| 2020 | SdSV Challenge 2020: Large-Scale Evaluation of Short-Duration Speaker Verification
Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
INTERSPEECH | 4 |
| 2020 | 13 years of speaker recognition research at BUT, with longitudinal analysis of NIST SRE
Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Johan Rohdin, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Ondrej Novotný, Mireia Díez, Jan Cernocký |
Comput. Speech Lang. | 4 |
| 2020 | End-to-end DNN based text-independent speaker recognition for long and short utterances
Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget, Ondrej Glembek |
Comput. Speech Lang. | 6 |
| 2020 | Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice PriorsabstractIn our previous work, we introduced our Bayesian Hidden Markov Model with eigenvoice priors, which has been recently recognized as the state-of-the-art model for Speaker Diarization. In this article we present a more complete analysis of the Diarization system. The inference of the model is fully described and derivations of all update formulas are provided for a complete understanding of the algorithm. An extensive analysis on the effect, sensitivity and interactions of all model parameters is provided, which might be used as a guide for their optimal setting. The newly introduced speaker regularization coefficient allows us to control the number of speakers inferred in an utterance. A naive speaker model merging strategy is also presented, which allows to drive the variational inference out of local optima. Experiments for the different diarization scenarios are presented on CALLHOME and DIHARD datasets. Mireia Díez, Lukás Burget, Federico Landini, Jan Cernocký |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Learning Document Embeddings Along With Their UncertaintiesabstractMajority of the text modeling techniques yield only point-estimates of document embeddings and lack in capturing the uncertainty of the estimates. These uncertainties give a notion of how well the embeddings represent a document. We present Bayesian subspace multinomial model (Bayesian SMM), a generative log-linear model that learns to represent documents in the form of Gaussian distributions, thereby encoding the uncertainty in its covariance. Additionally, in the proposed Bayesian SMM, we address a commonly encountered problem of intractability that appears during variational inference in mixed-logit models. We also present a generative Gaussian linear classifier for topic identification that exploits the uncertainty in document embeddings. Our intrinsic evaluation using perplexity measure shows that the proposed Bayesian SMM fits the unseen test data better as compared to the state-of-the-art neural variational document model on (Fisher) speech and (20Newsgroups) text corpora. Our topic identification experiments show that the proposed systems are robust to over-fitting on unseen test data. The topic ID results show that the proposed model outperforms state-of-the-art unsupervised topic models and achieve comparable results to the state-of-the-art fully supervised discriminative models. Santosh Kesiraju, Oldrich Plchot, Lukás Burget, Suryakanth V. Gangashetty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Speaker Verification with Application-Aware BeamformingabstractMultichannel speech processing applications usually employ beamformers as means of speech enhancement through spatial filtering. Beamformers with learnable parameters require training to minimize a loss function that is not necessarily correlated with the final objective. In this paper, we present a framework employing recent neural network based generalized eigenvalue beamformer and application-specific model that allows for optimization of beamformer w.r.t. target application. In our case, the application is speaker verification which utilizes a speaker embedding (x-vector) extractor that conveniently comes with desired loss. We show that application-specific training of the beamformer brings performance improvements over a system trained in the standard way. We perform our analysis on the recently introduced VOiCES corpus which contains multichannel data and allows us to modify the evaluation trials such that enrollment recordings remain single-channel and test utterances are multichannel. Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Lukás Burget, Jan Cernocký |
ASRU | 4 |
| 2019 | A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: The Deepmine DatabaseabstractDeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers and 540 thousand recordings overall, more than 480 hours of speech are transcribed. It is the first public large-scale speaker verification database in Persian, the largest public text-dependent and text-prompted speaker verification database in English, and the largest public evaluation dataset for text-independent speaker verification. It has a good coverage of age, gender, and accents. We provide several evaluation protocols for each part of the database to allow for research on different aspects of speaker verification. We also provide the results of several experiments that can be considered as baselines: HMM-based i-vectors for text-dependent speaker verification, and HMM-based as well as state-of-the-art deep neural network based ASR. We demonstrate that the database can serve for training robust ASR models. Hossein Zeinali, Lukás Burget, Jan Cernocký |
ASRU | 2 |
| 2019 | Promising Accurate Prefix Boosting for Sequence-to-sequence ASRabstractIn this paper, we present promising accurate prefix boosting (PAPB), a discriminative training technique for attention based sequence-to-sequence (seq2seq) ASR. PAPB is devised to unify the training and testing scheme effectively. The training procedure involves maximizing the score of each partial correct sequence obtained during beam search compared to other hypotheses. The training objective also includes minimization of token (character) error rate. PAPB shows its efficacy by achieving 10.8% and 3.8% WER with and without external RNNLM respectively on Wall Street Journal dataset. Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Martin Karafiát, Takaaki Hori, Jan Cernocký |
ICASSP | 2 |
| 2019 | Discriminatively Re-trained I-vector Extractor for Speaker RecognitionabstractIn this work we revisit discriminative training of the i-vector extractor component in the standard speaker verification (SV) system. The motivation of our research lies in the robustness and stability of this large generative model, which we want to preserve, and focus its power towards any intended SV task. We show that after generative initialization of the i-vector extractor, we can further refine it with discriminative training and obtain i-vectors that lead to better performance on various benchmarks representing different acoustic domains. Ondrej Novotný, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Pavel Matejka |
ICASSP | 4 |
| 2019 | Speaker Verification Using End-to-end Adversarial Language AdaptationabstractIn this paper we investigate the use of adversarial domain adaptation for addressing the problem of language mismatch between speaker recognition corpora. In the context of speaker verification, adversarial domain adaptation methods aim at minimizing certain divergences between the distribution that the utterance-level features follow (i.e. speaker embeddings) when drawn from source and target domains (i.e. languages), while preserving their capacity in recognizing speakers. Neural architectures for extracting utterance-level representations enable us to apply adversarial adaptation methods in an end-to-end fashion and train the network jointly with the standard cross-entropy loss. We examine several configurations, such as the use of (pseudo-)labels on the target domain as well as domain labels in the feature extractor, and we demonstrate the effectiveness of our method on the challenging NIST SRE16 and SRE18 benchmarks. Johan Rohdin, Themos Stafylakis, Anna Silnova, Hossein Zeinali, Lukás Burget, Oldrich Plchot |
ICASSP | 5 |
| 2019 | How to Improve Your Speaker Embeddings Extractor in Generic ToolkitsabstractRecently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine several tricks in training, such as the effects of normalizing input features and pooled statistics, different methods for preventing overfitting as well as alternative non-linearities that can be used instead of Rectifier Linear Units. In addition, we investigate the difference in performance between TDNN and CNN, and between two types of attention mechanism. Experimental results on Speaker in the Wild, SRE 2016 and SRE 2018 datasets demonstrate the effectiveness of the proposed implementation. Hossein Zeinali, Lukás Burget, Johan Rohdin, Themos Stafylakis, Jan Cernocký |
ICASSP | 2 |
| 2019 | Semi-Supervised Sequence-to-Sequence ASR Using Unpaired Speech and TextabstractSequence-to-sequence automatic speech recognition (ASR) models require large quantities of data to attain high performance. For this reason, there has been a recent surge in interest for unsupervised and semi-supervised training in such models. This work builds upon recent results showing notable improvements in semi-supervised training using cycle-consistency and related techniques. Such techniques derive training procedures and losses able to leverage unpaired speech and/or text data by combining ASR with Text-to-Speech (TTS) models. In particular, this work proposes a new semi-supervised loss combining an end-to-end differentiable ASR$\rightarrow$TTS loss with TTS$\rightarrow$ASR loss. The method is able to leverage both unpaired speech and text data to outperform recently proposed related techniques in terms of \%WER. We provide extensive results analyzing the impact of data quantity and speech and text modalities and show consistent gains across WSJ and Librispeech corpora. Our code is provided in ESPnet to reproduce the experiments. Murali Karthick Baskar, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Takaaki Hori, Lukás Burget, Jan Cernocký |
INTERSPEECH | 5 |
| 2019 | Bayesian HMM Based x-Vector Clustering for Speaker Diarization
Mireia Díez, Lukás Burget, Shuai Wang 0016, Johan Rohdin, Jan Cernocký |
INTERSPEECH | 2 |
| 2019 | Analysis of BUT Submission in Far-Field Scenarios of VOiCES 2019 Challenge
Pavel Matejka, Oldrich Plchot, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Lukás Burget, Ondrej Novotný, Ondrej Glembek |
INTERSPEECH | 6 |
| 2019 | Analysis of BUT Submission in Far-Field Scenarios of VOiCES 2019 Challenge
Pavel Matejka, Oldrich Plchot, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Lukás Burget, Ondrej Novotný, Ondrej Glembek |
INTERSPEECH | 6 |
| 2019 | Factorization of Discriminatively Trained i-Vector Extractor for Speaker RecognitionabstractIn this work, we continue in our research on i-vector extractor for speaker verification (SV) and we optimize its architecture for fast and effective discriminative training. We were motivated by computational and memory requirements caused by the large number of parameters of the original generative i-vector model. Our aim is to preserve the power of the original generative model, and at the same time focus the model towards extraction of speaker-related information. We show that it is possible to represent a standard generative i-vector extractor by a model with significantly less parameters and obtain similar performance on SV tasks. We can further refine this compact model by discriminative training and obtain i-vectors that lead to better performance on various SV benchmarks representing different acoustic domains. Ondrej Novotný, Oldrich Plchot, Ondrej Glembek, Lukás Burget |
INTERSPEECH | 4 |
| 2019 | Bayesian Subspace Hidden Markov Model for Acoustic Unit DiscoveryabstractThis work tackles the problem of learning a set of language specific acoustic units from unlabeled speech recordings given a set of labeled recordings from other languages. Our approach may be described by the following two steps procedure: first the model learns the notion of acoustic units from the labelled data and then the model uses its knowledge to find new acoustic units on the target language. We implement this process with the Bayesian Subspace Hidden Markov Model (SHMM), a model akin to the Subspace Gaussian Mixture Model (SGMM) where each low dimensional embedding represents an acoustic unit rather than just a HMM's state. The subspace is trained on 3 languages from the GlobalPhone corpus (German, Polish and Spanish) and the AUs are discovered on the TIMIT corpus. Results, measured in equivalent Phone Error Rate, show that this approach significantly outperforms previous HMM based acoustic units discovery systems and compares favorably with the Variational Auto Encoder-HMM. Lucas Ondel Yang, Hari Krishna Vydana, Lukás Burget, Jan Cernocký |
INTERSPEECH | 3 |
| 2019 | Self-Supervised Speaker EmbeddingsabstractContrary to i-vectors, speaker embeddings such as x-vectors are incapable of leveraging unlabelled utterances, due to the classification loss over training speakers. In this paper, we explore an alternative training strategy to enable the use of unlabelled utterances in training. We propose to train speaker embedding extractors via reconstructing the frames of a target speech segment, given the inferred embedding of another speech segment of the same utterance. We do this by attaching to the standard speaker embedding extractor a decoder network, which we feed not merely with the speaker embedding, but also with the estimated phone sequence of the target frame sequence. The reconstruction loss can be used either as a single objective, or be combined with the standard speaker classification loss. In the latter case, it acts as a regularizer, encouraging generalizability to speakers unseen during training. In all cases, the proposed architectures are trained from scratch and in an end-to-end fashion. We demonstrate the benefits from the proposed approach on VoxCeleb and Speakers in the wild, and we report notable improvements over the baseline. Themos Stafylakis, Johan Rohdin, Oldrich Plchot, Petr Mizera, Lukás Burget |
INTERSPEECH | 5 |
| 2019 | On the Usage of Phonetic Information for Text-Independent Speaker Embedding Extraction
Shuai Wang 0016, Johan Rohdin, Lukás Burget, Oldrich Plchot, Yanmin Qian, Kai Yu 0004, Jan Cernocký |
INTERSPEECH | 3 |
| 2019 | Detecting Spoofing Attacks Using VGG and SincNet: BUT-Omilia Submission to ASVspoof 2019 ChallengeabstractIn this paper, we present the system description of the joint efforts of Brno University of Technology (BUT) and Omilia -- Conversational Intelligence for the ASVSpoof2019 Spoofing and Countermeasures Challenge. The primary submission for Physical access (PA) is a fusion of two VGG networks, trained on single and two-channels features. For Logical access (LA), our primary system is a fusion of VGG and the recently introduced SincNet architecture. The results on PA show that the proposed networks yield very competitive performance in all conditions and achieved 86\:\% relative improvement compared to the official baseline. On the other hand, the results on LA showed that although the proposed architecture and training strategy performs very well on certain spoofing attacks, it fails to generalize to certain attacks that are unseen during training. Hossein Zeinali, Themos Stafylakis, Georgia Athanasopoulou, Johan Rohdin, Ioannis Gkinis, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2019 | Analysis of DNN Speech Signal Enhancement for Robust Speaker Recognition
Ondrej Novotný, Oldrich Plchot, Ondrej Glembek, Jan Cernocký, Lukás Burget |
Comput. Speech Lang. | 5 |
| 2018 | Out-of-Vocabulary Word Recovery using FST-Based Subword Unit Clustering in a Hybrid ASR SystemabstractThe paper presents a new approach to extracting useful information from out-of-vocabulary (OOV) speech regions in ASR system output. The system makes use of a hybrid decoding network with both words and sub-word units. In the decoded lattices, candidates for OOV regions are identified as sub-graphs of sub-word units. To facilitate OOV word recovery, we search for recurring OOV s by clustering the detected candidate OOV s. The metrics for clustering is based on a comparison of the sub-graphs corresponding to the OOV candidates. The proposed method discovers repeating out-of-vocabulary words and finds their graphemic representation more robustly than more conventional techniques taking into account only one best sub-word string hypotheses. Ekaterina Egorova, Lukás Burget |
ICASSP | 2 |
| 2018 | Analysis of Multilingual Blstm Acoustic Model on Low and High Resource LanguagesabstractThe paper provides an analysis of automatic speech recognition systems (ASR) based on multilingual BLSTM, where we used multi-task training with separate classification layer for each language. The focus is on low resource languages, where only a limited amount of transcribed speech is available. In such scenario, we found it essential to train the ASR systems in a multilingual fashion and we report superior results obtained with pre-trained multilingual BLSTM on this task. The high resource languages are also taken into account and we show the importance of language richness for multilingual training. Next, we present the performance of this technique as a function of amount of target language data. The importance of including context information into BLSTM multilingual systems is also stressed, and we report increased resilience of large NNs to overtraining in case of multi-task training. Martin Karafiát, Murali Karthick Baskar, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký |
ICASSP | 5 |
| 2018 | Bayesian Models for Unit Discovery on a Very Low Resource LanguageabstractDeveloping speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus. Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur |
ICASSP | 8 |
| 2018 | End-to-End DNN Based Speaker Recognition Inspired by I-Vector and PLDAabstractRecently, several end-to-end speaker verification systems based on deep neural networks (DNNs) have been proposed. These systems have been proven to be competitive for text-dependent tasks as well as for text-independent tasks with short utterances. However, for text-independent tasks with longer utterances, end-to-end systems are still outperformed by standard i-vector + PLDA systems. In this work, we develop an end-to-end speaker verification system that is initialized to mimic an i-vector + PLDA baseline. The system is then further trained in an end-to-end manner but regularized so that it does not deviate too far from the initial system. In this way we mitigate overfitting which normally limits the performance of end-to-end systems. The proposed system outperforms the i-vector + PLDA baseline on both long and short duration utterances. Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget |
ICASSP | 6 |
| 2018 | i-Vectors in Language Modeling: An Efficient Way of Domain Adaptation for Feed-Forward Models
Karel Benes, Santosh Kesiraju, Lukás Burget |
INTERSPEECH | 3 |
| 2018 | BUT System for DIHARD Speech Diarization Challenge 2018
Mireia Díez, Federico Landini, Lukás Burget, Johan Rohdin, Anna Silnova, Katerina Zmolíková, Ondrej Novotný, Karel Veselý, Ondrej Glembek, Oldrich Plchot, Ladislav Mosner, Pavel Matejka |
INTERSPEECH | 3 |
| 2018 | BUT OpenSAT 2017 Speech Recognition System
Martin Karafiát, Murali Karthick Baskar, Igor Szöke, Vladimir Malenovsky, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký |
INTERSPEECH | 7 |
| 2018 | BUT System for Low Resource Indian Language ASR
Bhargav Pulugundla, Murali Karthick Baskar, Santosh Kesiraju, Ekaterina Egorova, Martin Karafiát, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2018 | Fast Variational Bayes for Heavy-tailed PLDA Applied to i-vectors and x-vectorsabstractThe standard state-of-the-art backend for text-independent speaker recognizers that use i-vectors or x-vectors, is Gaussian PLDA (G-PLDA), assisted by a Gaussianization step involving length normalization. G-PLDA can be trained with both generative or discriminative methods. It has long been known that heavy-tailed PLDA (HT-PLDA), applied without length normalization, gives similar accuracy, but at considerable extra computational cost. We have recently introduced a fast scoring algorithm for a discriminatively trained HT-PLDA backend. This paper extends that work by introducing a fast, variational Bayes, generative training algorithm. We compare old and new backends, with and without length-normalization, with i-vectors and x-vectors, on SRE'10, SRE'16 and SITW. Anna Silnova, Niko Brümmer, Daniel Garcia-Romero, David Snyder, Lukás Burget |
INTERSPEECH | 5 |
| 2017 | Residual memory networks: Feed-forward approach to learn long-term temporal dependenciesabstractTraining deep recurrent neural network (RNN) architectures is complicated due to the increased network complexity. This disrupts the learning of higher order abstracts using deep RNN. In case of feed-forward networks training deep structures is simple and faster while learning long-term temporal information is not possible. In this paper we propose a residual memory neural network (RMN) architecture to model short-time dependencies using deep feed-forward layers having residual and time delayed connections. The residual connection paves way to construct deeper networks by enabling unhindered flow of gradients and the time delay units capture temporal information with shared weights. The number of layers in RMN signifies both the hierarchical processing depth and temporal depth. The computational complexity in training RMN is significantly less when compared to deep recurrent networks. RMN is further extended as bi-directional RMN (BRMN) to capture both past and future information. Experimental analysis is done on AMI corpus to substantiate the capability of RMN in learning long-term information and hierarchical information. Recognition performance of RMN trained with 300 hours of Switchboard corpus is compared with various state-of-the-art LVCSR systems. The results indicate that RMN and BRMN gains 6 % and 3.8 % relative improvement over LSTM and BLSTM networks. Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Karel Veselý, Frantisek Grézl, Jan Cernocký |
ICASSP | 3 |
| 2017 | Bayesian joint-sequence models for grapheme-to-phoneme conversionabstractWe describe a fully Bayesian approach to grapheme-to-phoneme conversion based on the joint-sequence model (JSM). Usually, standard smoothed n-gram language models (LM, e.g. Kneser-Ney) are used with JSMs to model graphone sequences (joint grapheme-phoneme pairs). However, we take a Bayesian approach using a hierarchical Pitman-Yor-Process LM. This provides an elegant alternative to using smoothing techniques to avoid over-training. No held-out sets and complex parameter tuning is necessary, and several convergence problems encountered in the discounted Expectation-Maximization (as used in the smoothed JSMs) are avoided. Every step is modeled by weighted finite state transducers and implemented with standard operations from the OpenFST toolkit. We evaluate our model on a standard data set (CMUdict), where it gives comparable results to the previously reported smoothed JSMs in terms of phoneme-error rate while requiring a much smaller training/testing time. Most importantly, our model can be used in a Bayesian framework and for (partly) un-supervised training. Mirko Hannemann, Jan Trmal, Lucas Ondel Yang, Santosh Kesiraju, Lukás Burget |
ICASSP | 5 |
| 2017 | Topic identification of spoken documents using unsupervised acoustic unit discoveryabstractThis paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The discovered phone-like (acoustic) units are further fed into the conventional topic ID framework. Using multilingual bottleneck features for the acoustic unit discovery, we show that the proposed method outperforms other systems that are based on cross-lingual phoneme recognizer. Santosh Kesiraju, Raghavendra Pappagari, Lucas Ondel Yang, Lukás Burget, Najim Dehak, Sanjeev Khudanpur, Jan Cernocký, Suryakanth V. Gangashetty |
ICASSP | 4 |
| 2017 | An empirical evaluation of zero resource acoustic unit discoveryabstractAcoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provided linguistic knowledge and transcribed speech are unavailable. Therefore, to further facilitate zero-resource AUD process, in this paper, we demonstrate acoustic feature representations can be significantly improved by (i) performing linear discriminant analysis (LDA) in an unsupervised self-trained fashion, and (ii) leveraging resources of other languages through building a multilingual bottleneck (BN) feature extractor to give effective cross-lingual generalization. Moreover, we perform comprehensive evaluations of AUD efficacy on multiple downstream speech applications, and their correlated performance suggests that AUD evaluations are feasible using different alternative language resources when only a subset of these evaluation resources can be available in typical zero resource applications. Chunxi Liu, Jinyi Yang, Santosh Kesiraju, Alena Rott, Lucas Ondel Yang, Pegah Ghahremani, Najim Dehak, Lukás Burget, Sanjeev Khudanpur |
ICASSP | 9 |
| 2017 | Bayesian phonotactic Language Model for Acoustic Unit DiscoveryabstractRecent work on Acoustic Unit Discovery (AUD) has led to the development of a non-parametric Bayesian phone-loop model where the prior over the probability of the phone-like units is assumed to be sampled from a Dirichlet Process (DP). In this work, we propose to improve this model by incorporating a Hierarchical Pitman-Yor based bigram Language Model on top of the units' transitions. This new model makes use of the phonotactic context information but assumes a fixed number of units. To remedy this limitation we first train a DP phone-loop model to infer the number of units, then, the bigram phone-loop is initialized from the DP phone-loop and trained until convergence of its parameters. Results show an absolute improvement of 1–2%on the Normalized Mutual Information (NMI) metric. Furthermore, we show that, combined with Multilingual Bottleneck (MBN) features the model yields a same or higher NMI as an English phone recogniser trained on TIMIT. Lucas Ondel Yang, Lukás Burget, Jan Cernocký, Santosh Kesiraju |
ICASSP | 2 |
| 2017 | Residual Memory Networks in Language Modeling: Improving the Reputation of Feed-Forward Networks
Karel Benes, Murali Karthick Baskar, Lukás Burget |
INTERSPEECH | 3 |
| 2017 | 2016 BUT Babel System: Multilingual BLSTM Acoustic Model with i-Vector Based Adaptation
Martin Karafiát, Murali Karthick Baskar, Pavel Matejka, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2017 | Analysis of Score Normalization in Multilingual Speaker Recognition
Pavel Matejka, Ondrej Novotný, Oldrich Plchot, Lukás Burget, Mireia Díez, Jan Cernocký |
INTERSPEECH | 4 |
| 2017 | Team ELISA System for DARPA LORELEI Speech Evaluation 2016
Pavlos Papadopoulos, Ruchir Travadi, Colin Vaz, Nikos Malandrakis, Ulf Hermjakob, Nima Pourdamghani, Michael Pust, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Ondrej Glembek, Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Mark Hasegawa-Johnson, Heng Ji 0001, Jonathan May, Kevin Knight, Shri Narayanan |
INTERSPEECH | 15 |
| 2017 | Alternative Approaches to Neural Network Based Speaker Verification
Anna Silnova, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2017 | Semi-Supervised DNN Training with Word Selection for ASR
Karel Veselý, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2017 | Text-dependent speaker verification based on i-vectors, Neural Networks and Hidden Markov Models
Hossein Zeinali, Hossein Sameti, Lukás Burget, Jan Cernocký |
Comput. Speech Lang. | 3 |
| 2017 | HMM-Based Phrase-Independent i-Vector Extractor for Text-Dependent Speaker VerificationabstractThe low-dimensional i-vector representation of speech segments is used in the state-of-the-art text-independent speaker verification systems. However, i-vectors were deemed unsuitable for the text-dependent task, where simpler and older speaker recognition approaches were found more effective. In this work, we propose a straightforward hidden Markov model (HMM) based extension of the i-vector approach, which allows i-vectors to be successfully applied to text-dependent speaker verification. In our approach, the Universal Background Model (UBM) for training phrase-independent i-vector extractor is based on a set of monophone HMMs instead of the standard Gaussian Mixture Model (GMM). To compensate for the channel variability, we propose to precondition i-vectors using a regularized variant of within-class covariance normalization, which can be robustly estimated in a phrase-dependent fashion on the small datasets available for the text-dependent task. The verification scores are cosine similarities between the i-vectors normalized using phrase-dependent s-norm. The experimental results on RSR2015 and RedDots databases confirm the effectiveness of the proposed approach, especially in rejecting test utterances with a wrong phrase. A simple MFCC based i-vector/HMM system performs competitively when compared to very computationally expensive DNN-based approaches or the conventional relevance MAP GMM-UBM, which does not allow for compact speaker representations. To our knowledge, this paper presents the best published results obtained with a single system on both RSR2015 and RedDots dataset. Hossein Zeinali, Hossein Sameti, Lukás Burget |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Multilingual region-dependent transformsabstractIn recent years, trained feature extraction (FE) schemes based on neural networks have replaced or complemented traditional approaches in top performing systems. This paper deals with FE in multilingual scenarios with a target language with low amount of transcribed data. Continuing our previous work on multilingual training of Stacked Bottle-Neck Neural Network FE schemes, we concentrate on improving the discriminatively trained Region-Dependent Transforms. We show that multilingual training of RDT can be implemented by merging statistics from several languages. In our case we used up to 11 source languages to build a FE which generalize well for a new language. This allows us to build a strong bootstrapping model for the final ASR system. The results are produced on IARPA Babel data. Martin Karafiát, Lukás Burget, Frantisek Grézl, Karel Veselý, Jan Cernocký |
ICASSP | 2 |
| 2016 | Analysis of DNN approaches to speaker identificationabstractThis work studies the usage of the Deep Neural Network (DNN) Bottleneck (BN) features together with the traditional MFCC features in the task of i-vector-based speaker recognition. We decouple the sufficient statistics extraction by using separate GMM models for frame alignment, and for statistics normalization and we analyze the usage of BN and MFCC features (and their concatenation) in the two stages. We also show the effect of using full-covariance GMM models, and, as a contrast, we compare the result to the recent DNN-alignment approach. On the NIST SRE2010, telephone condition, we show 60% relative gain over the traditional MFCC baseline for EER (and similar for the NIST DCF metrics), resulting in 0.94% EER. Pavel Matejka, Ondrej Glembek, Ondrej Novotný, Oldrich Plchot, Frantisek Grézl, Lukás Burget, Jan Cernocký |
ICASSP | 6 |
| 2016 | Audio enhancing with DNN autoencoder for speaker recognitionabstractIn this paper we present a design of a DNN-based autoencoder for speech enhancement and its use for speaker recognition systems for distant microphones and noisy data. We started with augmenting the Fisher database with artificially noised and reverberated data and trained the autoencoder to map noisy and reverberated speech to its clean version. We use the autoencoder as a preprocessing step in the later stage of modelling in state-of-the-art text-dependent and text-independent speaker recognition systems. We report relative improvements up to 50% for the text-dependent system and up to 48% for the text-independent one. With text-independent system, we present a more detailed analysis on various conditions of NIST SRE 2010 and PRISM suggesting that the proposed preprocessig is a promising and efficient way to build a robust speaker recognition system for distant microphone and noisy data. Oldrich Plchot, Lukás Burget, Hagai Aronowitz, Pavel Matejka |
ICASSP | 2 |
| 2016 | Sequence summarizing neural network for speaker adaptationabstractIn this paper, we propose a DNN adaptation technique, where the i-vector extractor is replaced by a Sequence Summarizing Neural Network (SSNN). Similarly to i-vector extractor, the SSNN produces a "summary vector", representing an acoustic summary of an utterance. Such vector is then appended to the input of main network, while both networks are trained together optimizing single loss function. Both the i-vector and SSNN speaker adaptation methods are compared on AMI meeting data. The results show comparable performance of both techniques on FBANK system with frame-classification training. Moreover, appending both the i-vector and "summary vector" to the FBANK features leads to additional improvement comparable to the performance of FMLLR adapted DNN system. Karel Veselý, Shinji Watanabe 0001, Katerina Zmolíková, Martin Karafiát, Lukás Burget, Jan Cernocký |
ICASSP | 5 |
| 2016 | Learning Document Representations Using Subspace Multinomial Model
Santosh Kesiraju, Lukás Burget, Igor Szöke, Jan Cernocký |
INTERSPEECH | 2 |
| 2016 | Exploiting Hidden-Layer Responses of Deep Neural Networks for Language RecognitionabstractAbstract : The most popular way to apply Deep Neural Network (DNN) for Language IDentification (LID) involves the extraction of bottleneck features from a network that was trained on automatic speech recognition task. These features are modeled using a classical I-vector system. Recently, a more direct DNN approach was proposed, it consists of estimating the language posteriors directly from a stacked frames input. The final decision score is based on averaging the scores for all the frames for a given speech segment. In this paper, we extended the direct DNN approach by modeling all hidden-layer activations rather than just averaging the output scores. One super-vector per utterance is formed by concatenating all hidden-layer responses. The dimensionality of this vector is then reduced using a Principal Component Analysis (PCA). The obtained reduce vector summarizes the most discriminative features for language recognition based on the trained DNNs. We evaluated this approach in NIST 2015 language recognition evaluation. The performances achieved by the proposed approach are very competitive to the classical I-vector baseline. Sri Harish Reddy Mallidi, Lukás Burget, Oldrich Plchot, Najim Dehak |
INTERSPEECH | 3 |
| 2016 | Analysis of Speaker Recognition Systems in Realistic Scenarios of the SITW 2016 Challenge
Ondrej Novotný, Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Jan Cernocký |
INTERSPEECH | 5 |
| 2016 | Sequence Summarizing Neural Networks for Spoken Language Recognition
Jan Pesán, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2016 | i-Vector/HMM Based Text-Dependent Speaker Verification System for RedDots Challenge
Hossein Zeinali, Hossein Sameti, Lukás Burget, Jan Cernocký, Nooshin Maghsoodi, Pavel Matejka |
INTERSPEECH | 3 |
| 2016 | Data Selection by Sequence Summarizing Neural Network in Mismatch Condition Training
Katerina Zmolíková, Martin Karafiát, Karel Veselý, Marc Delcroix, Shinji Watanabe 0001, Lukás Burget, Jan Cernocký |
INTERSPEECH | 6 |
| 2016 | Analysis of the DNN-based SRE systems in multi-language conditionsabstractThis paper analyzes the behavior of our state-of-the-art Deep Neural Network/i-vector/PLDA-based speaker recognition systems in multi-language conditions. On the “Language Pack” of the PRISM set, we evaluate the systems' performance using the NIST's standard metrics. We show that not only the gain from using DNNs vanishes, nor using dedicated DNNs for target conditions helps, but also the DNN-based systems tend to produce de-calibrated scores under the studied conditions. This work gives suggestions for directions of future research rather than any particular solutions to these issues. Ondrej Novotný, Pavel Matejka, Ondrej Glembek, Oldrich Plchot, Frantisek Grézl, Lukás Burget, Jan Cernocký |
SLT | 6 |
| 2015 | Robust speech recognition in unknown reverberant and noisy conditionsabstractIn this paper, we describe our work on the ASpIRE (Automatic Speech recognition In Reverberant Environments) challenge, which aims to assess the robustness of automatic speech recognition (ASR) systems. The main characteristic of the challenge is developing a high-performance system without access to matched training and development data. While the evaluation data are recorded with far-field microphones in noisy and reverberant rooms, the training data are telephone speech and close talking. Our approach to this challenge includes speech enhancement, neural network methods and acoustic model adaptation, We show that these techniques can successfully alleviate the performance degradation due to noisy audio and data mismatch. Roger Hsiao, Jeff Z. Ma, William Hartmann, Martin Karafiát, Frantisek Grézl, Lukás Burget, Igor Szöke, Jan Cernocký, Shinji Watanabe 0001, Zhuo Chen 0006, Sri Harish Reddy Mallidi, Hynek Hermansky, Stavros Tsakalidis, Richard M. Schwartz |
ASRU | 6 |
| 2015 | Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial WorkshopabstractA group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which the estimator was trained. The paper describes the problem and summarizes approaches that were taken by the group1. Hynek Hermansky, Lukás Burget, Jordan Cohen, Emmanuel Dupoux, Naomi Feldman, John Godfrey, Sanjeev Khudanpur, Matthew Maciejewski, Sri Harish Reddy Mallidi, Anjali Menon, Tetsuji Ogawa, Vijayaditya Peddinti, Richard C. Rose, Richard M. Stern, Matthew Wiesner, Karel Veselý |
ICASSP | 2 |
| 2015 | Employment of Subspace Gaussian Mixture Models in speaker recognitionabstractThis paper presents Subspace Gaussian Mixture Model (SGMM) approach employed as a probabilistic generative model to estimate speaker vector representations to be subsequently used in the speaker verification task. SGMMs have already been shown to significantly outperform traditional HMM/GMMs in Automatic Speech Recognition (ASR) applications. An extension to the basic SGMM framework allows to robustly estimate low-dimensional speaker vectors and exploit them for speaker adaptation. We propose a speaker verification framework based on low-dimensional speaker vectors estimated using SGMMs, trained in ASR manner using manual transcriptions. To test the robustness of the system, we evaluate the proposed approach with respect to the state-of-the-art i-vector extractor on the NIST SRE 2010 evaluation set and on four different length-utterance conditions: 3sec-10sec, 10 sec-30 sec, 30 sec-60 sec and full (untruncated) utterances. Experimental results reveal that while i-vector system performs better on truncated 3sec to 10sec and 10 sec to 30 sec utterances, noticeable improvements are observed with SGMMs especially on full length-utterance durations. Eventually, the proposed SGMM approach exhibits complementary properties and can thus be efficiently fused with i-vector based speaker verification system. Petr Motlícek, Subhadeep Dey, Srikanth R. Madikeri, Lukás Burget |
ICASSP | 4 |
| 2015 | Copingwith channel mismatch in Query-by-Example - But QUESST 2014abstractThe paper investigates into Query by Example (QbE) - a spoken term detection technique with queries entered by voice. It describes BUT QbE system that achieved the best accuracy in MediaEval QUESST2014 evaluations. This evaluation was challenging because of severe mismatch between queries and utterances, and introduction of new types of queries. The paper provides an analysis of DTW sub-system's in mismatched conditions (especially targeting DTW metrics) and discusses approaches investigated for QUESST2014: generation of calibration side-information by a language identification system, and handling T2 and T3 queries relaxing the constraints of an exact match. All results are provided on QUESST2014 development and evaluation data. Igor Szöke, Miroslav Skácel, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2015 | Migrating i-vectors between speaker recognition systems using regression neural networks
Ondrej Glembek, Pavel Matejka, Oldrich Plchot, Jan Pesán, Lukás Burget, Petr Schwarz |
INTERSPEECH | 5 |
| 2015 | Three ways to adapt a CTS recognizer to unseen reverberated speech in BUT system for the ASpIRE challenge
Martin Karafiát, Frantisek Grézl, Lukás Burget, Igor Szöke, Jan Cernocký |
INTERSPEECH | 3 |
| 2015 | DNN derived filters for processing of modulation spectrum of speechabstractWe propose a novel approach to design modulation frequency filters for the first stage processing of critical band spectrum of speech using deep neural network (DNN). These filters replace conventional modulation frequency filters currently used in state-of-the-art BUT speech recognition system and yield about 10% relative improvement in phoneme recognition accuracy. The resulting filters are consistent with some known temporal properties of higher levels of mammalian auditory processing and suggest more efficient scheme for pre-processing of speech for ASR. Index Terms: deep neural network, convolutive layer, modulation filters, mammalian auditory processing Jan Pesán, Lukás Burget, Hynek Hermansky, Karel Veselý |
INTERSPEECH | 2 |
| 2014 | Domain adaptation via within-class covariance correction in I-vector based speaker recognition systemsabstractIn this paper we propose a technique of Within-Class Covariance Correction (WCC) for Linear Discriminant Analysis (LDA) in Speaker Recognition to perform an unsupervised adaptation of LDA to an unseen data domain, and/or to compensate for speaker population difference among different portions of LDA training dataset. The paper follows on the study of source-normalization and inter-database variability compensation techniques which deal with multimodal distribution of i-vectors. On the DARPA RATS (Robust Automatic Transcription of Speech) task, we show that, with two hours of unsupervised data, we improve the Equal-Error Rate (EER) by 17.5%, and 36% relative on the unmatched and semi-matched conditions, respectively. On the Domain Adaptation Challenge we show up to 70% relative EER reduction and we propose a data clustering procedure to identify the directions of the domain-based variability in the adaptation data. Ondrej Glembek, Jeff Z. Ma, Pavel Matejka, Bing Zhang 0004, Oldrich Plchot, Lukás Burget, Spyridon Matsoukas |
ICASSP | 6 |
| 2014 | Unscented transform for ivector-based noisy speaker recognitionabstractRecently, a new version of the iVector modelling has been proposed for noise robust speaker recognition, where the nonlinear function that relates clean and noisy cepstral coefficients is approximated by a first order vector Taylor series (VTS). In this paper, it is proposed to substitute the first order VTS by an unscented transform, where unlike VTS, the nonlinear function is not applied over the clean model parameters directly, but over a set of sampled points. The resulting points in the transformed space are then used to calculate the model parameters. For very low signal-to-noise ratio improvements in equal error rate of about 7% for a clean backend and of 14.50% for a multistyle backend are obtained. David Martínez González, Lukás Burget, Themos Stafylakis, Yun Lei, Patrick Kenny, Eduardo Lleida |
ICASSP | 2 |
| 2014 | Calibration and fusion of query-by-example systems - But SWS 2013abstractThis paper summarizes our work for MediaEval 2013 Spoken Web Search task evaluations. The task was Query-by-Example (search of spoken queries within spoken data). We submitted a system composed of 26 subsystems, of which 13 are based on Acoustic Keyword Spotting and 13 on Dynamic Time Warping. All of them use three-state phoneme posteriors as input features. Our main contribution was m-norm normalization of particular subsystems together with the fusion based on binary logistic regression. The results, including per-language analysis, are provided on MediaEval 2013 dataset. Igor Szöke, Lukás Burget, Frantisek Grézl, Jan Cernocký, Lucas Ondel Yang |
ICASSP | 2 |
| 2014 | PLLR features in language recognition system for RATSabstractIn this paper, we study the use of features based on frame-byframe phone posteriors (PLLRs) for language recognition. The results are reported on the datasets developed for the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We show that systems based on the PLLRs outperform the standard acoustic system based on PLP2 features. By experimenting with the system combinations, we also demonstrate that the PLLR-based systems contain complementary information with respect to the PLP2 system. Finally we make a comparison between the PLLR and phonotactic systems with the outcome favorable to the PLLR. Oldrich Plchot, Mireia Díez, Mehdi Soufifar, Lukás Burget |
INTERSPEECH | 4 |
| 2014 | But ASR system for BABEL Surprise evaluation 2014abstractThe paper describes Brno University of Technology (BUT) ASR system for 2014 BABEL Surprise language evaluation (Tamil). While being largely based on our previous work, two original contributions were brought: (1) speaker-adapted bottle-neck neural network (BN) features were investigated as an input to DNN recognizer and semi-supervised training was found effective. (2) Adding of noise to training data outperformed a classical de-noising technique while dealing with noisy test data was found beneficial, and the performance of this approach was verified on a relatively clean training/test data setup from a different language. All results are reported on BABEL 2014 Tamil data. Martin Karafiát, Karel Veselý, Igor Szöke, Lukás Burget, Frantisek Grézl, Mirko Hannemann, Jan Cernocký |
SLT | 4 |
| 2014 | Non-Negative Factor Analysis of Gaussian Mixture Model Weight Adaptation for Language and Dialect RecognitionabstractRecent studies show that Gaussian mixture model (GMM) weights carry less, yet complimentary, information to GMM means for language and dialect recognition. However, state-of-the-art language recognition systems usually do not use this information. In this research, a non-negative factor analysis (NFA) approach is developed for GMM weight decomposition and adaptation. This modeling, which is conceptually simple and computationally inexpensive, suggests a new low-dimensional utterance representation method using a factor analysis similar to that of the i-vector framework. The obtained subspace vectors are then applied in conjunction with i-vectors to the language/dialect recognition problem. The suggested approach is evaluated on the NIST 2011 and RATS language recognition evaluation (LRE) corpora and on the QCRI Arabic dialect recognition evaluation (DRE) corpus. The assessment results show that the proposed adaptation method yields more accurate recognition results compared to three conventional weight adaptation approaches, namely maximum likelihood re-estimation, non-negative matrix factorization, and a subspace multinomial model. Experimental results also show that the intermediate-level fusion of i-vectors and NFA subspace vectors improves the performance of the state-of-the-art i-vector framework especially for the case of short utterances. Mohamad Hasan Bahari, Najim Dehak, Hugo Van hamme, Lukás Burget, Ahmed Ali 0002, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Semi-supervised training of Deep Neural NetworksabstractIn this paper we search for an optimal strategy for semi-supervised Deep Neural Network (DNN) training. We assume that a small part of the data is transcribed, while the majority of the data is untranscribed. We explore self-training strategies with data selection based on both the utterance-level and frame-level confidences. Further on, we study the interactions between semi-supervised frame-discriminative training and sequence-discriminative sMBR training. We found it beneficial to reduce the disproportion in amounts of transcribed and untranscribed data by including the transcribed data several times, as well as to do a frame-selection based on per-frame confidences derived from confusion in a lattice. For the experiments, we used the Limited language pack condition for the Surprise language task (Vietnamese) from the IARPA Babel program. The absolute Word Error Rate (WER) improvement for frame cross-entropy training is 2.2%, this corresponds to WER recovery of 36% when compared to the identical system, where the DNN is built on the fully transcribed data. Karel Veselý, Mirko Hannemann, Lukás Burget |
ASRU | 3 |
| 2013 | Rich system combination for keyword spotting in noisy and acoustically heterogeneous audio streamsabstractWe address the problem of retrieving spoken information from noisy and heterogeneous audio archives using system combination with a rich and diverse set of noise-robust modules. Audio search applications so far have focused on constrained domains or genres and not-so-noisy and heterogeneous acoustic or channel conditions. In this paper, our focus is to improve the accuracy of a keyword spotting system in highly degraded and diverse channel conditions by employing multiple recognition systems in parallel with different robust frontends and modeling choices, as well as different representations during audio indexing and search (words vs. subword units). After aligning keyword hits from different systems, we employ system combination at the score level using a logistic-regression-based classifier. Side information such as the output of an acoustic condition identification module is used to guide system combination system that is trained on a held-out dataset. Lattice-based indexing and search is used in all keyword spotting systems. We present improvements in probability-miss at a fixed probability-false-alarm by employing our proposed rich system combination approach on DARPA Robust Automatic Transcription of Speech (RATS) Phase-I evaluation data that contains highly degraded channel recordings (signal-to-noise ratio levels as low as 0 dB) and different channel characteristics. Murat Akbacak, Lukás Burget, Wen Wang 0001, Julien van Hout |
ICASSP | 2 |
| 2013 | A noise robust i-vector extractor using vector taylor series for speaker recognitionabstractWe propose a novel approach for noise-robust speaker recognition, where the model of distortions caused by additive and convolutive noises is integrated into the i-vector extraction framework. The model is based on a vector taylor series (VTS) approximation widely successful in noise robust speech recognition. The model allows for extracting “cleaned-up” i-vectors which can be used in a standard i-vector back end. We evaluate the proposed framework on the PRISM corpus, a NIST-SRE like corpus, where noisy conditions were created by artificially adding babble noises to clean speech segments. Results show that using VTS i-vectors present significant improvements in all noisy conditions compared to a state-of-the-art baseline speaker recognition. More importantly, the proposed framework is robust to noise, as improvements are maintained when the system is trained on clean data. Yun Lei, Lukás Burget, Nicolas Scheffer |
ICASSP | 2 |
| 2013 | A region-specific feature-space transformation for speaker adaptation and singularity analysis of jacobian matrix
Shakti P. Rath, Lukás Burget, Martin Karafiát, Ondrej Glembek, Jan Cernocký |
INTERSPEECH | 2 |
| 2013 | Regularized subspace n-gram model for phonotactic ivector extractionabstractPhonotactic language identification (LID) by means of n-gram statistics and discriminative classifiers is a popular approach for the LID problem. Low-dimensional representation of the n-gram statistics leads to the use of more diverse and efficient machine learning techniques in the LID. Recently, we proposed phototactic iVector as a low-dimensional representation of the n-gram statistics. In this work, an enhanced modeling of the n-gram probabilities along with regularized parameter estimation is proposed. The proposed model consistently improves the LID system performance over all conditions up to 15% relative to the previous state of the art system. The new model also alleviates memory requirement of the iVector extraction and helps to speed up subspace training. Results are presented in terms of Cavg over NIST LRE2009 evaluation set. Mehdi Soufifar, Lukás Burget, Oldrich Plchot, Sandro Cumani, Jan Cernocký |
INTERSPEECH | 2 |
| 2013 | Sequence-discriminative training of deep neural networksabstractSequence-discriminative training of deep neural networks (DNNs) is investigated on a 300 hour American English conversational telephone speech task. Different sequence-discriminative criteria ndash;- maximum mutual information (MMI), minimum phone error (MPE), state-level minimum Bayes risk (sMBR), and boosted MMI ndash;- are compared. Two different heuristics are investigated to improve the performance of the DNNs trained using sequence-based criteria ndash;- lattices are re-generated after the first iteration of training; and, for MMI and BMMI, the frames where the numerator and denominator hypotheses are disjoint are removed from the gradient computation. Starting from a competitive DNN baseline trained using cross-entropy, different sequence-discriminative criteria are shown to lower word error rates by 8-9% relative, on average. Little difference is noticed between the different sequence-based criteria that are investigated. The experiments are done using the open-source Kaldi toolkit, which makes it possible for the wider community to reproduce these results. Karel Veselý, Arnab Ghoshal, Lukás Burget, Daniel Povey |
INTERSPEECH | 3 |
| 2013 | Pairwise Discriminative Speaker Verification in the 𝕀-Vector SpaceabstractThis work presents a new and efficient approach to discriminative speaker verification in the${\rm i}$–vector space. We illustrate the development of a linear discriminative classifier that is trained to discriminate between the hypothesis that a pair of feature vectors in a trial belong to the same speaker or to different speakers. This approach is alternative to the usual discriminative setup that discriminates between a speaker and all the other speakers. We use a discriminative classifier based on a Support Vector Machine (SVM) that is trained to estimate the parameters of a symmetric quadratic function approximating a log–likelihood ratio score without explicit modeling of the${\rm i}$–vector distributions as in the generative Probabilistic Linear Discriminant Analysis (PLDA) models. Training these models is feasible because it is not necessary to expand the${\rm i}$–vector pairs, which would be expensive or even impossible even for medium sized training sets. The results of experiments performed on the tel-tel extended core condition of the NIST 2010 Speaker Recognition Evaluation are competitive with the ones obtained by generative models, in terms of normalized Detection Cost Function and Equal Error Rate. Moreover, we show that it is possible to train a gender–independent discriminative model that achieves state–of–the–art accuracy, comparable to the one of a gender–dependent system, saving memory and execution time both in training and in testing. Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface, Oldrich Plchot, Vasileios Vasilakakis |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | iVector-based prosodic system for language identificationabstractProsody is the part of speech where rhythm, stress, and intonation are reflected. In language identification tasks, these characteristics are assumed to be language dependent, and thus the language can be identified from them. In this paper, an automatic language recognition system that extracts prosody information from speech and makes decisions about the language with a generative classifier based on iVectors is built. The system is tested on the NIST LRE09 dataset. The results are still not comparable to state-of-the-art acoustic and phonotactic systems. However, they are promising and the fusion of the new approach with an iVector-based acoustic system is found to bring further improvements over the latter. David Martínez González, Lukás Burget, Luciana Ferrer, Nicolas Scheffer |
ICASSP | 2 |
| 2012 | Region dependent linear transforms in multilingual speech recognitionabstractIn today's speech recognition systems, linear or nonlinear transformations are usually applied to post-process speech features forming input to HMM based acoustic models. In this work, we experiment with three popular transforms: HLDA, MPE-HLDA and Region Dependent Linear Transforms (RDLT), which are trained jointly with the acoustic model to extract maximum of the discriminative information from the raw features and to represent it in a form suitable for the following GMM-HMM based acoustic model. We focus on multi-lingual environments, where limited resources are available for training recognizers of many languages. Using data from GlobalPhone database, we show that, under such restrictive conditions, the feature transformations can be advantageously shared across languages and robustly trained using data from several languages. Martin Karafiát, Milos Janda, Jan Cernocký, Lukás Burget |
ICASSP | 4 |
| 2012 | Improving language models for ASR using translated in-domain dataabstractAcquisition of in-domain training data to build speech recognition systems for under-resourced languages can be a costly, time-demanding and tedious process. In this work, we propose the use of machine translation to translate English transcripts of telephone speech into Czech language in order to improve a Czech CTS speech recognition system. The translated transcripts are used as additional language model training data in a scenario where the baseline language model is trained on off- and close-domain data only. We report perplexities, OOV and word error rates and examine different data sets and translators on their suitability for the described task. Stefan Kombrink, Tomás Mikolov, Martin Karafiát, Lukás Burget |
ICASSP | 4 |
| 2012 | Towards noise-robust speaker recognition using probabilistic linear discriminant analysisabstractThis work addresses the problem of speaker verification where additive noise is present in the enrollment and testing utterances. We show how the current state-of-the-art framework can be effectively used to mitigate this effect. We first look at the degradation a standard speaker verification system is subjected to when presented with noisy speech waveforms. We designed and generated a corpus with noisy conditions, based on the NIST SRE 2008 and 2010 data, built using open-source tools and freely available noise samples. We then show how adding noisy training data in the current i-vector-based approach followed by probabilistic linear discriminant analysis (PLDA) can bring significant gains in accuracy at various signal-to-noise ratio (SNR) levels. We demonstrate that this improvement is not feature-specific as we present positive results for three disparate sets of features: standard mel frequency cepstral coefficients, prosodic polynomial co-efficients and maximum likelihood linear regression (MLLR) transforms. Yun Lei, Lukás Burget, Luciana Ferrer, Martin Graciarena, Nicolas Scheffer |
ICASSP | 2 |
| 2012 | Generating exact lattices in the WFST frameworkabstractWe describe a lattice generation method that is exact, i.e. it satisfies all the natural properties we would want from a lattice of alternative transcriptions of an utterance. This method does not introduce substantial overhead above one-best decoding. Our method is most directly applicable when using WFST decoders where the WFST is “fully expanded”, i.e. where the arcs correspond to HMM transitions. It outputs lattices that include HMM-state-level alignments as well as word labels. The general idea is to create a state-level lattice during decoding, and to do a special form of determinization that retains only the best-scoring path for each word sequence. This special determinization algorithm is a solution to the following problem: Given a WFST A, compute a WFST B that, for each input-symbol-sequence of A, contains just the lowest-cost path through A. Daniel Povey, Mirko Hannemann, Gilles Boulianne, Lukás Burget, Arnab Ghoshal, Milos Janda, Martin Karafiát, Stefan Kombrink, Petr Motlícek, Yanmin Qian, Korbinian Riedhammer, Karel Veselý, Ngoc Thang Vu |
ICASSP | 4 |
| 2012 | Discriminative classifiers for phonotactic language recognition with iVectorsabstractPhonotactic models based on bags of n-grams representations and discriminative classifiers are a popular approach to the language recognition problem. However, the large size of n-gram count vectors brings about some difficulties in discriminative classifiers. The subspace Multinomial model was recently proposed to effectively represent information contained in the n-grams using low-dimensional iVectors. The availability of a low-dimensional feature vector allows investigating different post-processing techniques and different classifiers to improve recognition performance. In this work, we analyze a set of discriminative classifiers based on Support Vector Machines and Logistic Regression and we propose an iVector post-processing technique which allows to improve recognition performance. The proposed systems are evaluated on the NIST LRE 2009 task. Mehdi Soufifar, Sandro Cumani, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2012 | Discriminatively trained phoneme confusion model for keyword spottingabstractKeyword Spotting (KWS) aims at detecting speech segments that contain a given query within large amounts of audio data. Typically, a speech recognizer is involved in a first indexing step. One of the challenges of KWS is how to handle recognition errors and out-of-vocabulary (OOV) terms. This work proposes the use of discriminative training to construct a phoneme confusion model, which expands the phonemic index of a KWS system by adding phonemic variation to handle the abovementioned problems. The objective function that is optimized is the Figure of Merit (FOM), which is directly related to the KWS performance. The experiments conducted on English data sets show some improvement on the FOM and are promising for the use of such technique. Panagiota Karanasou, Lukás Burget, Dimitra Vergyri, Murat Akbacak, Arindam Mandal |
INTERSPEECH | 2 |
| 2012 | Bilinear Factor Analysis for iVector Based Speaker Verification
Yun Lei, Lukás Burget, Nicolas Scheffer |
INTERSPEECH | 2 |
| 2012 | Transcribing Meetings With the AMIDA SystemsabstractIn this paper, we give an overview of the AMIDA systems for transcription of conference and lecture room meetings. The systems were developed for participation in the Rich Transcription evaluations conducted by the National Institute for Standards and Technology in the years 2007 and 2009 and can process close talking and far field microphone recordings. The paper first discusses fundamental properties of meeting data with special focus on the AMI/AMIDA corpora. This is followed by a description and analysis of improved processing and modeling, with focus on techniques specifically addressing meeting transcription issues such as multi-room recordings or domain variability. In 2007 and 2009, two different strategies of systems building were followed. While in 2007 we used our traditional style system design based on cross adaptation, the 2009 systems were constructed semi-automatically, supported by improved decoders and a new method for system representation. Overall these changes gave a 6%-13% relative reduction in word error rate compared to our 2007 results while at the same time requiring less training material and reducing the real-time factor by five times. The meeting transcription systems are available at www.webasr.org. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Frantisek Grézl, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | iVector-based discriminative adaptation for automatic speech recognitionabstractWe presented a novel technique for discriminative feature-level adaptation of automatic speech recognition system. The concept of iVectors popular in Speaker Recognition is used to extract information about speaker or acoustic environment from speech segment. iVector is a low-dimensional fixed-length representing such information. To utilized iVectors for adaptation, Region Dependent Linear Transforms (RDLT) are discriminatively trained using MPE criterion on large amount of annotated data to extract the relevant information from iVectors and to compensate speech feature. The approach was tested on standard CTS data. We found it to be complementary to common adaptation techniques. On a well tuned RDLT system with standard CMLLR adaptation we reached 0.8% additive absolute WER improvement. Martin Karafiát, Lukás Burget, Pavel Matejka, Ondrej Glembek, Jan Cernocký |
ASRU | 2 |
| 2011 | Strategies for training large scale neural network language modelsabstractWe describe how to effectively train neural network based language models on large data sets. Fast convergence during training and better overall performance is observed when the training data are sorted by their relevance. We introduce hash-based implementation of a maximum entropy model, that can be trained as a part of the neural network model. This leads to significant reduction of computational complexity. We achieved around 10% relative reduction of word error rate on English Broadcast News speech recognition task, against large 4-gram model trained on 400M tokens. Tomás Mikolov, Anoop Deoras, Daniel Povey, Lukás Burget, Jan Cernocký |
ASRU | 4 |
| 2011 | Discriminatively trained Probabilistic Linear Discriminant Analysis for speaker verificationabstractRecently, i-vector extraction and Probabilistic Linear Discriminant Analysis (PLDA) have proven to provide state-of-the-art speaker verification performance. In this paper, the speaker verification score for a pair of i-vectors representing a trial is computed with a functional form derived from the successful PLDA generative model. In our case, however, parameters of this function are estimated based on a discriminative training criterion. We propose to use the objective function to directly address the task in speaker verification: discrimination between same-speaker and different-speaker trials. Compared with a baseline which uses a generatively trained PLDA model, discriminative training provides up to 40% relative improvement on the NIST S RE 2010 evaluation task. Lukás Burget, Oldrich Plchot, Sandro Cumani, Ondrej Glembek, Pavel Matejka, Niko Brümmer |
ICASSP | 1 |
| 2011 | Fast discriminative speaker verification in the i-vector spaceabstractThis work presents a new approach to discriminative speaker verification. Rather than estimating speaker models, or a model that discriminates between a speaker class and the class of all the other speakers, we directly solve the problem of classifying pairs of utterances as belonging to the same speaker or not. Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface |
ICASSP | 3 |
| 2011 | Simplification and optimization of i-vector extractionabstractThis paper introduces some simplifications to the i-vector speaker recognition systems. I-vector extraction as well as training of the i-vector extractor can be an expensive task both in terms of memory and speed. Under certain assumptions, the formulas for i-vector extraction—also used in i-vector extractor training—can be simplified and lead to a faster and memory more efficient code. The first assumption is that the GMM component alignment is constant across utterances and is given by the UBM GMM weights. The second assumption is that the i-vector extractor matrix can be linearly transformed so that its per-Gaussian components are orthogonal. We use PCA and HLDA to estimate this transform. Ondrej Glembek, Lukás Burget, Pavel Matejka, Martin Karafiát, Patrick Kenny |
ICASSP | 2 |
| 2011 | Recent progress in prosodic speaker verificationabstractWe describe recent progress in the field of prosodic modeling for speaker verification. In a previous paper, we proposed a technique for modeling syllable-based prosodic features that uses a multinomial subspace model for feature extraction and within-class covariance normalization or linear discriminant analysis for session variability compensation. In this paper, we show that performance can be significantly improved with the use of probabilistic linear discriminant analysis (PLDA) for session variability compensation. This system does not require score normalization. We report an equal error rate below 7% on a NIST 2008 task. To our knowledge, this is the best reported result to date for a prosodic system for speaker recognition. Fusion of this system with a state-of-the-art acoustic baseline system yields 10% relative improvement in the new detection cost function (DCF) as defined by NIST. Marcel Kockmann, Luciana Ferrer, Lukás Burget, Elizabeth Shriberg, Jan Cernocký |
ICASSP | 3 |
| 2011 | Full-covariance UBM and heavy-tailed PLDA in i-vector speaker verificationabstractIn this paper, we describe recent progress in i-vector based speaker verification. The use of universal background models (UBM) with full-covariance matrices is suggested and thoroughly experimentally tested. The i-vectors are scored using a simple cosine distance and advanced techniques such as Probabilistic Linear Discriminant Analysis (PLDA) and heavy-tailed variant of PLDA (PLDA-HT). Finally, we investigate into dimensionality reduction of i-vectors before entering the PLDA-HT modeling. The results are very competitive: on NIST 2010 SRE task, the results of a single full-covariance LDA-PLDA-HT system approach those of complex fused system. Pavel Matejka, Ondrej Glembek, Fabio Castaldo, Jahangir Alam 0001, Oldrich Plchot, Patrick Kenny, Lukás Burget, Jan Cernocký |
ICASSP | 7 |
| 2011 | Extensions of recurrent neural network language modelabstractWe present several modifications of the original recurrent neural network language model (RNN LM).While this model has been shown to significantly outperform many competitive language modeling techniques in terms of accuracy, the remaining problem is the computational complexity. In this work, we show approaches that lead to more than 15 times speedup for both training and testing phases. Next, we show importance of using a backpropagation through time algorithm. An empirical comparison with feedforward networks is also provided. In the end, we discuss possibilities how to reduce the amount of parameters in the model. The resulting RNN model can thus be smaller, faster both during training and testing, and more accurate than the basic one. Tomás Mikolov, Stefan Kombrink, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur |
ICASSP | 3 |
| 2011 | Discriminatively Trained i-vector Extractor for Speaker VerificationabstractWe propose a strategy for discriminative training of the ivector extractor in speaker recognition. The original i-vector extractor training was based on the maximum-likelihood generative modeling, where the EM algorithm was used. In our approach, the i-vector extractor parameters are numerically optimized to minimize the discriminative cross-entropy error function. Two versions of the i-vector extraction are studied—the original approach as defined for Joint Factor Analysis, and the simplified version, where orthogonalization of the i-vector extractor matrix is performed. Index Terms: speaker verification, i-vectors, PLDA, discriminative training Ondrej Glembek, Lukás Burget, Niko Brümmer, Oldrich Plchot, Pavel Matejka |
INTERSPEECH | 2 |
| 2011 | iVector Fusion of Prosodic and Cepstral Features for Speaker VerificationabstractIn this paper we apply the promising iVector extraction technique followed by PLDA modeling to simple prosodic contour features. With this procedure we achieve results comparable to a system that models much more complex prosodic features using our recently proposed SMM-based iVector modeling technique. We then propose a combination of both prosodic iVectors by joint PLDA modeling that leads to significant improvements over individual systems with an EER of 5.4% on NIST SRE 2008 telephone data. Finally, we can combine these two prosodic iVector front ends with a baseline cepstral iVector system to achieve up to 21% relative reduction in new DCF. Index Terms: speaker verification, prosody, JFA, iVector, SMM, fusion Marcel Kockmann, Luciana Ferrer, Lukás Burget, Jan Cernocký |
INTERSPEECH | 3 |
| 2011 | Recurrent Neural Network Based Language Modeling in Meeting RecognitionabstractWe use recurrent neural network (RNN) based language models to improve the BUT English meeting recognizer. On the baseline setup using the original language models we decrease word error rate (WER) more than 1% absolute by n-best list rescoring and language model adaptation. When n-gram language models are trained on the same moderately sized data set as the RNN models, improvements are higher yielding a system which performs comparable to the baseline. A noticeable improvement was observed with unsupervised adaptation of RNN models. Furthermore, we examine the influence of word history on WER and show how to speed-up rescoring by caching common prefix strings. Index Terms: automatic speech recognition, language modeling, recurrent neural networks, rescoring, adaptation Stefan Kombrink, Tomás Mikolov, Martin Karafiát, Lukás Burget |
INTERSPEECH | 4 |
| 2011 | Language Recognition in iVectors SpaceabstractThe concept of so called iVectors, where each utterance is represented by fixed-length low-dimensional feature vector, has recently become very successfully in speaker verification. In this work, we apply the same idea in the context of Language Recognition (LR). To recognize language in the iVector space, we experiment with three different linear classifiers: one based on a generative model, where classes are modeled by Gaussian distributions with shared covariance matrix, and two discriminative classifiers, namely linear Support Vector Machine and Logistic Regression. The tests were performed on the NIST LRE 2009 dataset and the results were compared with stateof-the-art LR based on Joint Factor Analysis (JFA). While the iVector system offers better performance, it also seems to be complementary to JFA, as their fusion shows another improvement. David Martínez González, Oldrich Plchot, Lukás Burget, Ondrej Glembek, Pavel Matejka |
INTERSPEECH | 3 |
| 2011 | Empirical Evaluation and Combination of Advanced Language Modeling TechniquesabstractWe present results obtained with several advanced language modeling techniques, including class based model, cache model, maximum entropy model, structured language model, random forest language model and several types of neural network based language models. We show results obtained after combining all these models by using linear interpolation. We conclude that for both small and moderately sized tasks, we obtain new state of the art results with combination of models, that is significantly better than performance of any individual model. Obtained perplexity reductions against Good-Turing trigram baseline are over 50% and against modified Kneser-Ney smoothed 5-gram over 40%. Index Terms: language modeling, neural networks, model combination, speech recognition Tomás Mikolov, Anoop Deoras, Stefan Kombrink, Lukás Burget, Jan Cernocký |
INTERSPEECH | 4 |
| 2011 | iVector Approach to Phonotactic Language RecognitionabstractThis paper addresses a novel technique for representation and processing of n-gram counts in phonotactic language recognition (LRE): subspace multinomial modelling represents the vectors of n-gram counts by low dimensional vectors of coordinates in total variability subspace, called iVector. Two techniques for iVector scoring are tested: support vector machines (SVM), and logistic regression (LR). Using standard NIST LRE 2009 task as our evaluation set, the latter scoring approach was shown to outperform phonotactic LRE system based on direct SVM classification of n-gram count vectors. The proposed iVector paradigm also shows comparable results to previously proposed PCA-based phonotactic feature extraction. Index Terms: language recognition, subspace modeling, multinomial distribution. Mehdi Soufifar, Marcel Kockmann, Lukás Burget, Oldrich Plchot, Ondrej Glembek, Torbjørn Svendsen |
INTERSPEECH | 3 |
| 2011 | The subspace Gaussian mixture model - A structured model for speech recognition
Daniel Povey, Lukás Burget, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
Comput. Speech Lang. | 2 |
| 2011 | Application of speaker- and language identification state-of-the-art techniques for emotion recognition
Marcel Kockmann, Lukás Burget, Jan Cernocký |
Speech Commun. | 2 |
| 2010 | Multilingual acoustic modeling for speech recognition based on subspace Gaussian Mixture ModelsabstractAlthough research has previously been done on multilingual speech recognition, it has been found to be very difficult to improve over separately trained systems. The usual approach has been to use some kind of “universal phone set” that covers multiple languages. We report experiments on a different approach to multilingual speech recognition, in which the phone sets are entirely distinct but the model has parameters not tied to specific states that are shared across languages. We use a model called a “Subspace Gaussian Mixture Model” where states' distributions are Gaussian Mixture Models with a common structure, constrained to lie in a subspace of the total parameter space. The parameters that define this subspace can be shared across languages. We obtain substantial WER improvements with this approach, especially with very small amounts of in-language training data. Lukás Burget, Petr Schwarz, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Daniel Povey, Ariya Rastrow, Richard C. Rose, Samuel Thomas 0001 |
ICASSP | 1 |
| 2010 | A novel estimation of feature-space MLLR for full-covariance modelsabstractIn this paper we present a novel approach for estimating feature-space maximum likelihood linear regression (fMLLR) transforms for full-covariance Gaussian models by directly maximizing the likelihood function by repeated line search in the direction of the gradient. We do this in a pre-transformed parameter space such that an approximation to the expected Hessian is proportional to the unit matrix. The proposed algorithm is as efficient or more efficient than standard approaches, and is more flexible because it can naturally be combined with sets of basis transforms and with full covariance and subspace precision and mean (SPAM) models. Arnab Ghoshal, Daniel Povey, Mohit Agarwal 0005, Pinar Akyazi, Lukás Burget, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
ICASSP | 5 |
| 2010 | Approaches to automatic lexicon learning with limited training examplesabstractPreparation of a lexicon for speech recognition systems can be a significant effort in languages where the written form is not exactly phonetic. On the other hand, in languages where the written form is quite phonetic, some common words are often mispronounced. In this paper, we use a combination of lexicon learning techniques to explore whether a lexicon can be learned when only a small lexicon is available for boot-strapping. We discover that for a phonetic language such as Spanish, it is possible to do that better than what is possible from generic rules or hand-crafted pronunciations. For a more complex language such as English, we find that it is still possible but with some loss of accuracy. Nagendra K. Goel, Samuel Thomas 0001, Mohit Agarwal 0005, Pinar Akyazi, Lukás Burget, Arnab Ghoshal, Ondrej Glembek, Martin Karafiát, Daniel Povey, Ariya Rastrow, Richard C. Rose, Petr Schwarz |
ICASSP | 5 |
| 2010 | Investigations into prosodic syllable contour features for speaker recognitionabstractWe investigate various ways of generating prosodic syllable contour features that have recently been applied to enhance systems for speaker recognition. We compare different approaches for segmentation of speech into syllable-like units, techniques for contour modeling and the extraction of pitch and energy, taking into account the computational complexity and gender dependence. We show that the performance is especially affected by the segmentation and the quality of the pitch tracking algorithm and that the features are highly gender dependent. Still, computationally simple ways of segmentation of speech can be used to achieve good results, as experiments on 2006 NIST speaker recognition evaluation task indicate. Marcel Kockmann, Lukás Burget, Jan Cernocký |
ICASSP | 2 |
| 2010 | Tuning phone decoders for language identificationabstractPhonotactic approach, phone recognition to be followed by language modeling, is one of the most popular approaches to language identification (LID). In this work, we explore how language identification accuracy of a phone decoder can be enhanced by varying acoustic resolution of the phone decoder, and subsequently how multiresolution versions of the same decoder can be integrated to improve the LID accuracy. We use mutual information to select the optimum set of phones for a specific acoustic resolution. Further, we propose strategies for building multilingual systems suitable for LID applications, and subsequently fine tune these systems to enhance the overall accuracy. C. Santhosh Kumar, Haizhou Li 0001, Rong Tong, Pavel Matejka, Lukás Burget, Jan Cernocký |
ICASSP | 5 |
| 2010 | Subspace Gaussian Mixture Models for speech recognitionabstractWe describe an acoustic modeling approach in which all phonetic states share a common Gaussian Mixture Model structure, and the means and mixture weights vary in a subspace of the total parameter space. We call this a Subspace Gaussian Mixture Model (SGMM). Globally shared parameters define the subspace. This style of acoustic model allows for a much more compact representation and gives better results than a conventional modeling approach, particularly with smaller amounts of training data. Daniel Povey, Lukás Burget, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
ICASSP | 2 |
| 2010 | The AMIDA 2009 meeting transcription systemabstractWe present the AMIDA 2009 system for participation in the NIST RT’2009 STT evaluations. Systems for close-talking, far field and speaker attributed STT conditions are described. Im- provements to our previous systems are: segmentation and diar- isation; stacked bottle-neck posterior feature extraction; fMPE training of acoustic models; adaptation on complete meetings; improvements to WFST decoding; automatic optimisation of decoders and system graphs. Overall these changes gave a 6- 13% relative reduction in word error rate while at the same time reducing the real-time factor by a factor of five and using con- siderably less data for acoustic model training. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
INTERSPEECH | 2 |
| 2010 | Similarity scoring for recognizing repeated out-of-vocabulary wordsabstractWe develop a similarity measure to detect repeatedly occurring Out-of-Vocabulary words (OOV), since these carry important information. Sub-word sequences in the recognition output from a hybrid word/sub-word recognizer are taken as detected OOVs and are aligned to each other with the help of an alignment error model. This model is able to deal with partial OOV detections and tries to reveal more complex word relations such as compound words. We apply the model to a selection of conversational phone calls to retrieve other examples of the same OOV, and to obtain a higher-level description of it such as being a derivation of a known word. Mirko Hannemann, Stefan Kombrink, Martin Karafiát, Lukás Burget |
INTERSPEECH | 4 |
| 2010 | Brno university of technology system for interspeech 2010 paralinguistic challengeabstractThis paper describes Brno University of Technology (BUT) system for the Interspeech 2010 Paralinguistic Challenge. Our submitted systems for the Ageand Gender-Sub-Challenges employ fusions of several sub-systems. We make use of our own acoustic frame-based feature sets, as well as the provided utterance-based acoustic, prosodic and voice quality features. Modeling is based on Gaussian Mixture Models (GMM) and Support Vector Machines (SVM), followed by linear Gaussian backends and logistic regression-based fusion. For a single subsystem, we obtain improvement of about 2% absolute, for both tasks, on the development-set. Our final fusion results in nearly 9% absolute improvement for the Age task and about 4.5% for the Gender task on the development set. On the final test set we obtain 3.5% and 2% absolute improvement, respectively. Marcel Kockmann, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2010 | Prosodic speaker verification using subspace multinomial models with intersession compensationabstractWe propose a novel approach to modeling prosodic features. Inspired by Joint Factor Analysis model (JFA), our model is based on the same idea of introducing subspace of model parameters. However, the underlying Gaussian Mixture distribution of JFA is replaced by multinomial distribution to model sequences of discrete units rather than continuous features. In this work, we use the subspace model as a feature extractor for support vector machines (SVMs), similar to the recently proposed JFA in total variability space. We can show the capability to reduce high-dimensional count vectors to low dimension while keeping system performance stable. With additional intersession compensation, we can improve 30 % relative to the baseline system and reach an equal error rate of 8.8 % on the NIST 2006 SRE dataset. Index Terms: speaker verification, prosody, JFA, multinomial model Marcel Kockmann, Lukás Burget, Ondrej Glembek, Luciana Ferrer, Jan Cernocký |
INTERSPEECH | 2 |
| 2010 | Recurrent neural network based language modelabstractA new recurrent neural network based language model (RNN LM) with applications to speech recognition is presented. Results indicate that it is possible to obtain around 50% reduction of perplexity by using mixture of several RNN LMs, compared to a state of the art backoff language model. Speech recognition experiments show around 18% reduction of word error rate on the Wall Street Journal task when comparing models trained on the same amount of data, and around 5% on the much harder NIST RT05 task, even when the backoff model is trained on much more data than the RNN LM. We provide ample empirical evidence to suggest that connectionist language models are superior to standard n-gram techniques, except their high computational (training) complexity. Index Terms: language modeling, recurrent neural networks, speech recognition Tomás Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2010 | Parallel training of neural networks for speech recognitionabstractThe feed-forward multi-layer neural networks have significant importance in speech recognition. A new parallel-training tool TNet was designed and optimized for multiprocessor computers. The training acceleration rates are reported on a phoneme-state classification task. Karel Veselý, Lukás Burget, Frantisek Grézl |
INTERSPEECH | 2 |
| 2009 | Support vector machines and Joint Factor Analysis for speaker verificationabstractThis article presents several techniques to combine between support vector machines (SVM) and joint factor analysis (JFA) model for speaker verification. In this combination, the SVMs are applied to different sources of information produced by the JFA. These informations are the Gaussian mixture model supervectors and speakers and common factors. We found that using SVM in JFA factors gave the best results especially when within class covariance normalization method is applied in order to compensate for the channel effect. The new combination results are comparable to other classical JFA scoring techniques. Najim Dehak, Patrick Kenny, Réda Dehak, Ondrej Glembek, Pierre Dumouchel, Lukás Burget, Valiantsina Hubeika, Fabio Castaldo |
ICASSP | 6 |
| 2009 | Comparison of scoring methods used in speaker recognition with Joint Factor AnalysisabstractThe aim of this paper is to compare different log-likelihood scoring methods, that different sites used in the latest state-of-the-art Joint Factor Analysis (JFA) Speaker Recognition systems. The algorithms use various assumptions and have been derived from various approximations of the objective functions of JFA. We compare the techniques in terms of speed and performance. We show, that approximations of the true log-likelihood ratio (LLR) may lead to significant speedup without any loss in performance. Ondrej Glembek, Lukás Burget, Najim Dehak, Niko Brümmer, Patrick Kenny |
ICASSP | 2 |
| 2009 | Neural network based language models for highly inflective languagesabstractSpeech recognition of inflectional and morphologically rich languages like Czech is currently quite a challenging task, because simple n-gram techniques are unable to capture important regularities in the data. Several possible solutions were proposed, namely class based models, factored models, decision trees and neural networks. This paper describes improvements obtained in recognition of spoken Czech lectures using language models based on neural networks. Relative reductions in word error rate are more than 15% over baseline obtained with adapted 4-gram backoff language model using modified Kneser-Ney smoothing. Tomás Mikolov, Jirí Kopecký, Lukás Burget, Ondrej Glembek, Jan Cernocký |
ICASSP | 3 |
| 2009 | Discriminative acoustic language recognition via channel-compensated GMM statisticsabstractWe propose a novel design for acoustic feature-based automatic spoken language recognizers. Our design is inspired by recent advances in text-independent speaker recognition, where intraclass variability is modeled by factor analysis in Gaussian mixture model (GMM) space. We use approximations to GMMlikelihoods which allow variable-length data sequences to be represented as statistics of fixed size. Our experiments on NIST LRE’07 show that variability-compensation of these statistics can reduce error-rates by a factor of three. Finally, we show that further improvements are possible with discriminative logistic regression training. Index Terms: acoustic language recognition, intersession variability compensation, discriminative training Niko Brümmer, Albert Strasheim, Valiantsina Hubeika, Pavel Matejka, Lukás Burget, Ondrej Glembek |
INTERSPEECH | 5 |
| 2009 | BUT system for NIST 2008 speaker recognition evaluationabstractThis paper presents BUT system submitted to NIST 2008 SRE. It includes two subsystems based on Joint Factor Analysis (JFA) GMM/UBM and one based on SVM-GMM. The systems were developed on NIST SRE2006 data, and the results arepresented on NIST SRE 2008 evaluation data. We concentrate on the influence of side information in the calibration. Index Terms: speaker recognition, joint factor analysis, NIST SRE 2008. Lukás Burget, Michal Fapso, Valiantsina Hubeika, Ondrej Glembek, Martin Karafiát, Marcel Kockmann, Pavel Matejka, Petr Schwarz, Jan Cernocký |
INTERSPEECH | 1 |
| 2009 | Investigation into variants of joint factor analysis for speaker recognitionabstractIn this paper, we have investigated into JFA used for speaker recognition. First, we performed systematic comparison of full JFA with its simplified variants and confirmed superior performance of the full JFA with both eigenchannels and eigenvoices. We investigated into sensitivity of JFA on the number of eigenvoices both for the full one and simplified variants. We studied the importance of normalization and found that genderdependent zt-norm was crucial. The results are reported on NIST 2006 and 2008 SRE evaluation data. Index Terms: speaker recognition, joint factor analysis. Lukás Burget, Pavel Matejka, Valiantsina Hubeika, Jan Cernocký |
INTERSPEECH | 1 |
| 2009 | Investigation into bottle-neck features for meeting speech recognitionabstractThis work investigates into recently proposed Bottle-Neck features for ASR. The bottle-neck ANN structure is imported into Split Context architecture gaining significant WER reduction. Further, Universal Context architecture was developed which simplifies the system by using only one universal ANN for all temporal splits. Significant WER reduction can be obtained by applying fMPE on top of our BN features as a technique for discriminative feature extraction and further gain is also obtained by retraining model parameters using MPE criterion. The results are reported on meeting data from RT07 evaluation. Frantisek Grézl, Martin Karafiát, Lukás Burget |
INTERSPEECH | 3 |
| 2009 | Brno University of Technology system for Interspeech 2009 emotion challengeabstractThis paper describes Brno University of Technology (BUT) system for the Interspeech 2009 Emotion Challenge. Our submitted system for the Open Performance Sub-Challenge uses acoustic frame based features as a front-end and Gaussian Mixture Models as a back-end. Different feature types and modeling approaches successfully applied in speakerand language recognition are investigated and we can achieve an 16% and 9% relative improvement over the best dynamic and static baseline system on the 5-class task, respectively. Marcel Kockmann, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2009 | Posterior-based out of vocabulary word detection in telephone speechabstractIn this paper we present an out-of-vocabulary word detector suitable for English conversational and read speech. We use an approach based on phone posteriors created by a Large Vocab-ulary Continuous Speech Recognition system and an additional phone recognizer, that allows detection of OOV and misrecog-nized words. In addition, the recognized word output can be transcribed more detailed using several classes. Reported re-sults are on CallHome English and Wall Street Journal data. Index Terms: confidence measures, out-of-vocabulary word detection, phone posteriors, neural net, OOV Stefan Kombrink, Lukás Burget, Pavel Matejka, Martin Karafiát, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Combination of strongly and weakly constrained recognizers for reliable detection of OOVSabstractThis paper addresses the detection of OOV segments in the output of a large vocabulary continuous speech recognition (LVCSR) system. First, standard confidence measures from frame-based wordand phone- posteriors are investigated. Substantial improvement is obtained when posteriors from two systems — strongly constrained (LVCSR) and weakly constrained (phone posterior estimator) are combined. We show that this approach is also suitable for detection of general recognition errors. All results are presented on WSJ task with reduced recognition vocabulary. Lukás Burget, Petr Schwarz, Pavel Matejka, Mirko Hannemann, Ariya Rastrow, Christopher M. White, Sanjeev Khudanpur, Hynek Hermansky, Jan Cernocký |
ICASSP | 1 |
| 2008 | Confidence estimation, OOV detection and language ID using phone-to-word transduction and phone-level alignmentsabstractAutomatic speech recognition (ASR) systems continue to make errors during search when handling various phenomena including noise, pronunciation variation, and out of vocabulary (OOV) words. Predicting the probability that a word is incorrect can prevent the error from propagating and perhaps allow the system to recover. This paper addresses the problem of detecting errors and OOVs for read Wall Street Journal speech when the word error rate (WER) is very low. It augments a traditional confidence estimate by introducing two novel methods: phone-level comparison using multi-string alignment (MSA) and word-level comparison using phone-to-word transduction. We show that features from phone and word string comparisons can be added to a standard maximum entropy framework thereby substantially improving performance in detecting both errors and OOVs. Additionally we show an extension to detecting English and accented English for the language identification (LID) task. Christopher M. White, Geoffrey Zweig, Lukás Burget, Petr Schwarz, Hynek Hermansky |
ICASSP | 3 |
| 2008 | Advances in phonotactic language recognitionabstractThis paper summarizes recent advances in PRLM language recognition within the context of the NIST 2007 LR evaluations (LRE). We present a comparison of binary decision tree (BT) vs. N -gram models when adaptation from a universal (background) model (UBM) is used, we introduce multi-models— anchor-model-like approach to scoring, and we adopt the framework of intersession variation using factor analysis. Ondrej Glembek, Pavel Matejka, Lukás Burget, Tomás Mikolov |
INTERSPEECH | 3 |
| 2008 | Discriminative training and channel compensation for acoustic language recognitionabstractThis paper describes the acoustic language recognition subsystems of Brno University of Technology (BUT) which contributed to the BUT main submission to the NIST LRE 2007. Two main techniques are employed in the subsystems discriminative training in terms of Maximum Mutual Information, and channel compensation in terms of eigenchannel adaptation in both, model and feature domain. The complementarity of the approaches is analyzed. Valiantsina Hubeika, Lukás Burget, Pavel Matejka, Petr Schwarz |
INTERSPEECH | 2 |
| 2008 | Discrimininative training of narrow band - wide band adapted systems for meeting recognitionabstractThe amount of training data has a crucial effect on the accuracy of HMM based meeting recognition systems. One of the largest collections of speech data is conversational telephone speech which was found to match speech in meetings well. However it is naturally recorded with limited bandwidth. In previous work we presented a scheme that allows to transform wide-band meeting data into the same space for improved model training. In this paper we focused on integration of discriminative adaptation into this scheme. This integration is not straightforward and we present the complexity of this process. The models are tested on the NIST RT’05 meeting evaluation where a relative reduction in word error rate of 5.6% against non-adapted meeting system was achieved. Martin Karafiát, Lukás Burget, Thomas Hain, Jan Cernocký |
INTERSPEECH | 2 |
| 2008 | BUT language recognition system for NIST 2007 evaluationsabstractThis paper describes Brno University of Technology (BUT) system for 2007 NIST Language recognition (LRE) evaluation. The system is a fusion of 4 acoustic and 9 phonotactic subsystems. We have investigated several new topics such as discriminatively trained language models in phonotactic systems, and eigen-channel adaptation in model and feature domain in acoustic systems. We also point out the importance of calibration and fusion. All results are presented on NIST 2007 LRE data. Pavel Matejka, Lukás Burget, Ondrej Glembek, Petr Schwarz, Valiantsina Hubeika, Michal Fapso, Tomás Mikolov, Oldrich Plchot, Jan Cernocký |
INTERSPEECH | 2 |
| 2008 | Contour modeling of prosodic and acoustic features for speaker recognitionabstractIn this paper we use acoustic and prosodic features jointly in a long-temporal lexical context for automatic speaker recognition from speech. The contours of pitch, energy and cepstral coefficients are continuously modeled over the time span of a syllable to capture the speaking style on phonetic level. As these features are affected by session variability, established channel compensation techniques are examined. Results for the combination of different features on a syllable-level as well as for channel compensation are presented for the NIST SRE 2006 speaker identification task. To show the complementary character of the features, the proposed system is fused with an acoustic short-time system, leading to a relative improvement of 10.4%. Marcel Kockmann, Lukás Burget |
SLT | 2 |
| 2008 | Morphological random forests for language modeling of inflectional languagesabstractIn this paper, we are concerned with using decision trees (DT) and random forests (RF) in language modeling for Czech LVCSR. We show that the RF approach can be successfully implemented for language modeling of an inflectional language. Performance of word-based and morphological DTs and RFs was evaluated on lecture recognition task. We show that while DTs perform worse than conventional trigram language models (LM), RFs of both kind outperform the latter. WER (up to 3.4% relative) and perplexity (10%) reduction over the trigram model can be gained with morphological RFs. Further improvement is obtained after interpolation of DT and RF LMs with the trigram one (up to 15.6% perplexity and 4.8% WER relative reduction). In this paper we also investigate distribution of morphological feature types chosen for splitting data at different levels of DTs. Ilya Oparin, Ondrej Glembek, Lukás Burget, Jan Cernocký |
SLT | 3 |
| 2008 | Sub-word modeling of out of vocabulary words in spoken term detectionabstractThis paper deals with comparison of sub-word based methods for spoken term detection (STD) task and phone recognition. The sub-word units are needed for search for out-of-vocabulary words. We compared words, phones and multigrams. The maximal length and pruning of multigrams were investigated first. Then two constrained methods of multigram training were proposed. We evaluated on the NIST STD06 dev-set CTS data. The conclusion is that the proposed method improves the phone accuracy more than 9% relative and STD accuracy more than 7% relative. Igor Szöke, Lukás Burget, Jan Cernocký, Michal Fapso |
SLT | 2 |
| 2007 | The AMI System for the Transcription of Speech in MeetingsabstractThis paper describes the AMI transcription system for speech in meetings developed in collaboration by five research groups. The system includes generic techniques such as discriminative and speaker adaptive training, vocal tract length normalisation, heteroscedastic linear discriminant analysis, maximum likelihood linear regression, and phone posterior based features, as well as techniques specifically designed for meeting data. These include segmentation and cross-talk suppression, beam-forming, domain adaptation, Web-data collection, and channel adaptive training. The system was improved by more than 20% relative in word error rate compared to our previous system and was used in the NIST RT106 evaluations where it was found to yield competitive performance. Thomas Hain, Vincent Wan, Lukás Burget, Martin Karafiát, John Dines, Jithendra Vepa, Giulia Garau, Mike Lincoln |
ICASSP (4) | 3 |
| 2007 | STBU System for the NIST 2006 Speaker Recognition EvaluationabstractThis paper describes STBU 2006 speaker recognition system, which performed well in the NIST 2006 speaker recognition evaluation. STBU is consortium of 4 partners: Spescom DataVoice (South Africa), TNO (Netherlands), BUT (Czech Republic) and University of Stellenbosch (South Africa). The primary system is a combination of three main kinds of systems: (1) GMM, with short-time MFCC or PLP features, (2) GMM-SVM, using GMM mean supervectors as input and (3) MLLR-SVM, using MLLR speaker adaptation coefficients derived from English LVCSR system. In this paper, we describe these sub-systems and present results for each system alone and in combination on the NIST Speaker Recognition Evaluation (SRE) 2006 development and evaluation data sets. Pavel Matejka, Lukás Burget, Petr Schwarz, Ondrej Glembek, Martin Karafiát, Frantisek Grézl, Jan Cernocký, David A. van Leeuwen, Niko Brümmer, Albert Strasheim |
ICASSP (4) | 2 |
| 2007 | Application of CMLLR in narrow band wide band adapted systems
Martin Karafiát, Lukás Burget, Jan Cernocký, Thomas Hain |
INTERSPEECH | 2 |
| 2007 | Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006abstractThis paper describes and discusses the "STBU" speaker recognition system, which performed well in the NIST Speaker Recognition Evaluation 2006 (SRE). STBU is a consortium of four partners: Spescom DataVoice (Stellenbosch, South Africa), TNO (Soesterberg, The Netherlands), BUT (Brno, Czech Republic), and the University of Stellenbosch (Stellenbosch, South Africa). The STBU system was a combination of three main kinds of subsystems: 1) GMM, with short-time Mel frequency cepstral coefficient (MFCC) or perceptual linear prediction (PLP) features, 2) Gaussian mixture model-support vector machine (GMM-SVM), using GMM mean supervectors as input to an SVM, and 3) maximum-likelihood linear regression-support vector machine (MLLR-SVM), using MLLR speaker adaptation coefficients derived from an English large vocabulary continuous speech recognition (LVCSR) system. All subsystems made use of supervector subspace channel compensation methods-either eigenchannel adaptation or nuisance attribute projection. We document the design and performance of all subsystems, as well as their fusion and calibration via logistic regression. Finally, we also present a cross-site fusion that was done with several additional systems from other NIST SRE-2006 participants. Niko Brümmer, Lukás Burget, Jan Cernocký, Ondrej Glembek, Frantisek Grézl, Martin Karafiát, David A. van Leeuwen, Pavel Matejka, Petr Schwarz, Albert Strasheim |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Analysis of Feature Extraction and Channel Compensation in a GMM Speaker Recognition SystemabstractIn this paper, several feature extraction and channel compensation techniques found in state-of-the-art speaker verification systems are analyzed and discussed. For the NIST SRE 2006 submission, cepstral mean subtraction, feature warping, RelAtive SpecTrAl (RASTA) filtering, heteroscedastic linear discriminant analysis (HLDA), feature mapping, and eigenchannel adaptation were incrementally added to minimize the system's error rate. This paper deals with eigenchannel adaptation in more detail and includes its theoretical background and implementation issues. The key part of the paper is, however, the post-evaluation analysis, undermining a common myth that “the more boxes in the scheme, the better the system.” All results are presented on NIST Speaker Recognition Evaluation (SRE) 2005 and 2006 data. Lukás Burget, Pavel Matejka, Petr Schwarz, Ondrej Glembek, Jan Cernocký |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Information Retrieval from Spoken Documents
Michal Fapso, Pavel Smrz, Petr Schwarz, Igor Szöke, Milan Schwarz, Jan Cernocký, Martin Karafiát, Lukás Burget |
CICLing | 8 |
| 2006 | Discriminative Training Techniques for Acoustic Language IdentificationabstractThis paper presents comparison of Maximum Likelihood (ML) and discriminative Maximum Mutual Information (MMI) training for acoustic modeling in language identification (LID). Both approaches are compared on state-of- the-art shifted delta-cepstra features, the results are reported on data from NIST 2003 evaluations. Clear advantage of MMI over ML training is shown. Further improvements of acoustic LID are discussed: Heteroscedastic Linear Discriminant Analysis (HLDA) for feature de-correlation and dimensionality reduction and Ergodic Hidden Markov models (EHMM) for better modeling of dynamics in the acoustic space. The final error rate compares favorably to other results published on NIST 2003 data. Lukás Burget, Pavel Matejka, Jan Cernocký |
ICASSP (1) | 1 |
| 2006 | Use of Anti-Models to Further Improve State-of-the-Art PRLM Language Recognition SystemabstractThis paper concentrates on PRLM (phoneme recognizer followed by language model) approach to language recognition. It elaborates on our prior work concerning the quality of phoneme recognition and amounts of training data for phoneme recognizer training. It reports improvements brought to our PRLM system by better phoneme recognition and Witten-Bell discounting in LM-modeling. The paper then concentrates on the use of phoneme lattices and anti-models. Training and scoring on phoneme lattices brought significant improvement in language recognition accuracy. The antimodels are simple, yet powerful technique to improve the discrimination between target and non-target languages. All results are reported on standard NIST 2003 data; comparison with other published results is favorable to our system. Pavel Matejka, Petr Schwarz, Lukás Burget, Jan Cernocký |
ICASSP (1) | 3 |
| 2005 | Non-parametric speaker turn segmentation of meeting dataabstractAn extension of conventional speaker segmentation framework is presented for a scenario in which a number of microphones record the activity of speakers present at a meeting (one microphone per speaker). Although each microphone can receive speech from both the participant wearing the microphone (local speech) and other participants (cross-talk), the recorded audio can be broadly classified in three ways: local speech, cross-talk, and silence. This paper proposes a technique which takes into account cross-correlations, values of its maxima, and energy differences as features to identify and segment speaker turns. In particular, we have used classical cross-correlation functions, time smoothing and in part temporal constraints to sharpen and disambiguate timing differences between microphone channels that may be dominated by noise and reverberation. Experimental results show that proposed technique can be successively used for speaker segmentation of data collected from a number of different setups. 1. Petr Motlícek, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2005 | Comparison of keyword spotting approaches for informal continuous speechabstractThis paper describes several approaches to keyword spotting (KWS) for informal continuous speech. We compare acoustic keyword spotting, spotting in word lattices generated by large vocabulary continuous speech recognition and a hybrid approach making use of phoneme lattices generated by a phoneme recognizer. The systems are compared on carefully defined test data extracted from ICSI meeting database. The acoustic and phoneme-lattice based KWS are based on a phoneme recognizer making use of temporal-pattern (TRAP) feature extraction and posterior estimation using neural nets. We show its superiority over traditional HMM/GMM systems. The advantages and drawbacks of different approaches are discussed. 1. Igor Szöke, Petr Schwarz, Pavel Matejka, Lukás Burget, Martin Karafiát, Michal Fapso, Jan Cernocký |
INTERSPEECH | 4 |
| 2004 | Combination of speech features using smoothed heteroscedastic linear discriminant analysis
Lukás Burget |
INTERSPEECH | 1 |
| 2002 | Qualcomm-ICSI-OGI features for ASRabstractOur feature extraction module for the Aurora task is based on a combination of a conventional noise supression technique (Wiener filtering) with our temporal processing technigues (linear discriminant RASTA filtering and nonlinear TempoRAl Pattern (TRAP) classifier). We observe better than 58% relative error improvement on the prescribed Aurora Digit Task, a performance level that is somewhat better than the new ETSI Advanced Feature standard. Further- more, to test generalization of our approach to an independent test set not available during development, we evaluate performance on American English SpeechDatCar digits and show 10.54% relative improvement over the new ETSI stan- dard. André Adami, Lukás Burget, Stéphane Dupont, Harinath Garudadri, Frantisek Grézl, Hynek Hermansky, Pratibha Jain, Sachin S. Kajarekar, Nelson Morgan, Sunil Sivadas |
INTERSPEECH | 2 |
| 2002 | Noise estimation for efficient speech enhancement and robust speech recognition
Petr Motlícek, Lukás Burget |
INTERSPEECH | 2 |
| 2001 | Robust ASR front-end using spectral-based and discriminant features: experiments on the Aurora tasksabstractThis paper describes an automatic speech recognition frontend that combines low-level robust ASR feature extraction techniques, and higher-level linear and non-linear feature transformations. The low-level algorithms use data-derived filters, mean and variance normalization of the feature vectors, and dropping of noise frames. The feature vectors are then linearly transformed using Principal Components Analysis (PCA). An Artificial Neural Network (ANN) is also used to compute features that are useful for classification of speech sounds. It is trained for phoneme probability estimation on a large corpus of noisy speech. These transformations lead to two feature streams whose vectors are concatenated and then used for speech recognition. This method was tested on the set of speech corpora used for the “Aurora” evaluation. Using the feature stream generated without the ANN yields an overall 41% reduction of the error rate over Mel-Frequency Cepstral Coefficients (MFCC) reference features. Adding the ANN stream further reduces the error rate yielding a 46% reduction over the reference features. M. Carmen Benítez, Lukás Burget, Barry Y. Chen, Stéphane Dupont, Harinath Garudadri, Hynek Hermansky, Pratibha Jain, Sachin S. Kajarekar, Nelson Morgan, Sunil Sivadas |
INTERSPEECH | 2 |