Jan Cernocký

dblp:93/2509 · also Honza Cernocký, Jan Honza Cernocký · DBLP profile ↗
← Back
150ranked-venue papers
3as first author
42since 2021 · last 2026
0000-0002-8800-0210ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 137 · 3 first-author · 39 since 2021Artificial intelligence and machine learning · 89 · 1 first-author · 24 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Trainable multi-channel front-ends for joint beamforming and speaker embedding extraction
abstract
Multi-channel speaker verification (SV), employing numerous microphones for capturing enrollment and/or test recordings, gained attention for its benefits in far-field scenarios. While some studies approach the problem by designing multi-channel embedding extractors, we focus on building and thoroughly analyzing a framework integrating beamforming pre-processing paired with single-channel embedding extraction. This strategy benefits from accommodating both multi-channel and single-channel inputs. Furthermore, it provides human-interpretable intermediate output — enhanced speech — that can be independently evaluated and related to SV performance. We first focus on the front-end, taking advantage of deep-learning source separation for direct or indirect mask estimation required by the beamformer. We alternate single-channel network architectures, subsequently extended to multi-channel ones by reference channel attention (RCA). We also analyze the impact of beamformer and network output fusion. Finally, we show improvements brought by end-to-end fine-tuning the entire architecture facilitated by our newly designed multi-channel corpus, MultiSV2, extending our previous MultiSV dataset.
Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký, Meng Yu 0003
Comput. Speech Lang.5
2026 DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition
Alexander Polok, Dominik Klement, Martin Kocour, Jiangyu Han, Federico Landini, Bolaji Yusuf, Matthew Wiesner, Sanjeev Khudanpur, Jan Cernocký, Lukás Burget
Comput. Speech Lang.9
2026 HPQ: A Hybrid Framework for Joint Pruning and Quantization of Self-Supervised Speech Models
Junyi Peng, Lin Zhang 0054, Jiangyu Han, Oldrich Plchot, Shuai Wang 0016, Jan Cernocký
IEEE Signal Process. Lett.6
2025 State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
abstract
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments be-longing to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation.
Sara Barahona, Ladislav Mosner, Themos Stafylakis, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký
ASRU7
2025 DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
abstract
This paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and generalization in both in- and out-of-domain settings. The proposed method, Decoder-Centric Regularization in EncoderDecoder (DeCRED), adds auxiliary classifiers to the decoder, enabling next token prediction via intermediate logits. Empirically, DeCRED reduces the mean internal LM BPE perplexity by 36.6% relative to 11 test sets. Furthermore, this translates into actual WER improvements over the baseline in 5 of 7 in-domain and 3 of 4 out-of-domain test sets, reducing macro WER from 6.4% to 6.3% and 18.2 % to 16.2 %, respectively. On TEDLIUM3, DeCRED achieves 7.0 % WER, surpassing the baseline and encoder-centric InterCTC regularization by 0.6 % and 0.5%, respectively. Finally, we compare DeCRED with OWSM v3.1 and Whisper-medium, showing competitive WERs despite training on much less data with fewer parameters.
Alexander Polok, Santosh Kesiraju, Karel Benes, Bolaji Yusuf, Lukás Burget, Jan Cernocký
ASRU6
2025 Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
abstract
Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
Sathvik Udupa, Shinji Watanabe 0001, Petr Schwarz, Jan Cernocký
ASRU4
2025 TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
abstract
Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions—a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.1
Junyi Peng, Takanori Ashihara, Marc Delcroix, Tsubasa Ochiai, Oldrich Plchot, Shoko Araki, Jan Cernocký
ICASSP7
2025 CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
abstract
Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42%, 0.48%, and 0.96% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.1
Junyi Peng, Ladislav Mosner, Lin Zhang 0054, Oldrich Plchot, Themos Stafylakis, Lukás Burget, Jan Cernocký
ICASSP7
2025 Target Speaker ASR with Whisper
abstract
We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single-speaker ASR models into target-speaker ASR models. Our approach also supports speaker-attributed ASR by sequentially generating transcripts for each speaker in a diarization output. This simplified method outperforms baseline speech separation and diarization cascade by 12.9 % absolute ORC-WER on the NOTSOFAR-1 dataset.
Alexander Polok, Dominik Klement, Matthew Wiesner, Sanjeev Khudanpur, Jan Cernocký, Lukás Burget
ICASSP5
2025 Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Jan Cernocký, Lukás Burget
INTERSPEECH6
2025 Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
Pradyoth Hegde, Santosh Kesiraju, Jan Svec, Simon Sedlácek, Bolaji Yusuf, Oldrich Plchot, Deepak K. T, Jan Cernocký
INTERSPEECH8
2025 Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Simon Sedlácek, Bolaji Yusuf, Jan Svec, Pradyoth Hegde, Santosh Kesiraju, Oldrich Plchot, Jan Cernocký
INTERSPEECH7
2024 Diacorrect: Error Correction Back-End for Speaker Diarization
abstract
In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and a transformer-based decoder. By exploiting the interactions between the input recording and the initial system’s outputs, DiaCorrect can automatically correct the initial speaker activities to minimize the diarization errors. Experiments on 2-speaker telephony data show that the proposed DiaCorrect can effectively improve the initial model’s results. Our source code is publicly available at https://github.com/BUTSpeechFIT/diacorrect.
Jiangyu Han, Federico Landini, Johan Rohdin, Mireia Díez, Lukás Burget, Yuhang Cao, Jan Cernocký
ICASSP8
2024 Target Speech Extraction with Pre-Trained Self-Supervised Learning Models
abstract
Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exploit pre-trained SSL models for two purposes within a TSE framework, i.e., to process the input mixture and to derive speaker embeddings from the enrollment. In this paper, we focus on how to effectively use SSL models for TSE. We first introduce a novel TSE downstream task following the SUPERB principles. This simple experiment shows the potential of SSL models for TSE, but extraction performance remains far behind the state-of-the-art. We then extend a powerful TSE architecture by incorporating two SSL-based modules: an Adaptive Input Enhancer (AIE) and a speaker encoder. Specifically, the proposed AIE utilizes intermediate representations from the CNN encoder by adjusting the time resolution of CNN encoder and transformer blocks through progressive upsampling, capturing both fine-grained and hierarchical features. Our method outperforms current TSE systems achieving a SI-SDR improvement of 14.0 dB on LibriMix. Moreover, we can further improve performance by 0.7 dB by fine-tuning the whole model including the SSL model parameters.
Junyi Peng, Marc Delcroix, Tsubasa Ochiai, Oldrich Plchot, Shoko Araki, Jan Cernocký
ICASSP6
2024 Multi-Channel Extension of Pre-trained Models for Speaker Verification
abstract
International audience
Ladislav Mosner, Romain Serizel, Lukás Burget, Oldrich Plchot, Emmanuel Vincent 0001, Junyi Peng, Jan Cernocký
INTERSPEECH7
2024 BESST Dataset: A Multimodal Resource for Speech-based Stress Detection and Analysis
Jan Pesán, Vojtech Jurík, Martin Karafiát, Jan Cernocký
INTERSPEECH4
2024 Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
Bolaji Yusuf, Jan Cernocký, Murat Saraclar
INTERSPEECH2
2023 Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters
abstract
Recently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model size grows and sometimes results in over-fitting on small datasets. In this paper, we conduct a comprehensive analysis of applying parameter-efficient transfer learning (PETL) methods to reduce the required learnable parameters for adapting to speaker verification tasks. Specifically, during the fine-tuning process, the pre-trained models are frozen, and only lightweight modules inserted in each Transformer block are trainable (a method known as adapters). Moreover, to boost the performance in a cross-language low-resource scenario, the Transformer model is further tuned on a large intermediate dataset before directly fine-tuning it on a small dataset. With updating fewer than 4% of parameters, (our proposed) PETL-based methods achieve comparable performances with full fine-tuning methods (Vox1-O: 0.55%, Vox1-E: 0.82%, Vox1-H:1.73%).
Junyi Peng, Themos Stafylakis, Rongzhi Gu, Oldrich Plchot, Ladislav Mosner, Lukás Burget, Jan Cernocký
ICASSP7
2023 Multi-Channel Speech Separation with Cross-Attention and Beamforming
Ladislav Mosner, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký
INTERSPEECH5
2023 Improving Speaker Verification with Self-Pretrained Transformer Models
Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký
INTERSPEECH6
2023 End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations
abstract
Conventional keyword search systems operate on automatic speech recognition (ASR) outputs, which causes them to have a complex indexing and search pipeline. This has led to interest in ASR-free approaches to simplify the search procedure. We recently proposed a neural ASR-free keyword search model which achieves competitive performance while maintaining an efficient and simplified pipeline, where queries and documents are encoded with a pair of recurrent neural network encoders and the encodings are combined with a dot-product. In this paper, we extend this work with multilingual pretraining and detailed analysis of the model. Our experiments show that the proposed multilingual training significantly improves the model performance and that despite not matching a strong ASR-based conventional keyword search system for short queries and queries comprising in-vocabulary words, the proposed model outperforms the ASR-based system for long queries and queries that do not appear in the training data.
Bolaji Yusuf, Jan Cernocký, Murat Saraclar
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction
abstract
In recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks. In this paper, from the time-frequency domain perspective, we propose a densely-connected pyramid complex convolutional network, termed DPCCN, to improve the robustness of speech separation under complicated conditions. Furthermore, we generalize the DPCCN to target speech extraction (TSE) by integrating a new specially designed speaker encoder. Moreover, we also investigate the robustness of DPCCN to unsupervised cross-domain TSE tasks. A Mixture-Remix approach is proposed to adapt the target domain acoustic characteristics for fine-tuning the source model. We evaluate the proposed methods not only under noisy and reverberant in-domain condition, but also in clean but cross-domain conditions. Results show that for both speech separation and extraction, the DPCCN-based systems achieve significantly better performance and robustness than the currently dominating time-domain methods, especially for the cross-domain tasks. Particularly, we find that the Mixture-Remix fine-tuning with DPCCN significantly outperforms the TD-SpeakerBeam for unsupervised cross-domain TSE, with around 3.5 dB SISNR improvement on target domain test set, without any source domain performance degradation.
Jiangyu Han, Yanhua Long, Lukás Burget, Jan Cernocký
ICASSP4
2022 Multisv: Dataset for Far-Field Multi-Channel Speaker Verification
abstract
Motivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experiments with dereverberation, denoising, and speech enhancement. We tackled the ever-present problem of the lack of multi-channel training data by utilizing data simulation on top of clean parts of the Voxceleb corpus. The development and evaluation trials are based on a retransmitted Voices Obscured in Complex Environmental Settings (VOiCES) corpus, which we modified to provide multi-channel trials. We publish full recipes that create the dataset from public sources as the MultiSV dataset, and we provide results with two of our multi-channel speaker verification systems with neural network-based beamforming based either on predicting ideal binary masks or the more recent Conv-TasNet.
Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký
ICASSP4
2022 Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer
abstract
We focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, and we adapt it to perform speech enhancement by forcing it to separate speech and additive noise. We experiment with using the STFT of Conv-TasNet outputs to obtain SCMs of speech and noise, and finally, we fine-tune this multi-channel frontend w.r.t. speaker verification objective. We successfully tackle the problem of the lack of a realistic multichannel training set by using simulated data of MultiSV corpus. The analysis is performed on its retransmitted and simulated test parts. We achieve consistent improvements with a 2.7 times smaller model than the baseline based on a scheme with mask estimating NN.
Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký
ICASSP4
2022 Speaker adaptation for Wav2vec2 based dysarthric ASR
Murali Karthick Baskar, Tim Herzig, Diana Nguyen, Mireia Díez, Tim Polzehl, Lukás Burget, Jan Cernocký
INTERSPEECH7
2022 Revisiting joint decoding based multi-talker speech recognition with DNN acoustic model
abstract
In typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker decoder which is applied on each speaker-specific output stream separately. In this work, we argue that such a scheme is sub-optimal and propose a principled solution that decodes all speakers jointly. We modify the acoustic model to predict joint state posteriors for all speakers, enabling the network to express uncertainty about the attribution of parts of the speech signal to the speakers. We employ a joint decoder that can make use of this uncertainty together with higher-level language information. For this, we revisit decoding algorithms used in factorial generative models in early multi-talker speech recognition systems. In contrast with these early works, we replace the GMM acoustic model with DNN, which provides greater modeling power and simplifies part of the inference. We demonstrate the advantage of joint decoding in proof of concept experiments on a mixed-TIDIGITS dataset.
Martin Kocour, Katerina Zmolíková, Lucas Ondel Yang, Jan Svec, Marc Delcroix, Tsubasa Ochiai, Lukás Burget, Jan Cernocký
INTERSPEECH8
2022 Learnable Sparse Filterbank for Speaker Verification
Junyi Peng, Rongzhi Gu, Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký
INTERSPEECH6
2022 Training speaker embedding extractors using multi-speaker audio with unknown speaker boundaries
abstract
In this paper, we demonstrate a method for training speaker embedding extractors using weak annotation.More specifically, we are using the full VoxCeleb recordings and the name of the celebrities appearing on each video without knowledge of the time intervals the celebrities appear in the video.We show that by combining a baseline speaker diarization algorithm that requires no training or parameter tuning, a modified loss with aggregation over segments, and a two-stage training approach, we are able to train a competitive ResNet-based embedding extractor.Finally, we experiment with two different aggregation functions and analyze their behaviour in terms of their gradients.
Themos Stafylakis, Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Anna Silnova, Lukás Burget, Jan Cernocký
INTERSPEECH7
2022 An Attention-Based Backend Allowing Efficient Fine-Tuning of Transformer Models for Speaker Verification
abstract
In recent years, self-supervised learning paradigm has received extensive attention due to its great success in various down-stream tasks. However, the fine-tuning strategies for adapting those pre-trained models to speaker verification task have yet to be fully explored. In this paper, we analyze several feature extraction approaches built on top of a pre-trained model, as well as regularization and a learning rate scheduler to stabilize the fine-tuning process and further boost performance: multi-head factorized attentive pooling is proposed to factorize the comparison of speaker representations into multiple phonetic clusters. We regularize towards the parameters of the pre-trained model and we set different learning rates for each layer of the pre-trained model during fine-tuning. The experimental results show our method can significantly shorten the training time to 4 hours and achieve SOTA performance: 0.59%, 0.79% and 1.77% EER on Vox1-O, Vox1-E and Vox1-H, respectively.11Code is available at https://github.com/JunyiPeng00/IEEE-SLT22-Pretrained-Model-for-SV.
Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký
SLT6
2022 Extracting Speaker and Emotion Information from Self-Supervised Speech Models via Channel-Wise Correlations
abstract
Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first - and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation.
Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot, Lukás Burget, Jan Cernocký
SLT6
2022 Spelling-Aware Word-Based End-to-End ASR
abstract
We propose a new end-to-end architecture for automatic speech recognition that expands the “listen, attend and spell” (LAS) paradigm. While the main word-predicting network is trained to predict words, the secondary, speller network, is optimized to predict word spellings from inner representations of the main network (e.g. word embeddings or context vectors from the attention module). We show that this joint training improves the word error rate of a word-based system and enables solving additional tasks, such as out-of-vocabulary word detection and recovery. The tests are conducted on LibriSpeech dataset consisting of 1000 h of read speech.
Ekaterina Egorova, Hari Krishna Vydana, Lukás Burget, Jan Cernocký
IEEE Signal Process. Lett.4
2021 Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition
abstract
Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the ASR hypotheses before forwarding it to TTS. 2) In the TTS→ASR direction, a hyper-parameter is introduced to scale the attention context from synthesized speech before sending it to ASR to handle out-of-domain data. Training strategies and the effectiveness of the EAT model are explored under out-of-domain data conditions. The results show that EAT reduces the performance gap between supervised and self-supervised training significantly by absolute 2.6% and 2.7% on Librispeech and BABEL respectively.
Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Jan Cernocký
ICASSP5
2021 Analysis of X-Vectors for Low-Resource Speech Recognition
abstract
The paper presents a study of usability of x-vectors for adaptation of automatic speech recognition (ASR) systems. X-vectors are Neural Network (NN)-based speaker embeddings recently proposed in speaker recognition (SR). They quickly replaced common i-vectors and became new state-of-the-art technique. Here, the same approach is adopted for ASR with the hope of similar outcome. All experiments were done on ASR for the latest IARPA MATERIAL evaluation running on Pashto language. Over 1% absolute improvement was observed with x-vectors over traditional i-vectors, even when the x-vector extractor was not trained on target Pashto data.
Martin Karafiát, Karel Veselý, Jan Cernocký, Ján Profant, Jirí Nytra, Miroslav Hlavácek, Tomás Pavlícek
ICASSP3
2021 Jointly Trained Transformers Models for Spoken Language Translation
abstract
End-to-End and cascade (ASR-MT) spoken language translation (SLT) systems are reaching comparable performances, however, a large degradation is observed when translating the ASR hypothesis in comparison to using oracle input text. In this work, degradation in performance is reduced by creating an End-to-End differentiable pipeline between the ASR and MT systems. In this work, we train SLT systems with ASR objective as an auxiliary loss and both the networks are connected through the neural hidden representations. This training has an End-to-End differentiable path with respect to the final objective function and utilizes the ASR objective for better optimization. This architecture has improved the BLEU score from 41.21 to 44.69. Ensembling the proposed architecture with independently trained ASR and MT systems further improved the BLEU score from 44.69 to 46.9. All the experiments are reported on English-Portuguese speech translation task using the How2 corpus. The final BLEU score is on-par with the best speech translation system on How2 dataset without using any additional training data and language model and using fewer parameters.
Hari Krishna Vydana, Martin Karafiát, Katerina Zmolíková, Lukás Burget, Jan Cernocký
ICASSP5
2021 A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
abstract
In this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach, we frame the task as one of learning embeddings on a low-dimensional phonetic subspace, and simultaneously specify the subspace itself as an embedding on a hyper-subspace. We train the hyper-subspace on a set of transcribed languages and transfer it to the target language. In the target language, we infer both the language and unit embeddings in an unsupervised manner, and in so doing, we simultaneously learn a subspace of units specific to that language and the units that dwell on it. We conduct experiments on TIMIT and two low-resource languages: Mboshi and Yoruba. Results show that our model outperforms major acoustic unit discovery techniques, both in terms of clustering quality and segmentation accuracy.
Bolaji Yusuf, Lucas Ondel Yang, Lukás Burget, Jan Cernocký, Murat Saraclar
ICASSP4
2021 Out-of-Vocabulary Words Detection with Attention and CTC Alignments in an End-to-End ASR System
Ekaterina Egorova, Hari Krishna Vydana, Lukás Burget, Jan Cernocký
Interspeech4
2021 Boosting of Contextual Information in ASR for Air-Traffic Call-Sign Recognition
abstract
Contextual adaptation of ASR can be very beneficial for multi-accent and often noisy Air-Traffic Control (ATC) speech. Our focus is call-sign recognition, which can be used to track conversations of ATC operators with individual airplanes. We developed a two-stage boosting strategy, consisting of HCLG boosting and Lattice boosting. Both are implemented as WFST compositions and the contextual information is specific to each utterance. In HCLG boosting we give score discounts to individual words, while in Lattice boosting the score discounts are given to word sequences. The context data have origin in surveillance database of OpenSky Network. From this, we obtain lists of call-signs that are made more likely to appear in the best hypothesis of ASR. This also improves the accuracy of the NLU module that recognizes the call-signs from the best hypothesis of ASR.
Martin Kocour, Karel Veselý, Alexander Blatt, Juan Zuluaga-Gomez, Igor Szöke, Jan Cernocký, Dietrich Klakow, Petr Motlícek
Interspeech6
2021 Effective Phase Encoding for End-To-End Speaker Verification
Junyi Peng, Xiaoyang Qu, Rongzhi Gu, Jianzong Wang, Jing Xiao 0006, Lukás Burget, Jan Cernocký
Interspeech7
2021 ICSpk: Interpretable Complex Speaker Embedding Extractor from Raw Waveform
Junyi Peng, Xiaoyang Qu, Jianzong Wang, Rongzhi Gu, Jing Xiao 0006, Lukás Burget, Jan Cernocký
Interspeech7
2021 Detecting English Speech in the Air Traffic Control Voice Communication
abstract
We launched a community platform for collecting the ATC speech world-wide in the ATCO2 project. Filtering out unseen non-English speech is one of the main components in the data processing pipeline. The proposed English Language Detection (ELD) system is based on the embeddings from Bayesian subspace multinomial model. It is trained on the word confusion network from an ASR system. It is robust, easy to train, and light weighted. We achieved 0.0439 equal-error-rate (EER), a 50% relative reduction as compared to the state-of-the-art acoustic ELD system based on x-vectors, in the in-domain scenario. Further, we achieved an EER of 0.1352, a 33% relative reduction as compared to the acoustic ELD, in the unseen language (out-of-domain) condition. We plan to publish the evaluation dataset from the ATCO2 project.
Igor Szöke, Santosh Kesiraju, Ondrej Novotný, Martin Kocour, Karel Veselý, Jan Cernocký
Interspeech6
2021 Auxiliary Loss Function for Target Speech Extraction and Recognition with Weak Supervision Based on Speaker Characteristics
Katerina Zmolíková, Marc Delcroix, Desh Raj, Shinji Watanabe 0001, Jan Cernocký
Interspeech5
2021 Integration of Variational Autoencoder and Spatial Clustering for Adaptive Multi-Channel Neural Speech Separation
abstract
In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in several works. As the spectral model, previous works used either factorial generative models of the mixed speech or discriminative neural networks. In our work, we combine the strengths of both approaches, by building a factorial model based on a generative neural network, a variational autoencoder. By doing so, we can exploit the modeling power of neural networks, but at the same time, keep a structured model. Such a model can be advantageous when adapting to new noise conditions as only the noise part of the model needs to be modified. We show experimentally, that our model significantly outperforms previous factorial model based on Gaussian mixture model (DOLPHIN), performs comparably to integration of permutation invariant training with spatial clustering, and enables us to easily adapt to new noise conditions.
Katerina Zmolíková, Marc Delcroix, Lukás Burget, Tomohiro Nakatani, Jan Cernocký
SLT5
2020 Investigation of Specaugment for Deep Speaker Embedding Learning
abstract
SpecAugment is a newly proposed data augmentation method for speech recognition. By randomly masking bands in the log Mel spectogram this method leads to impressive performance improvements. In this paper, we investigate the usage of SpecAugment for speaker verification tasks. Two different models, namely 1-D convolutional TDNN and 2-D convolutional ResNet34, trained with either Softmax or AAM-Softmax loss, are used to analyze SpecAugment's effectiveness. Experiments are carried out on the Voxceleb and NIST SRE 2016 dataset. By applying SpecAugment to the original clean data in an on-the-fly manner without complex off-line data augmentation methods, we obtained 3.72% and 11.49% EER for NIST SRE 2016 Cantonese and Tagalog, respectively. For Voxceleb1 evaluation set, we obtained 1.47% EER.
Shuai Wang 0016, Johan Rohdin, Oldrich Plchot, Lukás Burget, Kai Yu 0004, Jan Cernocký
ICASSP6
2020 Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge
abstract
This paper presents an analysis of our diarization system winning the second DIHARD speech diarization challenge, track 1. This system is based on clustering x-vector speaker embeddings extracted every 0.25s from short segments of the input recording. In this paper, we focus on the two x-vector clustering methods employed, namely Agglomerative Hierarchical Clustering followed by a clustering based on Bayesian Hidden Markov Model (BHMM). Even though the system submitted to the challenge had further post-processing steps, we will show that using this BHMM solely is enough to achieve the best performance in the challenge. The analysis will show improvements achieved by optimizing individual processing steps, including a simple procedure to effectively perform "domain adaptation" by Probabilistic Linear Discriminant Analysis model interpolation. All experiments are performed in the DIHARD II evaluation framework.
Mireia Díez, Lukás Burget, Federico Landini, Shuai Wang 0016, Jan Cernocký
ICASSP5
2020 13 years of speaker recognition research at BUT, with longitudinal analysis of NIST SRE
Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Johan Rohdin, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Ondrej Novotný, Mireia Díez, Jan Cernocký
Comput. Speech Lang.11
2020 Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors
abstract
In our previous work, we introduced our Bayesian Hidden Markov Model with eigenvoice priors, which has been recently recognized as the state-of-the-art model for Speaker Diarization. In this article we present a more complete analysis of the Diarization system. The inference of the model is fully described and derivations of all update formulas are provided for a complete understanding of the algorithm. An extensive analysis on the effect, sensitivity and interactions of all model parameters is provided, which might be used as a guide for their optimal setting. The newly introduced speaker regularization coefficient allows us to control the number of speakers inferred in an utterance. A naive speaker model merging strategy is also presented, which allows to drive the variational inference out of local optima. Experiments for the different diarization scenarios are presented on CALLHOME and DIHARD datasets.
Mireia Díez, Lukás Burget, Federico Landini, Jan Cernocký
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Speaker Verification with Application-Aware Beamforming
abstract
Multichannel speech processing applications usually employ beamformers as means of speech enhancement through spatial filtering. Beamformers with learnable parameters require training to minimize a loss function that is not necessarily correlated with the final objective. In this paper, we present a framework employing recent neural network based generalized eigenvalue beamformer and application-specific model that allows for optimization of beamformer w.r.t. target application. In our case, the application is speaker verification which utilizes a speaker embedding (x-vector) extractor that conveniently comes with desired loss. We show that application-specific training of the beamformer brings performance improvements over a system trained in the standard way. We perform our analysis on the recently introduced VOiCES corpus which contains multichannel data and allows us to modify the evaluation trials such that enrollment recordings remain single-channel and test utterances are multichannel.
Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Lukás Burget, Jan Cernocký
ASRU5
2019 A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: The Deepmine Database
abstract
DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers and 540 thousand recordings overall, more than 480 hours of speech are transcribed. It is the first public large-scale speaker verification database in Persian, the largest public text-dependent and text-prompted speaker verification database in English, and the largest public evaluation dataset for text-independent speaker verification. It has a good coverage of age, gender, and accents. We provide several evaluation protocols for each part of the database to allow for research on different aspects of speaker verification. We also provide the results of several experiments that can be considered as baselines: HMM-based i-vectors for text-dependent speaker verification, and HMM-based as well as state-of-the-art deep neural network based ASR. We demonstrate that the database can serve for training robust ASR models.
Hossein Zeinali, Lukás Burget, Jan Cernocký
ASRU3
2019 Promising Accurate Prefix Boosting for Sequence-to-sequence ASR
abstract
In this paper, we present promising accurate prefix boosting (PAPB), a discriminative training technique for attention based sequence-to-sequence (seq2seq) ASR. PAPB is devised to unify the training and testing scheme effectively. The training procedure involves maximizing the score of each partial correct sequence obtained during beam search compared to other hypotheses. The training objective also includes minimization of token (character) error rate. PAPB shows its efficacy by achieving 10.8% and 3.8% WER with and without external RNNLM respectively on Wall Street Journal dataset.
Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Martin Karafiát, Takaaki Hori, Jan Cernocký
ICASSP6
2019 How to Improve Your Speaker Embeddings Extractor in Generic Toolkits
abstract
Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine several tricks in training, such as the effects of normalizing input features and pooled statistics, different methods for preventing overfitting as well as alternative non-linearities that can be used instead of Rectifier Linear Units. In addition, we investigate the difference in performance between TDNN and CNN, and between two types of attention mechanism. Experimental results on Speaker in the Wild, SRE 2016 and SRE 2018 datasets demonstrate the effectiveness of the proposed implementation.
Hossein Zeinali, Lukás Burget, Johan Rohdin, Themos Stafylakis, Jan Cernocký
ICASSP5
2019 Semi-Supervised Sequence-to-Sequence ASR Using Unpaired Speech and Text
abstract
Sequence-to-sequence automatic speech recognition (ASR) models require large quantities of data to attain high performance. For this reason, there has been a recent surge in interest for unsupervised and semi-supervised training in such models. This work builds upon recent results showing notable improvements in semi-supervised training using cycle-consistency and related techniques. Such techniques derive training procedures and losses able to leverage unpaired speech and/or text data by combining ASR with Text-to-Speech (TTS) models. In particular, this work proposes a new semi-supervised loss combining an end-to-end differentiable ASR$\rightarrow$TTS loss with TTS$\rightarrow$ASR loss. The method is able to leverage both unpaired speech and text data to outperform recently proposed related techniques in terms of \%WER. We provide extensive results analyzing the impact of data quantity and speech and text modalities and show consistent gains across WSJ and Librispeech corpora. Our code is provided in ESPnet to reproduce the experiments.
Murali Karthick Baskar, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Takaaki Hori, Lukás Burget, Jan Cernocký
INTERSPEECH6
2019 Bayesian HMM Based x-Vector Clustering for Speaker Diarization
Mireia Díez, Lukás Burget, Shuai Wang 0016, Johan Rohdin, Jan Cernocký
INTERSPEECH5
2019 Analysis of Multilingual Sequence-to-Sequence Speech Recognition Systems
abstract
This paper investigates the applications of various multilingual approaches developed in conventional hidden Markov model (HMM) systems to sequence-to-sequence (seq2seq) automatic speech recognition (ASR). On a set composed of Babel data, we first show the effectiveness of multi-lingual training with stacked bottle-neck (SBN) features. Then we explore various architectures and training strategies of multi-lingual seq2seq models based on CTC-attention networks including combinations of output layer, CTC and/or attention component re-training. We also investigate the effectiveness of language-transfer learning in a very low resource scenario when the target language is not included in the original multi-lingual training data. Interestingly, we found multilingual features superior to multilingual models, and this finding suggests that we can efficiently combine the benefits of the HMM system with the seq2seq system through these multilingual feature techniques.
Martin Karafiát, Murali Karthick Baskar, Shinji Watanabe 0001, Takaaki Hori, Matthew Wiesner, Jan Cernocký
INTERSPEECH6
2019 Bayesian Subspace Hidden Markov Model for Acoustic Unit Discovery
abstract
This work tackles the problem of learning a set of language specific acoustic units from unlabeled speech recordings given a set of labeled recordings from other languages. Our approach may be described by the following two steps procedure: first the model learns the notion of acoustic units from the labelled data and then the model uses its knowledge to find new acoustic units on the target language. We implement this process with the Bayesian Subspace Hidden Markov Model (SHMM), a model akin to the Subspace Gaussian Mixture Model (SGMM) where each low dimensional embedding represents an acoustic unit rather than just a HMM's state. The subspace is trained on 3 languages from the GlobalPhone corpus (German, Polish and Spanish) and the AUs are discovered on the TIMIT corpus. Results, measured in equivalent Phone Error Rate, show that this approach significantly outperforms previous HMM based acoustic units discovery systems and compares favorably with the Variational Auto Encoder-HMM.
Lucas Ondel Yang, Hari Krishna Vydana, Lukás Burget, Jan Cernocký
INTERSPEECH4
2019 On the Usage of Phonetic Information for Text-Independent Speaker Embedding Extraction
Shuai Wang 0016, Johan Rohdin, Lukás Burget, Oldrich Plchot, Yanmin Qian, Kai Yu 0004, Jan Cernocký
INTERSPEECH7
2019 Detecting Spoofing Attacks Using VGG and SincNet: BUT-Omilia Submission to ASVspoof 2019 Challenge
abstract
In this paper, we present the system description of the joint efforts of Brno University of Technology (BUT) and Omilia -- Conversational Intelligence for the ASVSpoof2019 Spoofing and Countermeasures Challenge. The primary submission for Physical access (PA) is a fusion of two VGG networks, trained on single and two-channels features. For Logical access (LA), our primary system is a fusion of VGG and the recently introduced SincNet architecture. The results on PA show that the proposed networks yield very competitive performance in all conditions and achieved 86\:\% relative improvement compared to the official baseline. On the other hand, the results on LA showed that although the proposed architecture and training strategy performs very well on certain spoofing attacks, it fails to generalize to certain attacks that are unseen during training.
Hossein Zeinali, Themos Stafylakis, Georgia Athanasopoulou, Johan Rohdin, Ioannis Gkinis, Lukás Burget, Jan Cernocký
INTERSPEECH7
2019 Analysis of DNN Speech Signal Enhancement for Robust Speaker Recognition
Ondrej Novotný, Oldrich Plchot, Ondrej Glembek, Jan Cernocký, Lukás Burget
Comput. Speech Lang.4
2018 Analysis of Multilingual Blstm Acoustic Model on Low and High Resource Languages
abstract
The paper provides an analysis of automatic speech recognition systems (ASR) based on multilingual BLSTM, where we used multi-task training with separate classification layer for each language. The focus is on low resource languages, where only a limited amount of transcribed speech is available. In such scenario, we found it essential to train the ASR systems in a multilingual fashion and we report superior results obtained with pre-trained multilingual BLSTM on this task. The high resource languages are also taken into account and we show the importance of language richness for multilingual training. Next, we present the performance of this technique as a function of amount of target language data. The importance of including context information into BLSTM multilingual systems is also stressed, and we report increased resilience of large NNs to overtraining in case of multi-task training.
Martin Karafiát, Murali Karthick Baskar, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký
ICASSP6
2018 Dereverberation and Beamforming in Far-Field Speaker Recognition
abstract
This paper deals with far-field speaker recognition. On a corpus of NIST SRE 2010 data retransmitted in a real room with multiple microphones, we first demonstrate how room acoustics cause significant degradation of state-of-the-art i-vector based speaker recognition system. We then investigate several techniques to improve the performances ranging from probabilistic linear discriminant analysis (PLDA) re-training, through dereverberation, to beamforming. We found that weighted prediction error (WPE) based dereverberation combined with generalized eigenvalue beamformer with power-spectral density (PSD) weighting masks generated by neural networks (NN) provides results approaching the clean close-microphone setup. Further improvement was obtained by re-training PLDA or the mask-generating NNs on simulated target data. The work shows that a speaker recognition system working robustly in the far-field scenario can be developed.
Ladislav Mosner, Pavel Matejka, Ondrej Novotný, Jan Cernocký
ICASSP4
2018 Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion
abstract
This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters. Following our previous work, we inform the neural network about the target speaker using information extracted from an adaptation utterance’ enabling the network to track the target speaker. While in the previous work, this method was used to separately extract the speaker and then pass such preprocessed speech to a speech recognition system, here we explore training both systems jointly with a common speech recognition criterion. We show that integrating the two systems and training for the final objective improves the performance. In addition, the integration enables further sharing of information between the acoustic model and the speaker extraction system, by making use of the predicted HMM-state posteriors to refine the masks used for beamforming.
Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Tomohiro Nakatani, Jan Cernocký
ICASSP6
2018 BUT OpenSAT 2017 Speech Recognition System
Martin Karafiát, Murali Karthick Baskar, Igor Szöke, Vladimir Malenovsky, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký
INTERSPEECH8
2018 Dereverberation and Beamforming in Robust Far-Field Speaker Recognition
Ladislav Mosner, Oldrich Plchot, Pavel Matejka, Ondrej Novotný, Jan Cernocký
INTERSPEECH5
2018 BUT System for Low Resource Indian Language ASR
Bhargav Pulugundla, Murali Karthick Baskar, Santosh Kesiraju, Ekaterina Egorova, Martin Karafiát, Lukás Burget, Jan Cernocký
INTERSPEECH7
2018 Lightly Supervised vs. Semi-supervised Training of Acoustic Model on Luxembourgish for Low-resource Automatic Speech Recognition
Karel Veselý, Carlos Segura, Igor Szöke, Jordi Luque, Jan Cernocký
INTERSPEECH5
2017 Residual memory networks: Feed-forward approach to learn long-term temporal dependencies
abstract
Training deep recurrent neural network (RNN) architectures is complicated due to the increased network complexity. This disrupts the learning of higher order abstracts using deep RNN. In case of feed-forward networks training deep structures is simple and faster while learning long-term temporal information is not possible. In this paper we propose a residual memory neural network (RMN) architecture to model short-time dependencies using deep feed-forward layers having residual and time delayed connections. The residual connection paves way to construct deeper networks by enabling unhindered flow of gradients and the time delay units capture temporal information with shared weights. The number of layers in RMN signifies both the hierarchical processing depth and temporal depth. The computational complexity in training RMN is significantly less when compared to deep recurrent networks. RMN is further extended as bi-directional RMN (BRMN) to capture both past and future information. Experimental analysis is done on AMI corpus to substantiate the capability of RMN in learning long-term information and hierarchical information. Recognition performance of RMN trained with 300 hours of Switchboard corpus is compared with various state-of-the-art LVCSR systems. The results indicate that RMN and BRMN gains 6 % and 3.8 % relative improvement over LSTM and BLSTM networks.
Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Karel Veselý, Frantisek Grézl, Jan Cernocký
ICASSP6
2017 Topic identification of spoken documents using unsupervised acoustic unit discovery
abstract
This paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The discovered phone-like (acoustic) units are further fed into the conventional topic ID framework. Using multilingual bottleneck features for the acoustic unit discovery, we show that the proposed method outperforms other systems that are based on cross-lingual phoneme recognizer.
Santosh Kesiraju, Raghavendra Pappagari, Lucas Ondel Yang, Lukás Burget, Najim Dehak, Sanjeev Khudanpur, Jan Cernocký, Suryakanth V. Gangashetty
ICASSP7
2017 Bayesian phonotactic Language Model for Acoustic Unit Discovery
abstract
Recent work on Acoustic Unit Discovery (AUD) has led to the development of a non-parametric Bayesian phone-loop model where the prior over the probability of the phone-like units is assumed to be sampled from a Dirichlet Process (DP). In this work, we propose to improve this model by incorporating a Hierarchical Pitman-Yor based bigram Language Model on top of the units' transitions. This new model makes use of the phonotactic context information but assumes a fixed number of units. To remedy this limitation we first train a DP phone-loop model to infer the number of units, then, the bigram phone-loop is initialized from the DP phone-loop and trained until convergence of its parameters. Results show an absolute improvement of 1–2%on the Normalized Mutual Information (NMI) metric. Furthermore, we show that, combined with Multilingual Bottleneck (MBN) features the model yields a same or higher NMI as an English phone recogniser trained on TIMIT.
Lucas Ondel Yang, Lukás Burget, Jan Cernocký, Santosh Kesiraju
ICASSP3
2017 2016 BUT Babel System: Multilingual BLSTM Acoustic Model with i-Vector Based Adaptation
Martin Karafiát, Murali Karthick Baskar, Pavel Matejka, Karel Veselý, Frantisek Grézl, Lukás Burget, Jan Cernocký
INTERSPEECH7
2017 Analysis of Score Normalization in Multilingual Speaker Recognition
Pavel Matejka, Ondrej Novotný, Oldrich Plchot, Lukás Burget, Mireia Díez, Jan Cernocký
INTERSPEECH6
2017 Alternative Approaches to Neural Network Based Speaker Verification
Anna Silnova, Lukás Burget, Jan Cernocký
INTERSPEECH3
2017 Semi-Supervised DNN Training with Word Selection for ASR
Karel Veselý, Lukás Burget, Jan Cernocký
INTERSPEECH3
2017 Multilingually trained bottleneck features in spoken language recognition
Radek Fér, Pavel Matejka, Frantisek Grézl, Oldrich Plchot, Karel Veselý, Jan Cernocký
Comput. Speech Lang.6
2017 Text-dependent speaker verification based on i-vectors, Neural Networks and Hidden Markov Models
Hossein Zeinali, Hossein Sameti, Lukás Burget, Jan Cernocký
Comput. Speech Lang.4
2016 Multilingual region-dependent transforms
abstract
In recent years, trained feature extraction (FE) schemes based on neural networks have replaced or complemented traditional approaches in top performing systems. This paper deals with FE in multilingual scenarios with a target language with low amount of transcribed data. Continuing our previous work on multilingual training of Stacked Bottle-Neck Neural Network FE schemes, we concentrate on improving the discriminatively trained Region-Dependent Transforms. We show that multilingual training of RDT can be implemented by merging statistics from several languages. In our case we used up to 11 source languages to build a FE which generalize well for a new language. This allows us to build a strong bootstrapping model for the final ASR system. The results are produced on IARPA Babel data.
Martin Karafiát, Lukás Burget, Frantisek Grézl, Karel Veselý, Jan Cernocký
ICASSP5
2016 Analysis of DNN approaches to speaker identification
abstract
This work studies the usage of the Deep Neural Network (DNN) Bottleneck (BN) features together with the traditional MFCC features in the task of i-vector-based speaker recognition. We decouple the sufficient statistics extraction by using separate GMM models for frame alignment, and for statistics normalization and we analyze the usage of BN and MFCC features (and their concatenation) in the two stages. We also show the effect of using full-covariance GMM models, and, as a contrast, we compare the result to the recent DNN-alignment approach. On the NIST SRE2010, telephone condition, we show 60% relative gain over the traditional MFCC baseline for EER (and similar for the NIST DCF metrics), resulting in 0.94% EER.
Pavel Matejka, Ondrej Glembek, Ondrej Novotný, Oldrich Plchot, Frantisek Grézl, Lukás Burget, Jan Cernocký
ICASSP7
2016 Sequence summarizing neural network for speaker adaptation
abstract
In this paper, we propose a DNN adaptation technique, where the i-vector extractor is replaced by a Sequence Summarizing Neural Network (SSNN). Similarly to i-vector extractor, the SSNN produces a "summary vector", representing an acoustic summary of an utterance. Such vector is then appended to the input of main network, while both networks are trained together optimizing single loss function. Both the i-vector and SSNN speaker adaptation methods are compared on AMI meeting data. The results show comparable performance of both techniques on FBANK system with frame-classification training. Moreover, appending both the i-vector and "summary vector" to the FBANK features leads to additional improvement comparable to the performance of FMLLR adapted DNN system.
Karel Veselý, Shinji Watanabe 0001, Katerina Zmolíková, Martin Karafiát, Lukás Burget, Jan Cernocký
ICASSP6
2016 Learning Document Representations Using Subspace Multinomial Model
Santosh Kesiraju, Lukás Burget, Igor Szöke, Jan Cernocký
INTERSPEECH4
2016 Analysis of Speaker Recognition Systems in Realistic Scenarios of the SITW 2016 Challenge
Ondrej Novotný, Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Jan Cernocký
INTERSPEECH6
2016 Sequence Summarizing Neural Networks for Spoken Language Recognition
Jan Pesán, Lukás Burget, Jan Cernocký
INTERSPEECH3
2016 i-Vector/HMM Based Text-Dependent Speaker Verification System for RedDots Challenge
Hossein Zeinali, Hossein Sameti, Lukás Burget, Jan Cernocký, Nooshin Maghsoodi, Pavel Matejka
INTERSPEECH4
2016 Data Selection by Sequence Summarizing Neural Network in Mismatch Condition Training
Katerina Zmolíková, Martin Karafiát, Karel Veselý, Marc Delcroix, Shinji Watanabe 0001, Lukás Burget, Jan Cernocký
INTERSPEECH7
2016 Multilingual BLSTM and speaker-specific vector adaptation in 2016 but babel system
abstract
This paper provides an extensive summary of BUT 2016 system for the last IARPA Babel evaluations. It concentrates on multi-lingual training of both deep neural network (DNN)-based feature extraction and acoustic models including multilingual training of bidirectional Long Short Term memory networks. Next, two low-dimensional vector approaches to speaker adaptation are investigated: i-vectors and sequence-summarizing neural networks (SSNN). The results provided on three Babel Year 4 languages show clear advantage of both approaches in case limited amount of training data is available. The time necessary for the development of a new system is addressed too, as some of the investigated techniques do not require extensive re-training of the whole system.
Martin Karafiát, Murali Karthick Baskar, Pavel Matejka, Karel Veselý, Frantisek Grézl, Jan Cernocký
SLT6
2016 Analysis of the DNN-based SRE systems in multi-language conditions
abstract
This paper analyzes the behavior of our state-of-the-art Deep Neural Network/i-vector/PLDA-based speaker recognition systems in multi-language conditions. On the “Language Pack” of the PRISM set, we evaluate the systems' performance using the NIST's standard metrics. We show that not only the gain from using DNNs vanishes, nor using dedicated DNNs for target conditions helps, but also the DNN-based systems tend to produce de-calibrated scores under the studied conditions. This work gives suggestions for directions of future research rather than any particular solutions to these issues.
Ondrej Novotný, Pavel Matejka, Ondrej Glembek, Oldrich Plchot, Frantisek Grézl, Lukás Burget, Jan Cernocký
SLT7
2015 Robust speech recognition in unknown reverberant and noisy conditions
abstract
In this paper, we describe our work on the ASpIRE (Automatic Speech recognition In Reverberant Environments) challenge, which aims to assess the robustness of automatic speech recognition (ASR) systems. The main characteristic of the challenge is developing a high-performance system without access to matched training and development data. While the evaluation data are recorded with far-field microphones in noisy and reverberant rooms, the training data are telephone speech and close talking. Our approach to this challenge includes speech enhancement, neural network methods and acoustic model adaptation, We show that these techniques can successfully alleviate the performance degradation due to noisy audio and data mismatch.
Roger Hsiao, Jeff Z. Ma, William Hartmann, Martin Karafiát, Frantisek Grézl, Lukás Burget, Igor Szöke, Jan Cernocký, Shinji Watanabe 0001, Zhuo Chen 0006, Sri Harish Reddy Mallidi, Hynek Hermansky, Stavros Tsakalidis, Richard M. Schwartz
ASRU8
2015 Copingwith channel mismatch in Query-by-Example - But QUESST 2014
abstract
The paper investigates into Query by Example (QbE) - a spoken term detection technique with queries entered by voice. It describes BUT QbE system that achieved the best accuracy in MediaEval QUESST2014 evaluations. This evaluation was challenging because of severe mismatch between queries and utterances, and introduction of new types of queries. The paper provides an analysis of DTW sub-system's in mismatched conditions (especially targeting DTW metrics) and discusses approaches investigated for QUESST2014: generation of calibration side-information by a language identification system, and handling T2 and T3 queries relaxing the constraints of an exact match. All results are provided on QUESST2014 development and evaluation data.
Igor Szöke, Miroslav Skácel, Lukás Burget, Jan Cernocký
ICASSP4
2015 Multilingual bottleneck features for language recognition
Radek Fér, Pavel Matejka, Frantisek Grézl, Oldrich Plchot, Jan Cernocký
INTERSPEECH5
2015 Three ways to adapt a CTS recognizer to unseen reverberated speech in BUT system for the ASpIRE challenge
Martin Karafiát, Frantisek Grézl, Lukás Burget, Igor Szöke, Jan Cernocký
INTERSPEECH5
2014 But neural network features for spontaneous Vietnamese in BABEL
abstract
This paper presents our work on speech recognition of Vietnamese spontaneous telephone conversations. It focuses on feature extraction by Stacked Bottle-Neck neural networks: several improvements such as semi-supervised training on untranscribed data, increasing of precision of state targets, and CMLLR adaptations were investigated. We have also tested speaker adaptive training of this architecture and significant gain was found. The results are reported on BABEL Vietnamese data.
Martin Karafiát, Frantisek Grézl, Mirko Hannemann, Jan Cernocký
ICASSP4
2014 Calibration and fusion of query-by-example systems - But SWS 2013
abstract
This paper summarizes our work for MediaEval 2013 Spoken Web Search task evaluations. The task was Query-by-Example (search of spoken queries within spoken data). We submitted a system composed of 26 subsystems, of which 13 are based on Acoustic Keyword Spotting and 13 on Dynamic Time Warping. All of them use three-state phoneme posteriors as input features. Our main contribution was m-norm normalization of particular subsystems together with the fusion based on binary logistic regression. The results, including per-language analysis, are provided on MediaEval 2013 dataset.
Igor Szöke, Lukás Burget, Frantisek Grézl, Jan Cernocký, Lucas Ondel Yang
ICASSP4
2014 BUT 2014 Babel system: analysis of adaptation in NN based systems
Martin Karafiát, Frantisek Grézl, Karel Veselý, Mirko Hannemann, Igor Szöke, Jan Cernocký
INTERSPEECH6
2014 But ASR system for BABEL Surprise evaluation 2014
abstract
The paper describes Brno University of Technology (BUT) ASR system for 2014 BABEL Surprise language evaluation (Tamil). While being largely based on our previous work, two original contributions were brought: (1) speaker-adapted bottle-neck neural network (BN) features were investigated as an input to DNN recognizer and semi-supervised training was found effective. (2) Adding of noise to training data outperformed a classical de-noising technique while dealing with noisy test data was found beneficial, and the performance of this approach was verified on a relatively clean training/test data setup from a different language. All results are reported on BABEL 2014 Tamil data.
Martin Karafiát, Karel Veselý, Igor Szöke, Lukás Burget, Frantisek Grézl, Mirko Hannemann, Jan Cernocký
SLT7
2013 Manual and semi-automatic approaches to building a multilingual phoneme set
abstract
The paper addresses manual and semi-automatic approaches to building a multilingual phoneme set for automatic speech recognition. The first approach involves mapping and reduction of the phoneme set based on IPA and expert knowledge, the later one involves phoneme confusion matrix generated by a neural network. The comparison is done for 8 languages selected from GlobalPhone on three scenarios: 1) multilingual system with abundant data for all the languages, 2) multilingual systems excluding target language 3) multilingual systems with small amount of data for target languages. For 3), the multilingual system brought improvement for languages close enough to the others in the set.
Ekaterina Egorova, Karel Veselý, Martin Karafiát, Milos Janda, Jan Cernocký
ICASSP5
2013 BUT BABEL system for spontaneous Cantonese
abstract
This paper presents our work on speech recognition of Cantonese spontaneous telephone conversations. The key-points include feature extraction by 6-layer Stacked Bottle-Neck neural network and using fundamental frequency information at its input. We have also investigated into robustness of SBN training (silence, normalization) and shown an efficient combination with PLP using Region-Dependent transforms. A combination of RDT with another popular adaptation technique (SAT) was shown beneficial. The results are reported on BABEL Cantonese data. Index Terms: speech recognition, discriminative training, bottle-neck neural networks, region-dependent transforms
Martin Karafiát, Frantisek Grézl, Mirko Hannemann, Karel Veselý, Jan Cernocký
INTERSPEECH5
2013 Frequency warping and robust speaker verification: a comparison of alternative mel-scale representations
abstract
Accuracy of speaker verification is high under controlled condi-tions but falls off rapidly in the presence of interfering sounds. This is because spectral features, such as Mel-frequency cep-stral coefficients (MFCCs), are sensitive to additive noise. MFCCs are a particular realization of warped-frequency rep-resentation with low-frequency focus. But there are several alternative, potentially more robust, warped-frequency repre-sentations. We provide an experimental comparison of five warped-frequency features. They use exactly the same fre-quency warping function, the same number of coefficients and postprocessing, but differ in their internal computations. The compared variants are (1) conventional MFCCs from discrete Fourier transform (DFT), followed by Mel-scaled filterbank, (2) MFCCs via direct warping of DFT, followed by linear-scale fil-terbank, (3) warped linear prediction features, (4) perceptual minimum variance distortionless features and (5) recently pro-posed sparse Mel-scale histogram features. Experiments car-ried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases. Index Terms: speaker recognition, noise, frequency warping 1.
Tomi Kinnunen, Jahangir Alam 0001, Pavel Matejka, Patrick Kenny, Jan Cernocký, Douglas D. O'Shaughnessy
INTERSPEECH5
2013 A region-specific feature-space transformation for speaker adaptation and singularity analysis of jacobian matrix
Shakti P. Rath, Lukás Burget, Martin Karafiát, Ondrej Glembek, Jan Cernocký
INTERSPEECH5
2013 Improved feature processing for deep neural networks
abstract
In this paper, we investigate alternative ways of processing MFCC-based features to use as the input to Deep Neural Networks (DNNs). Our baseline is a conventional feature pipeline that involves splicing the 13-dimensional front-end MFCCs across 9 frames, followed by applying LDA to reduce the dimension to 40 and then further decorrelation using MLLT. Confirming the results of other groups, we show that speaker adaptation applied on the top of these features using feature-space MLLR is helpful. The fact that the number of parameters of a DNN is not strongly sensitive to the input feature dimension (unlike GMM-based systems) motivated us to investigate ways to increase the dimension of the features. In this paper, we investigate several approaches to derive higher-dimensional features and verify their performance with DNN. Our best result is obtained from splicing our baseline 40-dimensional speaker adapted features again across 9 frames, followed by reducing the dimension to 200 or 300 using another LDA. Our final result is about 3% absolute better than our best GMM system, which is a discriminatively trained model.
Shakti P. Rath, Daniel Povey, Karel Veselý, Jan Cernocký
INTERSPEECH4
2013 Regularized subspace n-gram model for phonotactic ivector extraction
abstract
Phonotactic language identification (LID) by means of n-gram statistics and discriminative classifiers is a popular approach for the LID problem. Low-dimensional representation of the n-gram statistics leads to the use of more diverse and efficient machine learning techniques in the LID. Recently, we proposed phototactic iVector as a low-dimensional representation of the n-gram statistics. In this work, an enhanced modeling of the n-gram probabilities along with regularized parameter estimation is proposed. The proposed model consistently improves the LID system performance over all conditions up to 15% relative to the previous state of the art system. The new model also alleviates memory requirement of the iVector extraction and helps to speed up subspace training. Results are presented in terms of Cavg over NIST LRE2009 evaluation set.
Mehdi Soufifar, Lukás Burget, Oldrich Plchot, Sandro Cumani, Jan Cernocký
INTERSPEECH5
2012 Region dependent linear transforms in multilingual speech recognition
abstract
In today's speech recognition systems, linear or nonlinear transformations are usually applied to post-process speech features forming input to HMM based acoustic models. In this work, we experiment with three popular transforms: HLDA, MPE-HLDA and Region Dependent Linear Transforms (RDLT), which are trained jointly with the acoustic model to extract maximum of the discriminative information from the raw features and to represent it in a form suitable for the following GMM-HMM based acoustic model. We focus on multi-lingual environments, where limited resources are available for training recognizers of many languages. Using data from GlobalPhone database, we show that, under such restrictive conditions, the feature transformations can be advantageously shared across languages and robustly trained using data from several languages.
Martin Karafiát, Milos Janda, Jan Cernocký, Lukás Burget
ICASSP3
2012 Discriminative classifiers for phonotactic language recognition with iVectors
abstract
Phonotactic models based on bags of n-grams representations and discriminative classifiers are a popular approach to the language recognition problem. However, the large size of n-gram count vectors brings about some difficulties in discriminative classifiers. The subspace Multinomial model was recently proposed to effectively represent information contained in the n-grams using low-dimensional iVectors. The availability of a low-dimensional feature vector allows investigating different post-processing techniques and different classifiers to improve recognition performance. In this work, we analyze a set of discriminative classifiers based on Support Vector Machines and Logistic Regression and we propose an iVector post-processing technique which allows to improve recognition performance. The proposed systems are evaluated on the NIST LRE 2009 task.
Mehdi Soufifar, Sandro Cumani, Lukás Burget, Jan Cernocký
ICASSP4
2012 Phonotactic Language Recognition using i-vectors and Phoneme Posteriogram Counts
Luis Fernando D'Haro, Ondrej Glembek, Oldrich Plchot, Pavel Matejka, Mehdi Soufifar, Ricardo de Córdoba, Jan Cernocký
INTERSPEECH7
2012 A factorized representation of FMLLR transform based on QR-decomposition
Shakti P. Rath, Martin Karafiát, Ondrej Glembek, Jan Cernocký
INTERSPEECH4
2012 Comparison of methods for language-dependent and language-independent query-by-example spoken term detection
abstract
This article investigates query-by-example (QbE) spoken term detection (STD), in which the query is not entered as text, but selected in speech data or spoken. Two feature extractors based on neural networks (NN) are introduced: the first producing phone-state posteriors and the second making use of a compressive NN layer. They are combined with three different QbE detectors: while the Gaussian mixture model/hidden Markov model (GMM/HMM) and dynamic time warping (DTW) both work on continuous feature vectors, the third one, based on weighted finite-state transducers (WFST), processes phone lattices. QbE STD is compared to two standard STD systems with text queries: acoustic keyword spotting and WFST-based search of phone strings in phone lattices. The results are reported on four languages (Czech, English, Hungarian, and Levantine Arabic) using standard metrics: equal error rate (EER) and two versions of popular figure-of-merit (FOM). Language-dependent and language-independent cases are investigated; the latter being particularly interesting for scenarios lacking standard resources to train speech recognition systems. While the DTW and GMM/HMM approaches produce the best results for a language-dependent setup depending on the target language, the GMM/HMM approach performs the best dealing with a language-independent setup. As far as WFSTs are concerned, they are promising as they allow for indexing and fast search.
Javier Tejedor, Michal Fapso, Igor Szöke, Jan Cernocký, Frantisek Grézl
ACM Trans. Inf. Syst.4
2011 iVector-based discriminative adaptation for automatic speech recognition
abstract
We presented a novel technique for discriminative feature-level adaptation of automatic speech recognition system. The concept of iVectors popular in Speaker Recognition is used to extract information about speaker or acoustic environment from speech segment. iVector is a low-dimensional fixed-length representing such information. To utilized iVectors for adaptation, Region Dependent Linear Transforms (RDLT) are discriminatively trained using MPE criterion on large amount of annotated data to extract the relevant information from iVectors and to compensate speech feature. The approach was tested on standard CTS data. We found it to be complementary to common adaptation techniques. On a well tuned RDLT system with standard CMLLR adaptation we reached 0.8% additive absolute WER improvement.
Martin Karafiát, Lukás Burget, Pavel Matejka, Ondrej Glembek, Jan Cernocký
ASRU5
2011 Strategies for training large scale neural network language models
abstract
We describe how to effectively train neural network based language models on large data sets. Fast convergence during training and better overall performance is observed when the training data are sorted by their relevance. We introduce hash-based implementation of a maximum entropy model, that can be trained as a part of the neural network model. This leads to significant reduction of computational complexity. We achieved around 10% relative reduction of word error rate on English Broadcast News speech recognition task, against large 4-gram model trained on 400M tokens.
Tomás Mikolov, Anoop Deoras, Daniel Povey, Lukás Burget, Jan Cernocký
ASRU5
2011 Recent progress in prosodic speaker verification
abstract
We describe recent progress in the field of prosodic modeling for speaker verification. In a previous paper, we proposed a technique for modeling syllable-based prosodic features that uses a multinomial subspace model for feature extraction and within-class covariance normalization or linear discriminant analysis for session variability compensation. In this paper, we show that performance can be significantly improved with the use of probabilistic linear discriminant analysis (PLDA) for session variability compensation. This system does not require score normalization. We report an equal error rate below 7% on a NIST 2008 task. To our knowledge, this is the best reported result to date for a prosodic system for speaker recognition. Fusion of this system with a state-of-the-art acoustic baseline system yields 10% relative improvement in the new detection cost function (DCF) as defined by NIST.
Marcel Kockmann, Luciana Ferrer, Lukás Burget, Elizabeth Shriberg, Jan Cernocký
ICASSP5
2011 Full-covariance UBM and heavy-tailed PLDA in i-vector speaker verification
abstract
In this paper, we describe recent progress in i-vector based speaker verification. The use of universal background models (UBM) with full-covariance matrices is suggested and thoroughly experimentally tested. The i-vectors are scored using a simple cosine distance and advanced techniques such as Probabilistic Linear Discriminant Analysis (PLDA) and heavy-tailed variant of PLDA (PLDA-HT). Finally, we investigate into dimensionality reduction of i-vectors before entering the PLDA-HT modeling. The results are very competitive: on NIST 2010 SRE task, the results of a single full-covariance LDA-PLDA-HT system approach those of complex fused system.
Pavel Matejka, Ondrej Glembek, Fabio Castaldo, Jahangir Alam 0001, Oldrich Plchot, Patrick Kenny, Lukás Burget, Jan Cernocký
ICASSP8
2011 Extensions of recurrent neural network language model
abstract
We present several modifications of the original recurrent neural network language model (RNN LM).While this model has been shown to significantly outperform many competitive language modeling techniques in terms of accuracy, the remaining problem is the computational complexity. In this work, we show approaches that lead to more than 15 times speedup for both training and testing phases. Next, we show importance of using a backpropagation through time algorithm. An empirical comparison with feedforward networks is also provided. In the end, we discuss possibilities how to reduce the amount of parameters in the model. The resulting RNN model can thus be smaller, faster both during training and testing, and more accurate than the basic one.
Tomás Mikolov, Stefan Kombrink, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur
ICASSP4
2011 General chair's message
abstract
The organizing committee of ICASSP 2011 is delighted to welcome you to the 36th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), which is being held at the Prague Congress Centre, May 22–27, 2011. This is the flagship conference for the IEEE Signal Processing Society. In 1997, ICASSP was held in Munich, Germany, and now, in 2011, ICASSP is back in Central Europe. During the intervening years, the event has been held across the continents and next year will return to Asia, when Kyoto in Japan will be the venue. The ICASSP conference is the world's largest and most comprehensive technical event, focused on signal processing and its applications. The conference includes overview lectures, detailed scientific sessions and exhibitions, together with social and cultural events.
Petr Tichavský, Jan Cernocký, Ales Procházka
ICASSP2
2011 iVector Fusion of Prosodic and Cepstral Features for Speaker Verification
abstract
In this paper we apply the promising iVector extraction technique followed by PLDA modeling to simple prosodic contour features. With this procedure we achieve results comparable to a system that models much more complex prosodic features using our recently proposed SMM-based iVector modeling technique. We then propose a combination of both prosodic iVectors by joint PLDA modeling that leads to significant improvements over individual systems with an EER of 5.4% on NIST SRE 2008 telephone data. Finally, we can combine these two prosodic iVector front ends with a baseline cepstral iVector system to achieve up to 21% relative reduction in new DCF. Index Terms: speaker verification, prosody, JFA, iVector, SMM, fusion
Marcel Kockmann, Luciana Ferrer, Lukás Burget, Jan Cernocký
INTERSPEECH4
2011 Empirical Evaluation and Combination of Advanced Language Modeling Techniques
abstract
We present results obtained with several advanced language modeling techniques, including class based model, cache model, maximum entropy model, structured language model, random forest language model and several types of neural network based language models. We show results obtained after combining all these models by using linear interpolation. We conclude that for both small and moderately sized tasks, we obtain new state of the art results with combination of models, that is significantly better than performance of any individual model. Obtained perplexity reductions against Good-Turing trigram baseline are over 50% and against modified Kneser-Ney smoothed 5-gram over 40%. Index Terms: language modeling, neural networks, model combination, speech recognition
Tomás Mikolov, Anoop Deoras, Stefan Kombrink, Lukás Burget, Jan Cernocký
INTERSPEECH5
2011 Application of speaker- and language identification state-of-the-art techniques for emotion recognition
Marcel Kockmann, Lukás Burget, Jan Cernocký
Speech Commun.3
2010 Investigations into prosodic syllable contour features for speaker recognition
abstract
We investigate various ways of generating prosodic syllable contour features that have recently been applied to enhance systems for speaker recognition. We compare different approaches for segmentation of speech into syllable-like units, techniques for contour modeling and the extraction of pitch and energy, taking into account the computational complexity and gender dependence. We show that the performance is especially affected by the segmentation and the quality of the pitch tracking algorithm and that the features are highly gender dependent. Still, computationally simple ways of segmentation of speech can be used to achieve good results, as experiments on 2006 NIST speaker recognition evaluation task indicate.
Marcel Kockmann, Lukás Burget, Jan Cernocký
ICASSP3
2010 Tuning phone decoders for language identification
abstract
Phonotactic approach, phone recognition to be followed by language modeling, is one of the most popular approaches to language identification (LID). In this work, we explore how language identification accuracy of a phone decoder can be enhanced by varying acoustic resolution of the phone decoder, and subsequently how multiresolution versions of the same decoder can be integrated to improve the LID accuracy. We use mutual information to select the optimum set of phones for a specific acoustic resolution. Further, we propose strategies for building multilingual systems suitable for LID applications, and subsequently fine tune these systems to enhance the overall accuracy.
C. Santhosh Kumar, Haizhou Li 0001, Rong Tong, Pavel Matejka, Lukás Burget, Jan Cernocký
ICASSP6
2010 Brno university of technology system for interspeech 2010 paralinguistic challenge
abstract
This paper describes Brno University of Technology (BUT) system for the Interspeech 2010 Paralinguistic Challenge. Our submitted systems for the Ageand Gender-Sub-Challenges employ fusions of several sub-systems. We make use of our own acoustic frame-based feature sets, as well as the provided utterance-based acoustic, prosodic and voice quality features. Modeling is based on Gaussian Mixture Models (GMM) and Support Vector Machines (SVM), followed by linear Gaussian backends and logistic regression-based fusion. For a single subsystem, we obtain improvement of about 2% absolute, for both tasks, on the development-set. Our final fusion results in nearly 9% absolute improvement for the Age task and about 4.5% for the Gender task on the development set. On the final test set we obtain 3.5% and 2% absolute improvement, respectively.
Marcel Kockmann, Lukás Burget, Jan Cernocký
INTERSPEECH3
2010 Prosodic speaker verification using subspace multinomial models with intersession compensation
abstract
We propose a novel approach to modeling prosodic features. Inspired by Joint Factor Analysis model (JFA), our model is based on the same idea of introducing subspace of model parameters. However, the underlying Gaussian Mixture distribution of JFA is replaced by multinomial distribution to model sequences of discrete units rather than continuous features. In this work, we use the subspace model as a feature extractor for support vector machines (SVMs), similar to the recently proposed JFA in total variability space. We can show the capability to reduce high-dimensional count vectors to low dimension while keeping system performance stable. With additional intersession compensation, we can improve 30 % relative to the baseline system and reach an equal error rate of 8.8 % on the NIST 2006 SRE dataset. Index Terms: speaker verification, prosody, JFA, multinomial model
Marcel Kockmann, Lukás Burget, Ondrej Glembek, Luciana Ferrer, Jan Cernocký
INTERSPEECH5
2010 Recurrent neural network based language model
abstract
A new recurrent neural network based language model (RNN LM) with applications to speech recognition is presented. Results indicate that it is possible to obtain around 50% reduction of perplexity by using mixture of several RNN LMs, compared to a state of the art backoff language model. Speech recognition experiments show around 18% reduction of word error rate on the Wall Street Journal task when comparing models trained on the same amount of data, and around 5% on the much harder NIST RT05 task, even when the backoff model is trained on much more data than the RNN LM. We provide ample empirical evidence to suggest that connectionist language models are superior to standard n-gram techniques, except their high computational (training) complexity. Index Terms: language modeling, recurrent neural networks, speech recognition
Tomás Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur
INTERSPEECH4
2010 Speech@FIT lecture browser
abstract
This paper describes an innovative web-based browser used for video recordings of lectures that is built on speech and image processing technologies. The aim of this project is to simplify the access to information that is spread across video recordings. This is mainly achieved by coupling to the speech search engine and due to a possibility to quickly navigate through an automatically generated list of slides presented. The reader is briefly acquainted with the technological background of the browser; the emphasis is laid on the use of the browser from the user point of view.
Igor Szöke, Jan Cernocký, Michal Fapso, Josef Zizka
SLT2
2010 Acoustic keyword spotter - optimization from end-user perspective
abstract
The paper deals with the development of acoustic keyword spotter (KWS) meeting requirements of a real user from the security community. While the basic scheme of the KWS is relatively standard, it uses novel features derived by a hierarchy of neural networks, and score normalization trained to maximize a user-like evaluation metric. The results are reported on a selection of Czech conversational telephone speech (CTS), radio and read data.
Igor Szöke, Frantisek Grézl, Jan Cernocký, Michal Fapso, Tomás Cipr
SLT3
2009 Neural network based language models for highly inflective languages
abstract
Speech recognition of inflectional and morphologically rich languages like Czech is currently quite a challenging task, because simple n-gram techniques are unable to capture important regularities in the data. Several possible solutions were proposed, namely class based models, factored models, decision trees and neural networks. This paper describes improvements obtained in recognition of spoken Czech lectures using language models based on neural networks. Relative reductions in word error rate are more than 15% over baseline obtained with adapted 4-gram backoff language model using modified Kneser-Ney smoothing.
Tomás Mikolov, Jirí Kopecký, Lukás Burget, Ondrej Glembek, Jan Cernocký
ICASSP5
2009 BUT system for NIST 2008 speaker recognition evaluation
abstract
This paper presents BUT system submitted to NIST 2008 SRE. It includes two subsystems based on Joint Factor Analysis (JFA) GMM/UBM and one based on SVM-GMM. The systems were developed on NIST SRE2006 data, and the results arepresented on NIST SRE 2008 evaluation data. We concentrate on the influence of side information in the calibration. Index Terms: speaker recognition, joint factor analysis, NIST SRE 2008.
Lukás Burget, Michal Fapso, Valiantsina Hubeika, Ondrej Glembek, Martin Karafiát, Marcel Kockmann, Pavel Matejka, Petr Schwarz, Jan Cernocký
INTERSPEECH9
2009 Investigation into variants of joint factor analysis for speaker recognition
abstract
In this paper, we have investigated into JFA used for speaker recognition. First, we performed systematic comparison of full JFA with its simplified variants and confirmed superior performance of the full JFA with both eigenchannels and eigenvoices. We investigated into sensitivity of JFA on the number of eigenvoices both for the full one and simplified variants. We studied the importance of normalization and found that genderdependent zt-norm was crucial. The results are reported on NIST 2006 and 2008 SRE evaluation data. Index Terms: speaker recognition, joint factor analysis.
Lukás Burget, Pavel Matejka, Valiantsina Hubeika, Jan Cernocký
INTERSPEECH4
2009 Brno University of Technology system for Interspeech 2009 emotion challenge
abstract
This paper describes Brno University of Technology (BUT) system for the Interspeech 2009 Emotion Challenge. Our submitted system for the Open Performance Sub-Challenge uses acoustic frame based features as a front-end and Gaussian Mixture Models as a back-end. Different feature types and modeling approaches successfully applied in speakerand language recognition are investigated and we can achieve an 16% and 9% relative improvement over the best dynamic and static baseline system on the 5-class task, respectively.
Marcel Kockmann, Lukás Burget, Jan Cernocký
INTERSPEECH3
2008 Combination of strongly and weakly constrained recognizers for reliable detection of OOVS
abstract
This paper addresses the detection of OOV segments in the output of a large vocabulary continuous speech recognition (LVCSR) system. First, standard confidence measures from frame-based wordand phone- posteriors are investigated. Substantial improvement is obtained when posteriors from two systems — strongly constrained (LVCSR) and weakly constrained (phone posterior estimator) are combined. We show that this approach is also suitable for detection of general recognition errors. All results are presented on WSJ task with reduced recognition vocabulary.
Lukás Burget, Petr Schwarz, Pavel Matejka, Mirko Hannemann, Ariya Rastrow, Christopher M. White, Sanjeev Khudanpur, Hynek Hermansky, Jan Cernocký
ICASSP9
2008 Discrimininative training of narrow band - wide band adapted systems for meeting recognition
abstract
The amount of training data has a crucial effect on the accuracy of HMM based meeting recognition systems. One of the largest collections of speech data is conversational telephone speech which was found to match speech in meetings well. However it is naturally recorded with limited bandwidth. In previous work we presented a scheme that allows to transform wide-band meeting data into the same space for improved model training. In this paper we focused on integration of discriminative adaptation into this scheme. This integration is not straightforward and we present the complexity of this process. The models are tested on the NIST RT’05 meeting evaluation where a relative reduction in word error rate of 5.6% against non-adapted meeting system was achieved.
Martin Karafiát, Lukás Burget, Thomas Hain, Jan Cernocký
INTERSPEECH4
2008 BUT language recognition system for NIST 2007 evaluations
abstract
This paper describes Brno University of Technology (BUT) system for 2007 NIST Language recognition (LRE) evaluation. The system is a fusion of 4 acoustic and 9 phonotactic subsystems. We have investigated several new topics such as discriminatively trained language models in phonotactic systems, and eigen-channel adaptation in model and feature domain in acoustic systems. We also point out the importance of calibration and fusion. All results are presented on NIST 2007 LRE data.
Pavel Matejka, Lukás Burget, Ondrej Glembek, Petr Schwarz, Valiantsina Hubeika, Michal Fapso, Tomás Mikolov, Oldrich Plchot, Jan Cernocký
INTERSPEECH9
2008 Morphological random forests for language modeling of inflectional languages
abstract
In this paper, we are concerned with using decision trees (DT) and random forests (RF) in language modeling for Czech LVCSR. We show that the RF approach can be successfully implemented for language modeling of an inflectional language. Performance of word-based and morphological DTs and RFs was evaluated on lecture recognition task. We show that while DTs perform worse than conventional trigram language models (LM), RFs of both kind outperform the latter. WER (up to 3.4% relative) and perplexity (10%) reduction over the trigram model can be gained with morphological RFs. Further improvement is obtained after interpolation of DT and RF LMs with the trigram one (up to 15.6% perplexity and 4.8% WER relative reduction). In this paper we also investigate distribution of morphological feature types chosen for splitting data at different levels of DTs.
Ilya Oparin, Ondrej Glembek, Lukás Burget, Jan Cernocký
SLT4
2008 Sub-word modeling of out of vocabulary words in spoken term detection
abstract
This paper deals with comparison of sub-word based methods for spoken term detection (STD) task and phone recognition. The sub-word units are needed for search for out-of-vocabulary words. We compared words, phones and multigrams. The maximal length and pruning of multigrams were investigated first. Then two constrained methods of multigram training were proposed. We evaluated on the NIST STD06 dev-set CTS data. The conclusion is that the proposed method improves the phone accuracy more than 9% relative and STD accuracy more than 7% relative.
Igor Szöke, Lukás Burget, Jan Cernocký, Michal Fapso
SLT3
2007 Probabilistic and Bottle-Neck Features for LVCSR of Meetings
abstract
In recent years, probabilistic features became an integral part of state-of-the-are LVCSR systems. In this work, we are exploring the possibility of obtaining the features directly from neural net without the necessity of converting output probabilities to features suitable for subsequent GMM-HMM system. We experimented with 5-layer MLP with bottle-neck in the middle layer. After training such a neural net, we used outputs of the bottle-neck as features for GMM-HMM recognition system. The benefits are twofold: first, improvement was gained when these features are used instead of the probabilistic features, second, the size of the system was reduced, as only part of the neural net is used. The experiments were performed on meetings recognition task defined in MST RT'05 evaluation.
Frantisek Grézl, Martin Karafiát, Stanislav Kontar, Jan Cernocký
ICASSP (4)4
2007 STBU System for the NIST 2006 Speaker Recognition Evaluation
abstract
This paper describes STBU 2006 speaker recognition system, which performed well in the NIST 2006 speaker recognition evaluation. STBU is consortium of 4 partners: Spescom DataVoice (South Africa), TNO (Netherlands), BUT (Czech Republic) and University of Stellenbosch (South Africa). The primary system is a combination of three main kinds of systems: (1) GMM, with short-time MFCC or PLP features, (2) GMM-SVM, using GMM mean supervectors as input and (3) MLLR-SVM, using MLLR speaker adaptation coefficients derived from English LVCSR system. In this paper, we describe these sub-systems and present results for each system alone and in combination on the NIST Speaker Recognition Evaluation (SRE) 2006 development and evaluation data sets.
Pavel Matejka, Lukás Burget, Petr Schwarz, Ondrej Glembek, Martin Karafiát, Frantisek Grézl, Jan Cernocký, David A. van Leeuwen, Niko Brümmer, Albert Strasheim
ICASSP (4)7
2007 Application of CMLLR in narrow band wide band adapted systems
Martin Karafiát, Lukás Burget, Jan Cernocký, Thomas Hain
INTERSPEECH3
2007 Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006
abstract
This paper describes and discusses the "STBU" speaker recognition system, which performed well in the NIST Speaker Recognition Evaluation 2006 (SRE). STBU is a consortium of four partners: Spescom DataVoice (Stellenbosch, South Africa), TNO (Soesterberg, The Netherlands), BUT (Brno, Czech Republic), and the University of Stellenbosch (Stellenbosch, South Africa). The STBU system was a combination of three main kinds of subsystems: 1) GMM, with short-time Mel frequency cepstral coefficient (MFCC) or perceptual linear prediction (PLP) features, 2) Gaussian mixture model-support vector machine (GMM-SVM), using GMM mean supervectors as input to an SVM, and 3) maximum-likelihood linear regression-support vector machine (MLLR-SVM), using MLLR speaker adaptation coefficients derived from an English large vocabulary continuous speech recognition (LVCSR) system. All subsystems made use of supervector subspace channel compensation methods-either eigenchannel adaptation or nuisance attribute projection. We document the design and performance of all subsystems, as well as their fusion and calibration via logistic regression. Finally, we also present a cross-site fusion that was done with several additional systems from other NIST SRE-2006 participants.
Niko Brümmer, Lukás Burget, Jan Cernocký, Ondrej Glembek, Frantisek Grézl, Martin Karafiát, David A. van Leeuwen, Pavel Matejka, Petr Schwarz, Albert Strasheim
IEEE Trans. Speech Audio Process.3
2007 Analysis of Feature Extraction and Channel Compensation in a GMM Speaker Recognition System
abstract
In this paper, several feature extraction and channel compensation techniques found in state-of-the-art speaker verification systems are analyzed and discussed. For the NIST SRE 2006 submission, cepstral mean subtraction, feature warping, RelAtive SpecTrAl (RASTA) filtering, heteroscedastic linear discriminant analysis (HLDA), feature mapping, and eigenchannel adaptation were incrementally added to minimize the system's error rate. This paper deals with eigenchannel adaptation in more detail and includes its theoretical background and implementation issues. The key part of the paper is, however, the post-evaluation analysis, undermining a common myth that “the more boxes in the scheme, the better the system.” All results are presented on NIST Speaker Recognition Evaluation (SRE) 2005 and 2006 data.
Lukás Burget, Pavel Matejka, Petr Schwarz, Ondrej Glembek, Jan Cernocký
IEEE Trans. Speech Audio Process.5
2006 Information Retrieval from Spoken Documents
Michal Fapso, Pavel Smrz, Petr Schwarz, Igor Szöke, Milan Schwarz, Jan Cernocký, Martin Karafiát, Lukás Burget
CICLing6
2006 Discriminative Training Techniques for Acoustic Language Identification
abstract
This paper presents comparison of Maximum Likelihood (ML) and discriminative Maximum Mutual Information (MMI) training for acoustic modeling in language identification (LID). Both approaches are compared on state-of- the-art shifted delta-cepstra features, the results are reported on data from NIST 2003 evaluations. Clear advantage of MMI over ML training is shown. Further improvements of acoustic LID are discussed: Heteroscedastic Linear Discriminant Analysis (HLDA) for feature de-correlation and dimensionality reduction and Ergodic Hidden Markov models (EHMM) for better modeling of dynamics in the acoustic space. The final error rate compares favorably to other results published on NIST 2003 data.
Lukás Burget, Pavel Matejka, Jan Cernocký
ICASSP (1)3
2006 Use of Anti-Models to Further Improve State-of-the-Art PRLM Language Recognition System
abstract
This paper concentrates on PRLM (phoneme recognizer followed by language model) approach to language recognition. It elaborates on our prior work concerning the quality of phoneme recognition and amounts of training data for phoneme recognizer training. It reports improvements brought to our PRLM system by better phoneme recognition and Witten-Bell discounting in LM-modeling. The paper then concentrates on the use of phoneme lattices and anti-models. Training and scoring on phoneme lattices brought significant improvement in language recognition accuracy. The antimodels are simple, yet powerful technique to improve the discrimination between target and non-target languages. All results are reported on standard NIST 2003 data; comparison with other published results is favorable to our system.
Pavel Matejka, Petr Schwarz, Lukás Burget, Jan Cernocký
ICASSP (1)4
2006 Hierarchical Structures of Neural Networks for Phoneme Recognition
abstract
This paper deals with phoneme recognition based on neural networks (NN). First, several approaches to improve the phoneme error rate are suggested and discussed. In the experimental part, we concentrate on TempoRAl Patterns (TRAPs) and novel split temporal context (STC) phoneme recognizers. We also investigate into tandem NN architectures. The results of the final system reported on standard TIMIT database compare favorably to the best published results.
Petr Schwarz, Pavel Matejka, Jan Cernocký
ICASSP (1)3
2005 Phonotactic language identification using high quality phoneme recognition
abstract
Phoneme Recognizers followed by Language Modeling (PRLM) have consistently yielded top performance in language identification (LID) task. Parallel ordering of PRLMs (PPRLM) improves performance even more. Since tokenizer is the most important part of LID system the high quality phoneme recognizer is employed. Two different multilingual databases for training phoneme recognizers are compared and the amount of sufficient training data is studied. Reported results are on data from NIST 2003 LID evaluation. Our four PRLM systems have Equal Error Rate (EER) of 2.4 % on 12 languages task. This result compares favorably to the best known result from this task. 1.
Pavel Matejka, Petr Schwarz, Jan Cernocký, Pavel Chytil
INTERSPEECH3
2005 Non-parametric speaker turn segmentation of meeting data
abstract
An extension of conventional speaker segmentation framework is presented for a scenario in which a number of microphones record the activity of speakers present at a meeting (one microphone per speaker). Although each microphone can receive speech from both the participant wearing the microphone (local speech) and other participants (cross-talk), the recorded audio can be broadly classified in three ways: local speech, cross-talk, and silence. This paper proposes a technique which takes into account cross-correlations, values of its maxima, and energy differences as features to identify and segment speaker turns. In particular, we have used classical cross-correlation functions, time smoothing and in part temporal constraints to sharpen and disambiguate timing differences between microphone channels that may be dominated by noise and reverberation. Experimental results show that proposed technique can be successively used for speaker segmentation of data collected from a number of different setups. 1.
Petr Motlícek, Lukás Burget, Jan Cernocký
INTERSPEECH3
2005 Comparison of keyword spotting approaches for informal continuous speech
abstract
This paper describes several approaches to keyword spotting (KWS) for informal continuous speech. We compare acoustic keyword spotting, spotting in word lattices generated by large vocabulary continuous speech recognition and a hybrid approach making use of phoneme lattices generated by a phoneme recognizer. The systems are compared on carefully defined test data extracted from ICSI meeting database. The acoustic and phoneme-lattice based KWS are based on a phoneme recognizer making use of temporal-pattern (TRAP) feature extraction and posterior estimation using neural nets. We show its superiority over traditional HMM/GMM systems. The advantages and drawbacks of different approaches are discussed. 1.
Igor Szöke, Petr Schwarz, Pavel Matejka, Lukás Burget, Martin Karafiát, Michal Fapso, Jan Cernocký
INTERSPEECH7
2004 TRAP based features for LVCSR of meting data
Frantisek Grézl, Martin Karafiát, Jan Cernocký
INTERSPEECH3
2004 Orthographic and Phonetic Annotation of Very Large Czech Corpora with Quality Assessment
Petr Pollák, Jan Cernocký
LREC2
2003 Time-domain based temporal processing with application of orthogonal transformations
Petr Motlícek, Jan Cernocký
INTERSPEECH2
2003 Autoregressive modeling based feature extraction for Aurora3 DSR task
Petr Motlícek, Jan Cernocký
INTERSPEECH2
2003 Recognition of phoneme strings using TRAP technique
abstract
We investigate and compare several techniques for automatic recognition of unconstrained context-independent phoneme strings from TIMIT and NTIMIT databases. Among the com-pared techniques, the technique based on TempoRAl Patterns (TRAP) achieves the best results in the clean speech, it achieves about 10 % relative improovements against baseline system. Its advantage is also observed in the presence of mismatch be-tween training and testing conditions. Issues such as the op-timal length of temporal patterns in the TRAP technique and the effectiveness of mean and variance normalization of the pat-terns and the multi-band input the TRAP estimations, are also explored. 1.
Petr Schwarz, Pavel Matejka, Jan Cernocký
INTERSPEECH3
2001 Speechdat-e: five eastern european speech databases for voice-operated teleservices completed
abstract
\n Contains fulltext :\n 76438.pdf (author's version ) (Open Access)\n
Henk van den Heuvel, Jérôme Boudy, Zsolt Bakcsi, Jan Cernocký, Valery Galunov, Julia Kochanina, Wojciech Majewski, Petr Pollák, Milan Rusko, Jerzy Sadowski, Piotr Staroniewicz, Herbert S. Tropf
INTERSPEECH4
1999 A segmental approach to text-independent speaker verification
abstract
In this paper a new automatic speech recognition (ASR) CPU-based software, called AlfaNum, with the chosen few heuristics optimized for applications in heterogeneous conditions is described. AlfaNum is a discrete speaker-independent ASR product intended for application in the largest bank-by-phone interactive voice response (IVR) system in Yugoslavia, with a lot of customers all over Serbia. That means a large variety of dialects, telephone line quality, and microphones used. This system has been tested on 500 speakers and it achieved an average accuracy of 98,2% in real life conditions. The whole software is developed in C++ programming language. Object oriented programming gave the software an elegant look, and minimized all possible errors. On the other hand, the power of C++ language and its tight interaction with machine made the software fast and efficient.
Jan Cernocký, Dijana Petrovska-Delacrétaz, Stéphane Pigeon, Patrick Verlinde, Gérard Chollet
EUROSPEECH1
1998 Segmental vocoder-going beyond the phonetic approach
abstract
The problem of very low bit rate segmental speech coding is addressed. The basic units are found automatically in the training database using temporal decomposition, vector quantization and multigrams. They are modelled by HMMs. The coding is based on recognition and synthesis. In single speaker tests, we obtained intelligible and naturally sounding speech at a mean rate of 211.2 b/s. In the end, future extensions of our scheme (diphone-like synthesis and speaker adaptation) as well as possible use of automatically derived units in recognition are discussed.
Jan Cernocký, Geneviève Baudoin, Gérard Chollet
ICASSP1
1998 Text-independent speaker verification using automatically labelled acoustic segments
abstract
Keywords: Non-Linear Signal Processing ; Speech Processing Reference LANOS-CONF-1998-024 Record created on 2004-12-03, modified on 2017-05-12
Dijana Petrovska-Delacrétaz, Jan Cernocký, Jean Hennebert, Gérard Chollet
ICSLP2
1997 Speech spectrum representation and coding using multigrams with distance
abstract
The multigrams allow us to split a string of symbols into a stream of variable length sequences. The direct application of this method to vector-quantized speech spectra fails, we develop an extension of the method called modified multigrams or multigrams with distance. The algorithm for modified multigram dictionary training as well as experimental results are presented. We found a significant improvement of rate/distortion ratio in comparison to vector quantization with small codebooks. For precise spectrum representation, this method is less suitable and we see its application rather in speech segmentation or in very low bit rate coding.
Jan Cernocký, Geneviève Baudoin, Gérard Chollet
ICASSP1
1997 Quantization of spectral sequences using variable length spectral segments for speech coding at very low bit rate
Geneviève Baudoin, Jan Cernocký, Gérard Chollet
EUROSPEECH2