Jan Svec

dblp:21/3262 · DBLP profile ↗
← Back
34ranked-venue papers
12as first author
15since 2021 · last 2025
0000-0001-8362-5927ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 11 first-author · 12 since 2021Artificial intelligence and machine learning · 22 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Semantic Search and Filtering with AI Agents
Martin Bulín, Jan Svec, Filip Polák, Lubos Smídl
ECIR (5)2
2025 Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
Pradyoth Hegde, Santosh Kesiraju, Jan Svec, Simon Sedlácek, Bolaji Yusuf, Oldrich Plchot, Deepak K. T, Jan Cernocký
INTERSPEECH3
2025 Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Simon Sedlácek, Bolaji Yusuf, Jan Svec, Pradyoth Hegde, Santosh Kesiraju, Oldrich Plchot, Jan Cernocký
INTERSPEECH3
2024 Asking Questions Framework for Oral History Archives
Jan Svec, Martin Bulín, Adam Frémund, Filip Polák
ECIR (3)1
2024 Zero-shot Out-of-domain is No Joke: Lessons Learned in the VoiceMOS 2023 MOS Prediction Challenge
Marie Kunesová, Jan Lehecka, Josef Michalek, Jindrich Matousek, Jan Svec
INTERSPEECH5
2023 The System for Efficient Indexing and Search in the Large Archives of Scanned Historical Documents
Martin Bulín, Jan Svec, Pavel Ircing
ECIR (3)2
2023 Ensemble of Deep Neural Network Models for MOS Prediction
abstract
Automatic evaluation of the quality of synthetic speech has the potential to serve as a cheaper and less time-consuming alternative to standard listening tests. In this paper, we present our contribution to the ongoing research: a system for automatic prediction of the mean opinion score (MOS) given by human listeners. The system was specifically developed for the recent VoiceMOS Challenge. Following the success of fusion systems in similar challenges, our contribution is an ensemble that interpolates the outputs of seven different models: four different wav2vec models, a CNN-RNN model, QuartzNet, and the LDNet baseline. During the VoiceMOS challenge, our system achieved the second-best utterance-level MSE of 0.171 and ranged from 2nd to 8th place among all 22 participating teams in terms of other evaluation metrics.
Marie Kunesová, Jindrich Matousek, Jan Lehecka, Jan Svec, Josef Michalek, Daniel Tihelka, Martin Bulín, Zdenek Hanzlícek, Markéta Rezácková
ICASSP4
2023 Transformer-based Speech Recognition Models for Oral History Archives in English, German, and Czech
Jan Lehecka, Jan Svec, Josef V. Psutka, Pavel Ircing
INTERSPEECH2
2023 Asking Questions: an Innovative Way to Interact with Oral History Archives
Jan Svec, Martin Bulín, Adam Frémund, Filip Polák
INTERSPEECH1
2022 Revisiting joint decoding based multi-talker speech recognition with DNN acoustic model
abstract
In typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker decoder which is applied on each speaker-specific output stream separately. In this work, we argue that such a scheme is sub-optimal and propose a principled solution that decodes all speakers jointly. We modify the acoustic model to predict joint state posteriors for all speakers, enabling the network to express uncertainty about the attribution of parts of the speech signal to the speakers. We employ a joint decoder that can make use of this uncertainty together with higher-level language information. For this, we revisit decoding algorithms used in factorial generative models in early multi-talker speech recognition systems. In contrast with these early works, we replace the GMM acoustic model with DNN, which provides greater modeling power and simplifies part of the inference. We demonstrate the advantage of joint decoding in proof of concept experiments on a mixed-TIDIGITS dataset.
Martin Kocour, Katerina Zmolíková, Lucas Ondel Yang, Jan Svec, Marc Delcroix, Tsubasa Ochiai, Lukás Burget, Jan Cernocký
INTERSPEECH4
2022 Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of Czech
abstract
In this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automatic speech recognition tasks using a combination of in-domain data and almost 6 thousand hours of out-of-domain transcribed speech. We are presenting a large palette of experiments with various fine-tuning setups evaluated on two public datasets (CommonVoice and VoxPopuli) and one extremely challenging dataset from the MALACH project. Our results show that monolingual Wav2Vec 2.0 models are robust ASR systems, which can take advantage of large labeled and unlabeled datasets and successfully compete with state-of-the-art LVCSR systems. Moreover, Wav2Vec models proved to be good zero-shot learners when no training data are available for the target ASR task.
Jan Lehecka, Jan Svec, Ales Prazák, Josef Psutka
INTERSPEECH2
2022 Deep LSTM Spoken Term Detection using Wav2Vec 2.0 Recognizer
abstract
In recent years, the standard hybrid DNN-HMM speech recognizers are outperformed by the end-to-end speech recognition systems. One of the very promising approaches is the grapheme Wav2Vec 2.0 model, which uses the self-supervised pretraining approach combined with transfer learning of the fine-tuned speech recognizer. Since it lacks the pronunciation vocabulary and language model, the approach is suitable for tasks where obtaining such models is not easy or almost impossible. In this paper, we use the Wav2Vec speech recognizer in the task of spoken term detection over a large set of spoken documents. The method employs a deep LSTM network which maps the recognized hypothesis and the searched term into a shared pronunciation embedding space in which the term occurrences and the assigned scores are easily computed. The paper describes a bootstrapping approach that allows the transfer of the knowledge contained in traditional pronunciation vocabulary of DNN-HMM hybrid ASR into the context of grapheme-based Wav2Vec. The proposed method outperforms the previously published system based on the combination of the DNN-HMM hybrid ASR and phoneme recognizer by a large margin on the MALACH data in both English and Czech languages.
Jan Svec, Jan Lehecka, Lubos Smídl
INTERSPEECH1
2021 Live TV Subtitling Through Respeaking
Ales Prazák, Zdenek Loose, Josef V. Psutka, Vlasta Radová, Josef Psutka, Jan Svec
Interspeech6
2021 T5G2P: Using Text-to-Text Transfer Transformer for Grapheme-to-Phoneme Conversion
abstract
Despite the increasing popularity of end-to-end text-to-speech (TTS) systems, the correct grapheme-to-phoneme (G2P) module is still a crucial part of those relying on a phonetic input. In this paper, we, therefore, introduce a T5G2P model, a Text-to-Text Transfer Transformer (T5) neural network model which is able to convert an input text sentence into a phoneme sequence with a high accuracy. The evaluation of our trained T5 model is carried out on English and Czech, since there are different specific properties of G2P, including homograph disambiguation, cross-word assimilation and irregular pronunciation of loanwords. The paper also contains an analysis of a homographs issue in English and offers another approach to Czech phonetic transcription using the detection of pronunciation exceptions.
Markéta Rezácková, Jan Svec, Daniel Tihelka
Interspeech2
2021 Spoken Term Detection and Relevance Score Estimation Using Dot-Product of Pronunciation Embeddings
abstract
The paper describes a novel approach to Spoken Term Detection (STD) in large spoken archives using deep LSTM networks. The work is based on the previous approach of using Siamese neural networks for STD and naturally extends it to directly localize a spoken term and estimate its relevance score. The phoneme confusion network generated by a phoneme recognizer is processed by the deep LSTM network which projects each segment of the confusion network into an embedding space. The searched term is projected into the same embedding space using another deep LSTM network. The relevance score is then computed using a simple dot-product in the embedding space and calibrated using a sigmoid function to predict the probability of occurrence. The location of the searched term is then estimated from the sequence of output probabilities. The deep LSTM networks are trained in a self-supervised manner from paired recognition hypotheses on word and phoneme levels. The method is experimentally evaluated on MALACH data in English and Czech languages.
Jan Svec, Lubos Smídl, Josef V. Psutka, Ales Prazák
Interspeech1
2019 Multimodal Dialog with the MALACH Audiovisual Archive
Adam Chýlek, Lubos Smídl, Jan Svec
INTERSPEECH3
2018 On the Use of Grapheme Models for Searching in Large Spoken Archives
abstract
This paper explores the possibility to use grapheme-based word and sub-word models in the task of spoken term detection (STD). The usage of grapheme models eliminates the need for expert-prepared pronunciation lexicons (which are often far from complete) and/or trainable grapheme-to-phoneme (G2P) algorithms that are frequently rather inaccurate, especially for rare words (words coming from a different language). Moreover, the G2P conversion of the search terms that need to be performed on-line can substantially increase the response time of the STD system. Our results show that using various grapheme-based models, we can achieve STD performance (measured in terms of ATWV) comparable with phoneme-based models but without the additional burden of G2P conversion.
Jan Svec, Josef V. Psutka, Jan Trmal, Lubas Smfdl, Pavel Ircing, Jan Sedmidubský
ICASSP1
2018 Design and Development of Speech Corpora for Air Traffic Control Training
Lubos Smídl, Jan Svec, Daniel Tihelka, Jindrich Matousek, Jan Romportl, Pavel Ircing
LREC2
2018 Towards Processing of the Oral History Interviews and Related Printed Documents
Zbynek Zajíc, Lucie Skorkovská, Petr Neduchal, Pavel Ircing, Josef V. Psutka, Marek Hrúz, Ales Prazák, Daniel Soutner, Jan Svec, Lukás Bures, Ludek Müller
LREC9
2017 Fast Subsequence Matching in Motion Capture Data
Jan Sedmidubský, Pavel Zezula, Jan Svec
ADBIS3
2017 A Relevance Score Estimation for Spoken Term Detection Based on RNN-Generated Pronunciation Embeddings
Jan Svec, Josef V. Psutka, Lubos Smídl, Jan Trmal
INTERSPEECH1
2016 A study of different weighting schemes for spoken language understanding based on convolutional neural networks
abstract
This paper describes the development of a stateless spoken spoken language understanding (SLU) module based on artificial neural networks that is able to deal with the uncertainty of the automatic speech recognition (ASR) output. The work builds upon the concept of weighted neurons introduced by the authors previously and presents a generalized weighting term for such a neuron. The effect of different forms and parameter estimation methods of the weighting term is experimentally evaluated on the multi-task training corpus, created by merging two different semantically annotated corpora. The robustness of the best performing weighting schemes is then demonstrated by experiments involving hybrid word-semantic (WSE) lattices and also limited data scenario.
Jan Svec, Adam Chýlek, Lubos Smídl, Pavel Ircing
ICASSP1
2016 A Multimodal Dialogue System for Air Traffic Control Trainees Based on Discrete-Event Simulation
Lubos Smídl, Adam Chýlek, Jan Svec
INTERSPEECH3
2016 An Engine for Online Video Search in Large Archives of the Holocaust Testimonies
Petr Stanislav, Jan Svec, Pavel Ircing
INTERSPEECH2
2016 An Automatic Training Tool for Air Traffic Control Training
Petr Stanislav, Lubos Smídl, Jan Svec
INTERSPEECH3
2015 Word-semantic lattices for spoken language understanding
abstract
The paper presents a method for converting word-based automatic speech recognition (ASR) lattices into word-semantic (W-SE) lattices that contain original words together with a partial semantic information - so-called semantic entities. Semantic entity detection algorithm generates semantic entities based on the expert-defined knowledge. The generated W-SE lattices have smaller vocabulary and consequently reduce the sparsity of the training data. The format of the W-SE lattices also naturally preserves the inherent uncertainty of the ASR output that can be exploited in subsequent dialog modules. The presented technique employs the framework of weighted finite state transducers which allows for efficient optimization of word-semantic lattices. We have evaluated the method in two different spoken language understanding tasks and obtained more than 10% reduction of concept error rate in comparison with using 1-best word hypothesis in both of those tasks.
Jan Svec, Lubos Smídl, Tomás Valenta, Adam Chýlek, Pavel Ircing
ICASSP1
2015 Hierarchical discriminative model for spoken language understanding based on convolutional neural network
Jan Svec, Adam Chýlek, Lubos Smídl
INTERSPEECH1
2014 WISE 2014 Challenge: Multi-label Classification of Print Media Articles to Topics
Grigorios Tsoumakas, Apostolos N. Papadopoulos, Weining Qian, Stavros Vologiannidis, Alexander D'yakonov, Antti Puurula, Jesse Read, Jan Svec, Stanislav Semenov
WISE (2)8
2013 Semantic entity detection from multiple ASR hypotheses within the WFST framework
abstract
The paper presents a novel approach to named entity detection from ASR lattices. Since the described method not only detects the named entities but also assigns a detailed semantic interpretation to them, we call our approach the semantic entity detection. All the algorithms are designed to use automata operations defined within the framework of weighted finite state transducers (WFST) - the ASR lattices are nowadays frequently represented as weighted acceptors. The expert knowledge about the semantics of the task at hand can be first expressed in the form of a context free grammar and then converted to the FST form. We use a WFST optimization to obtain compact representation of the ASR lattice. The WFST framework also allows to use the word confusion networks as another representation of multiple ASR hypotheses. That way we can use the full power of composition and optimization operations implemented in the OpenFST toolkit for our semantic entity detection algorithm. The devised method also employs the concept of a factor automaton; this approach allows us to overcome the need for a filler model and consequently makes the method more general. The paper includes experimental evaluation of the proposed algorithm and compares the performance obtained by using the one-best word hypothesis, optimized lattices and word confusion networks.
Jan Svec, Pavel Ircing, Lubos Smídl
ASRU1
2013 Efficient algorithm for rational kernel evaluation in large lattice sets
abstract
This paper presents an effective method for evaluation of the rational kernels represented by finite-state automata. The described algorithm is optimized for processing speed and thus facilitates the usage of state-of-the-art machine learning techniques like Support Vector Machines even in the real-time application of speech and language processing, such as dialogue systems and speech retrieval engines. The performance of the devised algorithm was tested on a spoken language understanding task and the results suggest that it consistently outperforms the baseline algorithm presented in the related literature.
Jan Svec, Pavel Ircing
ICASSP1
2013 Hierarchical discriminative model for spoken language understanding
abstract
The paper presents a new discriminative model for statistical spoken language understanding designed for use in spoken dialog systems. The parsing algorithm uses lexicalized grammar derived from unaligned training data with probability estimates generated by multiclass classifiers. The generated semantic trees are partially aligned with the input sentence to provide lexical realisation of semantic concepts. The model was evaluated on two semantically annotated corpora and in both tasks it outperforms the baseline Hidden Vector State parser and Semantic Tuple Classifiers model. The experiments were performed using both transcribed data and recognized lattices. The innovative aspect of using phoneme lattices in the understanding process instead of word lattices is examined and described.
Jan Svec, Lubos Smídl, Pavel Ircing
ICASSP1
2008 Extension of HVS semantic parser by allowing left-right branching
abstract
The hidden vector state (HVS) parser is a popular method for semantic parsing. It is used in the language understanding module of the statistical based spoken dialog system. This paper presents an extension of the HVS semantic parser. It enables the parser to generate broader class of semantic trees. This modification can be used to improve the performance of the parser by generating not only the right-branching trees (like original HVS parser) but also limited left-branching trees and their combinations. The extension retains simplicity and properties of the original HVS parser. We tested the method on Czech human-human train timetable corpus. The modified HVS parser yields statistically significant improvement. The accuracy of the system increased from 50.4% to 58.3% absolutely.
Filip Jurcícek, Jan Svec, Ludek Müller
ICASSP2
2008 Structural Metadata Annotation of Speech Corpora: Comparing Broadcast News and Broadcast Conversations
Jáchym Kolár, Jan Svec
LREC2
2005 Czech spontaneous speech corpus with structural metadata
abstract
Tento článek popisuje český korpus spontánní řeči skládajícíse z nahrávek rozhlasových diskusních pořadů. Jako první kompletní neanglický MDE korpus byl anotován strukturálními metadaty, která zvyšují čitelnost přepisů člověkem a umožňují i další automatické zpracování. Anotace zahrnuje rozdělení přepisů do syntakticko-sémantických jednotek a identifikace výplní a neplynulostí. Mimo modifikací nutných pouze pro češtinu také navrhujeme některé modifikace nezávislé na jazyku, jako je například limitované prozodické značkování na hranicích syntakticko-sémantických jednotek.
Jáchym Kolár, Jan Svec, Stephanie M. Strassel, Christopher Walker, Dagmar Kozlíková, Josef Psutka
INTERSPEECH2