VLDB 2026 Research / reviewers in the wild / expert
Jan Lehecka
dblp:21/10057
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-3889-8069ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zero-shot Out-of-domain is No Joke: Lessons Learned in the VoiceMOS 2023 MOS Prediction Challenge
Marie Kunesová, Jan Lehecka, Josef Michalek, Jindrich Matousek, Jan Svec |
INTERSPEECH | 2 |
| 2024 | A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History ArchivesabstractIn this paper, we are comparing monolingual Wav2Vec 2.0 models with various multilingual models to see whether we could improve speech recognition performance on a unique oral history archive containing a lot of mixed-language sentences. Our main goal is to push forward research on this unique dataset, which is an extremely valuable part of our cultural heritage. Our results suggest that monolingual speech recognition models are, in most cases, superior to multilingual models, even when processing the oral history archive full of mixed-language sentences from non-native speakers. We also performed the same experiments on the public CommonVoice dataset to verify our results. We are contributing to the research community by releasing our pre-trained models to the public. Jan Lehecka, Josef V. Psutka, Lubos Smídl, Pavel Ircing, Josef Psutka |
INTERSPEECH | 1 |
| 2023 | Ensemble of Deep Neural Network Models for MOS PredictionabstractAutomatic evaluation of the quality of synthetic speech has the potential to serve as a cheaper and less time-consuming alternative to standard listening tests. In this paper, we present our contribution to the ongoing research: a system for automatic prediction of the mean opinion score (MOS) given by human listeners. The system was specifically developed for the recent VoiceMOS Challenge. Following the success of fusion systems in similar challenges, our contribution is an ensemble that interpolates the outputs of seven different models: four different wav2vec models, a CNN-RNN model, QuartzNet, and the LDNet baseline. During the VoiceMOS challenge, our system achieved the second-best utterance-level MSE of 0.171 and ranged from 2nd to 8th place among all 22 participating teams in terms of other evaluation metrics. Marie Kunesová, Jindrich Matousek, Jan Lehecka, Jan Svec, Josef Michalek, Daniel Tihelka, Martin Bulín, Zdenek Hanzlícek, Markéta Rezácková |
ICASSP | 3 |
| 2023 | Transformer-based Speech Recognition Models for Oral History Archives in English, German, and Czech
Jan Lehecka, Jan Svec, Josef V. Psutka, Pavel Ircing |
INTERSPEECH | 1 |
| 2022 | Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of CzechabstractIn this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automatic speech recognition tasks using a combination of in-domain data and almost 6 thousand hours of out-of-domain transcribed speech. We are presenting a large palette of experiments with various fine-tuning setups evaluated on two public datasets (CommonVoice and VoxPopuli) and one extremely challenging dataset from the MALACH project. Our results show that monolingual Wav2Vec 2.0 models are robust ASR systems, which can take advantage of large labeled and unlabeled datasets and successfully compete with state-of-the-art LVCSR systems. Moreover, Wav2Vec models proved to be good zero-shot learners when no training data are available for the target ASR task. Jan Lehecka, Jan Svec, Ales Prazák, Josef Psutka |
INTERSPEECH | 1 |
| 2022 | Deep LSTM Spoken Term Detection using Wav2Vec 2.0 RecognizerabstractIn recent years, the standard hybrid DNN-HMM speech recognizers are outperformed by the end-to-end speech recognition systems. One of the very promising approaches is the grapheme Wav2Vec 2.0 model, which uses the self-supervised pretraining approach combined with transfer learning of the fine-tuned speech recognizer. Since it lacks the pronunciation vocabulary and language model, the approach is suitable for tasks where obtaining such models is not easy or almost impossible. In this paper, we use the Wav2Vec speech recognizer in the task of spoken term detection over a large set of spoken documents. The method employs a deep LSTM network which maps the recognized hypothesis and the searched term into a shared pronunciation embedding space in which the term occurrences and the assigned scores are easily computed. The paper describes a bootstrapping approach that allows the transfer of the knowledge contained in traditional pronunciation vocabulary of DNN-HMM hybrid ASR into the context of grapheme-based Wav2Vec. The proposed method outperforms the previously published system based on the combination of the DNN-HMM hybrid ASR and phoneme recognizer by a large margin on the MALACH data in both English and Czech languages. Jan Svec, Jan Lehecka, Lubos Smídl |
INTERSPEECH | 2 |