VLDB 2026 Research / reviewers in the wild / expert
Petr Cerva
dblp:09/36
· DBLP profile ↗
30ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0003-0767-0106ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Lukás Mateju, Jan Nouza, Petr Cerva, Jindrich Zdánský |
Speech Commun. | 3 |
| 2025 | Lightweight online punctuation and capitalization restoration for streaming ASR systems
Martin Polácek, Petr Cerva, Jindrich Zdánský |
Speech Commun. | 2 |
| 2023 | Combining Multilingual Resources and Models to Develop State-of-the-Art E2E ASR for Swedish
Lukás Mateju, Jan Nouza, Petr Cerva, Jindrich Zdánský, Frantisek Kynych |
INTERSPEECH | 3 |
| 2023 | Online Punctuation Restoration using ELECTRA Model for streaming ASR Systems
Martin Polácek, Petr Cerva, Jindrich Zdánský, Lenka Weingartová |
INTERSPEECH | 2 |
| 2022 | Overlapped Speech Detection in Broadcast Streams Using X-vectors
Lukás Mateju, Frantisek Kynych, Petr Cerva, Jirí Málek, Jindrich Zdánský |
INTERSPEECH | 3 |
| 2021 | Using X-Vectors for Speech Activity Detection in Broadcast Streams
Lukás Mateju, Frantisek Kynych, Petr Cerva, Jindrich Zdánský, Jirí Málek |
Interspeech | 3 |
| 2021 | Identification of related languages from spoken data: Moving from off-line to on-line scenario
Petr Cerva, Lukás Mateju, Jindrich Zdánský, Radek Safarík, Jan Nouza |
Comput. Speech Lang. | 1 |
| 2019 | An Approach to Online Speaker Change Point Detection Using DNNs and WFSTs
Lukás Mateju, Petr Cerva, Jindrich Zdánský |
INTERSPEECH | 2 |
| 2018 | Robust Recognition of Speech with Background Music in Acoustically Under-Resourced ScenariosabstractThis paper addresses the task of Automatic Speech Recognition (ASR) with music in the background. We consider two different situations: 1) scenarios with very small amount of labeled training utterances (duration 1 hour) and 2) scenarios with large amount of labeled training utterances (duration 132 hours). In these situations, we aim to achieve robust recognition. To this end we investigate the following techniques: a) multi-condition training of the acoustic model, b) denoising autoencoders for feature enhancement and c) joint training of both above mentioned techniques. We demonstrate that the considered methods can be successfully trained with the small amount of labeled acoustic data. We present substantially improved performance compared to acoustic models trained on clean speech. Further, we show a significant increase of accuracy in the under-resourced scenario, when utilizing additional amount of non-labeled data. Here, the non-labeled dataset is used to improve the accuracy of the feature enhancement via autoencoders. Subsequently, the autoencoders are jointly fine-tuned along with the acoustic model using the small amount of labeled utterances. Jirí Málek, Jindrich Zdánský, Petr Cerva |
ICASSP | 3 |
| 2018 | Using Deep Neural Networks for Identification of Slavic Languages from Acoustic Signal
Lukás Mateju, Petr Cerva, Jindrich Zdánský, Radek Safarík |
INTERSPEECH | 2 |
| 2017 | Robust Automatic Recognition of Speech with background musicabstractThis paper addresses the task of Automatic Speech Recognition (ASR) with music in the background, where the accuracy of recognition may deteriorate significantly. To improve the robustness of ASR in this task, e.g. for broadcast news transcription or subtitles creation, we adopt two approaches: 1) multi-condition training of the acoustic models and 2) denoising autoencoders followed by acoustic model training on the preprocessed data. In the latter case, two types of autoencoders are considered: the fully connected and the convolutional network. Presented experimental results show that all the investigated techniques are able to improve the recognition of speech distorted by music significantly. For example, in the case of artificial mixtures of speech and electronic music (low Signal-to-Noise Ratio (SNR) of 0 dB), we achieved absolute improvement of accuracy by 35.8%. For real-world broadcast news and a high SNR (about 10 dB), we achieved improvement by 2.4%. The important advantage of the studied approaches is that they do not deteriorate the accuracy in scenarios with clean speech (the decrease is about 1%). Jirí Málek, Jindrich Zdánský, Petr Cerva |
ICASSP | 3 |
| 2017 | Speech Activity Detection in online broadcast transcription using Deep Neural Networks and Weighted Finite State TransducersabstractIn this paper, a new approach to online Speech Activity Detection (SAD) is proposed. This approach is designed for the use in a system that carries out 24/7 transcription of radio/TV broadcasts containing a large amount of non-speech segments, such as advertisements or music. To improve the robustness of detection, we adopt Deep Neural Networks (DNNs) trained on artificially created mixtures of speech and non-speech signals at desired levels of signal-to-noise ratio (SNR). An integral part of our approach is an online decoder based on Weighted Finite State Transducers (WFSTs); this decoder smooths the output from DNN. The employed transduction model is context-based, i.e., both speech and non-speech events are modeled using sequences of states. The presented experimental results show that our approach yields state-of-the-art results on standardized QUT-NOISE-TIMIT data set for SAD and, at the same time, it is capable of a) operating with low latency and b) reducing the computational demands and error rate of the target transcription system. Lukás Mateju, Petr Cerva, Jindrich Zdánský, Jirí Málek |
ICASSP | 2 |
| 2016 | ASR for South Slavic Languages Developed in Almost Automated Way
Jan Nouza, Radek Safarík, Petr Cerva |
INTERSPEECH | 3 |
| 2014 | Speech-to-text technology to transcribe and disclose 100, 000+ hours of bilingual documents from historical Czech and Czechoslovak radio archive
Jan Nouza, Petr Cerva, Jindrich Zdánský, Karel Blavka, Marek Bohac, Jan Silovský, Josef Chaloupka, Michaela Kucharová, Ladislav Seps, Jirí Málek, Michal Rott |
INTERSPEECH | 2 |
| 2014 | Investigation of deep neural networks for robust recognition of nonlinearly distorted speech
Ladislav Seps, Jirí Málek, Petr Cerva, Jan Nouza |
INTERSPEECH | 3 |
| 2013 | Adding controlled amount of noise to improve recognition of compressed and spectrally distorted speechabstractThis paper deals with the recognition of speech whose spectrum is notably distorted by lossy compression (namely MP3) or by some implementations of `speech enhancement' techniques. We show that these non-linear treatments can introduce gaps in spectrum that significantly change the distribution of MFCCs and degrade performance of ASR. We propose a method that measures the level of spectrum distortion and use it for adding a controlled amount of noise to the signal. It effectively masks the gaps and helps namely in situations where the source and parameters of the distortion are not known and hence we cannot use a properly matched acoustic model. In spite of its simplicity, the method can improve significantly speech recognition of highly compressed or spectrally distorted signals. We demonstrate it in several large experiments conducted on publicly available speech databases, in two languages and for two types of spectral distortion. Jan Nouza, Petr Cerva, Jan Silovský |
ICASSP | 2 |
| 2013 | Speaker-adaptive speech recognition using speaker diarization for improved transcription of large spoken archives
Petr Cerva, Jan Silovský, Jindrich Zdánský, Jan Nouza, Ladislav Seps |
Speech Commun. | 1 |
| 2012 | Real-Time Lecture Transcription using ASR for Czech Hearing Impaired or Deaf Students
Petr Cerva, Jan Silovský, Jindrich Zdánský, Jan Nouza, Jirí Málek |
INTERSPEECH | 1 |
| 2012 | Study on Integration of Speaker Diarization with Speaker Adaptive Speech Recognition for Broadcast Transcription
Jan Silovský, Petr Cerva, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 2 |
| 2012 | Browsing, indexing and automatic transcription of lectures for distance learningabstractThis paper presents a complex system developed to improve the quality of distance learning by allowing people to browse the content of various (academic) lectures. The system consists of several main modules. The first automatic speech recognition (ASR) module is designed to cope with inflective Czech language and provides time-aligned transcriptions of input audio-visual recordings of lectures. These transcriptions are generated off-line in two recognition passes using speaker adaptation methods and language models mixed from various text sources including transcriptions of broadcast programs, spontaneous telephone talks, web discussions, thesis, etc. Lecture recordings and their transcriptions are then indexed and stored in the database. The next module, client-server web lecture browser, allows to browse or play the indexed content and search in it. Petr Cerva, Jan Silovský, Jindrich Zdánský, Ondrej Smola, Karel Blavka, Karel Palecek, Jan Nouza, Jirí Málek |
MMSP | 1 |
| 2012 | Large-scale processing, indexing and search system for Czech audio-visual cultural heritage archivesabstractThis paper describes a complex system developed for processing, indexing and accessing data collected in large audio and audio-visual archives that make an important part of Czech cultural heritage. Recently, the system is being applied to the Czech Radio archive, namely to its oral history segment with more than 200.000 individual recordings covering almost ninety years of broadcasting in the Czech Republic and former Czechoslovakia. The ultimate goals are a) to transcribe a significant portion of the archive - with the support of speech, speaker and language recognition technology, b) index the transcriptions, and c) make the audio and text files fully searchable. So far, the system has processed and indexed over 75.000 spoken documents. Most of them come from the last two decades, but the recent demo collection includes also a series of presidential speeches since 1934. The full coverage of the archive should be available by the end of 2014. Jan Nouza, Karel Blavka, Jindrich Zdánský, Petr Cerva, Jan Silovský, Marek Bohac, Josef Chaloupka, Michaela Kucharová, Ladislav Seps |
MMSP | 4 |
| 2012 | Incorporation of the ASR output in speaker segmentation and clustering within the task of speaker diarization of broadcast streamsabstractIn this paper we study the effect of incorporation of automatic transcriptions in the speaker diarization process. We aim to improve both the diarization accuracy as evaluated by standard objective measures and quality of the diarization output from user's perspective. Although the presented approach relies on output of an automatic speech recognizer, it makes no use of lexical information. Instead, we use information about word boundaries and classification of non-speech events occurring in the processed stream. The former information is used as constraining condition for speaker change-point candidates and the latter facilitate to neglect various vocal noise sounds that carry no speaker-specific information (considering representation of the signal by cepstral features) and thus harm the speaker's representation. The experimental evaluation of the presented approach was carried out using the COST278 multilingual broadcast news database. We demonstrate that the approach yields improvement in terms of both speaker diarization and segmentation performance measures. Furthermore, we show that the number of change-points detected within words (and not at their boundaries) is significantly reduced. Jan Silovský, Jindrich Zdánský, Jan Nouza, Petr Cerva, Jan Prazak |
MMSP | 4 |
| 2011 | Using Unsupervised Feature-Based Speaker Adaptation for Improved Transcription of Spoken Archives
Petr Cerva, Karel Palecek, Jan Silovský, Jan Nouza |
INTERSPEECH | 1 |
| 2011 | PLDA-Based Clustering for Speaker Diarization of Broadcast Streams
Jan Silovský, Jan Prazak, Petr Cerva, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 3 |
| 2009 | Very large vocabulary voice dictation for mobile devices
Jan Nouza, Petr Cerva, Jindrich Zdánský |
INTERSPEECH | 2 |
| 2008 | Czech-to-slovak adapted broadcast news transcription system
Jan Nouza, Jan Silovský, Jindrich Zdánský, Petr Cerva, Martin Kroul, Josef Chaloupka |
INTERSPEECH | 4 |
| 2007 | Design and development of voice controlled aids for motor-handicapped persons
Petr Cerva, Jan Nouza |
INTERSPEECH | 1 |
| 2006 | Two-step unsupervised speaker adaptation based on speaker and gender recognition and HMM combination
Petr Cerva, Jan Nouza, Jan Silovský |
INTERSPEECH | 1 |
| 2006 | Continual on-line monitoring of Czech spoken broadcast programs
Jan Nouza, Jindrich Zdánský, Petr Cerva, Jan Kolorenc |
INTERSPEECH | 3 |
| 2005 | Fully automated system for Czech spoken broadcast transcription with very large (300k+) lexicon
Jan Nouza, Jindrich Zdánský, Petr David, Petr Cerva, Jan Kolorenc, Dana Nejedlová |
INTERSPEECH | 4 |