VLDB 2026 Research / reviewers in the wild / expert
Jindrich Zdánský
dblp:84/4504
· DBLP profile ↗
34ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-5591-7228ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Combining multilingual resources to enhance end-to-end speech recognition systems for Scandinavian languages
Lukás Mateju, Jan Nouza, Petr Cerva, Jindrich Zdánský |
Speech Commun. | 4 |
| 2025 | Lightweight online punctuation and capitalization restoration for streaming ASR systems
Martin Polácek, Petr Cerva, Jindrich Zdánský |
Speech Commun. | 3 |
| 2023 | Combining Multilingual Resources and Models to Develop State-of-the-Art E2E ASR for Swedish
Lukás Mateju, Jan Nouza, Petr Cerva, Jindrich Zdánský, Frantisek Kynych |
INTERSPEECH | 4 |
| 2023 | Online Punctuation Restoration using ELECTRA Model for streaming ASR Systems
Martin Polácek, Petr Cerva, Jindrich Zdánský, Lenka Weingartová |
INTERSPEECH | 3 |
| 2022 | Overlapped Speech Detection in Broadcast Streams Using X-vectors
Lukás Mateju, Frantisek Kynych, Petr Cerva, Jirí Málek, Jindrich Zdánský |
INTERSPEECH | 5 |
| 2022 | Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker IdentificationabstractThis manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since the blind IVE cannot distinguish the target source by itself, it is guided towards the SOI via frame-wise speaker identification based on deep learning. Still, an incorrect speaker can be extracted due to guidance failings, especially when processing challenging data. To identify such cases, we propose a criterion for non-intrusively assessing the estimated speaker. It utilizes the same model as the speaker identification, so no additional training is required. When incorrect extraction is detected, we propose a “deflation” step in which the incorrect source is subtracted from the mixture and, subsequently, another attempt to extract the SOI is performed. The process is repeated until successful extraction is achieved. The proposed procedure is experimentally tested on artificial and real-world datasets containing challenging phenomena: source movements, reverberation, transient noise, or microphone failures. The method is compared with state-of-the-art blind algorithms as well as with current fully supervised deep learning-based methods. Jirí Málek, Jakub Janský, Zbynek Koldovský, Tomás Kounovský, Jaroslav Cmejla, Jindrich Zdánský |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-VectorsabstractWe propose a novel approach for semi-supervised extraction of a moving audio source of interest (SOI) applicable in reverberant and noisy environments. The blind part of the method is based on independent vector extraction (IVE) and uses the recently proposed constant separating vector (CSV) mixing model. This model allows for changes of mixing parameters within the processed interval of the mixture, which potentially leads to higher accuracy of SOI estimation. The supervised part of the method concerns a pilot signal, which is related to the SOI and ensures the convergence of the blind method towards the SOI. The pilot is based on robust detection of frames where SOI is dominant via speaker embeddings called X-vectors. Robustness of the detection is achieved through augmentation of the data for the supervised training of the X-vectors. The pilot-supported extraction yields significantly better performance compared to its unsupervised counterpart identifying SOI solely using the initialization. Jirí Málek, Jakub Janský, Tomás Kounovský, Zbynek Koldovský, Jindrich Zdánský |
ICASSP | 5 |
| 2021 | Using X-Vectors for Speech Activity Detection in Broadcast Streams
Lukás Mateju, Frantisek Kynych, Petr Cerva, Jindrich Zdánský, Jirí Málek |
Interspeech | 4 |
| 2021 | Identification of related languages from spoken data: Moving from off-line to on-line scenario
Petr Cerva, Lukás Mateju, Jindrich Zdánský, Radek Safarík, Jan Nouza |
Comput. Speech Lang. | 3 |
| 2020 | Adaptive Blind Audio Source Extraction Supervised By Dominant Speaker Identification Using X-VectorsabstractWe propose a novel algorithm for adaptive blind audio source extraction. The proposed method is based on independent vector analysis and utilizes the auxiliary function optimization to achieve high convergence speed. The algorithm is partially supervised by a pilot signal related to the source of interest (SOI), which ensures that the method correctly extracts the utterance of the desired speaker. The pilot is based on the identification of a dominant speaker in the mixture using x-vectors. The properties of the x-vectors computed in the presence of cross-talk are experimentally analyzed. The proposed approach is verified in a scenario with a moving SOI, static interfering speaker and environmental noise. Jakub Janský, Jirí Málek, Jaroslav Cmejla, Tomás Kounovský, Zbynek Koldovský, Jindrich Zdánský |
ICASSP | 6 |
| 2019 | An Approach to Online Speaker Change Point Detection Using DNNs and WFSTs
Lukás Mateju, Petr Cerva, Jindrich Zdánský |
INTERSPEECH | 3 |
| 2018 | Robust Recognition of Speech with Background Music in Acoustically Under-Resourced ScenariosabstractThis paper addresses the task of Automatic Speech Recognition (ASR) with music in the background. We consider two different situations: 1) scenarios with very small amount of labeled training utterances (duration 1 hour) and 2) scenarios with large amount of labeled training utterances (duration 132 hours). In these situations, we aim to achieve robust recognition. To this end we investigate the following techniques: a) multi-condition training of the acoustic model, b) denoising autoencoders for feature enhancement and c) joint training of both above mentioned techniques. We demonstrate that the considered methods can be successfully trained with the small amount of labeled acoustic data. We present substantially improved performance compared to acoustic models trained on clean speech. Further, we show a significant increase of accuracy in the under-resourced scenario, when utilizing additional amount of non-labeled data. Here, the non-labeled dataset is used to improve the accuracy of the feature enhancement via autoencoders. Subsequently, the autoencoders are jointly fine-tuned along with the acoustic model using the small amount of labeled utterances. Jirí Málek, Jindrich Zdánský, Petr Cerva |
ICASSP | 2 |
| 2018 | Using Deep Neural Networks for Identification of Slavic Languages from Acoustic Signal
Lukás Mateju, Petr Cerva, Jindrich Zdánský, Radek Safarík |
INTERSPEECH | 3 |
| 2017 | Robust Automatic Recognition of Speech with background musicabstractThis paper addresses the task of Automatic Speech Recognition (ASR) with music in the background, where the accuracy of recognition may deteriorate significantly. To improve the robustness of ASR in this task, e.g. for broadcast news transcription or subtitles creation, we adopt two approaches: 1) multi-condition training of the acoustic models and 2) denoising autoencoders followed by acoustic model training on the preprocessed data. In the latter case, two types of autoencoders are considered: the fully connected and the convolutional network. Presented experimental results show that all the investigated techniques are able to improve the recognition of speech distorted by music significantly. For example, in the case of artificial mixtures of speech and electronic music (low Signal-to-Noise Ratio (SNR) of 0 dB), we achieved absolute improvement of accuracy by 35.8%. For real-world broadcast news and a high SNR (about 10 dB), we achieved improvement by 2.4%. The important advantage of the studied approaches is that they do not deteriorate the accuracy in scenarios with clean speech (the decrease is about 1%). Jirí Málek, Jindrich Zdánský, Petr Cerva |
ICASSP | 2 |
| 2017 | Speech Activity Detection in online broadcast transcription using Deep Neural Networks and Weighted Finite State TransducersabstractIn this paper, a new approach to online Speech Activity Detection (SAD) is proposed. This approach is designed for the use in a system that carries out 24/7 transcription of radio/TV broadcasts containing a large amount of non-speech segments, such as advertisements or music. To improve the robustness of detection, we adopt Deep Neural Networks (DNNs) trained on artificially created mixtures of speech and non-speech signals at desired levels of signal-to-noise ratio (SNR). An integral part of our approach is an online decoder based on Weighted Finite State Transducers (WFSTs); this decoder smooths the output from DNN. The employed transduction model is context-based, i.e., both speech and non-speech events are modeled using sequences of states. The presented experimental results show that our approach yields state-of-the-art results on standardized QUT-NOISE-TIMIT data set for SAD and, at the same time, it is capable of a) operating with low latency and b) reducing the computational demands and error rate of the target transcription system. Lukás Mateju, Petr Cerva, Jindrich Zdánský, Jirí Málek |
ICASSP | 3 |
| 2014 | Speech-to-text technology to transcribe and disclose 100, 000+ hours of bilingual documents from historical Czech and Czechoslovak radio archive
Jan Nouza, Petr Cerva, Jindrich Zdánský, Karel Blavka, Marek Bohac, Jan Silovský, Josef Chaloupka, Michaela Kucharová, Ladislav Seps, Jirí Málek, Michal Rott |
INTERSPEECH | 3 |
| 2013 | Speaker-adaptive speech recognition using speaker diarization for improved transcription of large spoken archives
Petr Cerva, Jan Silovský, Jindrich Zdánský, Jan Nouza, Ladislav Seps |
Speech Commun. | 3 |
| 2012 | Real-Time Lecture Transcription using ASR for Czech Hearing Impaired or Deaf Students
Petr Cerva, Jan Silovský, Jindrich Zdánský, Jan Nouza, Jirí Málek |
INTERSPEECH | 3 |
| 2012 | Study on Integration of Speaker Diarization with Speaker Adaptive Speech Recognition for Broadcast Transcription
Jan Silovský, Petr Cerva, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 3 |
| 2012 | Browsing, indexing and automatic transcription of lectures for distance learningabstractThis paper presents a complex system developed to improve the quality of distance learning by allowing people to browse the content of various (academic) lectures. The system consists of several main modules. The first automatic speech recognition (ASR) module is designed to cope with inflective Czech language and provides time-aligned transcriptions of input audio-visual recordings of lectures. These transcriptions are generated off-line in two recognition passes using speaker adaptation methods and language models mixed from various text sources including transcriptions of broadcast programs, spontaneous telephone talks, web discussions, thesis, etc. Lecture recordings and their transcriptions are then indexed and stored in the database. The next module, client-server web lecture browser, allows to browse or play the indexed content and search in it. Petr Cerva, Jan Silovský, Jindrich Zdánský, Ondrej Smola, Karel Blavka, Karel Palecek, Jan Nouza, Jirí Málek |
MMSP | 3 |
| 2012 | Large-scale processing, indexing and search system for Czech audio-visual cultural heritage archivesabstractThis paper describes a complex system developed for processing, indexing and accessing data collected in large audio and audio-visual archives that make an important part of Czech cultural heritage. Recently, the system is being applied to the Czech Radio archive, namely to its oral history segment with more than 200.000 individual recordings covering almost ninety years of broadcasting in the Czech Republic and former Czechoslovakia. The ultimate goals are a) to transcribe a significant portion of the archive - with the support of speech, speaker and language recognition technology, b) index the transcriptions, and c) make the audio and text files fully searchable. So far, the system has processed and indexed over 75.000 spoken documents. Most of them come from the last two decades, but the recent demo collection includes also a series of presidential speeches since 1934. The full coverage of the archive should be available by the end of 2014. Jan Nouza, Karel Blavka, Jindrich Zdánský, Petr Cerva, Jan Silovský, Marek Bohac, Josef Chaloupka, Michaela Kucharová, Ladislav Seps |
MMSP | 3 |
| 2012 | Incorporation of the ASR output in speaker segmentation and clustering within the task of speaker diarization of broadcast streamsabstractIn this paper we study the effect of incorporation of automatic transcriptions in the speaker diarization process. We aim to improve both the diarization accuracy as evaluated by standard objective measures and quality of the diarization output from user's perspective. Although the presented approach relies on output of an automatic speech recognizer, it makes no use of lexical information. Instead, we use information about word boundaries and classification of non-speech events occurring in the processed stream. The former information is used as constraining condition for speaker change-point candidates and the latter facilitate to neglect various vocal noise sounds that carry no speaker-specific information (considering representation of the signal by cepstral features) and thus harm the speaker's representation. The experimental evaluation of the presented approach was carried out using the COST278 multilingual broadcast news database. We demonstrate that the approach yields improvement in terms of both speaker diarization and segmentation performance measures. Furthermore, we show that the number of change-points detected within words (and not at their boundaries) is significantly reduced. Jan Silovský, Jindrich Zdánský, Jan Nouza, Petr Cerva, Jan Prazak |
MMSP | 2 |
| 2011 | PLDA-Based Clustering for Speaker Diarization of Broadcast Streams
Jan Silovský, Jan Prazak, Petr Cerva, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 4 |
| 2009 | Very large vocabulary voice dictation for mobile devices
Jan Nouza, Petr Cerva, Jindrich Zdánský |
INTERSPEECH | 3 |
| 2008 | Enhancement of noisy speech recordings via blind source separationabstractWe propose an improved time-domain Blind Source Separa-tion method and apply it to speech signal enhancement using multiple microphone recordings. The improvement consists in utilization of fuzzy clustering instead of a hard one, which is verified by experiments where real-world mixtures of two au-dio signals are separated from two microphones. Performance of the method is demonstrated by recognizing mixed and sepa-rated utterances from the Czech part of the European broadcast news database using our Czech LVCSR system. The separation allows significantly better recognition, e.g., by 32 % when the jammer signal is a Gaussian noise and the input signal-to-noise ratio is 10dB. Jirí Málek, Zbynek Koldovský, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 3 |
| 2008 | Czech-to-slovak adapted broadcast news transcription system
Jan Nouza, Jan Silovský, Jindrich Zdánský, Petr Cerva, Martin Kroul, Josef Chaloupka |
INTERSPEECH | 3 |
| 2008 | Joint audio-visual processing, representation and indexing of TV news programmesabstractIn the paper we present a complex platform for automatic processing of Czech TV news programmes. Its audio processing module provides text transcription in form of metadata that contain information about spoken content, speaker identities, used pronunciation, word positions and intonation. The video processing module provides pictures representing individual video scenes and information about detected and possibly recognized human faces. The audio and video data are merged into single XML files that are indexed and stored in a searchable database. A simple Web-based search engine can be used to retrieve information from the database that recently contain more than 1800 hours of transcribed programmes from Czech CT24 station. Jindrich Zdánský, Josef Chaloupka, Jan Nouza |
MMSP | 1 |
| 2006 | Continual on-line monitoring of Czech spoken broadcast programs
Jan Nouza, Jindrich Zdánský, Petr Cerva, Jan Kolorenc |
INTERSPEECH | 2 |
| 2006 | BINSEG: an efficient speaker-based segmentation techniqueabstractIn this paper we present a new efficient approach to speaker-based audio stream segmentation. It employs binary segmentation technique that is well-known from mathematical statistic. Because integral part of this technique is hypotheses testing, we compare two well-founded (Maximum Likelihood, Informational) and one commonly used (BIC difference) approach for deriving speakerchange test statistics. Based on results of this comparison we propose both off-line and on-line speaker change detection algorithms (including way of effective training) that have merits of high accuracy and low computational costs. In simulated tests with artificially mixed data the on-line algorithm identified 95.7% of all speaker changes with precision of 96.9%. In tests done with 30 hours of real broadcast news (in 9 languages) the average recall was 74.4% and precision 70.3%. Index Terms: speaker change detection, acoustic segmentation. Jindrich Zdánský |
INTERSPEECH | 1 |
| 2005 | Fully automated system for Czech spoken broadcast transcription with very large (300k+) lexicon
Jan Nouza, Jindrich Zdánský, Petr David, Petr Cerva, Jan Kolorenc, Dana Nejedlová |
INTERSPEECH | 2 |
| 2005 | Detection of acoustic change-points in audio records via global BIC maximization and dynamic programming
Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 1 |
| 2005 | The COST278 broadcast news segmentation and speaker clustering evaluation - overview, methodology, systems, resultsabstractThis paper describes a large scale experiment in which eight research institutions have tested their audio partitioning and labeling algorithms on the same data, a multi-lingual database of news broadcasts, using the same evaluation tools and protocols. The experiments have provide more insight in the cross-lingual robustness of the methods and they have demonstrated that by further collaborating in thedomains of speaker change detection and speaker clustering it should be possible to achieve further technological progress in the near future. Janez Zibert, France Mihelic, Jean-Pierre Martens, Hugo Meinedo, João Paulo da Silva Neto, Laura Docío Fernández, Carmen García-Mateo, Petr David, Jindrich Zdánský, Matús Pleva, Anton Cizmar, Andrej Zgank, Zdravko Kacic, Csaba Teleki, Klára Vicsi |
INTERSPEECH | 9 |
| 2004 | Very large vocabulary speech recognition system for automatic transcription of czech broadcast programsabstractThis paper describes the first speech recognition system capable of transcribing a wide range of spoken broadcast programs in Czech language with the OOV rate being below 3 per cent. To achieve that level we had to a) create an optimized 200k word vocabulary with multiple text and pronunciation forms, b) extract an appropriate language model from a 300M word text corpus and c) develop an own decoder specially designed for the lexicon of that size. The system was tested on various types of broadcast programs with the following results: the Czech part of the European COST278 database of TV news (71.5 % accuracy rate on complete news streams, 82.7 % on their clean parts), radio news (80.2 %), read commentaries (78.6 %), broadcast debates (74.3 %) and recordings of the state presidents’ speeches (85.8 %). Jan Nouza, Dana Nejedlová, Jindrich Zdánský, Jan Kolorenc |
INTERSPEECH | 3 |
| 2004 | An improved preprocessor for the automatic transcription of broadcast news audio streamabstractThis paper deals with the preprocessing of the broadcast news (BN) audio stream for the automatic transcription purposes. The preprocessing consists of the automatic segmentation followed by the broad-class segment identification. The former is capable of detecting speaker and/or acoustic changes in the BN audio stream with the precision being 82.75%. The latter acts as a filter that removes nonspeech parts. The performance of the proposed system was evaluated on the multi-lingual pan-European COST278 BN database containing data in 6 languages. The preprocessing and segmentation module operates in a near-real-time way, with the total delay of 12 seconds. Its practical functionality was evaluated on the Czech part of the BN database. The automatically segmented signal was directly sent to the large vocabulary speech recognition system operating with a 200K-word Czech lexicon. The difference in performance between automatically and manually segmented BN streams was only minimal- 1.12%. 1. Jindrich Zdánský, Petr David, Jan Nouza |
INTERSPEECH | 1 |