VLDB 2026 Research / reviewers in the wild / expert
Jindrich Matousek
dblp:88/999
· DBLP profile ↗
41ranked-venue papers
18as first author
8since 2021 · last 2024
0000-0002-7408-7730ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 16 first-author · 6 since 2021Artificial intelligence and machine learning · 33 · 15 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zero-shot Out-of-domain is No Joke: Lessons Learned in the VoiceMOS 2023 MOS Prediction Challenge
Marie Kunesová, Jan Lehecka, Josef Michalek, Jindrich Matousek, Jan Svec |
INTERSPEECH | 4 |
| 2024 | Homograph Disambiguation with Text-to-Text Transfer Transformer
Markéta Rezácková, Daniel Tihelka, Jindrich Matousek |
INTERSPEECH | 3 |
| 2024 | Using LSTM neural networks for cross-lingual phonetic speech segmentation with an iterative correction procedureabstractAbstract This article describes experiments on speech segmentation using long short‐term memory recurrent neural networks. The main part of the paper deals with multi‐lingual and cross‐lingual segmentation, that is, it is performed on a language different from the one on which the model was trained. The experimental data involves large Czech, English, German, and Russian speech corpora designated for speech synthesis. For optimal multi‐lingual modeling, a compact phonetic alphabet was proposed by sharing and clustering phones of particular languages. Many experiments were performed exploring various experimental conditions and data combinations. We proposed a simple procedure that iteratively adapts the inaccurate default model to the new voice/language. The segmentation accuracy was evaluated by comparison with reference segmentation created by a well‐tuned hidden Markov model‐based framework with additional manual corrections. The resulting segmentation was also employed in a unit selection text‐to‐speech system. The generated speech quality was compared with the reference segmentation by a preference listening test. Zdenek Hanzlícek, Jindrich Matousek, Jakub Vit |
Comput. Intell. | 2 |
| 2024 | T5G2P: Text-to-Text Transfer Transformer Based Grapheme-to-Phoneme ConversionabstractThe present paper explores the use of several deep neural network architectures to carry out a grapheme-to-phoneme (G2P) conversion, aiming to find a universal and language-independent approach to the task. The models explored are trained on whole sentences in order to automatically capture cross-word context (such as voicedness assimilation) if it exists in the given language. Four different languages, English, Czech, Russian, and German, were chosen due to their different nature and requirements for the G2P task. Ultimately, the Text-to-Text Transfer Transformer (T5) based model achieved very high conversion accuracy on all the tested languages. Also, it exceeded the accuracy reached by a similar system, when trained on a public LibriSpeech database. Markéta Rezácková, Daniel Tihelka, Jindrich Matousek |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Ensemble of Deep Neural Network Models for MOS PredictionabstractAutomatic evaluation of the quality of synthetic speech has the potential to serve as a cheaper and less time-consuming alternative to standard listening tests. In this paper, we present our contribution to the ongoing research: a system for automatic prediction of the mean opinion score (MOS) given by human listeners. The system was specifically developed for the recent VoiceMOS Challenge. Following the success of fusion systems in similar challenges, our contribution is an ensemble that interpolates the outputs of seven different models: four different wav2vec models, a CNN-RNN model, QuartzNet, and the LDNet baseline. During the VoiceMOS challenge, our system achieved the second-best utterance-level MSE of 0.171 and ranged from 2nd to 8th place among all 22 participating teams in terms of other evaluation metrics. Marie Kunesová, Jindrich Matousek, Jan Lehecka, Jan Svec, Josef Michalek, Daniel Tihelka, Martin Bulín, Zdenek Hanzlícek, Markéta Rezácková |
ICASSP | 2 |
| 2023 | Neural Speech Synthesis with Enriched Phrase Boundaries
Marie Kunesová, Jindrich Matousek |
INTERSPEECH | 2 |
| 2021 | A Comparison of Convolutional Neural Networks for Glottal Closure Instant Detection from Raw SpeechabstractIn this paper, we continue to investigate the use of machine learning for the automatic detection of glottal closure instants (GCIs) from raw speech. We compare several deep one-dimensional convolutional neural network architectures on the same data and show that the InceptionV3 model yields the best results on the test set. On publicly available databases, the proposed 1D InceptionV3 outperforms XGBoost, a non-deep machine learning model, as well as other traditional GCI detection algorithms. Jindrich Matousek, Daniel Tihelka |
ICASSP | 1 |
| 2021 | Save Your Voice: Voice Banking and TTS for Anyone
Daniel Tihelka, Markéta Rezácková, Martin Gruber, Zdenek Hanzlícek, Jakub Vit, Jindrich Matousek |
Interspeech | 6 |
| 2019 | Using Extreme Gradient Boosting to Detect Glottal Closure Instants in Speech SignalabstractIn this paper, we continue to investigate the use of classifiers for the automatic detection of glottal closure instants (GCIs) from the speech signal. We focus on extreme gradient boosting (XGB), a fast and powerful implementation of a gradient boosting algorithm. We show that XGB outperforms other classifiers, achieving GCI detection accuracy F 1 = 98.55% and AUC = 99.90%. The proposed XGB model is also shown to outperform other existing GCI detection algorithms on publicly available databases. Despite using much less training data, the performance of XGB is comparable to a deep convolutional neural network based approach, especially when it is tested on voices that were not included in the training data. Jindrich Matousek, Daniel Tihelka |
ICASSP | 1 |
| 2019 | Framework for Conducting Tasks Requiring Human Assessment
Martin Gruber, Adam Chýlek, Jindrich Matousek |
INTERSPEECH | 3 |
| 2019 | Web-Based Speech Synthesis Editor
Martin Gruber, Jakub Vit, Jindrich Matousek |
INTERSPEECH | 3 |
| 2018 | On the Analysis of Training Data for Wavenet-Based Speech SynthesisabstractIn this paper, we analyze how much, how consistent and how accurate data WaveNet-based speech synthesis method needs to be able to generate speech of good quality. We do this by adding artificial noise to the description of our training data and observing how well WaveNet trains and produces speech. More specifically, we add noise to both phonetic segmentation and annotation accuracy, and we also reduce the size of training data by using a fewer number of sentences during training of a WaveNet model. We conducted MUSHRA listening tests and used objective measures to track speech quality within the conducted experiments. We show that WaveNet retains high quality even after adding a small amount of noise (up to 10%) to phonetic segmentation and annotation. A small degradation of speech quality was observed for our WaveNet configuration when only 3 hours of training data were used. Jakub Vit, Zdenek Hanzlícek, Jindrich Matousek |
ICASSP | 3 |
| 2018 | Glottal Closure Instant Detection from Speech Signal Using Voting Classifier and Recursive Feature Elimination
Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2018 | Design and Development of Speech Corpora for Air Traffic Control Training
Lubos Smídl, Jan Svec, Daniel Tihelka, Jindrich Matousek, Jan Romportl, Pavel Ircing |
LREC | 4 |
| 2017 | WebSubDub - Experimental System for Creating High-Quality Alternative Audio Track for TV Broadcasting
Martin Gruber, Jindrich Matousek, Zdenek Hanzlícek, Jakub Vit, Daniel Tihelka |
INTERSPEECH | 2 |
| 2017 | Voice Conservation and TTS System for People Facing Total Laryngectomy
Markéta Juzová, Daniel Tihelka, Jindrich Matousek, Zdenek Hanzlícek |
INTERSPEECH | 3 |
| 2017 | Classification-Based Detection of Glottal Closure Instants from Speech Signals
Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2017 | Anomaly-based annotation error detection in speech-synthesis corpora
Jindrich Matousek, Daniel Tihelka |
Comput. Speech Lang. | 1 |
| 2016 | ARET - Automatic Reading of Educational Texts for Visually Impaired Students
Martin Gruber, Jindrich Matousek, Zdenek Hanzlícek, Zdenek Krnoul, Zbynek Zajíc |
INTERSPEECH | 2 |
| 2016 | Voting Detector: A Combination of Anomaly Detectors to Reveal Annotation Errors in TTS Corpora
Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2015 | Anomaly-based annotation errors detection in TTS corpora
Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2014 | Very fast unit selection using Viterbi search with zero-concatenation-cost chainsabstractThis paper introduces a very fast heuristic search algorithm for unit-selection speech synthesis. The algorithm modifies commonly used Viterbi search framework by introducing zero-concatenation-cost (ZCC) chains of unit candidates that immediately neighbored in a source speech corpus. ZCC chains are preferred as they represent perfect speech segment concatenations (so there is no need to compute concatenation costs inside the chains) unless a so-called target specification is violated. The number of ZCC chains is reduced based on statistics calculated upon the synthesis of a large number of utterances. ZCC chains are then combined with single unit candidates to fill possible gaps in the sequence of candidates. The proposed method reduces the computational load of a unit selection system up to hundreds of times. According to listening tests, the quality of synthetic speech was not deteriorated. Jirí Kala, Jindrich Matousek |
ICASSP | 2 |
| 2013 | Annotation errors detection in TTS corporaabstractWe investigate the problem of automatic detection of annotation errors in single-speaker read-speech corpora used for text-to-speech (TTS) synthesis. Various word-level feature sets were used, and the performance of several detection methods based on support vector machines, extremely randomized trees, k-nearest neighbors, and the performance of novelty and outlier detection are evaluated. We show that both word- and utterance-level annotation error detections perform very well with both high precision and recall scores and with F1 measure being almost 90%, or 97%, respectively. Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2012 | Improving automatic dubbing with subtitle timing optimisation using video cut detectionabstractThis paper presents improvements to an automatic dubbing system in which text-to-speech technology is used to synthesise speech from subtitles. Spring-based subtitle timing optimisation was proposed to reduce the need for speeding up synthetic speech to fit it into corresponding subtitle slots. Video cut detection algorithm was also introduced, and the cuts were then used to prevent stretching subtitles across the cuts. Results show that after the optimisation smaller speeding-up factors are applied on synthetic speech while keeping optimised subtitle start and end times close to original positions. Jindrich Matousek, Jakub Vit |
ICASSP | 1 |
| 2011 | On the detection of pitch marks using a robust multi-phase algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka |
Speech Commun. | 2 |
| 2010 | Enhancements of viterbi search for fast unit selection synthesisabstractThe paper describes the optimisation of Viterbi search used in unit selection TTS, since with a large speech corpus necessary to achieve a high level of naturalness, the performance still suffers. To improve the search speed, the combination of sophisticated stopping schemes and pruning thresholds is employed into the baseline search. The optimised search is, moreover, extremely flexible in configuration, requiring only three intuitively comprehensible coefficients to be set. This provides the means for tuning the search depending on device resources, while it allows reaching significant performance increase. To illustrate it, several configuration scenarios, with speed-up ranging from 6 to 58 times, are presented. Their impact on speech quality is verified by CCR listening test, taking into account only the phrases with the highest number of differences when compared to the baseline search. Daniel Tihelka, Jirí Kala, Jindrich Matousek |
INTERSPEECH | 3 |
| 2009 | Identification and automatic detection of parasitic speech soundsabstractThis paper presents initial experiments with the identification and automatic detection of parasitic sounds in speech signals. The main goal of this study is to identify such sounds in the source recordings for unit-selection-based speech synthesis systems and thus to avoid their unintended usage in synthesised speech. The first part of the paper describes the phonetic analysis and identification of parasitic phenomena in recordings of two Czech speakers. In the second part, experiments with the automatic detection of parasitic sounds using HMM-based and BVM classifiers are presented. The results are encouraging, especially those for glottalization phenomena. Jindrich Matousek, Radek Skarnitzl, Pavel Machac, Jan Trmal |
INTERSPEECH | 1 |
| 2008 | Automatic pitch-synchronous phonetic segmentationabstractThis paper deals with an HMM-based automatic phonetic segmentation (APS) system and proposes to increase its performance by employing a pitch-synchronous (PS) coding scheme. Such a coding scheme uses different frames of speech throughout voiced and unvoiced speech regions and enables thus better modelling of each individual phone. The PS coding scheme is shown to outperform the traditionally utilised pitch-asynchronous (PA) coding scheme for two corpora of Czech speech (one female and one male) both in the case of a base (not-refined) APS and in the case of a CART-refined APS. Better results were observed for each of the voicing-dependent boundary types (unvoiced-unvoiced, unvoiced-voiced, voiced-unvoiced and voiced-voiced). Jindrich Matousek, Jan Romportl |
INTERSPEECH | 1 |
| 2008 | Building of a Speech Corpus Optimised for Unit Selection TTS Synthesis
Jindrich Matousek, Daniel Tihelka, Jan Romportl |
LREC | 1 |
| 2007 | F0 transformation within the voice conversion frameworkabstractV tomto článku je prezentováno několik experimentů ohledně transformace F0 v úloze konverze hlasu. Konverzní systém je založený na pravděpodobnostní transformaci parametrů a reziduální predikci. Celkem jsou zde porovnány 3 pravděpodobnostní metody pro transformaci F0. Zdenek Hanzlícek, Jindrich Matousek |
INTERSPEECH | 2 |
| 2007 | A robust multi-phase pitch-mark detection algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka |
INTERSPEECH | 2 |
| 2006 | Unit selection and its relation to symbolic prosody: a new approachabstractProtože je naším cílem zvyšovat kvalitu syntetické řeči generované TTS systémem ARTIC, adoptovali jsme přístuv dynamického výběru jednotek. Naše verze tohoto přístupu je řísená výhradně symbolickými prozodickými příznaky vysoké úrovně, které jsou v relaci s prozodií syntetické fráze skrze jevy prozodické synonymie a homonymie. Potvrdili jsme, že tento přístup generuje vysoce přirozenou řeč při zachování bohatosti prozodie. Kvalita výstupní řeči generovaná naším přístupem byla hodnoce velmi blízko přirozené řeči. Koncept prozodické synonymie a homonymie je v tomto článku tedy dále rozšířen a formálně popsán a je zde předvedena jeho důležitost pro přístup výběru jednotek. Také zde ukazujeme odlišnost našeho konceptu od konceptů převážně používaných. Navíc jsme provedli první formální experiment, který je v souladu s formálním popisem prezentovaném v tomto článku. Daniel Tihelka, Jindrich Matousek |
INTERSPEECH | 2 |
| 2006 | Design, implementation and evaluation of the Czech realistic audio-visual speech synthesis
Milos Zelezný, Zdenek Krnoul, Petr Císar, Jindrich Matousek |
Signal Process. | 4 |
| 2005 | Hybrid syllable/triphone speech synthesisabstractIn this paper, the syllable, an alternative phonetic unit to the phone, is researched in the context of speech synthesis.Several approaches to syllable modelling within the statistical approach (using hidden Markov models) to the acoustic unit inventory creation are proposed and evaluated.To be able to synthesize an arbitrary text, the syllable inventories were supplemented with triphones resulting in hybrid syllable/triphone inventories.Listening tests were accomplished both to assess the quality of the resulting synthetic speech produced using the hybrid syllable/triphone inventories and to choose the best approach to syllable modelling.The resulting synthetic speech is highly intelligible and fluent.Although the synthetic speech generated using the baseline triphone inventory was assessed slightly better, the results of the very first experiments with syllable modelling are very promising. Jindrich Matousek, Zdenek Hanzlícek, Daniel Tihelka |
INTERSPEECH | 1 |
| 2004 | Recent improvements on ARTIC: czech text-to-speech systemabstractČlánek prezentuje nejnovější vylepšení systému ARTIC, moderního českého korpusově orientovaného TTS systému. Protože jsme použili statistický přístup (skryté Markovovy modely) k vytvoření inventáře akustických jednotek, vylepšení se týkala všech jeho komponent. Vylepšeným modelováním, shlukováním, a segmentací akustických jednotek jsme dosáhli zvýšené srozumitelnosti výsledné řeči. Navrhli jsme rovněž 2 přístupy ke generování prozodických charakteristik a získali tak vyšší přirozenost syntetické řeči. Abychom zvýšili i plynulost vytvářené řeči, navrhli jsme rovněž schéma využívající více realizací každé řečové jednotky s výběrem nejvhodnějšího kandidáta on-line. Zmíníme také alternativní metodu vytváření řeči využívající harmonický model a model šumu. Implementací německého a slovenského jazykového modulu (vedle 2 českých hlasů) jsme navíc vytvořili důležitý krok směrem k vícejazyčnosti našeho TTS systému ARTIC. Jindrich Matousek, Jan Romportl, Daniel Tihelka, Zbynek Tychtl |
INTERSPEECH | 1 |
| 2004 | The Design of Czech Language Formal Listening Tests for the Evaluation of TTS Systems
Daniel Tihelka, Jindrich Matousek |
LREC | 2 |
| 2003 | Automatic segmentation for czech concatenative speech synthesis using statistical approach with boundary-specific correctionabstractThis paper deals with the problems of automatic segmentation for the purposes of Czech concatenative speech synthesis. Statistical approach to speech segmentation using HMMs is applied in the baseline system. Several improvements of this system are then proposed to get more accurate segmentation results. These enhancements mainly concern the various strategies of HMM initialization (flat-start initialization, hand-labeled or speaker independent HMM bootstrapping). Since HTK was utilized in our work, a correction of the output boundary placements is proposed to reflect speech parameterization mechanism. An objective comparison of various automatic methods and manual segmentation is performed to find out the best method. The best results were obtained for boundary-specific statistical correction of the segmentation that resulted from bootstrapping with hand-labeled HMMs (96% segmentation accuracy in tolerance region 20ms). Jindrich Matousek, Daniel Tihelka, Josef Psutka |
INTERSPEECH | 1 |
| 2001 | Design of speech corpus for text-to-speech synthesisabstractThis paper deals with the design of a speech corpus for a\nconcatenation-based text-to-speech (TTS) synthesis. Several\naspects of the design process are discussed here. We propose a\nsentence selection algorithm to choose sentences (from a large\ntext corpus) which will be read and stored in a speech corpus.\nThe selected sentences should include all possible triphones in a\nsufficient number of occurrences. Some notes on recording the\nspeech are also discussed to ensure a quality speech corpus. As\nsome popular speech synthesis techniques require knowing the\nmoments of principal excitation of vocal tract during the speech,\npitch-mark detection is also a subject of our attention. Several\nautomatic pitch-mark detection methods are discussed here and\na comparison test is performed to find out the best method. Jindrich Matousek, Josef Psutka, Jiri Kruta |
INTERSPEECH | 1 |
| 2001 | Large broadcast news and read speech corpora of spoken czech
Josef Psutka, Vlasta Radová, Ludek Müller, Jindrich Matousek, Pavel Ircing, David Graff |
INTERSPEECH | 4 |
| 2000 | ARTIC: a new Czech text-to-speech system using statistical approach to speech segment database construction
Jindrich Matousek, Josef Psutka |
INTERSPEECH | 1 |
| 1999 | Speech synthesis using HMM-based acoustic unit inventory
Jindrich Matousek |
EUROSPEECH | 1 |