Jindrich Matousek

dblp:88/999 · DBLP profile ↗
← Back
41ranked-venue papers
18as first author
8since 2021 · last 2024
0000-0002-7408-7730ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 16 first-author · 6 since 2021Artificial intelligence and machine learning · 33 · 15 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Zero-shot Out-of-domain is No Joke: Lessons Learned in the VoiceMOS 2023 MOS Prediction Challenge
Marie Kunesová, Jan Lehecka, Josef Michalek, Jindrich Matousek, Jan Svec
INTERSPEECH4
2024 Homograph Disambiguation with Text-to-Text Transfer Transformer
Markéta Rezácková, Daniel Tihelka, Jindrich Matousek
INTERSPEECH3
2024 Using LSTM neural networks for cross-lingual phonetic speech segmentation with an iterative correction procedure
abstract
Abstract This article describes experiments on speech segmentation using long short‐term memory recurrent neural networks. The main part of the paper deals with multi‐lingual and cross‐lingual segmentation, that is, it is performed on a language different from the one on which the model was trained. The experimental data involves large Czech, English, German, and Russian speech corpora designated for speech synthesis. For optimal multi‐lingual modeling, a compact phonetic alphabet was proposed by sharing and clustering phones of particular languages. Many experiments were performed exploring various experimental conditions and data combinations. We proposed a simple procedure that iteratively adapts the inaccurate default model to the new voice/language. The segmentation accuracy was evaluated by comparison with reference segmentation created by a well‐tuned hidden Markov model‐based framework with additional manual corrections. The resulting segmentation was also employed in a unit selection text‐to‐speech system. The generated speech quality was compared with the reference segmentation by a preference listening test.
Zdenek Hanzlícek, Jindrich Matousek, Jakub Vit
Comput. Intell.2
2024 T5G2P: Text-to-Text Transfer Transformer Based Grapheme-to-Phoneme Conversion
abstract
The present paper explores the use of several deep neural network architectures to carry out a grapheme-to-phoneme (G2P) conversion, aiming to find a universal and language-independent approach to the task. The models explored are trained on whole sentences in order to automatically capture cross-word context (such as voicedness assimilation) if it exists in the given language. Four different languages, English, Czech, Russian, and German, were chosen due to their different nature and requirements for the G2P task. Ultimately, the Text-to-Text Transfer Transformer (T5) based model achieved very high conversion accuracy on all the tested languages. Also, it exceeded the accuracy reached by a similar system, when trained on a public LibriSpeech database.
Markéta Rezácková, Daniel Tihelka, Jindrich Matousek
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Ensemble of Deep Neural Network Models for MOS Prediction
abstract
Automatic evaluation of the quality of synthetic speech has the potential to serve as a cheaper and less time-consuming alternative to standard listening tests. In this paper, we present our contribution to the ongoing research: a system for automatic prediction of the mean opinion score (MOS) given by human listeners. The system was specifically developed for the recent VoiceMOS Challenge. Following the success of fusion systems in similar challenges, our contribution is an ensemble that interpolates the outputs of seven different models: four different wav2vec models, a CNN-RNN model, QuartzNet, and the LDNet baseline. During the VoiceMOS challenge, our system achieved the second-best utterance-level MSE of 0.171 and ranged from 2nd to 8th place among all 22 participating teams in terms of other evaluation metrics.
Marie Kunesová, Jindrich Matousek, Jan Lehecka, Jan Svec, Josef Michalek, Daniel Tihelka, Martin Bulín, Zdenek Hanzlícek, Markéta Rezácková
ICASSP2
2023 Neural Speech Synthesis with Enriched Phrase Boundaries
Marie Kunesová, Jindrich Matousek
INTERSPEECH2
2021 A Comparison of Convolutional Neural Networks for Glottal Closure Instant Detection from Raw Speech
abstract
In this paper, we continue to investigate the use of machine learning for the automatic detection of glottal closure instants (GCIs) from raw speech. We compare several deep one-dimensional convolutional neural network architectures on the same data and show that the InceptionV3 model yields the best results on the test set. On publicly available databases, the proposed 1D InceptionV3 outperforms XGBoost, a non-deep machine learning model, as well as other traditional GCI detection algorithms.
Jindrich Matousek, Daniel Tihelka
ICASSP1
2021 Save Your Voice: Voice Banking and TTS for Anyone
Daniel Tihelka, Markéta Rezácková, Martin Gruber, Zdenek Hanzlícek, Jakub Vit, Jindrich Matousek
Interspeech6
2019 Using Extreme Gradient Boosting to Detect Glottal Closure Instants in Speech Signal
abstract
In this paper, we continue to investigate the use of classifiers for the automatic detection of glottal closure instants (GCIs) from the speech signal. We focus on extreme gradient boosting (XGB), a fast and powerful implementation of a gradient boosting algorithm. We show that XGB outperforms other classifiers, achieving GCI detection accuracy F 1 = 98.55% and AUC = 99.90%. The proposed XGB model is also shown to outperform other existing GCI detection algorithms on publicly available databases. Despite using much less training data, the performance of XGB is comparable to a deep convolutional neural network based approach, especially when it is tested on voices that were not included in the training data.
Jindrich Matousek, Daniel Tihelka
ICASSP1
2019 Framework for Conducting Tasks Requiring Human Assessment
Martin Gruber, Adam Chýlek, Jindrich Matousek
INTERSPEECH3
2019 Web-Based Speech Synthesis Editor
Martin Gruber, Jakub Vit, Jindrich Matousek
INTERSPEECH3
2018 On the Analysis of Training Data for Wavenet-Based Speech Synthesis
abstract
In this paper, we analyze how much, how consistent and how accurate data WaveNet-based speech synthesis method needs to be able to generate speech of good quality. We do this by adding artificial noise to the description of our training data and observing how well WaveNet trains and produces speech. More specifically, we add noise to both phonetic segmentation and annotation accuracy, and we also reduce the size of training data by using a fewer number of sentences during training of a WaveNet model. We conducted MUSHRA listening tests and used objective measures to track speech quality within the conducted experiments. We show that WaveNet retains high quality even after adding a small amount of noise (up to 10%) to phonetic segmentation and annotation. A small degradation of speech quality was observed for our WaveNet configuration when only 3 hours of training data were used.
Jakub Vit, Zdenek Hanzlícek, Jindrich Matousek
ICASSP3
2018 Glottal Closure Instant Detection from Speech Signal Using Voting Classifier and Recursive Feature Elimination
Jindrich Matousek, Daniel Tihelka
INTERSPEECH1
2018 Design and Development of Speech Corpora for Air Traffic Control Training
Lubos Smídl, Jan Svec, Daniel Tihelka, Jindrich Matousek, Jan Romportl, Pavel Ircing
LREC4
2017 WebSubDub - Experimental System for Creating High-Quality Alternative Audio Track for TV Broadcasting
Martin Gruber, Jindrich Matousek, Zdenek Hanzlícek, Jakub Vit, Daniel Tihelka
INTERSPEECH2
2017 Voice Conservation and TTS System for People Facing Total Laryngectomy
Markéta Juzová, Daniel Tihelka, Jindrich Matousek, Zdenek Hanzlícek
INTERSPEECH3
2017 Classification-Based Detection of Glottal Closure Instants from Speech Signals
Jindrich Matousek, Daniel Tihelka
INTERSPEECH1
2017 Anomaly-based annotation error detection in speech-synthesis corpora
Jindrich Matousek, Daniel Tihelka
Comput. Speech Lang.1
2016 ARET - Automatic Reading of Educational Texts for Visually Impaired Students
Martin Gruber, Jindrich Matousek, Zdenek Hanzlícek, Zdenek Krnoul, Zbynek Zajíc
INTERSPEECH2
2016 Voting Detector: A Combination of Anomaly Detectors to Reveal Annotation Errors in TTS Corpora
Jindrich Matousek, Daniel Tihelka
INTERSPEECH1
2015 Anomaly-based annotation errors detection in TTS corpora
Jindrich Matousek, Daniel Tihelka
INTERSPEECH1
2014 Very fast unit selection using Viterbi search with zero-concatenation-cost chains
abstract
This paper introduces a very fast heuristic search algorithm for unit-selection speech synthesis. The algorithm modifies commonly used Viterbi search framework by introducing zero-concatenation-cost (ZCC) chains of unit candidates that immediately neighbored in a source speech corpus. ZCC chains are preferred as they represent perfect speech segment concatenations (so there is no need to compute concatenation costs inside the chains) unless a so-called target specification is violated. The number of ZCC chains is reduced based on statistics calculated upon the synthesis of a large number of utterances. ZCC chains are then combined with single unit candidates to fill possible gaps in the sequence of candidates. The proposed method reduces the computational load of a unit selection system up to hundreds of times. According to listening tests, the quality of synthetic speech was not deteriorated.
Jirí Kala, Jindrich Matousek
ICASSP2
2013 Annotation errors detection in TTS corpora
abstract
We investigate the problem of automatic detection of annotation errors in single-speaker read-speech corpora used for text-to-speech (TTS) synthesis. Various word-level feature sets were used, and the performance of several detection methods based on support vector machines, extremely randomized trees, k-nearest neighbors, and the performance of novelty and outlier detection are evaluated. We show that both word- and utterance-level annotation error detections perform very well with both high precision and recall scores and with F1 measure being almost 90%, or 97%, respectively.
Jindrich Matousek, Daniel Tihelka
INTERSPEECH1
2012 Improving automatic dubbing with subtitle timing optimisation using video cut detection
abstract
This paper presents improvements to an automatic dubbing system in which text-to-speech technology is used to synthesise speech from subtitles. Spring-based subtitle timing optimisation was proposed to reduce the need for speeding up synthetic speech to fit it into corresponding subtitle slots. Video cut detection algorithm was also introduced, and the cuts were then used to prevent stretching subtitles across the cuts. Results show that after the optimisation smaller speeding-up factors are applied on synthetic speech while keeping optimised subtitle start and end times close to original positions.
Jindrich Matousek, Jakub Vit
ICASSP1
2011 On the detection of pitch marks using a robust multi-phase algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka
Speech Commun.2
2010 Enhancements of viterbi search for fast unit selection synthesis
abstract
The paper describes the optimisation of Viterbi search used in unit selection TTS, since with a large speech corpus necessary to achieve a high level of naturalness, the performance still suffers. To improve the search speed, the combination of sophisticated stopping schemes and pruning thresholds is employed into the baseline search. The optimised search is, moreover, extremely flexible in configuration, requiring only three intuitively comprehensible coefficients to be set. This provides the means for tuning the search depending on device resources, while it allows reaching significant performance increase. To illustrate it, several configuration scenarios, with speed-up ranging from 6 to 58 times, are presented. Their impact on speech quality is verified by CCR listening test, taking into account only the phrases with the highest number of differences when compared to the baseline search.
Daniel Tihelka, Jirí Kala, Jindrich Matousek
INTERSPEECH3
2009 Identification and automatic detection of parasitic speech sounds
abstract
This paper presents initial experiments with the identification and automatic detection of parasitic sounds in speech signals. The main goal of this study is to identify such sounds in the source recordings for unit-selection-based speech synthesis systems and thus to avoid their unintended usage in synthesised speech. The first part of the paper describes the phonetic analysis and identification of parasitic phenomena in recordings of two Czech speakers. In the second part, experiments with the automatic detection of parasitic sounds using HMM-based and BVM classifiers are presented. The results are encouraging, especially those for glottalization phenomena.
Jindrich Matousek, Radek Skarnitzl, Pavel Machac, Jan Trmal
INTERSPEECH1
2008 Automatic pitch-synchronous phonetic segmentation
abstract
This paper deals with an HMM-based automatic phonetic segmentation (APS) system and proposes to increase its performance by employing a pitch-synchronous (PS) coding scheme. Such a coding scheme uses different frames of speech throughout voiced and unvoiced speech regions and enables thus better modelling of each individual phone. The PS coding scheme is shown to outperform the traditionally utilised pitch-asynchronous (PA) coding scheme for two corpora of Czech speech (one female and one male) both in the case of a base (not-refined) APS and in the case of a CART-refined APS. Better results were observed for each of the voicing-dependent boundary types (unvoiced-unvoiced, unvoiced-voiced, voiced-unvoiced and voiced-voiced).
Jindrich Matousek, Jan Romportl
INTERSPEECH1
2008 Building of a Speech Corpus Optimised for Unit Selection TTS Synthesis
Jindrich Matousek, Daniel Tihelka, Jan Romportl
LREC1
2007 F0 transformation within the voice conversion framework
abstract
V tomto článku je prezentováno několik experimentů ohledně transformace F0 v úloze konverze hlasu. Konverzní systém je založený na pravděpodobnostní transformaci parametrů a reziduální predikci. Celkem jsou zde porovnány 3 pravděpodobnostní metody pro transformaci F0.
Zdenek Hanzlícek, Jindrich Matousek
INTERSPEECH2
2007 A robust multi-phase pitch-mark detection algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2006 Unit selection and its relation to symbolic prosody: a new approach
abstract
Protože je naším cílem zvyšovat kvalitu syntetické řeči generované TTS systémem ARTIC, adoptovali jsme přístuv dynamického výběru jednotek. Naše verze tohoto přístupu je řísená výhradně symbolickými prozodickými příznaky vysoké úrovně, které jsou v relaci s prozodií syntetické fráze skrze jevy prozodické synonymie a homonymie. Potvrdili jsme, že tento přístup generuje vysoce přirozenou řeč při zachování bohatosti prozodie. Kvalita výstupní řeči generovaná naším přístupem byla hodnoce velmi blízko přirozené řeči. Koncept prozodické synonymie a homonymie je v tomto článku tedy dále rozšířen a formálně popsán a je zde předvedena jeho důležitost pro přístup výběru jednotek. Také zde ukazujeme odlišnost našeho konceptu od konceptů převážně používaných. Navíc jsme provedli první formální experiment, který je v souladu s formálním popisem prezentovaném v tomto článku.
Daniel Tihelka, Jindrich Matousek
INTERSPEECH2
2006 Design, implementation and evaluation of the Czech realistic audio-visual speech synthesis
Milos Zelezný, Zdenek Krnoul, Petr Císar, Jindrich Matousek
Signal Process.4
2005 Hybrid syllable/triphone speech synthesis
abstract
In this paper, the syllable, an alternative phonetic unit to the phone, is researched in the context of speech synthesis.Several approaches to syllable modelling within the statistical approach (using hidden Markov models) to the acoustic unit inventory creation are proposed and evaluated.To be able to synthesize an arbitrary text, the syllable inventories were supplemented with triphones resulting in hybrid syllable/triphone inventories.Listening tests were accomplished both to assess the quality of the resulting synthetic speech produced using the hybrid syllable/triphone inventories and to choose the best approach to syllable modelling.The resulting synthetic speech is highly intelligible and fluent.Although the synthetic speech generated using the baseline triphone inventory was assessed slightly better, the results of the very first experiments with syllable modelling are very promising.
Jindrich Matousek, Zdenek Hanzlícek, Daniel Tihelka
INTERSPEECH1
2004 Recent improvements on ARTIC: czech text-to-speech system
abstract
Článek prezentuje nejnovější vylepšení systému ARTIC, moderního českého korpusově orientovaného TTS systému. Protože jsme použili statistický přístup (skryté Markovovy modely) k vytvoření inventáře akustických jednotek, vylepšení se týkala všech jeho komponent. Vylepšeným modelováním, shlukováním, a segmentací akustických jednotek jsme dosáhli zvýšené srozumitelnosti výsledné řeči. Navrhli jsme rovněž 2 přístupy ke generování prozodických charakteristik a získali tak vyšší přirozenost syntetické řeči. Abychom zvýšili i plynulost vytvářené řeči, navrhli jsme rovněž schéma využívající více realizací každé řečové jednotky s výběrem nejvhodnějšího kandidáta on-line. Zmíníme také alternativní metodu vytváření řeči využívající harmonický model a model šumu. Implementací německého a slovenského jazykového modulu (vedle 2 českých hlasů) jsme navíc vytvořili důležitý krok směrem k vícejazyčnosti našeho TTS systému ARTIC.
Jindrich Matousek, Jan Romportl, Daniel Tihelka, Zbynek Tychtl
INTERSPEECH1
2004 The Design of Czech Language Formal Listening Tests for the Evaluation of TTS Systems
Daniel Tihelka, Jindrich Matousek
LREC2
2003 Automatic segmentation for czech concatenative speech synthesis using statistical approach with boundary-specific correction
abstract
This paper deals with the problems of automatic segmentation for the purposes of Czech concatenative speech synthesis. Statistical approach to speech segmentation using HMMs is applied in the baseline system. Several improvements of this system are then proposed to get more accurate segmentation results. These enhancements mainly concern the various strategies of HMM initialization (flat-start initialization, hand-labeled or speaker independent HMM bootstrapping). Since HTK was utilized in our work, a correction of the output boundary placements is proposed to reflect speech parameterization mechanism. An objective comparison of various automatic methods and manual segmentation is performed to find out the best method. The best results were obtained for boundary-specific statistical correction of the segmentation that resulted from bootstrapping with hand-labeled HMMs (96% segmentation accuracy in tolerance region 20ms).
Jindrich Matousek, Daniel Tihelka, Josef Psutka
INTERSPEECH1
2001 Design of speech corpus for text-to-speech synthesis
abstract
This paper deals with the design of a speech corpus for a\nconcatenation-based text-to-speech (TTS) synthesis. Several\naspects of the design process are discussed here. We propose a\nsentence selection algorithm to choose sentences (from a large\ntext corpus) which will be read and stored in a speech corpus.\nThe selected sentences should include all possible triphones in a\nsufficient number of occurrences. Some notes on recording the\nspeech are also discussed to ensure a quality speech corpus. As\nsome popular speech synthesis techniques require knowing the\nmoments of principal excitation of vocal tract during the speech,\npitch-mark detection is also a subject of our attention. Several\nautomatic pitch-mark detection methods are discussed here and\na comparison test is performed to find out the best method.
Jindrich Matousek, Josef Psutka, Jiri Kruta
INTERSPEECH1
2001 Large broadcast news and read speech corpora of spoken czech
Josef Psutka, Vlasta Radová, Ludek Müller, Jindrich Matousek, Pavel Ircing, David Graff
INTERSPEECH4
2000 ARTIC: a new Czech text-to-speech system using statistical approach to speech segment database construction
Jindrich Matousek, Josef Psutka
INTERSPEECH1
1999 Speech synthesis using HMM-based acoustic unit inventory
Jindrich Matousek
EUROSPEECH1