Daniel Tihelka

dblp:35/3557 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-3149-2330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Homograph Disambiguation with Text-to-Text Transfer Transformer
Markéta Rezácková, Daniel Tihelka, Jindrich Matousek
INTERSPEECH2
2024 T5G2P: Text-to-Text Transfer Transformer Based Grapheme-to-Phoneme Conversion
abstract
The present paper explores the use of several deep neural network architectures to carry out a grapheme-to-phoneme (G2P) conversion, aiming to find a universal and language-independent approach to the task. The models explored are trained on whole sentences in order to automatically capture cross-word context (such as voicedness assimilation) if it exists in the given language. Four different languages, English, Czech, Russian, and German, were chosen due to their different nature and requirements for the G2P task. Ultimately, the Text-to-Text Transfer Transformer (T5) based model achieved very high conversion accuracy on all the tested languages. Also, it exceeded the accuracy reached by a similar system, when trained on a public LibriSpeech database.
Markéta Rezácková, Daniel Tihelka, Jindrich Matousek
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Ensemble of Deep Neural Network Models for MOS Prediction
abstract
Automatic evaluation of the quality of synthetic speech has the potential to serve as a cheaper and less time-consuming alternative to standard listening tests. In this paper, we present our contribution to the ongoing research: a system for automatic prediction of the mean opinion score (MOS) given by human listeners. The system was specifically developed for the recent VoiceMOS Challenge. Following the success of fusion systems in similar challenges, our contribution is an ensemble that interpolates the outputs of seven different models: four different wav2vec models, a CNN-RNN model, QuartzNet, and the LDNet baseline. During the VoiceMOS challenge, our system achieved the second-best utterance-level MSE of 0.171 and ranged from 2nd to 8th place among all 22 participating teams in terms of other evaluation metrics.
Marie Kunesová, Jindrich Matousek, Jan Lehecka, Jan Svec, Josef Michalek, Daniel Tihelka, Martin Bulín, Zdenek Hanzlícek, Markéta Rezácková
ICASSP6
2021 A Comparison of Convolutional Neural Networks for Glottal Closure Instant Detection from Raw Speech
abstract
In this paper, we continue to investigate the use of machine learning for the automatic detection of glottal closure instants (GCIs) from raw speech. We compare several deep one-dimensional convolutional neural network architectures on the same data and show that the InceptionV3 model yields the best results on the test set. On publicly available databases, the proposed 1D InceptionV3 outperforms XGBoost, a non-deep machine learning model, as well as other traditional GCI detection algorithms.
Jindrich Matousek, Daniel Tihelka
ICASSP2
2021 T5G2P: Using Text-to-Text Transfer Transformer for Grapheme-to-Phoneme Conversion
abstract
Despite the increasing popularity of end-to-end text-to-speech (TTS) systems, the correct grapheme-to-phoneme (G2P) module is still a crucial part of those relying on a phonetic input. In this paper, we, therefore, introduce a T5G2P model, a Text-to-Text Transfer Transformer (T5) neural network model which is able to convert an input text sentence into a phoneme sequence with a high accuracy. The evaluation of our trained T5 model is carried out on English and Czech, since there are different specific properties of G2P, including homograph disambiguation, cross-word assimilation and irregular pronunciation of loanwords. The paper also contains an analysis of a homographs issue in English and offers another approach to Czech phonetic transcription using the detection of pronunciation exceptions.
Markéta Rezácková, Jan Svec, Daniel Tihelka
Interspeech3
2021 Save Your Voice: Voice Banking and TTS for Anyone
Daniel Tihelka, Markéta Rezácková, Martin Gruber, Zdenek Hanzlícek, Jakub Vit, Jindrich Matousek
Interspeech1
2019 Using Extreme Gradient Boosting to Detect Glottal Closure Instants in Speech Signal
abstract
In this paper, we continue to investigate the use of classifiers for the automatic detection of glottal closure instants (GCIs) from the speech signal. We focus on extreme gradient boosting (XGB), a fast and powerful implementation of a gradient boosting algorithm. We show that XGB outperforms other classifiers, achieving GCI detection accuracy F 1 = 98.55% and AUC = 99.90%. The proposed XGB model is also shown to outperform other existing GCI detection algorithms on publicly available databases. Despite using much less training data, the performance of XGB is comparable to a deep convolutional neural network based approach, especially when it is tested on voices that were not included in the training data.
Jindrich Matousek, Daniel Tihelka
ICASSP2
2019 Unified Language-Independent DNN-Based G2P Converter
abstract
Představujeme jednotný model pro převod grafémů na fonémy založený na hlubokých neuronových sítích. Na rozdíl od obvyklých přístupů, které používají pro trénovaní slovník, používáme celé fráze, což nám umožňuje zachytit různé jazykové vlastnosti, např. spodobu znělosti přes hranici slov, bez nutnosti definovat specifika pro konkrétní jazyk. Vyhodnocení přístupu probíhá na třech různých jazycích - angličtině, češtině a ruštině. Každý z nich vyžaduje řešení specifických vlastností, a proto to obvykle vede k použití odlišných přístupů. První výsledky použití navrhovaného modelu prokazují, že je schopný se specifika jednotlivých jazyků naučit. Považujeme tedy model za jazykově nezávislý pro širokou škálu jazyků.
Markéta Juzová, Daniel Tihelka, Jakub Vit
INTERSPEECH2
2018 Glottal Closure Instant Detection from Speech Signal Using Voting Classifier and Recursive Feature Elimination
Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2018 Design and Development of Speech Corpora for Air Traffic Control Training
Lubos Smídl, Jan Svec, Daniel Tihelka, Jindrich Matousek, Jan Romportl, Pavel Ircing
LREC3
2017 WebSubDub - Experimental System for Creating High-Quality Alternative Audio Track for TV Broadcasting
Martin Gruber, Jindrich Matousek, Zdenek Hanzlícek, Jakub Vit, Daniel Tihelka
INTERSPEECH5
2017 Voice Conservation and TTS System for People Facing Total Laryngectomy
Markéta Juzová, Daniel Tihelka, Jindrich Matousek, Zdenek Hanzlícek
INTERSPEECH2
2017 Classification-Based Detection of Glottal Closure Instants from Speech Signals
Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2017 Anomaly-based annotation error detection in speech-synthesis corpora
Jindrich Matousek, Daniel Tihelka
Comput. Speech Lang.2
2016 Voting Detector: A Combination of Anomaly Detectors to Reveal Annotation Errors in TTS Corpora
Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2015 Anomaly-based annotation errors detection in TTS corpora
Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2013 Annotation errors detection in TTS corpora
abstract
We investigate the problem of automatic detection of annotation errors in single-speaker read-speech corpora used for text-to-speech (TTS) synthesis. Various word-level feature sets were used, and the performance of several detection methods based on support vector machines, extremely randomized trees, k-nearest neighbors, and the performance of novelty and outlier detection are evaluated. We show that both word- and utterance-level annotation error detections perform very well with both high precision and recall scores and with F1 measure being almost 90%, or 97%, respectively.
Jindrich Matousek, Daniel Tihelka
INTERSPEECH2
2011 On the detection of pitch marks using a robust multi-phase algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka
Speech Commun.3
2010 Enhancements of viterbi search for fast unit selection synthesis
abstract
The paper describes the optimisation of Viterbi search used in unit selection TTS, since with a large speech corpus necessary to achieve a high level of naturalness, the performance still suffers. To improve the search speed, the combination of sophisticated stopping schemes and pruning thresholds is employed into the baseline search. The optimised search is, moreover, extremely flexible in configuration, requiring only three intuitively comprehensible coefficients to be set. This provides the means for tuning the search depending on device resources, while it allows reaching significant performance increase. To illustrate it, several configuration scenarios, with speed-up ranging from 6 to 58 times, are presented. Their impact on speech quality is verified by CCR listening test, taking into account only the phrases with the highest number of differences when compared to the baseline search.
Daniel Tihelka, Jirí Kala, Jindrich Matousek
INTERSPEECH1
2009 Exploring automatic similarity measures for unit selection tuning
abstract
The present paper focuses on the current handling of target fea-tures in the unit selection approach basically requiring huge cor-pora. In the paper there are outlined possible solutions based on measuring (dis)similarity among prosodic patterns. As the start of research, the feasibility of (dis)similarity estimation is ex-amined on several intuitively chosen measures of acoustic sig-nal which are correlated to perceived similarity obtained from a large-scale listening test. Index Terms: speech synthesis, unit selection, target features, prosodic patterns, perceived similarity, signal similarity, multi-
Daniel Tihelka, Jan Romportl
INTERSPEECH1
2008 Building of a Speech Corpus Optimised for Unit Selection TTS Synthesis
Jindrich Matousek, Daniel Tihelka, Jan Romportl
LREC2
2007 A robust multi-phase pitch-mark detection algorithm
Milan Legát, Jindrich Matousek, Daniel Tihelka
INTERSPEECH3
2006 Unit selection and its relation to symbolic prosody: a new approach
abstract
Protože je naším cílem zvyšovat kvalitu syntetické řeči generované TTS systémem ARTIC, adoptovali jsme přístuv dynamického výběru jednotek. Naše verze tohoto přístupu je řísená výhradně symbolickými prozodickými příznaky vysoké úrovně, které jsou v relaci s prozodií syntetické fráze skrze jevy prozodické synonymie a homonymie. Potvrdili jsme, že tento přístup generuje vysoce přirozenou řeč při zachování bohatosti prozodie. Kvalita výstupní řeči generovaná naším přístupem byla hodnoce velmi blízko přirozené řeči. Koncept prozodické synonymie a homonymie je v tomto článku tedy dále rozšířen a formálně popsán a je zde předvedena jeho důležitost pro přístup výběru jednotek. Také zde ukazujeme odlišnost našeho konceptu od konceptů převážně používaných. Navíc jsme provedli první formální experiment, který je v souladu s formálním popisem prezentovaném v tomto článku.
Daniel Tihelka, Jindrich Matousek
INTERSPEECH1
2005 Hybrid syllable/triphone speech synthesis
abstract
In this paper, the syllable, an alternative phonetic unit to the phone, is researched in the context of speech synthesis.Several approaches to syllable modelling within the statistical approach (using hidden Markov models) to the acoustic unit inventory creation are proposed and evaluated.To be able to synthesize an arbitrary text, the syllable inventories were supplemented with triphones resulting in hybrid syllable/triphone inventories.Listening tests were accomplished both to assess the quality of the resulting synthetic speech produced using the hybrid syllable/triphone inventories and to choose the best approach to syllable modelling.The resulting synthetic speech is highly intelligible and fluent.Although the synthetic speech generated using the baseline triphone inventory was assessed slightly better, the results of the very first experiments with syllable modelling are very promising.
Jindrich Matousek, Zdenek Hanzlícek, Daniel Tihelka
INTERSPEECH3
2005 Symbolic prosody driven unit selection for highly natural synthetic speech
Daniel Tihelka
INTERSPEECH1
2004 Recent improvements on ARTIC: czech text-to-speech system
abstract
Článek prezentuje nejnovější vylepšení systému ARTIC, moderního českého korpusově orientovaného TTS systému. Protože jsme použili statistický přístup (skryté Markovovy modely) k vytvoření inventáře akustických jednotek, vylepšení se týkala všech jeho komponent. Vylepšeným modelováním, shlukováním, a segmentací akustických jednotek jsme dosáhli zvýšené srozumitelnosti výsledné řeči. Navrhli jsme rovněž 2 přístupy ke generování prozodických charakteristik a získali tak vyšší přirozenost syntetické řeči. Abychom zvýšili i plynulost vytvářené řeči, navrhli jsme rovněž schéma využívající více realizací každé řečové jednotky s výběrem nejvhodnějšího kandidáta on-line. Zmíníme také alternativní metodu vytváření řeči využívající harmonický model a model šumu. Implementací německého a slovenského jazykového modulu (vedle 2 českých hlasů) jsme navíc vytvořili důležitý krok směrem k vícejazyčnosti našeho TTS systému ARTIC.
Jindrich Matousek, Jan Romportl, Daniel Tihelka, Zbynek Tychtl
INTERSPEECH3
2004 The Design of Czech Language Formal Listening Tests for the Evaluation of TTS Systems
Daniel Tihelka, Jindrich Matousek
LREC1
2003 Automatic segmentation for czech concatenative speech synthesis using statistical approach with boundary-specific correction
abstract
This paper deals with the problems of automatic segmentation for the purposes of Czech concatenative speech synthesis. Statistical approach to speech segmentation using HMMs is applied in the baseline system. Several improvements of this system are then proposed to get more accurate segmentation results. These enhancements mainly concern the various strategies of HMM initialization (flat-start initialization, hand-labeled or speaker independent HMM bootstrapping). Since HTK was utilized in our work, a correction of the output boundary placements is proposed to reflect speech parameterization mechanism. An objective comparison of various automatic methods and manual segmentation is performed to find out the best method. The best results were obtained for boundary-specific statistical correction of the segmentation that resulted from bootstrapping with hand-labeled HMMs (96% segmentation accuracy in tolerance region 20ms).
Jindrich Matousek, Daniel Tihelka, Josef Psutka
INTERSPEECH2