EDBT 2026 Demo / reviewers in the wild / expert
Tamás Gábor Csapó
dblp:24/8420
· DBLP profile ↗
30ranked-venue papers
10as first author
7since 2021 · last 2023
0000-0003-4375-7524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 24 · 10 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Raw Ultrasound-Based Phonetic Segments Classification Via Mask ModelingabstractUltrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep models rely on a large amount of manually labeled data, which is often difficult to obtain in practical settings. To address this issue, this paper focuses on how to utilize a large amount of unlabeled UTI data to improve the performance of UTI classification task. Specifically, we explore self-supervised learning with masking modeling strategy. By predicting the masked part, our pre-trained model enables the neural network to infer contextual information. Then, we fine-tune the pre-trained model with a small amount of labeled data. Compared with the previous competing algorithms, our method can improve the classification accuracy by an average of 13.33% in four different scenarios. Kang You, Bo Liu 0014, Kele Xu, Yunsheng Xiong, Qisheng Xu, Ming Feng, Tamás Gábor Csapó, Boqing Zhu |
ICASSP | 7 |
| 2023 | Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks
László Tóth 0001, Amin Honarmandi Shandiz, Gábor Gosztolya, Tamás Gábor Csapó |
INTERSPEECH | 4 |
| 2023 | Towards Ultrasound Tongue Image prediction from EEG during speech productionabstractPrevious initial research has already been carried out to propose speech-based BCI using brain signals (e.g.non-invasive EEG and invasive sEEG / ECoG), but there is a lack of combined methods that investigate non-invasive brain, articulation, and speech signals together and analyze the cognitive processes in the brain, the kinematics of the articulatory movement and the resulting speech signal.In this paper, we describe our multimodal (electroencephalography, ultrasound tongue imaging, and speech) analysis and synthesis experiments, as a feasibility study.We extend the analysis of brain signals recorded during speech production with ultrasound-based articulation data.From the brain signal measured with EEG, we predict ultrasound images of the tongue with a fully connected deep neural network.The results show that there is a weak but noticeable relationship between EEG and ultrasound tongue images, i.e. the network can differentiate articulated speech and neutral tongue position. Tamás Gábor Csapó, Frigyes Viktor Arthur, Péter Nagy, Ádám Boncz |
INTERSPEECH | 1 |
| 2023 | Investigations on speaker adaptation using a continuous vocoder within recurrent neural network based text-to-speech synthesisabstractAbstract This paper presents an investigation of speaker adaptation using a continuous vocoder for parametric text-to-speech (TTS) synthesis. In purposes that demand low computational complexity, conventional vocoder-based statistical parametric speech synthesis can be preferable. While capable of remarkable naturalness, recent neural vocoders nonetheless fall short of the criteria for real-time synthesis. We investigate our former continuous vocoder, in which the excitation is characterized employing two one-dimensional parameters: Maximum Voiced Frequency and continuous fundamental frequency (F0). We show that an average voice can be trained for deep neural network-based TTS utilizing data from nine English speakers. We did speaker adaptation experiments for each target speaker with 400 utterances (approximately 14 minutes). We showed an apparent enhancement in the quality and naturalness of synthesized speech compared to our previous work by utilizing the recurrent neural network topologies. According to the objective studies (Mel-Cepstral Distortion and F0 correlation), the quality of speaker adaptation using Continuous Vocoder-based DNN-TTS is slightly better than the WORLD Vocoder-based baseline. The subjective MUSHRA-like test results also showed that our speaker adaptation technique is almost as natural as the WORLD vocoder using Gated Recurrent Unit and Long Short Term Memory networks. The proposed vocoder, being capable of real-time synthesis, can be used for applications which need fast synthesis speed. Ali Raheem Mandeel, Mohammed Salah Al-Radhi, Tamás Gábor Csapó |
Multim. Tools Appl. | 3 |
| 2021 | Continuous Wavelet Vocoder-Based Decomposition of Parametric Speech Waveform SynthesisabstractTo date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models that nearly resembles the human voice, has to generate a waveform in a time consuming sequential manner with an extremely complex structure of its neural networks. Mohammed Salah Al-Radhi, Tamás Gábor Csapó, Csaba Zainkó, Géza Németh |
Interspeech | 2 |
| 2021 | Neural Speaker Embeddings for Ultrasound-Based Silent Speech InterfacesabstractArticulatory-to-acoustic mapping seeks to reconstruct speech from a recording of the articulatory movements, for example, an ultrasound video. Just like speech signals, these recordings represent not only the linguistic content, but are also highly specific to the actual speaker. Hence, due to the lack of multi-speaker data sets, researchers have so far concentrated on speaker-dependent modeling. Here, we present multi-speaker experiments using the recently published TaL80 corpus. To model speaker characteristics, we adjusted the x-vector framework popular in speech processing to operate with ultrasound tongue videos. Next, we performed speaker recognition experiments using 50 speakers from the corpus. Then, we created speaker embedding vectors and evaluated them on the remaining speakers. Finally, we examined how the embedding vector influences the accuracy of our ultrasound-to-speech conversion network in a multi-speaker scenario. In the experiments we attained speaker recognition error rates below 3%, and we also found that the embedding vectors generalize nicely to unseen speakers. Our first attempt to apply them in a multi-speaker silent speech framework brought about a marginal reduction in the error rate of the spectral estimation step. Amin Honarmandi Shandiz, László Tóth 0001, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó |
Interspeech | 5 |
| 2021 | Noise and acoustic modeling with waveform generator in text-to-speech and neutral speech conversionabstractAbstract This article focuses on developing a system for high-quality synthesized and converted speech by addressing three fundamental principles. Although the noise-like component in the state-of-the-art parametric vocoders (for example, STRAIGHT) is often not accurate enough, a novel analytical approach for modeling unvoiced excitations using a temporal envelope is proposed. Discrete All Pole, Frequency Domain Linear Prediction, Low Pass Filter, and True envelopes are firstly studied and applied to the noise excitation signal in our continuous vocoder. Second, we build a deep learning model based text–to–speech (TTS) which converts written text into human-like speech with a feed-forward and several sequence-to-sequence models (long short-term memory, gated recurrent unit, and hybrid model). Third, a new voice conversion system is proposed using a continuous fundamental frequency to provide accurate time-aligned voiced segments. The results have been evaluated in terms of objective measures and subjective listening tests. Experimental results showed that the proposed models achieved the highest speaker similarity and better quality compared with the other conventional methods. Mohammed Salah Al-Radhi, Tamás Gábor Csapó, Géza Németh |
Multim. Tools Appl. | 2 |
| 2020 | Speaker Dependent Articulatory-to-Acoustic Mapping Using Real-Time MRI of the Vocal TractabstractArticulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for this purpose. The advantage of MRI is that it has a high `relative' spatial resolution: it can capture not only lingual, labial and jaw motion, but also the velum and the pharyngeal region, which is typically not possible with other techniques. In the current paper, we train various DNNs (fully connected, convolutional and recurrent neural networks) for articulatory-to-speech conversion, using rtMRI as input, in a speaker-specific way. We use two male and two female speakers of the USC-TIMIT articulatory database, each of them uttering 460 sentences. We evaluate the results with objective (Normalized MSE and MCD) and subjective measures (perceptual test) and show that CNN-LSTM networks are preferred which take multiple images as input, and achieve MCD scores between 2.8-4.5 dB. In the experiments, we find that the predictions of speaker `m1' are significantly weaker than other speakers. We show that this is caused by the fact that 74% of the recordings of speaker `m1' are out of sync. Tamás Gábor Csapó |
INTERSPEECH | 1 |
| 2020 | Speaker Dependent Acoustic-to-Articulatory Inversion Using Real-Time MRI of the Vocal TractabstractAcoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring.Most earlier inversion studies are based on point-tracking articulatory techniques (e.g.EMA or XRMB).The advantage of rtMRI is that it provides dynamic information about the full midsagittal plane of the upper airway, with a high 'relative' spatial resolution.In this work, we estimated midsagittal rtMRI images of the vocal tract for speaker dependent AAI, using MGC-LSP spectral features as input.We applied FC-DNNs, CNNs and recurrent neural networks, and have shown that LSTMs are the most suitable for this task.As objective evaluation we measured normalized MSE, Structural Similarity Index (SSIM) and its complex wavelet version (CW-SSIM).The results indicate that the combination of FC-DNNs and LSTMs can achieve smooth generated MR images of the vocal tract, which are similar to the original MRI recordings (average CW-SSIM: 0.94). Tamás Gábor Csapó |
INTERSPEECH | 1 |
| 2020 | Quantification of Transducer Misalignment in Ultrasound Tongue ImagingabstractIn speech production research, different imaging modalities have been employed to obtain accurate information about the movement and shaping of the vocal tract. Ultrasound is an affordable and non-invasive imaging modality with relatively high temporal and spatial resolution to study the dynamic behavior of tongue during speech production. However, a long-standing problem for ultrasound tongue imaging is the transducer misalignment during longer data recording sessions. In this paper, we propose a simple, yet effective, misalignment quantification approach. The analysis employs MSE distance and two similarity measurement metrics to identify the relative displacement between the chin and the transducer. We visualize these measures as a function of the timestamp of the utterances. Extensive experiments are conducted on a Hungarian and Scottish English child dataset. The results suggest that large values of Mean Square Error (MSE) and small values of Structural Similarity Index (SSIM) and Complex Wavelet SSIM indicate corruptions or issues during the data recordings, which can either be caused by transducer misalignment or lack of gel. Tamás Gábor Csapó, Kele Xu |
INTERSPEECH | 1 |
| 2020 | Ultrasound-Based Articulatory-to-Acoustic Mapping with WaveGlow Speech SynthesisabstractFor articulatory-to-acoustic mapping using deep neural networks, typically spectral and excitation parameters of vocoders have been used as the training targets. However, vocoding often results in buzzy and muffled final speech quality. Therefore, in this paper on ultrasound-based articulatory-to-acoustic conversion, we use a flow-based neural vocoder (WaveGlow) pre-trained on a large amount of English and Hungarian speech data. The inputs of the convolutional neural network are ultrasound tongue images. The training target is the 80-dimensional mel-spectrogram, which results in a finer detailed spectral representation than the previously used 25-dimensional Mel-Generalized Cepstrum. From the output of the ultrasound-to-mel-spectrogram prediction, WaveGlow inference results in synthesized speech. We compare the proposed WaveGlow-based system with a continuous vocoder which does not use strict voiced/unvoiced decision when predicting F0. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the WaveGlow neural vocoder produces significantly more natural synthesized speech than the baseline system. Besides, the advantage of WaveGlow is that F0 is included in the mel-spectrogram representation, and it is not necessary to predict the excitation separately. Tamás Gábor Csapó, Csaba Zainkó, László Tóth 0001, Gábor Gosztolya, Alexandra Markó |
INTERSPEECH | 1 |
| 2020 | A continuous vocoder for statistical parametric speech synthesis and its evaluation using an audio-visual phonetically annotated Arabic corpus
Mohammed Salah Al-Radhi, Omnia Abdo, Tamás Gábor Csapó, Sherif M. Abdou, Géza Németh, Mervat Fashal |
Comput. Speech Lang. | 3 |
| 2019 | RNN-based speech synthesis using a continuous sinusoidal modelabstractRecently in statistical parametric speech synthesis, we proposed a continuous sinusoidal model (CSM) using continuous F0 (contF0) in combination with Maximum Voiced Frequency (MVF), which was successfully giving state-of-the-art vocoders performance (e.g. similar to STRAIGHT) in synthesized speech. In this paper, we address the use of sequence-to-sequence modeling with recurrent neural networks (RNNs). Bidirectional long short-term memory (Bi-LSTM) is investigated and applied using our CSM to model contF0, MVF, and Mel-Generalized Cepstrum (MGC) for more natural sounding synthesized speech. For refining the output of the contF0 estimation, post-processing based on time-warping approach is applied to reduce the unwanted voiced component of the unvoiced speech sounds, resulting in an enhanced contF0 track. The overall conclusion is covered by objective evaluation and subjective listening test, showing that the proposed framework provides satisfactory results in terms of naturalness and intelligibility, and is comparable to the high-quality WORLD model based RNNs. Mohammed Salah Al-Radhi, Tamás Gábor Csapó, Géza Németh |
IJCNN | 2 |
| 2019 | Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech InterfacesabstractWhen using ultrasound video as input, Deep Neural Network-based Silent Speech Interfaces usually rely on the whole image to estimate the spectral parameters required for the speech synthesis step. Although this approach is quite straightforward, and it permits the synthesis of understandable speech, it has several disadvantages as well. Besides the inability to capture the relations between close regions (i.e. pixels) of the image, this pixelby-pixel representation of the image is also quite uneconomical. It is easy to see that a significant part of the image is irrelevant for the spectral parameter estimation task as the information stored by the neighbouring pixels is redundant, and the neural network is quite large due to the large number of input features. To resolve these issues, in this study we train an autoencoder neural network on the ultrasound image; the estimation of the spectral speech parameters is done by a second DNN, using the activations of the bottleneck layer of the autoencoder network as features. In our experiments, the proposed method proved to be more efficient than the standard approach: the measured normalized mean squared error scores were lower, while the correlation values were higher in each case. Based on the result of a listening test, the synthesized utterances also sounded more natural to native speakers. A further advantage of our proposed approach is that, due to the (relatively) small size of the bottleneck layer, we can utilize several consecutive ultrasound images during estimation without a significant increase in the network size, while significantly increasing the accuracy of parameter estimation. Gábor Gosztolya, Ádám Pintér, László Tóth 0001, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
IJCNN | 6 |
| 2019 | DNN-based Acoustic-to-Articulatory Inversion using Ultrasound Tongue ImagingabstractSpeech sounds are produced as the coordinated movement of the speaking organs. There are several available methods to model the relation of articulatory movements and the resulting speech signal. The reverse problem is often called as acoustic-to-articulatory inversion (AAI). In this paper we have implemented several different Deep Neural Networks (DNNs) to estimate the articulatory information from the acoustic signal. There are several previous works related to performing this task, but most of them are using ElectroMagnetic Articulography (EMA) for tracking the articulatory movement. Compared to EMA, Ultrasound Tongue Imaging (UTI) is a technique of higher cost-benefit if we take into account equipment cost, portability, safety and visualized structures. Seeing that, our goal is to train a DNN to obtain UT images, when using speech as input. We also test two approaches to represent the articulatory information: 1) the EigenTongue space and 2) the raw ultrasound image. As an objective quality measure for the reconstructed UT images, we use MSE, Structural Similarity Index (SSIM) and Complex- Wavelet SSIM (CW-SSIM). Our experimental results show that CW-SSIM is the most useful error measure in the UTI context. We tested three different system configurations: a) simple DNN composed of 2 hidden layers with 64x64 pixels of an UTI file as target; b) the same simple DNN but with ultrasound images projected to the EigenTongue space as the target; c) and a more complex DNN composed of 5 hidden layers with UTI files projected to the EigenTongue space. In a subjective experiment the subjects found that the neural networks with two hidden layers were more suitable for this inversion task. Dagoberto Porras, Alexander Sepúlveda, Tamás Gábor Csapó |
IJCNN | 3 |
| 2019 | Ultrasound-Based Silent Speech Interface Built on a Continuous VocoderabstractRecently it was shown that within the Silent Speech Interface (SSI) field, the prediction of F0 is possible from Ultrasound Tongue Images (UTI) as the articulatory input, using Deep Neural Networks for articulatory-to-acoustic mapping. Moreover, text-to-speech synthesizers were shown to produce higher quality speech when using a continuous pitch estimate, which takes non-zero pitch values even when voicing is not present. Therefore, in this paper on UTI-based SSI, we use a simple continuous F0 tracker which does not apply a strict voiced / unvoiced decision. Continuous vocoder parameters (ContF0, Maximum Voiced Frequency and Mel-Generalized Cepstrum) are predicted using a convolutional neural network, with UTI as input. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the continuous F0 is predicted with lower error, and the continuous vocoder produces slightly more natural synthesized speech than the baseline vocoder using standard discontinuous F0. Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 1 |
| 2019 | V-to-V Coarticulation Induced Acoustic and Articulatory Variability of Vowels: The Effect of Pitch-AccentabstractIn the present study we analyzed vowel variation induced by carryover V-to-V coarticulation under the effect of pitch-accent as a function of vowel quality (using a minimally constrained intervening consonant to maximize V-to-V effects). We tested if /i/ is more resistant to coarticulation than /u/, and if both vowels show increased coarticulatory resistance in pitch-accented syllables. Our approach was unprecedented in the sense that it involved the analysis of parallel acoustic (F2) and articulatory (x-axis dorsum position) data in a great number of speakers (9 speaker), and real words of Hungarian. To analyze the degree of coarticulation, we adopted the locus equation approach, and fitted linear models on vowel onset and midpoint data, and calculated the differences between coarticulated and non-coarticulated vowels in both domains. To measure variability, we calculated standard deviations of midpoint F2 values and dorsum positions. \nThe results showed that accent clearly exerted an effect on the phonetic realization of vowels, but the effect we found was dependent on both the vowel quality, and the domain (articulation/acoustics) at hand. Observation of the patterns we found in parallel acoustic and articulatory data warrants for reconsideration of the term ‘coarticulatory resistance’, and how it should be conceptualized. Andrea Deme, Márton Bartók, Tekla Etelka Gráczi, Tamás Gábor Csapó, Alexandra Markó |
INTERSPEECH | 4 |
| 2019 | Articulatory Analysis of Transparent Vowel /iː/ in Harmonic and Antiharmonic Hungarian Stems: Is There a Difference?abstractThe aim of our study is to analyse the articulatory characteristics of /iː/ occurring in Hungarian monosyllabic harmonic and antiharmonic stems.In their frequently cited work, based on 3 speakers' data, Beňuš and Gafos (2007) [1] claimed that the tongue position in transparent vowels of antiharmonic Hungarian stems is less advanced than that of the phonemically identical vowels in harmonic stems.In their study, the authors compared different harmonic and antiharmonic stems (even if the consonantal context was more or less controlled).In the present study, we analysed two homophonous pairs of words /siːv/ and /ɲiːr/, which are antiharmonic in their verbal usage, but are harmonic as nouns.The words were produced by 4 speakers both (i) in isolation and (ii) in sentence-initial position, where they were followed by front and back vowels, in a well-controlled manner.The experiment was carried out using electromagnetic articulography.We compared the sequence of the horizontal position of four receiver coils (ttip, tbl, tbo1, tbo2) across the conditions with Generalized Additive Models.The results showed that the horizontal positions of the receivers did not vary as a function of the harmonicity of the stem in either the isolated or the coarticulated condition. Alexandra Markó, Márton Bartók, Tamás Gábor Csapó, Tekla Etelka Gráczi, Andrea Deme |
INTERSPEECH | 3 |
| 2019 | Continuous vocoder applied in deep neural network based voice conversionabstractAbstract In this paper, a novel vocoder is proposed for a Statistical Voice Conversion (SVC) framework using deep neural network, where multiple features from the speech of two speakers (source and target) are converted acoustically. Traditional conversion methods focus on the prosodic feature represented by the discontinuous fundamental frequency (F0) and the spectral envelope. Studies have shown that speech analysis/synthesis solutions play an important role in the overall quality of the converted voice. Recently, we have proposed a new continuous vocoder, originally for statistical parametric speech synthesis, in which all parameters are continuous. Therefore, this work introduces a new method by using a continuous F0 (contF0) in SVC to avoid alignment errors that may happen in voiced and unvoiced segments and can degrade the converted speech. Our contribution includes the following. (1) We integrate into the SVC framework the continuous vocoder, which provides an advanced model of the excitation signal, by converting its contF0, maximum voiced frequency, and spectral features. (2) We show that the feed-forward deep neural network (FF-DNN) using our vocoder yields high quality conversion. (3) We apply a geometric approach to spectral subtraction (GA-SS) in the final stage of the proposed framework, to improve the signal-to-noise ratio of the converted speech. Our experimental results, using two male and one female speakers, have shown that the resulting converted speech with the proposed SVC technique is similar to the target speaker and gives state-of-the-art performance as measured by objective evaluation and subjective listening tests. Mohammed Salah Al-Radhi, Tamás Gábor Csapó, Géza Németh |
Multim. Tools Appl. | 2 |
| 2018 | F0 Estimation for DNN-Based Ultrasound Silent Speech InterfacesabstractState-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solution is based on the assumption that the F0 curve is unpredictable from articulatory data that does not contain direct measurements of the vocal fold vibration. Here, we experimented with deep neural networks to perform articulatory-to-acoustic conversion from ultrasound images, with an emphasis on estimating the voicing feature and the F0 curve from the ultrasound input. Contrary to the common belief that F0 is unpredictable, we attained a correlation rate of 0.74 between the original and the predicted F0 curve. What is more, the listening tests revealed that our subjects could not distinguish the sentences synthesized using the DNN-estimated and the original F0 curve, and ranked them as having the same quality. Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Tamás Gábor Csapó, Alexandra Markó |
ICASSP | 4 |
| 2018 | Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces
László Tóth 0001, Gábor Gosztolya, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
INTERSPEECH | 5 |
| 2017 | Time-Domain Envelope Modulating the Noise Component of Excitation in a Continuous Residual-Based Vocoder for Statistical Parametric Speech Synthesis
Mohammed Salah Al-Radhi, Tamás Gábor Csapó, Géza Németh |
INTERSPEECH | 2 |
| 2017 | DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech InterfaceabstractIn this paper we present our initial results in articulatory-toacoustic conversion based on tongue movement recordings using Deep Neural Networks (DNNs).Despite the fact that deep learning has revolutionized several fields, so far only a few researchers have applied DNNs for this task.Here, we compare various possible feature representation approaches combined with DNN-based regression.As the input, we recorded synchronized 2D ultrasound images and speech signals.The task of the DNN was to estimate Mel-Generalized Cepstrum-based Line Spectral Pair (MGC-LSP) coefficients, which then served as input to a standard pulse-noise vocoder for speech synthesis.As the raw ultrasound images have a relatively high resolution, we experimented with various feature selection and transformation approaches to reduce the size of the feature vectors.The synthetic speech signals resulting from the various DNN configurations were evaluated both using objective measures and a subjective listening test.We found that the representation that used several neighboring image frames in combination with a feature selection method was preferred both by the subjects taking part in the listening experiments, and in terms of the Normalized Mean Squared Error.Our results may be useful for creating Silent Speech Interface applications in the future. Tamás Gábor Csapó, Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 1 |
| 2015 | From text to formants - indirect model for trajectory prediction based on a multi-speaker parallel speech databaseabstractAn indirect model is presented, capable of estimating formant trajectories from text only (Text-to-Formants, TTF). The result is a phonetically correct formant trajectory flow of any virtual speech signal, i.e. one that has never been uttered. The focus is on the pattern forms inside the given sound, taking into account the sound environment (up to quinphone), and not on individual formant value measurements. The model is based on a multi-speaker parallel speech database with precise manual corrections and a HMM-based formant trajectory predictor. The validation of the TTF model shows that formant trajectories can be predicted with good accuracy from text. The model indirectly gives information about a theoretically possible articulation flow of the sentence. Thus it gives a general ‘formantprint’ of the language. Kálmán Abari, Tamás Gábor Csapó, Bálint Tóth, Gábor Olaszy |
INTERSPEECH | 2 |
| 2015 | Error analysis of extracted tongue contours from 2d ultrasound imagesabstractThe goal of this study was to characterize errors involved in obtaining midsagittal tongue contours from two-dimensional ultrasound image sequences. Toward that end, two basic experiments were conducted. First, manual tongue contours were obtained from 1,145 tongue ultrasound images recorded from four speakers during production of the sentence ‘I owe you a yoyo’, and the uncertainty associated with the contours was quantified. Second, tongue contours from the same images were obtained using the EdgeTrak, TongueTrack, and AutoTrace algorithms, and these were compared quantitatively with the manual tongue contours. Three basic error types associated with the tongue contours are identified, indicating areas in need of improvement in future algorithmic developments. Depending on the speaker, RMS errors for the algorithmically obtained contours ranged from 1.76 to 7.11 mm, and the standard deviation of manual contours ranged from 0.97 to 2.07 mm. Tamás Gábor Csapó, Steven M. Lulich |
INTERSPEECH | 1 |
| 2015 | Automatic transformation of irregular to regular voice by residual analysis and synthesisabstractThis paper presents an automatic speech transformation method of non-ideal phonation of speech (irregular or creaky voice). The irregular-to-regular transformation is performed by analyzing and resynthesizing the residual. A recent continuous pitch estimation algorithm is used for interpolating F0 in regions of irregular voice. The linear prediction residual of irregular sections of speech is replaced by overlap-added frames from a codebook of pitch-synchronous residuals. Finally, speech is reconstructed from the residual. A listening experiment showed that by transforming natural speech samples containing irregular voice, the perceived roughness of the transformed speech is decreased. Tamás Gábor Csapó, Géza Németh |
INTERSPEECH | 1 |
| 2012 | Synthesizing expressive speech from amateur audiobook recordingsabstractFreely available audiobooks are a rich resource of expressive speech recordings that can be used for the purposes of speech synthesis. Natural sounding, expressive synthetic voices have previously been built from audiobooks that contained large amounts of highly expressive speech recorded from a professionally trained speaker. The majority of freely available audiobooks, however, are read by amateur speakers, are shorter and contain less expressive (less emphatic, less emotional, etc.) speech both in terms of quality and quantity. Synthesizing expressive speech from a typical online audiobook therefore poses many challenges. In this work we address these challenges by applying a method consisting of minimally supervised techniques to align the text with the recorded speech, select groups of expressive speech segments and build expressive voices for hidden Markov-model based synthesis using speaker adaptation. Subjective listening tests have shown that the expressive synthetic speech generated with this method is often able to produce utterances suited to an emotional message. We used a restricted amount of speech data in our experiment, in order to show that the method is generally applicable to most typical audiobooks widely available online. Éva Székely, Tamás Gábor Csapó, Bálint Tóth, Péter Mihajlik, Julie Carson-Berndsen |
SLT | 2 |
| 2011 | Context and Speaker Dependency in the Relation of Vowel Formants and Subglottal Resonances - Evidence from HungarianabstractSubglottal resonances are claimed to divide front/back vowels and low/high vowels in several languages, including Hungarian. However, some ‘recalcitrant’ vowels appear to resist this mould. We therefore performed a careful analysis of the role coarticulation and speaker-dependent effects might play in the recalcitrance of these vowels in Hungarian. The present analyzis focused on various stop contexts in order to see the place of articulation triggered effects. It is shown that the subglottal resonances indeed divide the vowel space as claimed, and that the recalcitrance of certain vowels is due to coarticulation with specific consonants. The magnitude of the coarticulation effect is speaker dependent. Tekla Etelka Gráczi, Steven M. Lulich, Tamás Gábor Csapó, András Beke |
INTERSPEECH | 3 |
| 2009 | Relation of formants and subglottal resonances in Hungarian vowelsabstractThe relation between vowel formants and subglottal resonances (SGRs) has previously been explored in English, German, and Korean. Results from these studies indicate that vowel classes are categorically separated by SGRs. We extended this work to Hungarian vowels, which have not been related to SGRs before. The Hungarian vowel system contains paired long and short vowels as well as a series of front rounded vowels, similar to German but more complex than English and Korean. Results indicate that SGRs separate vowel classes in Hungarian as in English, German, and Korean, and uncover additional patterns of vowel formants relative to the third subglottal resonance (Sg3). These results have implications for understanding phonological distinctive features, and applications in automatic speech technologies. Index Terms: subglottal resonances, Hungarian, quantal theory, vowels Tamás Gábor Csapó, Zsuzsanna Bárkányi, Tekla Etelka Gráczi, Tamás Bohm, Steven M. Lulich |
INTERSPEECH | 1 |
| 2007 | Increasing prosodic variability of text-to-speech synthesizersabstractThe lack of prosody variation in text-to-speech systems contributes to their perceived unnaturalness when synthesizing extended passages. In this paper, we present a method to improve prosody generation in this direction. A database of natural sample sentences is searched for sentences having similar word and syllable structure to the input. One sentence is selected randomly from the similar sentences found. The prosody of the randomly selected natural sentence is used as a target to generate the prosody of the synthetic one. An experiment was conducted to determine the potential of the proposed method. The rule-based pitch contour generation of a Hungarian concatenative synthesizer was replaced by a semi-automatic implementation of the proposed method. A listening test showed that subjects preferred sentences synthesized by the proposed method over a rule-based solution. Géza Németh, Márk Fék, Tamás Gábor Csapó |
INTERSPEECH | 3 |