László Tóth 0001

dblp:69/3781-1 · DBLP profile ↗
← Back
56ranked-venue papers
15as first author
11since 2021 · last 2025
0000-0003-0161-1375ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 12 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Voice Reconstruction through Large-Scale TTS Models: Comparing Zero-Shot and Fine-tuning Approaches to Personalise TTS in Assistive Communication
Éva Székely, Péter Mihajlik, Mate Kadar, László Tóth 0001
INTERSPEECH4
2024 Combining Acoustic Feature Sets for Detecting Mild Cognitive Impairment in the Interspeech'24 TAUKADIAL Challenge
abstract
Shared tasks or challenges provide valuable opportunities for the machine learning community, as they offer a chance to compare the performance of machine learning approaches without peeking (due to the hidden test set).We present the approach of our team for the Interspeech'24 TAUKADIAL Challenge, where the task is to distinguish patients of Mild Cognitive Impairment (MCI) from healthy controls based on their speech.Our workflow focuses entirely on the acoustics, mixing standard feature sets (ComParE functionals and wav2vec2 embeddings) and custom attributes focusing on the amount of silent and filled pause segments.By training dedicated SVM classifiers on the three speech tasks and combining the predictions over the different speech tasks and feature sets, we obtained F1 values of up to 0.76 for the MCI identification task using crossvalidation, while our RMSE scores for the MMSE estimation task were as low as 2.769 (cross-validation) and 2.608 (test).
Gábor Gosztolya, László Tóth 0001
INTERSPEECH2
2023 Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks
László Tóth 0001, Amin Honarmandi Shandiz, Gábor Gosztolya, Tamás Gábor Csapó
INTERSPEECH1
2022 Using Acoustic Deep Neural Network Embeddings to Detect Multiple Sclerosis From Speech
abstract
Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system. It affects cognitive and motor functions, and the limitation of executive functions can also manifest itself in speech production. Due to this, automatic speech analysis might serve as an effective technique for assessing MS, or for monitoring the status of the patient. However, choosing the features to be extracted from the recordings is not straightforward. In the past few years, general feature extractors such as i-vectors, d-vectors and x-vectors have found their way into automatic speech analysis. In this study we show that there is no need to employ a special neural network architecture such as x-vectors to calculate effective features, but (even more) indicative features can be derived on the basis of a standard Deep Neural Network acoustic model. From our results, these features could effectively be used to distinguish MS subjects from healthy controls, as we measured AUC scores up to 0.935. We found that classification performance depended only slightly on the choice of the hid-den layer used to extract our features, but the speech task per-formed by the subject turned out to be an important factor.
Gábor Gosztolya, László Tóth 0001, Veronika Svindt, Judit Bóna, Ildikó Hoffmann
ICASSP2
2022 Using Spectral Sequence-to-Sequence Autoencoders to Assess Mild Cognitive Impairment
abstract
Dementia is a chronic or progressive clinical syndrome, mainly characterized by the deterioration of memory, thinking, reasoning and language. In Mild cognitive impairment (MCI), often considered as the prodromal stage of dementia, there is also a subtle deterioration of these functions, but they do not affect the daily life of the patient. However, due to the slight nature of the changes, it is quite hard to diagnose MCI. In this study, we employ sequence-to-sequence deep autoencoders in order to extract compact, robust and efficient attributes from the spontaneous speech of 25 MCI subjects and 25 healthy controls. From our results, this approach gives a competitive performance, as we significantly outperformed x-vectors even though they were trained on more data. Our additional efforts to identify mild Alzheimer’s (mAD) subjects as well were less successful; but since the focus is on the early detection of dementia, this is not a limitation of the methodology from a practical point of view.
Mercedes Vetráb, José Vicente Egas López, Réka Balogh, Nóra Imre, Ildikó Hoffmann, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya
ICASSP6
2022 Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks
Amin Honarmandi Shandiz, László Tóth 0001
IEA/AIE2
2022 Linguistic Parameters of Spontaneous Speech for Identifying Mild Cognitive Impairment and Alzheimer Disease
abstract
Abstract In this article, we seek to automatically identify Hungarian patients suffering from mild cognitive impairment (MCI) or mild Alzheimer disease (mAD) based on their speech transcripts, focusing only on linguistic features. In addition to the features examined in our earlier study, we introduce syntactic, semantic, and pragmatic features of spontaneous speech that might affect the detection of dementia. In order to ascertain the most useful features for distinguishing healthy controls, MCI patients, and mAD patients, we carry out a statistical analysis of the data and investigate the significance level of the extracted features among various speaker group pairs and for various speaking tasks. In the second part of the article, we use this rich feature set as a basis for an effective discrimination among the three speaker groups. In our machine learning experiments, we analyze the efficacy of each feature group separately. Our model that uses all the features achieves competitive scores, either with or without demographic information (3-class accuracy values: 68%–70%, 2-class accuracy values: 77.3%–80%). We also analyze how different data recording scenarios affect linguistic features and how they can be productively used when distinguishing MCI patients from healthy controls.
Veronika Vincze, Martina Katalin Szabó, Ildikó Hoffmann, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya
Comput. Linguistics4
2022 Automatic screening of mild cognitive impairment and Alzheimer's disease by means of posterior-thresholding hesitation representation
José Vicente Egas López, Réka Balogh, Nóra Imre, Ildikó Hoffmann, Martina Katalin Szabó, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya
Comput. Speech Lang.6
2021 Identifying Conflict Escalation and Primates by Using Ensemble X-Vectors and Fisher Vector Features
abstract
Computational paralinguistics is concerned with the automatic identification of non-verbal information in human speech.The Interspeech ComParE challenge features new paralinguistic tasks each year; this time, among others, a cross-corpus conflict escalation task and the identification of primates based solely on audio are the actual problems set.In our entry to ComParE 2021, we utilize x-vectors and Fisher vectors as features.To improve the robustness of the predictions, we also experiment with building an ensemble of classifiers from the x-vectors.Lastly, we exploit the fact that the Escalation Sub-Challenge is a conflict detection task, and incorporate the SSPNet Conflict Corpus in our training workflow.Using these approaches, at the time of writing, we had already surpassed the official Challenge baselines on both tasks, which demonstrates the efficiency of the employed techniques.
José Vicente Egas López, Mercedes Vetráb, László Tóth 0001, Gábor Gosztolya
Interspeech3
2021 Neural Speaker Embeddings for Ultrasound-Based Silent Speech Interfaces
abstract
Articulatory-to-acoustic mapping seeks to reconstruct speech from a recording of the articulatory movements, for example, an ultrasound video. Just like speech signals, these recordings represent not only the linguistic content, but are also highly specific to the actual speaker. Hence, due to the lack of multi-speaker data sets, researchers have so far concentrated on speaker-dependent modeling. Here, we present multi-speaker experiments using the recently published TaL80 corpus. To model speaker characteristics, we adjusted the x-vector framework popular in speech processing to operate with ultrasound tongue videos. Next, we performed speaker recognition experiments using 50 speakers from the corpus. Then, we created speaker embedding vectors and evaluated them on the remaining speakers. Finally, we examined how the embedding vector influences the accuracy of our ultrasound-to-speech conversion network in a multi-speaker scenario. In the experiments we attained speaker recognition error rates below 3%, and we also found that the embedding vectors generalize nicely to unseen speakers. Our first attempt to apply them in a multi-speaker silent speech framework brought about a marginal reduction in the error rate of the spectral estimation step.
Amin Honarmandi Shandiz, László Tóth 0001, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó
Interspeech2
2021 Cross-lingual detection of mild cognitive impairment based on temporal parameters of spontaneous speech
Gábor Gosztolya, Réka Balogh, Nóra Imre, José Vicente Egas López, Ildikó Hoffmann, Veronika Vincze, László Tóth 0001, Davangere P. Devanand, Magdolna Pákáski, János Kálmán
Comput. Speech Lang.7
2020 Ultrasound-Based Articulatory-to-Acoustic Mapping with WaveGlow Speech Synthesis
abstract
For articulatory-to-acoustic mapping using deep neural networks, typically spectral and excitation parameters of vocoders have been used as the training targets. However, vocoding often results in buzzy and muffled final speech quality. Therefore, in this paper on ultrasound-based articulatory-to-acoustic conversion, we use a flow-based neural vocoder (WaveGlow) pre-trained on a large amount of English and Hungarian speech data. The inputs of the convolutional neural network are ultrasound tongue images. The training target is the 80-dimensional mel-spectrogram, which results in a finer detailed spectral representation than the previously used 25-dimensional Mel-Generalized Cepstrum. From the output of the ultrasound-to-mel-spectrogram prediction, WaveGlow inference results in synthesized speech. We compare the proposed WaveGlow-based system with a continuous vocoder which does not use strict voiced/unvoiced decision when predicting F0. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the WaveGlow neural vocoder produces significantly more natural synthesized speech than the baseline system. Besides, the advantage of WaveGlow is that F0 is included in the mel-spectrogram representation, and it is not necessary to predict the excitation separately.
Tamás Gábor Csapó, Csaba Zainkó, László Tóth 0001, Gábor Gosztolya, Alexandra Markó
INTERSPEECH3
2020 Social Signal Detection by Probabilistic Sampling DNN Training
abstract
When our task is to detect social signals such as laughter and filler events in an audio recording, the most straightforward way is to apply a Hidden Markov Model-or a Hidden Markov Model/Deep Neural Network (HMM/DNN) hybrid, which is considered state-of-the-art nowadays. In this hybrid model, the DNN component is trained on frame-level samples of the classes we are looking for. In such event detection tasks, however, the training labels are seriously imbalanced, as typically only a small fraction of the training data corresponds to these social signals, while the bulk of the utterances consists of speech segments or silence. A strong imbalance of the training classes is known to cause difficulties during DNN training. To alleviate these problems, here we apply the technique called probabilistic sampling, which seeks to balance the class distribution. Probabilistic sampling is a mathematically well-founded combination of upsampling and downsampling, which was found to outperform both of these simple resampling approaches. With this strategy, we managed to achieve a 7-8 percent relative error reduction both at the segment level and frame level, and we efficiently reduced the DNN training times as well.
Gábor Gosztolya, Tamás Grósz, László Tóth 0001
IEEE Trans. Affect. Comput.3
2019 Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces
abstract
When using ultrasound video as input, Deep Neural Network-based Silent Speech Interfaces usually rely on the whole image to estimate the spectral parameters required for the speech synthesis step. Although this approach is quite straightforward, and it permits the synthesis of understandable speech, it has several disadvantages as well. Besides the inability to capture the relations between close regions (i.e. pixels) of the image, this pixelby-pixel representation of the image is also quite uneconomical. It is easy to see that a significant part of the image is irrelevant for the spectral parameter estimation task as the information stored by the neighbouring pixels is redundant, and the neural network is quite large due to the large number of input features. To resolve these issues, in this study we train an autoencoder neural network on the ultrasound image; the estimation of the spectral speech parameters is done by a second DNN, using the activations of the bottleneck layer of the autoencoder network as features. In our experiments, the proposed method proved to be more efficient than the standard approach: the measured normalized mean squared error scores were lower, while the correlation values were higher in each case. Based on the result of a listening test, the synthesized utterances also sounded more natural to native speakers. A further advantage of our proposed approach is that, due to the (relatively) small size of the bottleneck layer, we can utilize several consecutive ultrasound images during estimation without a significant increase in the network size, while significantly increasing the accuracy of parameter estimation.
Gábor Gosztolya, Ádám Pintér, László Tóth 0001, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó
IJCNN3
2019 Ultrasound-Based Silent Speech Interface Built on a Continuous Vocoder
abstract
Recently it was shown that within the Silent Speech Interface (SSI) field, the prediction of F0 is possible from Ultrasound Tongue Images (UTI) as the articulatory input, using Deep Neural Networks for articulatory-to-acoustic mapping. Moreover, text-to-speech synthesizers were shown to produce higher quality speech when using a continuous pitch estimate, which takes non-zero pitch values even when voicing is not present. Therefore, in this paper on UTI-based SSI, we use a simple continuous F0 tracker which does not apply a strict voiced / unvoiced decision. Continuous vocoder parameters (ContF0, Maximum Voiced Frequency and Mel-Generalized Cepstrum) are predicted using a convolutional neural network, with UTI as input. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the continuous F0 is predicted with lower error, and the continuous vocoder produces slightly more natural synthesized speech than the baseline vocoder using standard discontinuous F0.
Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth 0001, Alexandra Markó
INTERSPEECH6
2019 Calibrating DNN Posterior Probability Estimates of HMM/DNN Models to Improve Social Signal Detection from Audio Data
abstract
To detect social signals such as laughter or filler events from audio data, a straightforward choice is to apply a Hidden Markov Model (HMM) in combination with a Deep Neural Network (DNN) that supplies the local class posterior estimates (HMM/DNN hybrid model).However, the posterior estimates of the DNN may be suboptimal due to a mismatch between the cost function used during training (e.g.frame-level crossentropy) and the actual evaluation metric (e.g.segment-level F1 score).In this study, we show experimentally that by employing a simple posterior probability calibration technique on the DNN outputs, the performance of the HMM/DNN workflow can be significantly improved.Specifically, we apply a linear transformation on the activations of the output layer right before using the softmax function, and fine-tune the parameters of this transformation.Out of the calibration approaches tested, we got the best F1 scores when the posterior calibration process was adjusted so as to maximize the actual HMM-based evaluation metric.
Gábor Gosztolya, László Tóth 0001
INTERSPEECH2
2019 Examining the Combination of Multi-Band Processing and Channel Dropout for Robust Speech Recognition
abstract
sponsorship: Laszlo Toth was supported by the J ' anos Bolyai Research Scholarship of the Hungarian Academy of Sciences and the UNKP19-4 New Excellence Program of the Hungarian Ministry of Innovation and Technology. (Hungarian Academy of Sciences, New Excellence Program of the Hungarian Ministry of Innovation and Technology|UNKP19-4)
György Kovács 0001, László Tóth 0001, Dirk Van Compernolle, Marcus Liwicki
INTERSPEECH2
2019 Identifying Mild Cognitive Impairment and mild Alzheimer's disease based on spontaneous speech using ASR and linguistic features
Gábor Gosztolya, Veronika Vincze, László Tóth 0001, Magdolna Pákáski, János Kálmán, Ildikó Hoffmann
Comput. Speech Lang.3
2018 F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces
abstract
State-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solution is based on the assumption that the F0 curve is unpredictable from articulatory data that does not contain direct measurements of the vocal fold vibration. Here, we experimented with deep neural networks to perform articulatory-to-acoustic conversion from ultrasound images, with an emphasis on estimating the voicing feature and the F0 curve from the ultrasound input. Contrary to the common belief that F0 is unpredictable, we attained a correlation rate of 0.74 between the original and the predicted F0 curve. What is more, the listening tests revealed that our subjects could not distinguish the sentences synthesized using the DNN-estimated and the original F0 curve, and ranked them as having the same quality.
Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Tamás Gábor Csapó, Alexandra Markó
ICASSP3
2018 General Utterance-Level Feature Extraction for Classifying Crying Sounds, Atypical & Self-Assessed Affect and Heart Beats
abstract
In the area of computational paralinguistics, there is a growing need for general techniques that can be applied in a variety of tasks, and which can be easily realized using standard and publicly available tools.In our contribution to the 2018 Interspeech Computational Paralinguistic Challenge (ComParE), we test four general ways of extracting features.Besides the standard ComParE feature set consisting of 6373 diverse attributes, we experiment with two variations of Bag-of-Audio-Words representations, and define a simple feature set inspired by Gaussian Mixture Models.Our results indicate that the UAR scores obtained via the different approaches vary among the tasks.In our view, this is mainly because most feature sets tested were local by nature, and they could not properly represent the utterances of the Atypical Affect and Self-Assessed Affect Sub-Challenges.On the Crying Sub-Challenge, however, a simple combination of all four feature sets proved to be effective.
Gábor Gosztolya, Tamás Grósz, László Tóth 0001
INTERSPEECH3
2018 Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces
László Tóth 0001, Gábor Gosztolya, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó
INTERSPEECH1
2018 Multi-Band Processing With Gabor Filters and Time Delay Neural Nets for Noise Robust Speech Recognition
abstract
Spectro-temporal feature extraction and multi-band processing were both invented with the goal of increasing the robustness of speech recognisers. However, although these methods have been in use for a long time now, and they are evidently compatible, few attempts have been made to combine them. This is why here we investigate the combination of multi-band processing with the use of spectro-temporal Gabor filters. First, based on the TIMIT corpus, we optimise their meta-parameters like the overlap, and the number of bands. Then we verify the cross-corpus viability of our multi-band processing approach on the Aurora-4 corpus. Lastly, we combine our method with the recently proposed channel dropout method. Our results show that this combination not only leads to lower error rates than those got using either multi-band processing or channel dropout, but these results compare favourably to those recently reported for the clean training scenario on the Aurora-4 corpus.
György Kovács 0001, László Tóth 0001, Gábor Gosztolya
SLT2
2018 Efficient visual code localization with neural networks
Péter Bodnár, Tamás Grósz, László Tóth 0001, László G. Nyúl
Pattern Anal. Appl.3
2018 A feature selection-based speaker clustering method for paralinguistic tasks
Gábor Gosztolya, László Tóth 0001
Pattern Anal. Appl.2
2017 DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech Interface
abstract
In this paper we present our initial results in articulatory-toacoustic conversion based on tongue movement recordings using Deep Neural Networks (DNNs).Despite the fact that deep learning has revolutionized several fields, so far only a few researchers have applied DNNs for this task.Here, we compare various possible feature representation approaches combined with DNN-based regression.As the input, we recorded synchronized 2D ultrasound images and speech signals.The task of the DNN was to estimate Mel-Generalized Cepstrum-based Line Spectral Pair (MGC-LSP) coefficients, which then served as input to a standard pulse-noise vocoder for speech synthesis.As the raw ultrasound images have a relatively high resolution, we experimented with various feature selection and transformation approaches to reduce the size of the feature vectors.The synthetic speech signals resulting from the various DNN configurations were evaluated both using objective measures and a subjective listening test.We found that the representation that used several neighboring image frames in combination with a feature selection method was preferred both by the subjects taking part in the listening experiments, and in terms of the Normalized Mean Squared Error.Our results may be useful for creating Silent Speech Interface applications in the future.
Tamás Gábor Csapó, Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Alexandra Markó
INTERSPEECH4
2017 DNN-Based Feature Extraction and Classifier Combination for Child-Directed Speech, Cold and Snoring Identification
abstract
In this study we deal with the three sub-challenges of the Interspeech ComParE Challenge 2017, where the goal is to identify child-directed speech, speakers having a cold, and different types of snoring sounds.For the first two sub-challenges we propose a simple, two-step feature extraction and classification scheme: first we perform frame-level classification via Deep Neural Networks (DNNs), and then we extract utterancelevel features from the DNN outputs.By utilizing these features for classification, we were able to match the performance of the standard paralinguistic approach (which involves extracting thousands of features, many of them being completely irrelevant to the actual task).As for the Snoring Sub-Challenge, we divided the recordings into segments, and averaged out some frame-level features segment-wise, which were then used for utterance-level classification.When combining the predictions of the proposed approaches with those got by the standard paralinguistic approach, we managed to outperform the baseline values of the Cold and Snoring sub-challenges on the hidden test sets.
Gábor Gosztolya, Róbert Busa-Fekete, Tamás Grósz, László Tóth 0001
INTERSPEECH4
2017 Training Context-Dependent DNN Acoustic Models Using Probabilistic Sampling
abstract
In current HMM/DNN speech recognition systems, the purpose of the DNN component is to estimate the posterior probabilities of tied triphone states.In most cases the distribution of these states is uneven, meaning that we have a markedly different number of training samples for the various states.This imbalance of the training data is a source of suboptimality for most machine learning algorithms, and DNNs are no exception.A straightforward solution is to re-sample the data, either by upsampling the rarer classes or by dowsampling the more common classes.Here, we experiment with the so-called probabilistic sampling method that applies downsampling and upsampling at the same time.For this, it defines a new class distribution for the training data, which is a linear combination of the original and the uniform class distributions.As an extension to previous studies, we propose a new method to re-estimate the class priors, which is required to remedy the mismatch between the training and the test data distributions introduced by re-sampling.Using probabilistic sampling and the proposed modification we report 5% and 6% relative error rate reductions on the TED-LIUM and on the AMI corpora, respectively.
Tamás Grósz, Gábor Gosztolya, László Tóth 0001
INTERSPEECH3
2017 A Comparative Evaluation of GMM-Free State Tying Methods for ASR
Tamás Grósz, Gábor Gosztolya, László Tóth 0001
INTERSPEECH3
2017 Increasing the robustness of CNN acoustic models using autoregressive moving average spectrogram features and channel dropout
György Kovács 0001, László Tóth 0001, Dirk Van Compernolle, Sriram Ganapathy
Pattern Recognit. Lett.2
2017 DNN-Based Feature Extraction for Conflict Intensity Estimation From Speech
abstract
Over past few years, there has been an increasing need to extract nonlinguistic information from audio sources. This trend has created a new area in speech technology known as computational paralinguistics. A task belonging to this area is to estimate the intensity of conflicts arising in speech recordings, based only on the audio information. It was shown that the human comprehension of conflict intensity is closely related to speaker overlap; that is, when multiple persons are speaking at the same time. This type of information can also aid automated conflict intensity estimation. In this study, we propose a simple, DNN-based feature extraction step, and show that this approach is superior to those introduced in the literature so far: By combining our results with an efficient greedy feature selection algorithm, we were able to outperform all previous results on the SSPNet conflict dataset, achieving a correlation coefficient of 0.856 on the test set.
Gábor Gosztolya, László Tóth 0001
IEEE Signal Process. Lett.2
2016 Determining Native Language and Deception Using Phonetic Features and Classifier Combination
abstract
For several years, the Interspeech ComParE Challenge has focused on paralinguistic tasks of various kinds.In this paper we focus on the Native Language and the Deception subchallenges of ComParE 2016, where the goal is to identify the native language of the speaker, and to recognize deceptive speech.As both tasks can be treated as classification ones, we experiment with several state-of-the-art machine learning methods (Support-Vector Machines, AdaBoost.MH and Deep Neural Networks), and also test a simple-yet-robust combination method.Furthermore, we will assume that the native language of the speaker affects the pronunciation of specific phonemes in the language he is currently using.To exploit this, we extract phonetic features for the Native Language task.Moreover, for the Deception Sub-Challenge we compensate for the highly unbalanced class distribution by instance re-sampling.With these techniques we are able to significantly outperform the baseline SVM on the unpublished test set.
Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001
INTERSPEECH4
2016 Estimating the Sincerity of Apologies in Speech by DNN Rank Learning and Prosodic Analysis
abstract
In the Sincerity Sub-Challenge of the Interspeech ComParE 2016 Challenge, the task is to estimate user-annotated sincerity scores for speech samples.We interpret this challenge as a ranklearning regression task, since the evaluation metric (Spearman's correlation) is calculated from the rank of the instances.As a first approach, Deep Neural Networks are used by introducing a novel error criterion which maximizes the correlation metric directly.We obtained the best performance by combining the proposed error function with the conventional MSE error.This approach yielded results that outperform the baseline on the Challenge test set.Furthermore, we introduce a compact prosodic feature set based on a dynamic representation of F0, energy and sound duration.We extract syllable-based prosodic features which are used as the basis of another machine learning step.We show that a small set of prosodic features is capable of yielding a result very close to the baseline one and that by combining the predictions yielded by DNN and the prosodic feature set, further improvement can be reached, significantly outperforming the baseline SVR on the Challenge test set.
Gábor Gosztolya, Tamás Grósz, György Szaszák, László Tóth 0001
INTERSPEECH4
2016 GMM-Free Flat Start Sequence-Discriminative DNN Training
abstract
Recently, attempts have been made to remove Gaussian mixture models (GMM) from the training process of deep neural network-based hidden Markov models (HMM/DNN). For the GMM-free training of a HMM/DNN hybrid we have to solve two problems, namely the initial alignment of the frame-level state labels and the creation of context-dependent states. Although flat-start training via iteratively realigning and retraining the DNN using a frame-level error function is viable, it is quite cumbersome. Here, we propose to use a sequence-discriminative training criterion for flat start. While sequence-discriminative training is routinely applied only in the final phase of model training, we show that with proper caution it is also suitable for getting an alignment of context-independent DNN models. For the construction of tied states we apply a recently proposed KL-divergence-based state clustering method, hence our whole training process is GMM-free. In the experimental evaluation we found that the sequence-discriminative flat start training method is not only significantly faster than the straightforward approach of iterative retraining and realignment, but the word error rates attained are slightly better as well.
Gábor Gosztolya, Tamás Grósz, László Tóth 0001
INTERSPEECH3
2016 Detecting Mild Cognitive Impairment from Spontaneous Speech by Correlation-Based Phonetic Feature Selection
abstract
Mild Cognitive Impairment (MCI), sometimes regarded as a prodromal stage of Alzheimer's disease, is a mental disorder that is difficult to diagnose.Recent studies reported that MCI causes slight changes in the speech of the patient.Our previous studies showed that MCI can be efficiently classified by machine learning methods such as Support-Vector Machines and Random Forest, using features describing the amount of pause in the spontaneous speech of the subject.Furthermore, as hesitation is the most important indicator of MCI, we took special care when handling filled pauses, which usually correspond to hesitation.In contrast to our previous studies which employed manually constructed feature sets, we now employ (automatic) correlation-based feature selection methods to find the relevant feature subset for MCI classification.By analyzing the selected feature subsets we also show that features related to filled pauses are useful for MCI detection from speech samples.
Gábor Gosztolya, László Tóth 0001, Tamás Grósz, Veronika Vincze, Ildikó Hoffmann, Gréta Szatlóczki, Magdolna Pákáski, János Kálmán
INTERSPEECH2
2015 Building context-dependent DNN acoustic models using Kullback-Leibler divergence-based state tying
abstract
Deep neural network (DNN) based speech recognizers have recently replaced Gaussian mixture (GMM) based systems as the state-of-the-art. HMM/DNN systems have kept many refinements of the HMM/GMM framework, even though some of these may be suboptimal for them. One such example is the creation of context-dependent tied states, for which an efficient decision tree state tying method exists. The tied states used to train DNNs are usually obtained using the same tying algorithm, even though it is based on likelihoods of Gaussians. In this paper, we investigate an alternative state clustering method that uses the Kullback-Leibler (KL) divergence of DNN output vectors to build the decision tree. It has already been successfully applied within the framework of KL-HMM systems, and here we show that it is also beneficial for HMM/DNN hybrids. In a large vocabulary recognition task we report a 4% relative word error rate reduction using this state clustering method.
Gábor Gosztolya, Tamás Grósz, László Tóth 0001, David Imseng
ICASSP3
2015 Modeling long temporal contexts in convolutional neural network-based phone recognition
abstract
The deep neural network component of current hybrid speech recognizers is trained on a context of consecutive feature vectors. Here, we investigate whether the time span of this input can be extended by splitting it up and modeling it in smaller chunks. One method for this is to train a hierarchy of two networks, while the less well-known split temporal context (STC) method models the left and right contexts of a frame separately. Here, we evaluate these techniques within a convolutional neural network framework, and find that the two approaches can be nicely combined. With the combined model we can expand the time-span of our network to 69 frames, and we achieve a 7.5% relative error rate reduction compared to modeling this large context as one block. We report a phone error rate of 17.1% on the TIMIT core test set, which is one of the best scores published.
László Tóth 0001
ICASSP1
2015 Assessing the degree of nativeness and parkinson's condition using Gaussian processes and deep rectifier neural networks
abstract
The Interspeech 2015 Computational Paralinguistics Challenge includes two regression learning tasks, namely the Parkinson's Condition Sub-Challenge and the Degree of Nativeness Sub-Challenge.We evaluated two state-of-the-art machine learning methods on the tasks, namely Deep Neural Networks (DNN) and Gaussian Processes Regression (GPR).We also experiented with various classifier combination and feature selection methods.For the Degree of Nativeness sub-challenge we obtained a far better Spearman correlation value than the one presented in the baseline paper.As regards the Parkinson's Condition Sub-Challenge, we showed that both DNN and GPR are competitive with the baseline SVM, and that the results can be improved further by combining the classifiers.However, we obtained by far the best results when we applied a speaker clustering method to identify the files that belong to the same speaker.
Tamás Grósz, Róbert Busa-Fekete, Gábor Gosztolya, László Tóth 0001
INTERSPEECH4
2015 Automatic detection of mild cognitive impairment from spontaneous speech using ASR
abstract
Mild Cognitive Impairment (MCI), sometimes regarded as a prodromal stage of Alzheimer's disease, is a mental disorder that is difficult to diagnose.However, recent studies reported that MCI causes slight changes in the speech of the patient.Our starting point here is a study that found acoustic correlates of MCI, but extracted the proposed features manually.Here, we automate the extraction of the features by applying automatic speech recognition (ASR).Unlike earlier authors, we use ASR to extract only a phonetic level segmentation and annotation.While the phonetic output allows the calculation of features like the speech rate, it avoids the problems caused by the agrammatical speech frequently produced by the targeted patient group.Furthermore, as hesitation is the most important indicator of MCI, we take special care when handling filled pauses, which usually correspond to hesitation.Using the ASR-based features, we employ machine learning methods to separate the subjects with MCI from the control group.The classification results obtained with ASR-based feature extraction are just slightly worse that those got with the manual method.The F1 value achieved (85.3) is very promising regarding the creation of an automated MCI screening application.
László Tóth 0001, Gábor Gosztolya, Veronika Vincze, Ildikó Hoffmann, Gréta Szatlóczki, Edit Biró, Fruzsina Zsura, Magdolna Pákáski, János Kálmán
INTERSPEECH1
2014 Combining time- and frequency-domain convolution in convolutional neural network-based phone recognition
abstract
Convolutional neural networks have proved very successful in image recognition, thanks to their tolerance to small translations. They have recently been applied to speech recognition as well, using a spectral representation as input. However, in this case the translations along the two axes - time and frequency - should be handled quite differently. So far, most authors have focused on convolution along the frequency axis, which offers invariance to speaker and speaking style variations. Other researchers have developed a different network architecture that applies time-domain convolution in order to process a longer time-span of input in a hierarchical manner. These two approaches have different background motivations, and both offer significant gains over a standard fully connected network. Here we show that the two network architectures can be readily combined, like their advantages. With the combined model we report an error rate of 16.7% on the TIMIT phone recognition task, a new record on this dataset.
László Tóth 0001
ICASSP1
2014 Detecting the intensity of cognitive and physical load using AdaBoost and deep rectifier neural networks
abstract
The Interspeech ComParE 2014 Challenge consists of two machine learning tasks, which have quite a small number of examples. Due to our good results in ComParE 2013, we considered AdaBoost a suitable machine learning meta-algorithm for these tasks, besides we also experimented with Deep Rectifier Neural Networks. These differ from traditional neural networks in that the former have several hidden layers, and use rectifier neurons as hidden units. With AdaBoost we achieved competitive results, whereas with the neural networks we were able to outperform baseline SVM scores in both Sub-Challenges. Index Terms :s peech technology, AdaBoost, deep neural networks, rectifier activation function
Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001
INTERSPEECH4
2014 Convolutional deep maxout networks for phone recognition
abstract
Convolutional neural networks have recently been shown to outperform fully connected deep neural networks on several speech recognition tasks. Their superior performance is due to their convolutional structure that processes several, slightly shifted versions of the input window using the same weights, and then pools the resulting neural activations. This pooling operation makes the network less sensitive to translations. The convolutional network results published up till now used sigmoid or rectified linear neurons. However, quite recently a new type of activation function called the maxout activation has been proposed. Its operation is closely related to convolutional networks, as it applies a similar pooling step, but over different neurons evaluated on the same input. Here, we combine the two technologies, and experiment with deep convolutional neural networks built from maxout neurons. Phone recognition tests on the TIMIT database show that switching to maxout units from rectifier units decreases the phone error rate for each network configuration studied, and yields relative error rate reductions of between 2% and 6%.
László Tóth 0001
INTERSPEECH1
2013 Phone recognition with deep sparse rectifier neural networks
abstract
Rectifier neurons differ from standard ones only in that the sigmoid activation function is replaced by the rectifier function, max(0, x). This modification requires only minimal changes to any existing neural net implementation, but makes it more effective. In particular, we show that a deep architecture of rectifier neurons can attain the same recognition accuracy as deep neural networks, but without the need for pre-training. With 4-5 hidden layers of rectifier neurons we report 20.8% and 19.8% phone error rates on TIMIT (with CI and CD units, respectively), which are competitive with the best results on this database.
László Tóth 0001
ICASSP1
2013 Detecting autism, emotions and social signals using adaboost
abstract
In the area of speech technology, tasks that involve the extraction of non-lingustic information have been receiving more attention recently. The Computational Paralinguistics Challenge (ComParE 2013) sought to develop techniques to efficiently detect a number of paralinguistic events, including the detection of non-linguistic events (laughter and fillers) in speech recordings as well as categorizing whole (albeit short) recordings by speaker emotion, conflict or the presence of development disorders (autism). We treated these sub-challenges as general classification tasks and applied the general-purpose machine learning meta-algorithm, AdaBoost.MH, and its recently proposed variant, AdaBoost.MH.BA, to them. The results show that these new algorithms convincingly outperform baseline SVM scores.
Gábor Gosztolya, Róbert Busa-Fekete, László Tóth 0001
INTERSPEECH3
2013 Convolutional deep rectifier neural nets for phone recognition
abstract
Rectifier neurons differ from standard ones only in that the sigmoid activation function is replaced by the rectifier function, max(0, x). Several recent studies suggest that rectifier units may be more suitable building units for deep nets. For example, we found that with deep rectifier networks one can attain a similar speech recognition performance than that with sigmoid nets, but without the need for the time-consuming pre-training procedure. Here, we extend the previous results by modifying the rectifier network so that it has a convolutional structure. As convolutional networks are inherently deep, rectifier neurons seem to be an ideal choice as their building units. Indeed, on the TIMIT phone recognition task we report a 6% relative error reduction compared to our earlier results, giving an 18.6% error rate on the core test set. Then, with the application of the recently proposed ‘dropout’ training method we reduce the error rate further to 17.8%, which, to our knowledge, is the best result to date on this database.
László Tóth 0001
INTERSPEECH1
2011 A hierarchical, context-dependent neural network architecture for improved phone recognition
abstract
In this paper we combine three simple refinements proposed recently to improve HMM/ANN hybrid models. The first refinement is to apply a hierarchy of two nets, where the second net models the contextual relations of the state posteriors produced by the first network. The second idea is to train the network on context-dependent units (HMM states) instead of context-independent phones or phone states. As the latter refinement results in a lot of output neurons, combining the two methods directly would be problematic. Hence the third trick is to shrink the output layer of the first net using the bottleneck technique before applying the second net on top of it. The phone recognition results obtained on the TIMIT database demonstrate that both the context-dependent and the 2-stage modeling methods can bring about marked improvements. Using them in combination, however, results in a further significant gain in accuracy. With the bottleneck technique a further improvement can be obtained, especially when the number of context-dependent units is large.
László Tóth 0001
ICASSP1
2008 Cross-lingual portability of MLP-based tandem features - a case study for English and Hungarian
abstract
One promising approach for building ASR systems for lessresourced languages is cross-lingual adaptation.Tandem ASR is particularly well suited to such adaptation, as it includes two cascaded modelling steps: feature extraction using multi-layer perceptrons (MLPs), followed by modelling using a standard HMM.The language-specific tuning can be performed by adjusting the HMM only, leaving the MLP untouched.Here we examine the portability of feature extractor MLPs between an Indo-European (English) and a Finno-Ugric (Hungarian) language.We present experiments which use both conventional phone-posterior and articulatory feature (AF) detector MLPs, both trained on a much larger quantity of (English) data than the monolingual (Hungarian) system.We find that the cross-lingual configurations achieve similar performance to the monolingual system, and that, interestingly, the AF detectors lead to slightly worse performance, despite the expectation that they should be more language-independent than phone-based MLPs.However, the cross-lingual system outperforms all other configurations when the English phone MLP is adapted on the Hungarian data.
László Tóth 0001, Joe Frankel, Gábor Gosztolya, Simon King 0001
INTERSPEECH1
2007 Benchmarking human performance on the acoustic and linguistic subtasks of ASR systems
abstract
Many believe that comparisons of machine and human speech recognition could help determine both the room for and the direction of improvement for speech recognizers. Yet, such experiments are made quite rarely or over such complex domains where instructive conclusions are hard to draw. In this paper we attempt to measure human performance on the tasks of the acoustic and language models of ASR systems separately. To simulate the task of acoustic decoding, subjects were instructed to phonetically transcribe short nonsense sentences. Here, besides the well-known superior segment classification, we also observed a good performance in word segmentation. To imitate higher-level processing, the subjects had to correct deliberately corrupted texts. Here we found that humans can achieve a word accuracy of about 80% even when almost one third of the phonemes are incorrect, and that with word boundary position information the word error rate roughly halves.
László Tóth 0001
INTERSPEECH1
2007 A segment-based interpretation of HMM/ANN hybrids
László Tóth 0001, András Kocsor
Comput. Speech Lang.1
2005 Training HMM/ANN Hybrid Speech Recognizers by Probabilistic Sampling
László Tóth 0001, András Kocsor
ICANN (1)1
2005 Fundamental frequency estimation by least-squares harmonic model fitting
abstract
This paper proposes a pitch estimation algorithm that is based on optimal harmonic model fitting. The algorithm operates directly on the time-domain signal and has a relatively simple mathematical background. To increase its efficiency and accuracy, the algorithm is applied in combination with an autocorrelation-based initialization phase. For testing purposes we compare its performance on pitch-annotated corpora with several conventional time-domain pitch estimation algorithms, and also with a recently proposed one. The results show that even the autocorrelation-based first phase significantly outperforms the traditional methods, and also slightly the recently proposed yin algorithm. After applying the second phase – the harmonic approximation step – the amount of errors can be further reduced by about 20 % relative to the error obtained in the first phase. 1.
András Bánhalmi, Kornél Kovács, András Kocsor, László Tóth 0001
INTERSPEECH4
2004 Initialization of directions in projection pursuit learning
abstract
The Projection Pursuit Learner is a multi-class classifier that resembles a two-layer neutral network in which the sigmoid activation functions of the hidden neurons have been replaced by an interpolating polynomial. This modification increases the flexibility of the model but also makes it more inclined to get stuck in a local minimum during gradient-based training. This problem can be alleviated to a certain extent by replacing the random initialization the projection directions by means of feature space transformation methods such as independent component analysis (IDA), principal component analysis (PCA), linear discriminant analysis (LDA) and springy discriminant analysis (SDA). We find that with this refinement the number of processing units can be reduced by 10 - 40%.
Gabor Faddi, András Kocsor, László Tóth 0001
IJCNN3
2004 Replicator Neural Networks for Outlier Modeling in Segmental Speech Recognition
László Tóth 0001, Gábor Gosztolya
ISNN (1)1
2004 Application of Kernel-Based Feature Space Transformations and Learning Methods to Phoneme Classification
András Kocsor, László Tóth 0001
Appl. Intell.2
2003 Harmonic alternatives to sine-wave speech
László Tóth 0001, András Kocsor
INTERSPEECH1
2002 Hungarian Speech Synthesis Using a Phase Exact HNM Approach
Kornél Kovács, András Kocsor, László Tóth 0001
SOFSEM3
2001 Application of Feature Transformation and Learning Methods in Phoneme Classification
András Kocsor, László Tóth 0001, László Felföldi
IEA/AIE2