EDBT 2026 Demo / reviewers in the wild / expert
Gábor Gosztolya
dblp:51/1288
· DBLP profile ↗
63ranked-venue papers
38as first author
19since 2021 · last 2025
0000-0002-2864-6466ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 34 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 26 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conformer-based Ultrasound-to-Speech ConversionabstractDeep neural networks have shown promising potential for ultrasound-to-speech conversion task towards Silent Speech Interfaces. In this work, we applied two Conformer-based DNN architectures (Base and one with bi-LSTM) for this task. Speaker-specific models were trained on the data of four speakers from the Ultrasuite-Tal80 dataset, while the generated mel spectrograms were synthesized to audio waveform using a HiFi-GAN vocoder. Compared to a standard 2D-CNN baseline, objective measurements (MSE and mel cepstral distortion) showed no statistically significant improvement for either model. However, a MUSHRA listening test revealed that Conformer with bi-LSTM provided better perceptual quality, while Conformer Base matched the performance of the baseline along with a 3× faster training time due to its simpler architecture. These findings suggest that Conformer-based models, especially the Conformer with bi-LSTM, offer a promising alternative to CNNs for ultrasound-to-speech conversion. © 2025 Elsevier B.V., All rights reserved. Ibrahim Ibrahimov, Csaba Zainkó, Gábor Gosztolya |
INTERSPEECH | 3 |
| 2024 | Combining Acoustic Feature Sets for Detecting Mild Cognitive Impairment in the Interspeech'24 TAUKADIAL ChallengeabstractShared tasks or challenges provide valuable opportunities for the machine learning community, as they offer a chance to compare the performance of machine learning approaches without peeking (due to the hidden test set).We present the approach of our team for the Interspeech'24 TAUKADIAL Challenge, where the task is to distinguish patients of Mild Cognitive Impairment (MCI) from healthy controls based on their speech.Our workflow focuses entirely on the acoustics, mixing standard feature sets (ComParE functionals and wav2vec2 embeddings) and custom attributes focusing on the amount of silent and filled pause segments.By training dedicated SVM classifiers on the three speech tasks and combining the predictions over the different speech tasks and feature sets, we obtained F1 values of up to 0.76 for the MCI identification task using crossvalidation, while our RMSE scores for the MMSE estimation task were as low as 2.769 (cross-validation) and 2.608 (test). Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 1 |
| 2024 | Automatic Longitudinal Investigation of Multiple Sclerosis SubjectsabstractMultiple Sclerosis is a chronic inflammatory disease of the central nervous system. Over time, people with MS may experience significant changes in cognition, language and speech processes. In this study we investigate speech utterances recorded over the course of three years for 16 MS subjects and 12 healthy controls. Our examination is based on speaker category classification (healthy or MS) using wav2vec2 embeddings as features. We found that subject classification performance improved over time: the 0.745-0.844 AUC values from year one increased to 0.891-0.979 in the third year. By analyzing the posterior estimates, we measured a statistically significant improvement in the scores corresponding to the third year for the MS category, while for the control subjects there was no such tendency. This, in our view, indicates that the change is due to a subtle deterioration in the condition of MS patients, which was detected by our machine learning workflow. Gábor Gosztolya, Veronika Svindt, Judit Bóna, Ildikó Hoffmann |
INTERSPEECH | 1 |
| 2024 | Wav2vec 2.0 Embeddings Are No Swiss Army Knife - A Case Study for Multiple SclerosisabstractIn the past few years, self-supervised learning has revolutionalized automatic speech recognition.Self-supervised models such as wav2vec2, due to their generalization ability on huge unannotated audio corpora, were claimed to be state-ofthe-art feature extractors in paralinguistic and pathological applications as well.In this study we test embeddings extracted from a wav2vec 2.0 model fine-tuned on the target language as features on a multiple sclerosis audio corpus, using three speech tasks.After comparing the resulting classification performances with traditional features such as ComParE functionals, ECAPA-TDNN and activations of a HMM/DNN hybrid acoustic model, we found that wav2vec2-based models, surprisingly, only produced a mediocre classification performance.In contrast, the decade-old ComParE functionals feature set consistently led to high scores.Our results also indicate that the number of features correlates surprisingly well with classification performance. Gábor Gosztolya, Mercedes Vetráb, Veronika Svindt, Judit Bóna, Ildikó Hoffmann |
INTERSPEECH | 1 |
| 2023 | Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks
László Tóth 0001, Amin Honarmandi Shandiz, Gábor Gosztolya, Tamás Gábor Csapó |
INTERSPEECH | 3 |
| 2023 | Automated Multiple Sclerosis Screening Based on Encoded Speech Representations
José Vicente Egas López, Veronika Svindt, Judit Bóna, Ildikó Hoffmann, Gábor Gosztolya |
INTERSPEECH | 5 |
| 2022 | Using Acoustic Deep Neural Network Embeddings to Detect Multiple Sclerosis From SpeechabstractMultiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system. It affects cognitive and motor functions, and the limitation of executive functions can also manifest itself in speech production. Due to this, automatic speech analysis might serve as an effective technique for assessing MS, or for monitoring the status of the patient. However, choosing the features to be extracted from the recordings is not straightforward. In the past few years, general feature extractors such as i-vectors, d-vectors and x-vectors have found their way into automatic speech analysis. In this study we show that there is no need to employ a special neural network architecture such as x-vectors to calculate effective features, but (even more) indicative features can be derived on the basis of a standard Deep Neural Network acoustic model. From our results, these features could effectively be used to distinguish MS subjects from healthy controls, as we measured AUC scores up to 0.935. We found that classification performance depended only slightly on the choice of the hid-den layer used to extract our features, but the speech task per-formed by the subject turned out to be an important factor. Gábor Gosztolya, László Tóth 0001, Veronika Svindt, Judit Bóna, Ildikó Hoffmann |
ICASSP | 1 |
| 2022 | Automatic Assessment of the Degree of Clinical Depression from Speech Using X-VectorsabstractDepression is a frequent and curable psychiatric disorder, detrimentally affecting daily activities, harming both work-place productivity and personal relationships. Among many other symptoms, depression is associated with disordered speech production, which might permit its automatic screening by means of the speech of the subject. However, the choice of actual features extracted from the recordings is not trivial. In this study, we employ x-vectors, a DNN-based feature extractor technique, to detect depression from a Hungarian corpus. We experiment with training custom x-vector extractors, and we also explore the performance of an out-of-domain pre-trained one. Our findings confirm that x-vectors are able to capture meaningful speaker traits that contain information for depression discrimination. We also show that the language of the extractor is of secondary importance compared to the frame-level feature set: our best model, which achieved an AUC score of 0.940 and an RMSE score of 9.54, was trained on log-energies instead of MFCCs. José Vicente Egas López, Gábor Kiss, Dávid Sztahó, Gábor Gosztolya |
ICASSP | 4 |
| 2022 | Using Spectral Sequence-to-Sequence Autoencoders to Assess Mild Cognitive ImpairmentabstractDementia is a chronic or progressive clinical syndrome, mainly characterized by the deterioration of memory, thinking, reasoning and language. In Mild cognitive impairment (MCI), often considered as the prodromal stage of dementia, there is also a subtle deterioration of these functions, but they do not affect the daily life of the patient. However, due to the slight nature of the changes, it is quite hard to diagnose MCI. In this study, we employ sequence-to-sequence deep autoencoders in order to extract compact, robust and efficient attributes from the spontaneous speech of 25 MCI subjects and 25 healthy controls. From our results, this approach gives a competitive performance, as we significantly outperformed x-vectors even though they were trained on more data. Our additional efforts to identify mild Alzheimer’s (mAD) subjects as well were less successful; but since the focus is on the early detection of dementia, this is not a limitation of the methodology from a practical point of view. Mercedes Vetráb, José Vicente Egas López, Réka Balogh, Nóra Imre, Ildikó Hoffmann, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya |
ICASSP | 9 |
| 2022 | Identification of Subjects Wearing a Surgical Mask from Their Speech by Means of X-vectors and Fisher Vectors
José Vicente Egas López, Gábor Gosztolya |
MDAI | 2 |
| 2022 | Linguistic Parameters of Spontaneous Speech for Identifying Mild Cognitive Impairment and Alzheimer DiseaseabstractAbstract In this article, we seek to automatically identify Hungarian patients suffering from mild cognitive impairment (MCI) or mild Alzheimer disease (mAD) based on their speech transcripts, focusing only on linguistic features. In addition to the features examined in our earlier study, we introduce syntactic, semantic, and pragmatic features of spontaneous speech that might affect the detection of dementia. In order to ascertain the most useful features for distinguishing healthy controls, MCI patients, and mAD patients, we carry out a statistical analysis of the data and investigate the significance level of the extracted features among various speaker group pairs and for various speaking tasks. In the second part of the article, we use this rich feature set as a basis for an effective discrimination among the three speaker groups. In our machine learning experiments, we analyze the efficacy of each feature group separately. Our model that uses all the features achieves competitive scores, either with or without demographic information (3-class accuracy values: 68%–70%, 2-class accuracy values: 77.3%–80%). We also analyze how different data recording scenarios affect linguistic features and how they can be productively used when distinguishing MCI patients from healthy controls. Veronika Vincze, Martina Katalin Szabó, Ildikó Hoffmann, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya |
Comput. Linguistics | 7 |
| 2022 | Automatic screening of mild cognitive impairment and Alzheimer's disease by means of posterior-thresholding hesitation representation
José Vicente Egas López, Réka Balogh, Nóra Imre, Ildikó Hoffmann, Martina Katalin Szabó, László Tóth 0001, Magdolna Pákáski, János Kálmán, Gábor Gosztolya |
Comput. Speech Lang. | 9 |
| 2022 | Optimizing class priors to improve the detection of social signals in audio data
Gábor Gosztolya |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Estimating the degree of conflict in speech by employing Bag-of-Audio-Words and Fisher Vectors
Gábor Gosztolya |
Expert Syst. Appl. | 1 |
| 2021 | Deep Neural Network Embeddings for the Estimation of the Degree of SleepinessabstractEstimating the degree of sleepiness from the human speech is an emerging research problem with straightforward applications. In this study, we employ the x-vector approach, currently the state-of-the-art in speaker recognition, as a neural network feature extractor to detect the level of sleepiness of a speaker. Besides using different corpora for fitting the x- vector DNN, we also experiment with adding noise and reverberation to the training samples. According to our experimental results for the publicly available Dusseldorf Sleepy Language Corpus, utilizing x-vector embeddings as features for Support Vector Regression consistently leads to competitive performance scores in sleepiness detection. In particular, we present the highest Spearman's correlation coefficient on the public corpus that was achieved by a single method. José Vicente Egas López, Gábor Gosztolya |
ICASSP | 2 |
| 2021 | Identifying Conflict Escalation and Primates by Using Ensemble X-Vectors and Fisher Vector FeaturesabstractComputational paralinguistics is concerned with the automatic identification of non-verbal information in human speech.The Interspeech ComParE challenge features new paralinguistic tasks each year; this time, among others, a cross-corpus conflict escalation task and the identification of primates based solely on audio are the actual problems set.In our entry to ComParE 2021, we utilize x-vectors and Fisher vectors as features.To improve the robustness of the predictions, we also experiment with building an ensemble of classifiers from the x-vectors.Lastly, we exploit the fact that the Escalation Sub-Challenge is a conflict detection task, and incorporate the SSPNet Conflict Corpus in our training workflow.Using these approaches, at the time of writing, we had already surpassed the official Challenge baselines on both tasks, which demonstrates the efficiency of the employed techniques. José Vicente Egas López, Mercedes Vetráb, László Tóth 0001, Gábor Gosztolya |
Interspeech | 4 |
| 2021 | Neural Speaker Embeddings for Ultrasound-Based Silent Speech InterfacesabstractArticulatory-to-acoustic mapping seeks to reconstruct speech from a recording of the articulatory movements, for example, an ultrasound video. Just like speech signals, these recordings represent not only the linguistic content, but are also highly specific to the actual speaker. Hence, due to the lack of multi-speaker data sets, researchers have so far concentrated on speaker-dependent modeling. Here, we present multi-speaker experiments using the recently published TaL80 corpus. To model speaker characteristics, we adjusted the x-vector framework popular in speech processing to operate with ultrasound tongue videos. Next, we performed speaker recognition experiments using 50 speakers from the corpus. Then, we created speaker embedding vectors and evaluated them on the remaining speakers. Finally, we examined how the embedding vector influences the accuracy of our ultrasound-to-speech conversion network in a multi-speaker scenario. In the experiments we attained speaker recognition error rates below 3%, and we also found that the embedding vectors generalize nicely to unseen speakers. Our first attempt to apply them in a multi-speaker silent speech framework brought about a marginal reduction in the error rate of the spectral estimation step. Amin Honarmandi Shandiz, László Tóth 0001, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó |
Interspeech | 3 |
| 2021 | Cross-lingual detection of mild cognitive impairment based on temporal parameters of spontaneous speech
Gábor Gosztolya, Réka Balogh, Nóra Imre, José Vicente Egas López, Ildikó Hoffmann, Veronika Vincze, László Tóth 0001, Davangere P. Devanand, Magdolna Pákáski, János Kálmán |
Comput. Speech Lang. | 1 |
| 2021 | Ensemble Bag-of-Audio-Words Representation Improves Paralinguistic Classification AccuracyabstractA recently introduced, effective feature extraction technique for computational paralinguistics is that of Bag-of-Audio-Words (BoAW), where we cluster the frame-level training vectors, and represent each speech utterance based on the cluster of its frames. Over the past few years, several improvements have been proposed for the original BoAW approach, but none of them has examined the impact of the stochastic nature of the clustering step. In this study we demonstrate experimentally that the random factor present in the BoAW clustering step is indeed propagated into the next classification step, eventually leading to suboptimal classification performance. As a solution, we propose to train an ensemble of classifiers; that is, we repeat the BoAW codebook selection step several times, train separate classifier models for these BoAW representation versions and combine their predictions. Our results, obtained for three different paralinguistic datasets, demonstrate that this ensemble technique makes the whole paralinguistic classification process more robust, and it leads to improvements in the classification performance. We tested this technique on three different paralinguistic datasets, and achieved the highest Unweighted Average Recall score reported so far on the iHEARu-EAT corpus. Gábor Gosztolya, Róbert Busa-Fekete |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Ultrasound-Based Articulatory-to-Acoustic Mapping with WaveGlow Speech SynthesisabstractFor articulatory-to-acoustic mapping using deep neural networks, typically spectral and excitation parameters of vocoders have been used as the training targets. However, vocoding often results in buzzy and muffled final speech quality. Therefore, in this paper on ultrasound-based articulatory-to-acoustic conversion, we use a flow-based neural vocoder (WaveGlow) pre-trained on a large amount of English and Hungarian speech data. The inputs of the convolutional neural network are ultrasound tongue images. The training target is the 80-dimensional mel-spectrogram, which results in a finer detailed spectral representation than the previously used 25-dimensional Mel-Generalized Cepstrum. From the output of the ultrasound-to-mel-spectrogram prediction, WaveGlow inference results in synthesized speech. We compare the proposed WaveGlow-based system with a continuous vocoder which does not use strict voiced/unvoiced decision when predicting F0. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the WaveGlow neural vocoder produces significantly more natural synthesized speech than the baseline system. Besides, the advantage of WaveGlow is that F0 is included in the mel-spectrogram representation, and it is not necessary to predict the excitation separately. Tamás Gábor Csapó, Csaba Zainkó, László Tóth 0001, Gábor Gosztolya, Alexandra Markó |
INTERSPEECH | 4 |
| 2020 | Very Short-Term Conflict Intensity Estimation Using Fisher VectorsabstractThe automatic detection of conflict situations from human speech has several applications like obtaining feedback of employees in call centers, the surveillance of public spaces, and other roles in human-computer interactions.Although several methods have been developed to automatic conflict detection, they were designed to operate on relatively long utterances.In practice, however, it would be beneficial to process much shorter speech segments.With the traditional workflow of paralinguistic speech processing, this would require properly annotated training and testing material consisting of short clips.In this study we show that Support Vector Regression machine learning models using Fisher vectors as features, even when trained on longer utterances, allow us to efficiently and accurately detect conflict intensity from very short audio segments.Even without having reliable annotations of these such short chunks, the mean scores of the predictions corresponding to short segments of the same original, longer utterances correlate well to the reference manual annotation.We also verify the validity of this approach by comparing the SVM predictions of the chunks with a manual annotation for the full and the 5-second-long cases.Our findings allow the construction of conflict detection systems having smaller delay, therefore being more useful in practice. Gábor Gosztolya |
INTERSPEECH | 1 |
| 2020 | Making a Distinction Between Schizophrenia and Bipolar Disorder Based on Temporal Parameters in Spontaneous SpeechabstractSchizophrenia is a heterogeneous chronic and severe mental disorder.There are several different theories for the development of schizophrenia from an etiological point of view: neurochemical, neuroanatomical, psychological and genetic factors may also be present in the background of the disease.In this study, we examined spontaneous speech productions by patients suffering from schizophrenia (SCH) and bipolar disorder (BD).We extracted 15 temporal parameters from the speech excerpts and used machine learning techniques for distinguishing the SCH and BD groups, their subgroups (SCH-S and SCH-Z) and subtypes (BD-I and BD-II).Our results indicated, that there is a notable difference between spontaneous speech productions of certain subgroups, while some appears to be indistinguishable for the used classification model.Firstly, SCH and BD groups were found to be different.Secondly, the results of SCH-S subgroup were distinct from BD. Thirdly, the spontaneous speech of the SCH-Z subgroup was found to be very similar to the BD-I, however, it was sharply distinct from BD-II.Our detailed examination highlighted the indistinguishable subgroups and led to us to make our S and Z theory more clarified. Gábor Gosztolya, Anita Bagi, Szilvia Szalóki, István Szendi, Ildikó Hoffmann |
INTERSPEECH | 1 |
| 2020 | Social Signal Detection by Probabilistic Sampling DNN TrainingabstractWhen our task is to detect social signals such as laughter and filler events in an audio recording, the most straightforward way is to apply a Hidden Markov Model-or a Hidden Markov Model/Deep Neural Network (HMM/DNN) hybrid, which is considered state-of-the-art nowadays. In this hybrid model, the DNN component is trained on frame-level samples of the classes we are looking for. In such event detection tasks, however, the training labels are seriously imbalanced, as typically only a small fraction of the training data corresponds to these social signals, while the bulk of the utterances consists of speech segments or silence. A strong imbalance of the training classes is known to cause difficulties during DNN training. To alleviate these problems, here we apply the technique called probabilistic sampling, which seeks to balance the class distribution. Probabilistic sampling is a mathematically well-founded combination of upsampling and downsampling, which was found to outperform both of these simple resampling approaches. With this strategy, we managed to achieve a 7-8 percent relative error reduction both at the segment level and frame level, and we efficiently reduced the DNN training times as well. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech InterfacesabstractWhen using ultrasound video as input, Deep Neural Network-based Silent Speech Interfaces usually rely on the whole image to estimate the spectral parameters required for the speech synthesis step. Although this approach is quite straightforward, and it permits the synthesis of understandable speech, it has several disadvantages as well. Besides the inability to capture the relations between close regions (i.e. pixels) of the image, this pixelby-pixel representation of the image is also quite uneconomical. It is easy to see that a significant part of the image is irrelevant for the spectral parameter estimation task as the information stored by the neighbouring pixels is redundant, and the neural network is quite large due to the large number of input features. To resolve these issues, in this study we train an autoencoder neural network on the ultrasound image; the estimation of the spectral speech parameters is done by a second DNN, using the activations of the bottleneck layer of the autoencoder network as features. In our experiments, the proposed method proved to be more efficient than the standard approach: the measured normalized mean squared error scores were lower, while the correlation values were higher in each case. Based on the result of a listening test, the synthesized utterances also sounded more natural to native speakers. A further advantage of our proposed approach is that, due to the (relatively) small size of the bottleneck layer, we can utilize several consecutive ultrasound images during estimation without a significant increase in the network size, while significantly increasing the accuracy of parameter estimation. Gábor Gosztolya, Ádám Pintér, László Tóth 0001, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
IJCNN | 1 |
| 2019 | Ultrasound-Based Silent Speech Interface Built on a Continuous VocoderabstractRecently it was shown that within the Silent Speech Interface (SSI) field, the prediction of F0 is possible from Ultrasound Tongue Images (UTI) as the articulatory input, using Deep Neural Networks for articulatory-to-acoustic mapping. Moreover, text-to-speech synthesizers were shown to produce higher quality speech when using a continuous pitch estimate, which takes non-zero pitch values even when voicing is not present. Therefore, in this paper on UTI-based SSI, we use a simple continuous F0 tracker which does not apply a strict voiced / unvoiced decision. Continuous vocoder parameters (ContF0, Maximum Voiced Frequency and Mel-Generalized Cepstrum) are predicted using a convolutional neural network, with UTI as input. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the continuous F0 is predicted with lower error, and the continuous vocoder produces slightly more natural synthesized speech than the baseline vocoder using standard discontinuous F0. Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 4 |
| 2019 | Using Fisher Vector and Bag-of-Audio-Words Representations to Identify Styrian Dialects, Sleepiness, Baby & Orca SoundsabstractThe 2019 INTERSPEECH Computational Paralinguistics Challenge (ComParE) consists of four Sub-Challenges, where the tasks are to identify different German (Austrian) dialects, estimate how sleepy the speaker is, what type of sound a given baby uttered, and whether there is a sound of an orca (killer whale) present in the recording.Following our team's last year entry, we continue our research by looking for feature set types that might be employed on a wide variety of tasks without alteration.This year, besides the standard 6373-sized ComParE functionals, we experimented with the Fisher vector representation along with the Bag-of-Audio-Words technique.To adapt Fisher vectors from the field of image processing, we utilized them on standard MFCC features instead of the originally intended SIFT attributes (which describe local objects found in the image).Our results indicate that using these feature representation techniques was indeed beneficial, as we could outperform the baseline values in three of the four Sub-Challenges; the performance of our approach seems to be even higher if we consider that the baseline scores were obtained by combining different methods as well. Gábor Gosztolya |
INTERSPEECH | 1 |
| 2019 | Using the Bag-of-Audio-Word Feature Representation of ASR DNN Posteriors for Paralinguistic ClassificationabstractThe Bag-of-Audio-Word (or BoAW) representation is an utterance-level feature representation approach that was successfully applied in the past in various computational paralinguistic tasks.Here, we extend the BoAW feature extraction process with the use of Deep Neural Networks: first we train a DNN acoustic model on an acoustic dataset consisting of 22 hours of speech for phoneme identification, then we evaluate this DNN on a standard paralinguistic dataset.To construct utterance-level features from the frame-level posterior vectors, we calculate their BoAW representation.We found that this approach can be utilized even on its own, although the results obtained lag behind those of the standard paralinguistic approach, and the optimal size of the extracted feature vectors tends to be large.Our approach, however, can be easily and efficiently combined with the standard paralinguistic one, resulting in the highest Unweighted Average Recall (UAR) score achieved so far for a general paralinguistic dataset. Gábor Gosztolya |
INTERSPEECH | 1 |
| 2019 | Calibrating DNN Posterior Probability Estimates of HMM/DNN Models to Improve Social Signal Detection from Audio DataabstractTo detect social signals such as laughter or filler events from audio data, a straightforward choice is to apply a Hidden Markov Model (HMM) in combination with a Deep Neural Network (DNN) that supplies the local class posterior estimates (HMM/DNN hybrid model).However, the posterior estimates of the DNN may be suboptimal due to a mismatch between the cost function used during training (e.g.frame-level crossentropy) and the actual evaluation metric (e.g.segment-level F1 score).In this study, we show experimentally that by employing a simple posterior probability calibration technique on the DNN outputs, the performance of the HMM/DNN workflow can be significantly improved.Specifically, we apply a linear transformation on the activations of the output layer right before using the softmax function, and fine-tune the parameters of this transformation.Out of the calibration approaches tested, we got the best F1 scores when the posterior calibration process was adjusted so as to maximize the actual HMM-based evaluation metric. Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 1 |
| 2019 | Assessing Parkinson's Disease from Speech Using Fisher VectorsabstractParkinson's Disease (PD) is a neuro-degenerative disorder that affects primarily the motor system of the body.Besides other functions, the subject's speech also deteriorates during the disease, which allows for a non-invasive way of automatic screening.In this study, we represent the utterances of subjects having PD and those of healthy controls by means of the Fisher Vector approach.This technique is very common in the area of image recognition, where it provides a representation of the local image descriptors via frequency and high order statistics.In the present work, we used four frame-level feature sets as the input of the FV method, and applied (linear) Support Vector Machines (SVM) for classifying the speech of subjects.We found that our approach offers superior performance compared to classification based on the i-vector and cosine distance approach, and it also provides an efficient combination of machine learning models trained on different feature sets or on different speaker tasks. José Vicente Egas López, Juan Rafael Orozco-Arroyave, Gábor Gosztolya |
INTERSPEECH | 3 |
| 2019 | Identifying Mild Cognitive Impairment and mild Alzheimer's disease based on spontaneous speech using ASR and linguistic features
Gábor Gosztolya, Veronika Vincze, László Tóth 0001, Magdolna Pákáski, János Kálmán, Ildikó Hoffmann |
Comput. Speech Lang. | 1 |
| 2019 | Posterior-thresholding feature extraction for paralinguistic speech classification
Gábor Gosztolya |
Knowl. Based Syst. | 1 |
| 2019 | Calibrating AdaBoost for phoneme classification
Gábor Gosztolya, Róbert Busa-Fekete |
Soft Comput. | 1 |
| 2018 | F0 Estimation for DNN-Based Ultrasound Silent Speech InterfacesabstractState-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solution is based on the assumption that the F0 curve is unpredictable from articulatory data that does not contain direct measurements of the vocal fold vibration. Here, we experimented with deep neural networks to perform articulatory-to-acoustic conversion from ultrasound images, with an emphasis on estimating the voicing feature and the F0 curve from the ultrasound input. Contrary to the common belief that F0 is unpredictable, we attained a correlation rate of 0.74 between the original and the predicted F0 curve. What is more, the listening tests revealed that our subjects could not distinguish the sentences synthesized using the DNN-estimated and the original F0 curve, and ranked them as having the same quality. Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Tamás Gábor Csapó, Alexandra Markó |
ICASSP | 2 |
| 2018 | Identifying Schizophrenia Based on Temporal Parameters in Spontaneous SpeechabstractSchizophrenia is a neurodegenerative disease with spectrum disorder, consisting of groups of different deficits.It is, among other symptoms, characterized by reduced information processing speed and deficits in verbal fluency.In this study we focus on the speech production fluency of patients with schizophrenia compared to healthy controls.Our aim is to show that a temporal speech parameter set consisting of articulation tempo, speech tempo and various pause-related indicators, originally defined for the sake of early detection of various dementia types such as Mild Cognitive Impairment and early Alzheimer's Disease, is able to capture specific differences in the spontaneous speech of the two groups.We tested the applicability of the temporal indicators by machine learning (i.e. by using Support-Vector Machines).Our results show that members of the two speaker groups could be identified with classification accuracy scores of between 70 -80% and F-measure scores between 81% and 87%.Our detailed examination revealed that, among the pause-related temporal parameters, the most useful for distinguishing the two speaker groups were those which took into account both the silent and filled pauses. Gábor Gosztolya, Anita Bagi, Szilvia Szalóki, István Szendi, Ildikó Hoffmann |
INTERSPEECH | 1 |
| 2018 | General Utterance-Level Feature Extraction for Classifying Crying Sounds, Atypical & Self-Assessed Affect and Heart BeatsabstractIn the area of computational paralinguistics, there is a growing need for general techniques that can be applied in a variety of tasks, and which can be easily realized using standard and publicly available tools.In our contribution to the 2018 Interspeech Computational Paralinguistic Challenge (ComParE), we test four general ways of extracting features.Besides the standard ComParE feature set consisting of 6373 diverse attributes, we experiment with two variations of Bag-of-Audio-Words representations, and define a simple feature set inspired by Gaussian Mixture Models.Our results indicate that the UAR scores obtained via the different approaches vary among the tasks.In our view, this is mainly because most feature sets tested were local by nature, and they could not properly represent the utterances of the Atypical Affect and Self-Assessed Affect Sub-Challenges.On the Crying Sub-Challenge, however, a simple combination of all four feature sets proved to be effective. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 1 |
| 2018 | Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces
László Tóth 0001, Gábor Gosztolya, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
INTERSPEECH | 2 |
| 2018 | User-centric Evaluation of Automatic Punctuation in ASR Closed CaptioningabstractPunctuation of ASR-produced transcripts has received increasing attention in the recent years; RNN-based sequence modelling solutions which exploit textual and/or acoustic features show encouraging performance.Switching the focus from the technical side, qualifying and quantifying the benefits of such punctuation from end-user perspective have not been performed yet exhaustively.The ambition of the current paper is to explore to what extent automatic punctuation can improve human readability and understandability.The paper presents a user-centric evaluation of a real-time closed captioning system enhanced by a lightweight RNN-based punctuation module.Subjective tests involve both normal hearing and deaf or hard-of-hearing (DHH) subjects.Results confirm that automatic punctuation itself significantly increases understandability, even if several other factors interplay in subjective impression.The perceived improvement is even more pronounced in the DHH group.A statistical analysis is carried out to identify objectively measurable factors which are well reflected by subjective scores. Máté Ákos Tündik, György Szaszák, Gábor Gosztolya, András Beke |
INTERSPEECH | 3 |
| 2018 | Posterior Calibration for Multi-Class Paralinguistic ClassificationabstractComputational paralinguistics is an area which contains diverse classification tasks. In many cases the class distribution of these tasks is highly imbalanced by nature, as the phenomena needed to detect in human speech do not occur uniformly. To ignore this imbalance, it is common to measure the efficiency of classification approaches via the Unweighted Average Recall (UAR) metric in this area. However, general classification methods such as Support-Vector Machines (SVM) and Deep Neural Networks (DNNs) were shown to focus on traditional classification accuracy, which might lead to a suboptimal performance for imbalanced datasets. In this study we show that by performing posterior calibration, this effect can be countered and the UAR scores obtained might be improved. Our approach led to relative error reduction values of 4% and 14% on the test set of two multi-class paralinguistic datasets that had imbalanced class distributions, outperforming the traditional downsampling. Gábor Gosztolya, Róbert Busa-Fekete |
SLT | 1 |
| 2018 | Multi-Band Processing With Gabor Filters and Time Delay Neural Nets for Noise Robust Speech RecognitionabstractSpectro-temporal feature extraction and multi-band processing were both invented with the goal of increasing the robustness of speech recognisers. However, although these methods have been in use for a long time now, and they are evidently compatible, few attempts have been made to combine them. This is why here we investigate the combination of multi-band processing with the use of spectro-temporal Gabor filters. First, based on the TIMIT corpus, we optimise their meta-parameters like the overlap, and the number of bands. Then we verify the cross-corpus viability of our multi-band processing approach on the Aurora-4 corpus. Lastly, we combine our method with the recently proposed channel dropout method. Our results show that this combination not only leads to lower error rates than those got using either multi-band processing or channel dropout, but these results compare favourably to those recently reported for the clean training scenario on the Aurora-4 corpus. György Kovács 0001, László Tóth 0001, Gábor Gosztolya |
SLT | 3 |
| 2018 | A feature selection-based speaker clustering method for paralinguistic tasks
Gábor Gosztolya, László Tóth 0001 |
Pattern Anal. Appl. | 1 |
| 2017 | DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech InterfaceabstractIn this paper we present our initial results in articulatory-toacoustic conversion based on tongue movement recordings using Deep Neural Networks (DNNs).Despite the fact that deep learning has revolutionized several fields, so far only a few researchers have applied DNNs for this task.Here, we compare various possible feature representation approaches combined with DNN-based regression.As the input, we recorded synchronized 2D ultrasound images and speech signals.The task of the DNN was to estimate Mel-Generalized Cepstrum-based Line Spectral Pair (MGC-LSP) coefficients, which then served as input to a standard pulse-noise vocoder for speech synthesis.As the raw ultrasound images have a relatively high resolution, we experimented with various feature selection and transformation approaches to reduce the size of the feature vectors.The synthetic speech signals resulting from the various DNN configurations were evaluated both using objective measures and a subjective listening test.We found that the representation that used several neighboring image frames in combination with a feature selection method was preferred both by the subjects taking part in the listening experiments, and in terms of the Normalized Mean Squared Error.Our results may be useful for creating Silent Speech Interface applications in the future. Tamás Gábor Csapó, Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 3 |
| 2017 | Optimized Time Series Filters for Detecting Laughter and Filler Events
Gábor Gosztolya |
INTERSPEECH | 1 |
| 2017 | DNN-Based Feature Extraction and Classifier Combination for Child-Directed Speech, Cold and Snoring IdentificationabstractIn this study we deal with the three sub-challenges of the Interspeech ComParE Challenge 2017, where the goal is to identify child-directed speech, speakers having a cold, and different types of snoring sounds.For the first two sub-challenges we propose a simple, two-step feature extraction and classification scheme: first we perform frame-level classification via Deep Neural Networks (DNNs), and then we extract utterancelevel features from the DNN outputs.By utilizing these features for classification, we were able to match the performance of the standard paralinguistic approach (which involves extracting thousands of features, many of them being completely irrelevant to the actual task).As for the Snoring Sub-Challenge, we divided the recordings into segments, and averaged out some frame-level features segment-wise, which were then used for utterance-level classification.When combining the predictions of the proposed approaches with those got by the standard paralinguistic approach, we managed to outperform the baseline values of the Cold and Snoring sub-challenges on the hidden test sets. Gábor Gosztolya, Róbert Busa-Fekete, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 1 |
| 2017 | Training Context-Dependent DNN Acoustic Models Using Probabilistic SamplingabstractIn current HMM/DNN speech recognition systems, the purpose of the DNN component is to estimate the posterior probabilities of tied triphone states.In most cases the distribution of these states is uneven, meaning that we have a markedly different number of training samples for the various states.This imbalance of the training data is a source of suboptimality for most machine learning algorithms, and DNNs are no exception.A straightforward solution is to re-sample the data, either by upsampling the rarer classes or by dowsampling the more common classes.Here, we experiment with the so-called probabilistic sampling method that applies downsampling and upsampling at the same time.For this, it defines a new class distribution for the training data, which is a linear combination of the original and the uniform class distributions.As an extension to previous studies, we propose a new method to re-estimate the class priors, which is required to remedy the mismatch between the training and the test data distributions introduced by re-sampling.Using probabilistic sampling and the proposed modification we report 5% and 6% relative error rate reductions on the TED-LIUM and on the AMI corpora, respectively. Tamás Grósz, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 2 |
| 2017 | A Comparative Evaluation of GMM-Free State Tying Methods for ASR
Tamás Grósz, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 2 |
| 2017 | DNN-Based Feature Extraction for Conflict Intensity Estimation From SpeechabstractOver past few years, there has been an increasing need to extract nonlinguistic information from audio sources. This trend has created a new area in speech technology known as computational paralinguistics. A task belonging to this area is to estimate the intensity of conflicts arising in speech recordings, based only on the audio information. It was shown that the human comprehension of conflict intensity is closely related to speaker overlap; that is, when multiple persons are speaking at the same time. This type of information can also aid automated conflict intensity estimation. In this study, we propose a simple, DNN-based feature extraction step, and show that this approach is superior to those introduced in the literature so far: By combining our results with an efficient greedy feature selection algorithm, we were able to outperform all previous results on the SSPNet conflict dataset, achieving a correlation coefficient of 0.856 on the test set. Gábor Gosztolya, László Tóth 0001 |
IEEE Signal Process. Lett. | 1 |
| 2016 | Determining Native Language and Deception Using Phonetic Features and Classifier CombinationabstractFor several years, the Interspeech ComParE Challenge has focused on paralinguistic tasks of various kinds.In this paper we focus on the Native Language and the Deception subchallenges of ComParE 2016, where the goal is to identify the native language of the speaker, and to recognize deceptive speech.As both tasks can be treated as classification ones, we experiment with several state-of-the-art machine learning methods (Support-Vector Machines, AdaBoost.MH and Deep Neural Networks), and also test a simple-yet-robust combination method.Furthermore, we will assume that the native language of the speaker affects the pronunciation of specific phonemes in the language he is currently using.To exploit this, we extract phonetic features for the Native Language task.Moreover, for the Deception Sub-Challenge we compensate for the highly unbalanced class distribution by instance re-sampling.With these techniques we are able to significantly outperform the baseline SVM on the unpublished test set. Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001 |
INTERSPEECH | 1 |
| 2016 | Estimating the Sincerity of Apologies in Speech by DNN Rank Learning and Prosodic AnalysisabstractIn the Sincerity Sub-Challenge of the Interspeech ComParE 2016 Challenge, the task is to estimate user-annotated sincerity scores for speech samples.We interpret this challenge as a ranklearning regression task, since the evaluation metric (Spearman's correlation) is calculated from the rank of the instances.As a first approach, Deep Neural Networks are used by introducing a novel error criterion which maximizes the correlation metric directly.We obtained the best performance by combining the proposed error function with the conventional MSE error.This approach yielded results that outperform the baseline on the Challenge test set.Furthermore, we introduce a compact prosodic feature set based on a dynamic representation of F0, energy and sound duration.We extract syllable-based prosodic features which are used as the basis of another machine learning step.We show that a small set of prosodic features is capable of yielding a result very close to the baseline one and that by combining the predictions yielded by DNN and the prosodic feature set, further improvement can be reached, significantly outperforming the baseline SVR on the Challenge test set. Gábor Gosztolya, Tamás Grósz, György Szaszák, László Tóth 0001 |
INTERSPEECH | 1 |
| 2016 | GMM-Free Flat Start Sequence-Discriminative DNN TrainingabstractRecently, attempts have been made to remove Gaussian mixture models (GMM) from the training process of deep neural network-based hidden Markov models (HMM/DNN). For the GMM-free training of a HMM/DNN hybrid we have to solve two problems, namely the initial alignment of the frame-level state labels and the creation of context-dependent states. Although flat-start training via iteratively realigning and retraining the DNN using a frame-level error function is viable, it is quite cumbersome. Here, we propose to use a sequence-discriminative training criterion for flat start. While sequence-discriminative training is routinely applied only in the final phase of model training, we show that with proper caution it is also suitable for getting an alignment of context-independent DNN models. For the construction of tied states we apply a recently proposed KL-divergence-based state clustering method, hence our whole training process is GMM-free. In the experimental evaluation we found that the sequence-discriminative flat start training method is not only significantly faster than the straightforward approach of iterative retraining and realignment, but the word error rates attained are slightly better as well. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 1 |
| 2016 | Detecting Mild Cognitive Impairment from Spontaneous Speech by Correlation-Based Phonetic Feature SelectionabstractMild Cognitive Impairment (MCI), sometimes regarded as a prodromal stage of Alzheimer's disease, is a mental disorder that is difficult to diagnose.Recent studies reported that MCI causes slight changes in the speech of the patient.Our previous studies showed that MCI can be efficiently classified by machine learning methods such as Support-Vector Machines and Random Forest, using features describing the amount of pause in the spontaneous speech of the subject.Furthermore, as hesitation is the most important indicator of MCI, we took special care when handling filled pauses, which usually correspond to hesitation.In contrast to our previous studies which employed manually constructed feature sets, we now employ (automatic) correlation-based feature selection methods to find the relevant feature subset for MCI classification.By analyzing the selected feature subsets we also show that features related to filled pauses are useful for MCI detection from speech samples. Gábor Gosztolya, László Tóth 0001, Tamás Grósz, Veronika Vincze, Ildikó Hoffmann, Gréta Szatlóczki, Magdolna Pákáski, János Kálmán |
INTERSPEECH | 1 |
| 2015 | Building context-dependent DNN acoustic models using Kullback-Leibler divergence-based state tyingabstractDeep neural network (DNN) based speech recognizers have recently replaced Gaussian mixture (GMM) based systems as the state-of-the-art. HMM/DNN systems have kept many refinements of the HMM/GMM framework, even though some of these may be suboptimal for them. One such example is the creation of context-dependent tied states, for which an efficient decision tree state tying method exists. The tied states used to train DNNs are usually obtained using the same tying algorithm, even though it is based on likelihoods of Gaussians. In this paper, we investigate an alternative state clustering method that uses the Kullback-Leibler (KL) divergence of DNN output vectors to build the decision tree. It has already been successfully applied within the framework of KL-HMM systems, and here we show that it is also beneficial for HMM/DNN hybrids. In a large vocabulary recognition task we report a 4% relative word error rate reduction using this state clustering method. Gábor Gosztolya, Tamás Grósz, László Tóth 0001, David Imseng |
ICASSP | 1 |
| 2015 | Conflict intensity estimation from speech using Greedy forward-backward feature selectionabstractIn the recent years extracting non-trivial information from audio sources has become possible.The resulting data has induced a new area in speech technology known as computational paralinguistics.A task in this area was presented at the ComParE 2013 Challenge (using the SSPNet Conflict Corpus), where the task was to determine the intensity of conflicts arising in speech recordings, based only on the audio information.Most authors approached this task by following standard paralinguistic practice, where we extract a huge number of potential features and perform the actual classification or regression process in the hope that the machine learning method applied is able to completely ignore irrelevant features.Although current stateof-the-art methods can indeed handle an overcomplete feature set, studies show that they can still be aided by feature selection.We opted for a simple greedy feature selection algorithm, by which we were able to outperform all previous scores on the SSPNet Conflict dataset, achieving a UAR score of 85.6%. Gábor Gosztolya |
INTERSPEECH | 1 |
| 2015 | On evaluation metrics for social signal detectionabstractSocial signal detection is a task in speech technology which has recently became more popular.In the Interspeech 2013 Com-ParE Challenge one of the tasks was social signal detection, and since then, new results have been published on the dataset.These studies all used the Area Under Curve (AUC) metric to evaluate the performance; here we argue that this metric is not really suitable for social signals detection.Besides raising some serious theoretical objections, we will also demonstrate this unsuitability experimentally: we will show that applying a very simple smoothing function on the output of the framelevel scores of state-of-the-art classifiers can significantly improve the AUC scores, but perform poorly when employed in a Hidden Markov Model.As the latter is more like real-world applications, we suggest relying on utterance-level evaluation metrics in the future. Gábor Gosztolya |
INTERSPEECH | 1 |
| 2015 | Assessing the degree of nativeness and parkinson's condition using Gaussian processes and deep rectifier neural networksabstractThe Interspeech 2015 Computational Paralinguistics Challenge includes two regression learning tasks, namely the Parkinson's Condition Sub-Challenge and the Degree of Nativeness Sub-Challenge.We evaluated two state-of-the-art machine learning methods on the tasks, namely Deep Neural Networks (DNN) and Gaussian Processes Regression (GPR).We also experiented with various classifier combination and feature selection methods.For the Degree of Nativeness sub-challenge we obtained a far better Spearman correlation value than the one presented in the baseline paper.As regards the Parkinson's Condition Sub-Challenge, we showed that both DNN and GPR are competitive with the baseline SVM, and that the results can be improved further by combining the classifiers.However, we obtained by far the best results when we applied a speaker clustering method to identify the files that belong to the same speaker. Tamás Grósz, Róbert Busa-Fekete, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 3 |
| 2015 | Automatic detection of mild cognitive impairment from spontaneous speech using ASRabstractMild Cognitive Impairment (MCI), sometimes regarded as a prodromal stage of Alzheimer's disease, is a mental disorder that is difficult to diagnose.However, recent studies reported that MCI causes slight changes in the speech of the patient.Our starting point here is a study that found acoustic correlates of MCI, but extracted the proposed features manually.Here, we automate the extraction of the features by applying automatic speech recognition (ASR).Unlike earlier authors, we use ASR to extract only a phonetic level segmentation and annotation.While the phonetic output allows the calculation of features like the speech rate, it avoids the problems caused by the agrammatical speech frequently produced by the targeted patient group.Furthermore, as hesitation is the most important indicator of MCI, we take special care when handling filled pauses, which usually correspond to hesitation.Using the ASR-based features, we employ machine learning methods to separate the subjects with MCI from the control group.The classification results obtained with ASR-based feature extraction are just slightly worse that those got with the manual method.The F1 value achieved (85.3) is very promising regarding the creation of an automated MCI screening application. László Tóth 0001, Gábor Gosztolya, Veronika Vincze, Ildikó Hoffmann, Gréta Szatlóczki, Edit Biró, Fruzsina Zsura, Magdolna Pákáski, János Kálmán |
INTERSPEECH | 2 |
| 2014 | Detecting the intensity of cognitive and physical load using AdaBoost and deep rectifier neural networksabstractThe Interspeech ComParE 2014 Challenge consists of two machine learning tasks, which have quite a small number of examples. Due to our good results in ComParE 2013, we considered AdaBoost a suitable machine learning meta-algorithm for these tasks, besides we also experimented with Deep Rectifier Neural Networks. These differ from traditional neural networks in that the former have several hidden layers, and use rectifier neurons as hidden units. With AdaBoost we achieved competitive results, whereas with the neural networks we were able to outperform baseline SVM scores in both Sub-Challenges. Index Terms :s peech technology, AdaBoost, deep neural networks, rectifier activation function Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001 |
INTERSPEECH | 1 |
| 2014 | Applying Representative Uninorms for Phonetic Classifier Combination
Gábor Gosztolya, József Dombi 0001 |
MDAI | 1 |
| 2013 | Detecting autism, emotions and social signals using adaboostabstractIn the area of speech technology, tasks that involve the extraction of non-lingustic information have been receiving more attention recently. The Computational Paralinguistics Challenge (ComParE 2013) sought to develop techniques to efficiently detect a number of paralinguistic events, including the detection of non-linguistic events (laughter and fillers) in speech recordings as well as categorizing whole (albeit short) recordings by speaker emotion, conflict or the presence of development disorders (autism). We treated these sub-challenges as general classification tasks and applied the general-purpose machine learning meta-algorithm, AdaBoost.MH, and its recently proposed variant, AdaBoost.MH.BA, to them. The results show that these new algorithms convincingly outperform baseline SVM scores. Gábor Gosztolya, Róbert Busa-Fekete, László Tóth 0001 |
INTERSPEECH | 1 |
| 2013 | Using the Logarithmic Generator Function in the Spoken Term Detection Task
Gábor Gosztolya |
MDAI | 1 |
| 2008 | Cross-lingual portability of MLP-based tandem features - a case study for English and HungarianabstractOne promising approach for building ASR systems for lessresourced languages is cross-lingual adaptation.Tandem ASR is particularly well suited to such adaptation, as it includes two cascaded modelling steps: feature extraction using multi-layer perceptrons (MLPs), followed by modelling using a standard HMM.The language-specific tuning can be performed by adjusting the HMM only, leaving the MLP untouched.Here we examine the portability of feature extractor MLPs between an Indo-European (English) and a Finno-Ugric (Hungarian) language.We present experiments which use both conventional phone-posterior and articulatory feature (AF) detector MLPs, both trained on a much larger quantity of (English) data than the monolingual (Hungarian) system.We find that the cross-lingual configurations achieve similar performance to the monolingual system, and that, interestingly, the AF detectors lead to slightly worse performance, despite the expectation that they should be more language-independent than phone-based MLPs.However, the cross-lingual system outperforms all other configurations when the English phone MLP is adapted on the Hungarian data. László Tóth 0001, Joe Frankel, Gábor Gosztolya, Simon King 0001 |
INTERSPEECH | 3 |
| 2005 | Speeding Up Dynamic Search Methods in Speech Recognition
Gábor Gosztolya, András Kocsor |
IEA/AIE | 1 |
| 2004 | Replicator Neural Networks for Outlier Modeling in Segmental Speech Recognition
László Tóth 0001, Gábor Gosztolya |
ISNN (1) | 2 |
| 2003 | Improving the Multi-stack Decoding Algorithm in a Segment-Based Speech Recognizer
Gábor Gosztolya, András Kocsor |
IEA/AIE | 1 |