EDBT 2026 Demo / reviewers in the wild / expert
Tamás Grósz
dblp:133/1170
· DBLP profile ↗
37ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0001-7918-9579ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single model capable of multilingual SER by distilling knowledge from multiple teacher models. To address this, we introduce a novel language-aware multi-teacher knowledge distillation method to advance SER in English, Finnish, and French. It leverages Wav2Vec2.0 as the foundation of monolingual teacher models and then distills their knowledge into a single multilingual student model. The student model demonstrates state-of-the-art performance, with a weighted recall of 72.9 on the English dataset and an unweighted recall of 63.4 on the Finnish dataset, surpassing fine-tuning and knowledge distillation baselines. Our method excels in improving recall for sad and neutral emotions, although it still faces challenges in recognizing anger and happiness. Mehedi Hasan Bijoy, Dejan Porjazovski, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 3 |
| 2025 | Is your model big enough? Training and interpreting large-scale monolingual speech foundation models
Yaroslav Getman, Tamás Grósz, Tommi Lehtonen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2025 | Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland SwedishabstractMispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%). Nhan Phan, Mikko Kuronen, Maria Kautonen, Riikka Ullakonoja, Anna von Zansen, Yaroslav Getman, Ekaterina Voskoboinik, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 8 |
| 2024 | Collecting Linguistic Resources for Assessing Children's Pronunciation of Nordic LanguagesabstractThis paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD). The paper presents the challenges encountered recording and annotating data for Finnish, Swedish and Norwegian, as well as the ethical considerations related with making this data publicly available. We hope that sharing this experience will encourage others to collect similar data for other languages. Of the different data collections, we were able to make the Norwegian corpus publicly available in the hope that it will serve as a reference in pronunciation assessment research. Anne Marte Haug Olstad, Anna-Riikka Smolander, Sofia Strömbergsson, Sari Ylinen, Minna Lehtonen, Mikko Kurimo, Yaroslav Getman, Tamás Grósz, Xinwei Cao, Torbjørn Svendsen, Giampiero Salvi |
LREC/COLING | 8 |
| 2024 | Investigating the Clusters Discovered By Pre-Trained AV-HuBERTabstractSelf-supervised models, such as HuBERT and its audio-visual version AV-HuBERT, have demonstrated excellent performance on various tasks. The main factor for their success is the pre-training procedure, which requires only raw data without human transcription. During the self-supervised pre-training phase, HuBERT is trained to discover latent clusters in the training data, but these clusters are discarded, and only the last hidden layer is used by the conventional finetuning step. We investigate what latent information the AV-HuBERT model managed to uncover via its clusters and can we use them directly for speech recognition. To achieve this, we consider the sequence of cluster ids as a ’language’ developed by the AV-HuBERT and attempt to translate it to English text via small LSTM-based models. These translation models enable us to investigate the relations between the clusters and the English alphabet, shedding light on groups of latent clusters specialized to recognise specific phonetic groups. Our results demonstrate that using the pre-trained system as a quantizer, we are able to compress the video to as low as 275 bit/sec while maintaining acceptable speech recognition accuracy. Furthermore, compared to the conventional finetuning step, our solution has considerably lower computational cost. Anja Virkkunen, Marek Sarvas, Guangpu Huang, Tamás Grósz, Mikko Kurimo |
ICASSP | 4 |
| 2024 | Exploring adaptation techniques of large speech foundation models for low-resource ASR: a case study on Northern SámiabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Yaroslav Getman, Tamás Grósz, Katri Hiovain-Asikainen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2024 | What happens in continued pre-training? Analysis of self-supervised speech models with continued pre-training for colloquial Finnish ASRabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Yaroslav Getman, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 2 |
| 2024 | Oversampling, Augmentation and Curriculum Learning for Speaking Assessment with Limited Training DataabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Tin Mei Lun, Ekaterina Voskoboinik, Ragheb Al-Ghezi, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 4 |
| 2024 | CaptainA self-study mobile app for practising speaking: task completion assessment and feedback with generative AI
Nhan Phan, Anna von Zansen, Maria Kautonen, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 4 |
| 2024 | Automated content assessment and feedback for Finnish L2 learners in a picture description speaking taskabstractWe propose a framework to address several unsolved challenges in second language (L2) automatic speaking assessment (ASA) and feedback. The challenges include: 1. ASA of visual task completion, 2. automated content grading and explanation of spontaneous L2 speech, 3. corrective feedback generation for L2 learners, and 4. all the above for a language that has minimal speech data of L2 learners. The proposed solution combines visual natural language generation (NLG), automatic speech recognition (ASR) and prompting a large language model (LLM) for low-resource L2 learners. We describe the solution and the outcomes of our case study for a picture description task in Finnish. Our results indicate substantial agreement with human experts in grading, explanation and feedback. This framework has the potential for a significant impact in constructing next-generation computer-assisted language learning systems to provide automatic scoring with feedback for learners of low-resource languages. Nhan Phan, Anna von Zansen, Maria Kautonen, Ekaterina Voskoboinik, Tamás Grósz, Raili Hildén, Mikko Kurimo |
INTERSPEECH | 5 |
| 2024 | Comparison and analysis of new curriculum criteria for end-to-end ASRabstractTraditionally, teaching a human and a Machine Learning (ML) model is quite different, but organized and structured learning has the ability to enable faster and better understanding of the underlying concepts. For example, when humans learn to speak, they first learn how to utter basic phones and then slowly move towards more complex structures such as words and sentences. Motivated by this observation, researchers have started to adapt this approach for training ML models. Since the main concept, the gradual increase in difficulty, resembles the notion of the curriculum in education, the methodology became known as Curriculum Learning (CL). In this work, we design and test new CL approaches to train Automatic Speech Recognition systems, specifically focusing on the so-called end-to-end models. These models consist of a single, large-scale neural network that performs the recognition task, in contrast to the traditional way of having several specialized components focusing on different subtasks (e.g., acoustic and language modeling). We demonstrate that end-to-end models can achieve better performances if they are provided with an organized training set consisting of examples that exhibit an increasing level of difficulty. To impose structure on the training set and to define the notion of an easy example, we explored multiple solutions that use either external, static scoring methods or incorporate feedback from the model itself. In addition, we examined the effect of pacing functions that control how much data is presented to the network during each training epoch. Our proposed curriculum learning strategies were tested on the task of speech recognition on two data sets, one containing spontaneous Finnish speech where volunteers were asked to speak about a given topic, and one containing planned English speech. Empirical results showed that a good curriculum strategy can yield performance improvements and speed-up convergence. After a given number of epochs, our best strategy achieved a 5.6% and 3.4% decrease in terms of test set word error rate for the Finnish and English data sets, respectively. Georgios Karakasidis, Mikko Kurimo, Peter Bell 0001, Tamás Grósz |
Speech Commun. | 4 |
| 2024 | From Raw Speech to Fixed Representations: A Comprehensive Evaluation of Speech Embedding TechniquesabstractSpeech embeddings, fixed-size representations derived from raw audio data, play a crucial role in diverse machine learning applications. Despite the abundance of speech embedding techniques, selecting the most suitable one remains challenging. Existing studies often focus on intrinsic or extrinsic aspects, seldom exploring both simultaneously. Furthermore, comparing the state-of-the-art pre-trained models with prior speech embedding solutions is notably scarce in the literature. To address these gaps, we undertake a comprehensive evaluation of both small and large-scale speech embedding models, which, in our opinion, needs to incorporate both intrinsic and extrinsic assessments. The intrinsic experiments delve into the models' ability to pick speaker-related characteristics and assess their discriminative capacities, providing insights into their inherent capabilities and internal workings. Concurrently, the extrinsic experiments evaluate whether the models learned semantic cues during pre-training. The findings underscore the superior performance of the large-scale pre-trained models, albeit at an elevated computational cost. The base self-supervised models show comparable results to their large counterparts, making them a better choice for many applications. Furthermore, we show that by selecting the most crucial dimensions, the models' performance often does not suffer drastically and even improves in some cases. This research contributes valuable insights into the nuanced landscape of speech embeddings, aiding researchers and practitioners in making informed choices for various applications. Dejan Porjazovski, Tamás Grósz, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Principled Comparisons for End-to-End Speech Recognition: Attention vs Hybrid at the 1000-Hour ScaleabstractEnd-to-End speech recognition has become the center of attention for speech recognition research, but Hybrid Hidden Markov Model Deep Neural Network (HMM/DNN) -systems remain a competitive approach in terms of performance. End-to-End models may be better at very large data scales, and HMM / DNN-systems may have an advantage in low-resource scenarios, but the thousand-hour scale is particularly interesting for comparisons. At that scale experiments have not been able to conclusively demonstrate which approach is best, or if the heterogeneous approaches yield similar results. In this work, we work towards answering that question for Attention-based Encoder-Decoder models compared with HMM / DNN-systems. We present two simple experimental design principles, and how to build systems adhering to those principles. We demonstrate how those principles remove confounding variables related to both data, and neural architecture and training. We apply the principles in a set of experiments on three diverse thousand-hour-scale tasks. In our experiments, the HMM / DNN-systems yield equal or better results in almost all cases. Aku Rouhe, Tamás Grósz, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Investigating wav2vec2 context representations and the effects of fine-tuning, a case-study of a Finnish modelabstractSelf-supervised speech models, such as the wav2vec2, have become extremely popular in the past few years. Their main appeal is that after their pre-training on a large amount of audio, they require only a small amount of supervised, finetuning data to achieve outstanding results. Despite their immense success, very little is understood about the pre-trained models and how finetuning changes them. In this work, we take the first steps towards a better understanding of wav2vec2 systems using model interpretation tools such as visualization and latent embedding clustering. Through our analysis, we gain new insights into the abilities of the pre-trained networks and the effect that finetuning has on them. We demonstrate that the clusters learned by the pre-trained model are just as important a factor as the supervised training data distribution in determining the accuracy of the finetuned system, which could aid us in selecting the most suitable pre-trained model for the supervised data. Tamás Grósz, Yaroslav Getman, Ragheb Al-Ghezi, Aku Rouhe, Mikko Kurimo |
INTERSPEECH | 1 |
| 2023 | Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian InferenceabstractLarge pre-trained models are essential in paralinguistic systems, demonstrating effectiveness in tasks like emotion recognition and stuttering detection. In this paper, we employ large pre-trained models for the ACM Multimedia Computational Paralinguistics Challenge, addressing the Requests and Emotion Share tasks. We explore audio-only and hybrid solutions leveraging audio and text modalities. Our empirical results consistently show the superiority of the hybrid approaches over the audio-only models. Moreover, we introduce a Bayesian layer as an alternative to the standard linear output layer. The multimodal fusion approach achieves an 85.4% UAR on HC-Requests and 60.2% on HC-Complaints. The ensemble model for the Emotion Share task yields the best ρ value of .614. The Bayesian wav2vec2 approach, explored in this study, allows us to easily build ensembles, at the cost of fine-tuning only one model. Moreover, we can have usable confidence values instead of the usual overconfident posterior probabilities. Dejan Porjazovski, Yaroslav Getman, Tamás Grósz, Mikko Kurimo |
ACM Multimedia | 3 |
| 2022 | wav2vec2-based Speech Rating System for Children with Speech Sound DisorderabstractThe computational resources were provided by Aalto ScienceIT. This work was supported by NordForsk through the funding to Technology-enhanced foreign and second-language learning of Nordic languages, project number 103893. Yaroslav Getman, Ragheb Al-Ghezi, Katja Voskoboinik, Tamás Grósz, Mikko Kurimo, Giampiero Salvi, Torbjørn Svendsen, Sofia Strömbergsson |
INTERSPEECH | 4 |
| 2022 | Comparison and Analysis of New Curriculum Criteria for End-to-End ASRabstractThe computational resources were provided by Aalto ScienceIT. We are grateful for the Academy of Finland project funding number 345790 in ICT 2023 programme's project”Understanding speech and scene with ears and eyes” Georgios Karakasidis, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 2 |
| 2022 | Wav2vec2-based Paralinguistic Systems to Recognise Vocalised Emotions and StutteringabstractWith the rapid advancement in automatic speech recognition and natural language understanding, a complementary field (paralinguistics) emerged, focusing on the non-verbal content of speech. The ACM Multimedia 2022 Computational Paralinguistics Challenge introduced several exciting tasks of this field. In this work, we focus on tackling two Sub-Challenges using modern, pre-trained models called wav2vec2. Our experimental results demonstrated that wav2vec2 is an excellent tool for detecting the emotions behind vocalisations and recognising different types of stutterings. Albeit they achieve outstanding results on their own, our results demonstrated that wav2vec2-based systems could be further improved by ensembling them with other models. Our best systems outperformed the competition baselines by a considerable margin, achieving an unweighted average recall of 44.0 (absolute improvement of 6.6% over baseline) on the Vocalisation Sub-Challenge and 62.1 (absolute improvement of 21.7% over baseline) on the Stuttering Sub-Challenge. Tamás Grósz, Dejan Porjazovski, Yaroslav Getman, Sudarsana Reddy Kadiri, Mikko Kurimo |
ACM Multimedia | 1 |
| 2020 | Data Augmentation Using Prosody and False Starts to Recognize Non-Native Children's SpeechabstractThis paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from children of various age groups given a limited amount of speech. Moreover, the speech being spontaneous has false starts transcribed as partial words, which in the test transcriptions leads to unseen partial words. To cope with these two challenges, we investigate a data augmentation-based approach. Firstly, we apply the prosody-based data augmentation to supplement the audio data. Secondly, we simulate false starts by introducing partial-word noise in the language modeling corpora creating new words. Acoustic models trained on prosody-based augmented data outperform the models using the baseline recipe or the SpecAugment-based augmentation. The partial-word noise also helps to improve the baseline language model. Our ASR system, a combination of these schemes, is placed third in the evaluation period and achieves the word error rate of 18.71%. Post-evaluation period, we observe that increasing the amounts of prosody-based augmented data leads to better performance. Furthermore, removing low-confidence-score words from hypotheses can lead to further gains. These two improvements lower the ASR error rate to 17.99%. Hemant Kumar Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 3 |
| 2020 | Social Signal Detection by Probabilistic Sampling DNN TrainingabstractWhen our task is to detect social signals such as laughter and filler events in an audio recording, the most straightforward way is to apply a Hidden Markov Model-or a Hidden Markov Model/Deep Neural Network (HMM/DNN) hybrid, which is considered state-of-the-art nowadays. In this hybrid model, the DNN component is trained on frame-level samples of the classes we are looking for. In such event detection tasks, however, the training labels are seriously imbalanced, as typically only a small fraction of the training data corresponds to these social signals, while the bulk of the utterances consists of speech segments or silence. A strong imbalance of the training classes is known to cause difficulties during DNN training. To alleviate these problems, here we apply the technique called probabilistic sampling, which seeks to balance the class distribution. Probabilistic sampling is a mathematically well-founded combination of upsampling and downsampling, which was found to outperform both of these simple resampling approaches. With this strategy, we managed to achieve a 7-8 percent relative error reduction both at the segment level and frame level, and we efficiently reduced the DNN training times as well. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech InterfacesabstractWhen using ultrasound video as input, Deep Neural Network-based Silent Speech Interfaces usually rely on the whole image to estimate the spectral parameters required for the speech synthesis step. Although this approach is quite straightforward, and it permits the synthesis of understandable speech, it has several disadvantages as well. Besides the inability to capture the relations between close regions (i.e. pixels) of the image, this pixelby-pixel representation of the image is also quite uneconomical. It is easy to see that a significant part of the image is irrelevant for the spectral parameter estimation task as the information stored by the neighbouring pixels is redundant, and the neural network is quite large due to the large number of input features. To resolve these issues, in this study we train an autoencoder neural network on the ultrasound image; the estimation of the spectral speech parameters is done by a second DNN, using the activations of the bottleneck layer of the autoencoder network as features. In our experiments, the proposed method proved to be more efficient than the standard approach: the measured normalized mean squared error scores were lower, while the correlation values were higher in each case. Based on the result of a listening test, the synthesized utterances also sounded more natural to native speakers. A further advantage of our proposed approach is that, due to the (relatively) small size of the bottleneck layer, we can utilize several consecutive ultrasound images during estimation without a significant increase in the network size, while significantly increasing the accuracy of parameter estimation. Gábor Gosztolya, Ádám Pintér, László Tóth 0001, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
IJCNN | 4 |
| 2019 | Ultrasound-Based Silent Speech Interface Built on a Continuous VocoderabstractRecently it was shown that within the Silent Speech Interface (SSI) field, the prediction of F0 is possible from Ultrasound Tongue Images (UTI) as the articulatory input, using Deep Neural Networks for articulatory-to-acoustic mapping. Moreover, text-to-speech synthesizers were shown to produce higher quality speech when using a continuous pitch estimate, which takes non-zero pitch values even when voicing is not present. Therefore, in this paper on UTI-based SSI, we use a simple continuous F0 tracker which does not apply a strict voiced / unvoiced decision. Continuous vocoder parameters (ContF0, Maximum Voiced Frequency and Mel-Generalized Cepstrum) are predicted using a convolutional neural network, with UTI as input. The results demonstrate that during the articulatory-to-acoustic mapping experiments, the continuous F0 is predicted with lower error, and the continuous vocoder produces slightly more natural synthesized speech than the baseline vocoder using standard discontinuous F0. Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 5 |
| 2018 | F0 Estimation for DNN-Based Ultrasound Silent Speech InterfacesabstractState-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solution is based on the assumption that the F0 curve is unpredictable from articulatory data that does not contain direct measurements of the vocal fold vibration. Here, we experimented with deep neural networks to perform articulatory-to-acoustic conversion from ultrasound images, with an emphasis on estimating the voicing feature and the F0 curve from the ultrasound input. Contrary to the common belief that F0 is unpredictable, we attained a correlation rate of 0.74 between the original and the predicted F0 curve. What is more, the listening tests revealed that our subjects could not distinguish the sentences synthesized using the DNN-estimated and the original F0 curve, and ranked them as having the same quality. Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Tamás Gábor Csapó, Alexandra Markó |
ICASSP | 1 |
| 2018 | General Utterance-Level Feature Extraction for Classifying Crying Sounds, Atypical & Self-Assessed Affect and Heart BeatsabstractIn the area of computational paralinguistics, there is a growing need for general techniques that can be applied in a variety of tasks, and which can be easily realized using standard and publicly available tools.In our contribution to the 2018 Interspeech Computational Paralinguistic Challenge (ComParE), we test four general ways of extracting features.Besides the standard ComParE feature set consisting of 6373 diverse attributes, we experiment with two variations of Bag-of-Audio-Words representations, and define a simple feature set inspired by Gaussian Mixture Models.Our results indicate that the UAR scores obtained via the different approaches vary among the tasks.In our view, this is mainly because most feature sets tested were local by nature, and they could not properly represent the utterances of the Atypical Affect and Self-Assessed Affect Sub-Challenges.On the Crying Sub-Challenge, however, a simple combination of all four feature sets proved to be effective. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 2 |
| 2018 | Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces
László Tóth 0001, Gábor Gosztolya, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó |
INTERSPEECH | 3 |
| 2018 | Efficient visual code localization with neural networks
Péter Bodnár, Tamás Grósz, László Tóth 0001, László G. Nyúl |
Pattern Anal. Appl. | 2 |
| 2017 | DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech InterfaceabstractIn this paper we present our initial results in articulatory-toacoustic conversion based on tongue movement recordings using Deep Neural Networks (DNNs).Despite the fact that deep learning has revolutionized several fields, so far only a few researchers have applied DNNs for this task.Here, we compare various possible feature representation approaches combined with DNN-based regression.As the input, we recorded synchronized 2D ultrasound images and speech signals.The task of the DNN was to estimate Mel-Generalized Cepstrum-based Line Spectral Pair (MGC-LSP) coefficients, which then served as input to a standard pulse-noise vocoder for speech synthesis.As the raw ultrasound images have a relatively high resolution, we experimented with various feature selection and transformation approaches to reduce the size of the feature vectors.The synthetic speech signals resulting from the various DNN configurations were evaluated both using objective measures and a subjective listening test.We found that the representation that used several neighboring image frames in combination with a feature selection method was preferred both by the subjects taking part in the listening experiments, and in terms of the Normalized Mean Squared Error.Our results may be useful for creating Silent Speech Interface applications in the future. Tamás Gábor Csapó, Tamás Grósz, Gábor Gosztolya, László Tóth 0001, Alexandra Markó |
INTERSPEECH | 2 |
| 2017 | DNN-Based Feature Extraction and Classifier Combination for Child-Directed Speech, Cold and Snoring IdentificationabstractIn this study we deal with the three sub-challenges of the Interspeech ComParE Challenge 2017, where the goal is to identify child-directed speech, speakers having a cold, and different types of snoring sounds.For the first two sub-challenges we propose a simple, two-step feature extraction and classification scheme: first we perform frame-level classification via Deep Neural Networks (DNNs), and then we extract utterancelevel features from the DNN outputs.By utilizing these features for classification, we were able to match the performance of the standard paralinguistic approach (which involves extracting thousands of features, many of them being completely irrelevant to the actual task).As for the Snoring Sub-Challenge, we divided the recordings into segments, and averaged out some frame-level features segment-wise, which were then used for utterance-level classification.When combining the predictions of the proposed approaches with those got by the standard paralinguistic approach, we managed to outperform the baseline values of the Cold and Snoring sub-challenges on the hidden test sets. Gábor Gosztolya, Róbert Busa-Fekete, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 3 |
| 2017 | Training Context-Dependent DNN Acoustic Models Using Probabilistic SamplingabstractIn current HMM/DNN speech recognition systems, the purpose of the DNN component is to estimate the posterior probabilities of tied triphone states.In most cases the distribution of these states is uneven, meaning that we have a markedly different number of training samples for the various states.This imbalance of the training data is a source of suboptimality for most machine learning algorithms, and DNNs are no exception.A straightforward solution is to re-sample the data, either by upsampling the rarer classes or by dowsampling the more common classes.Here, we experiment with the so-called probabilistic sampling method that applies downsampling and upsampling at the same time.For this, it defines a new class distribution for the training data, which is a linear combination of the original and the uniform class distributions.As an extension to previous studies, we propose a new method to re-estimate the class priors, which is required to remedy the mismatch between the training and the test data distributions introduced by re-sampling.Using probabilistic sampling and the proposed modification we report 5% and 6% relative error rate reductions on the TED-LIUM and on the AMI corpora, respectively. Tamás Grósz, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 1 |
| 2017 | A Comparative Evaluation of GMM-Free State Tying Methods for ASR
Tamás Grósz, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 1 |
| 2016 | Determining Native Language and Deception Using Phonetic Features and Classifier CombinationabstractFor several years, the Interspeech ComParE Challenge has focused on paralinguistic tasks of various kinds.In this paper we focus on the Native Language and the Deception subchallenges of ComParE 2016, where the goal is to identify the native language of the speaker, and to recognize deceptive speech.As both tasks can be treated as classification ones, we experiment with several state-of-the-art machine learning methods (Support-Vector Machines, AdaBoost.MH and Deep Neural Networks), and also test a simple-yet-robust combination method.Furthermore, we will assume that the native language of the speaker affects the pronunciation of specific phonemes in the language he is currently using.To exploit this, we extract phonetic features for the Native Language task.Moreover, for the Deception Sub-Challenge we compensate for the highly unbalanced class distribution by instance re-sampling.With these techniques we are able to significantly outperform the baseline SVM on the unpublished test set. Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001 |
INTERSPEECH | 2 |
| 2016 | Estimating the Sincerity of Apologies in Speech by DNN Rank Learning and Prosodic AnalysisabstractIn the Sincerity Sub-Challenge of the Interspeech ComParE 2016 Challenge, the task is to estimate user-annotated sincerity scores for speech samples.We interpret this challenge as a ranklearning regression task, since the evaluation metric (Spearman's correlation) is calculated from the rank of the instances.As a first approach, Deep Neural Networks are used by introducing a novel error criterion which maximizes the correlation metric directly.We obtained the best performance by combining the proposed error function with the conventional MSE error.This approach yielded results that outperform the baseline on the Challenge test set.Furthermore, we introduce a compact prosodic feature set based on a dynamic representation of F0, energy and sound duration.We extract syllable-based prosodic features which are used as the basis of another machine learning step.We show that a small set of prosodic features is capable of yielding a result very close to the baseline one and that by combining the predictions yielded by DNN and the prosodic feature set, further improvement can be reached, significantly outperforming the baseline SVR on the Challenge test set. Gábor Gosztolya, Tamás Grósz, György Szaszák, László Tóth 0001 |
INTERSPEECH | 2 |
| 2016 | GMM-Free Flat Start Sequence-Discriminative DNN TrainingabstractRecently, attempts have been made to remove Gaussian mixture models (GMM) from the training process of deep neural network-based hidden Markov models (HMM/DNN). For the GMM-free training of a HMM/DNN hybrid we have to solve two problems, namely the initial alignment of the frame-level state labels and the creation of context-dependent states. Although flat-start training via iteratively realigning and retraining the DNN using a frame-level error function is viable, it is quite cumbersome. Here, we propose to use a sequence-discriminative training criterion for flat start. While sequence-discriminative training is routinely applied only in the final phase of model training, we show that with proper caution it is also suitable for getting an alignment of context-independent DNN models. For the construction of tied states we apply a recently proposed KL-divergence-based state clustering method, hence our whole training process is GMM-free. In the experimental evaluation we found that the sequence-discriminative flat start training method is not only significantly faster than the straightforward approach of iterative retraining and realignment, but the word error rates attained are slightly better as well. Gábor Gosztolya, Tamás Grósz, László Tóth 0001 |
INTERSPEECH | 2 |
| 2016 | Detecting Mild Cognitive Impairment from Spontaneous Speech by Correlation-Based Phonetic Feature SelectionabstractMild Cognitive Impairment (MCI), sometimes regarded as a prodromal stage of Alzheimer's disease, is a mental disorder that is difficult to diagnose.Recent studies reported that MCI causes slight changes in the speech of the patient.Our previous studies showed that MCI can be efficiently classified by machine learning methods such as Support-Vector Machines and Random Forest, using features describing the amount of pause in the spontaneous speech of the subject.Furthermore, as hesitation is the most important indicator of MCI, we took special care when handling filled pauses, which usually correspond to hesitation.In contrast to our previous studies which employed manually constructed feature sets, we now employ (automatic) correlation-based feature selection methods to find the relevant feature subset for MCI classification.By analyzing the selected feature subsets we also show that features related to filled pauses are useful for MCI detection from speech samples. Gábor Gosztolya, László Tóth 0001, Tamás Grósz, Veronika Vincze, Ildikó Hoffmann, Gréta Szatlóczki, Magdolna Pákáski, János Kálmán |
INTERSPEECH | 3 |
| 2015 | Building context-dependent DNN acoustic models using Kullback-Leibler divergence-based state tyingabstractDeep neural network (DNN) based speech recognizers have recently replaced Gaussian mixture (GMM) based systems as the state-of-the-art. HMM/DNN systems have kept many refinements of the HMM/GMM framework, even though some of these may be suboptimal for them. One such example is the creation of context-dependent tied states, for which an efficient decision tree state tying method exists. The tied states used to train DNNs are usually obtained using the same tying algorithm, even though it is based on likelihoods of Gaussians. In this paper, we investigate an alternative state clustering method that uses the Kullback-Leibler (KL) divergence of DNN output vectors to build the decision tree. It has already been successfully applied within the framework of KL-HMM systems, and here we show that it is also beneficial for HMM/DNN hybrids. In a large vocabulary recognition task we report a 4% relative word error rate reduction using this state clustering method. Gábor Gosztolya, Tamás Grósz, László Tóth 0001, David Imseng |
ICASSP | 2 |
| 2015 | Assessing the degree of nativeness and parkinson's condition using Gaussian processes and deep rectifier neural networksabstractThe Interspeech 2015 Computational Paralinguistics Challenge includes two regression learning tasks, namely the Parkinson's Condition Sub-Challenge and the Degree of Nativeness Sub-Challenge.We evaluated two state-of-the-art machine learning methods on the tasks, namely Deep Neural Networks (DNN) and Gaussian Processes Regression (GPR).We also experiented with various classifier combination and feature selection methods.For the Degree of Nativeness sub-challenge we obtained a far better Spearman correlation value than the one presented in the baseline paper.As regards the Parkinson's Condition Sub-Challenge, we showed that both DNN and GPR are competitive with the baseline SVM, and that the results can be improved further by combining the classifiers.However, we obtained by far the best results when we applied a speaker clustering method to identify the files that belong to the same speaker. Tamás Grósz, Róbert Busa-Fekete, Gábor Gosztolya, László Tóth 0001 |
INTERSPEECH | 1 |
| 2014 | Detecting the intensity of cognitive and physical load using AdaBoost and deep rectifier neural networksabstractThe Interspeech ComParE 2014 Challenge consists of two machine learning tasks, which have quite a small number of examples. Due to our good results in ComParE 2013, we considered AdaBoost a suitable machine learning meta-algorithm for these tasks, besides we also experimented with Deep Rectifier Neural Networks. These differ from traditional neural networks in that the former have several hidden layers, and use rectifier neurons as hidden units. With AdaBoost we achieved competitive results, whereas with the neural networks we were able to outperform baseline SVM scores in both Sub-Challenges. Index Terms :s peech technology, AdaBoost, deep neural networks, rectifier activation function Gábor Gosztolya, Tamás Grósz, Róbert Busa-Fekete, László Tóth 0001 |
INTERSPEECH | 2 |