Eesung Kim

dblp:243/6661 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
Eesung Kim, Vijendra Raj Apsingekar
INTERSPEECH2
2024 Joint End-to-End Spoken Language Understanding and Automatic Speech Recognition Training Based on Unified Speech-to-Text Pre-Training
abstract
Modern spoken language understanding (SLU) approaches optimize the system in an end-to-end (E2E) manner. This approach offers two key advantages. Firstly, it helps mitigate error propagation from upstream systems. Secondly, combining various information types and optimizing them towards the same objective is straightforward. In this study, we attempt to build an SLU system by integrating information from two modalities, i.e., speech and text, and concurrently optimizing the associated tasks. We leverage a pre-trained model built with speech and text data and fine-tune it for the E2E SLU tasks. The SLU model is jointly optimized with automatic speech recognition (ASR) and SLU tasks under single-mode and dual-mode schemes. In the single-mode model, ASR and SLU results are predicted sequentially, whereas the dualmode model predicts either ASR or SLU outputs based on the task tag. Our proposed method demonstrates its superiority through benchmarking against FSC, SLURP, and in-house datasets, exhibiting improved intent accuracy, SLU-F1, and Word Error Rate (WER).
Eesung Kim, Taeyeon Ki, Divya Neelagiri, Vijendra Raj Apsingekar
ICASSP1
2023 Efficient Adaptation of Spoken Language Understanding based on End-to-End Automatic Speech Recognition
Eesung Kim, Aditya Jajodia, Cindy Tseng, Divya Neelagiri, Taeyeon Ki, Vijendra Raj Apsingekar
INTERSPEECH1
2022 Automatic Pronunciation Assessment using Self-Supervised Speech Representation Learning
abstract
Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community.In particular, speech representations learned by SSL models have been shown to be effective for encoding various speech-related characteristics.In this context, we propose a novel automatic pronunciation assessment method based on SSL models.First, the proposed method fine-tunes the pre-trained SSL models with connectionist temporal classification to adapt the English pronunciation of English-as-a-second-language (ESL) learners in a data environment.Then, the layer-wise contextual representations are extracted from all across the transformer layers of the SSL models.Finally, the automatic pronunciation score is estimated using bidirectional long short-term memory with the layer-wise contextual representations and the corresponding text.We show that the proposed SSL model-based methods outperform the baselines, in terms of the Pearson correlation coefficient, on datasets of Korean ESL learner children and Speechocean762.Furthermore, we analyze how different representations of transformer layers in the SSL model affect the performance of the pronunciation assessment task.
Eesung Kim, Jae-Jin Jeon, Hyeji Seo
INTERSPEECH1
2022 JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech
abstract
In neural text-to-speech (TTS), two-stage system or a cascade of separately learned models have shown synthesis quality close to human speech.For example, FastSpeech2 transforms an input text to a mel-spectrogram and then HiFi-GAN generates a raw waveform from a mel-spectogram where they are called an acoustic feature generator and a neural vocoder respectively.However, their training pipeline is somewhat cumbersome in that it requires a fine-tuning and an accurate speech-text alignment for optimal performance.In this work, we present endto-end text-to-speech (E2E-TTS) model which has a simplified training pipeline and outperforms a cascade of separately learned models.Specifically, our proposed model is jointly trained FastSpeech2 and HiFi-GAN with an alignment module.Since there is no acoustic feature mismatch between training and inference, it does not requires fine-tuning.Furthermore, we remove dependency on an external speech-text alignment tool by adopting an alignment learning objective in our joint training framework.Experiments on LJSpeech corpus shows that the proposed model outperforms publicly available, stateof-the-art implementations of ESPNet2-TTS on subjective evaluation (MOS) and some objective evaluations.
Dan Lim, Sunghee Jung, Eesung Kim
INTERSPEECH3
2021 Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
abstract
Recently, several types of end-to-end speech recognition methods named transformer-transducer were introduced. According to those kinds of methods, transcription networks are generally modeled by transformer-based neural networks, while prediction networks could be modeled by either transformers or recurrent neural networks (RNN). In this paper, we propose novel multitask learning, joint optimization, and joint decoding methods for transformer-RNN-transducer systems. Our proposed methods have the main advantage in that the model can maintain information on the large text corpus eliminating the necessity of an external language model (LM). We prove their effectiveness by performing experiments utilizing the well-known ESPNET toolkit for the widely used Librispeech datasets. We also show that the proposed methods can reduce word error rate (WER) by 16.6 % and 13.3 % for test-clean and test-other datasets, respectively, without changing the overall model structure nor exploiting an external LM.
Jae-Jin Jeon, Eesung Kim
ICASSP2
2021 U-Convolution Based Residual Echo Suppression with Multiple Encoders
abstract
In this paper, we propose an efficient end-to-end neural network that can estimate near-end speech using a U-convolution block by exploiting various signals to achieve residual echo suppression (RES). Specifically, the proposed model employs multiple encoders and an integration block to utilize complete signal information in an acoustic echo cancellation system and also applies the U-convolution blocks to separate near-end speech efficiently. The proposed network affords an improvement in the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI), as compared to baselines, in scenarios involving smart audio devices. The experimental results show that the proposed method outperforms the baselines for various types of mismatched background noise and environmental reverberation, while requiring low computational resources.
Eesung Kim, Jae-Jin Jeon, Hyeji Seo
ICASSP1
2021 SE-Conformer: Time-Domain Speech Enhancement Using Conformer
Eesung Kim, Hyeji Seo
Interspeech1
2020 Accelerating RNN Transducer Inference via Adaptive Expansion Search
abstract
Recurrent neural network transducers (RNN-T) are a promising end-to-end speech recognition framework that transduce input acoustic frames to a character sequence. Best- and breadth-first searches have been used as decoding strategies for RNN-T. However, best-first search follows a sequential process for its expansion search, which slows down the decoding process. Although breadth-first search replaces the sequential process of best-first search with a parallel one, it unnecessarily conducts an expansion search for all decoding steps. As most of the decoding frames correspond to a blank symbol because the length of the character sequence is much shorter than that of the decoding frames, this induces computational overhead. To address these limitations, we introduce an adaptive expansion search (AES) to accelerate RNN-T inference. AES overcomes the aforementioned limitations by batching the hypotheses and adopting a decision-making process that decides whether to continue the expansion search; thus, AES can avoid unnecessary expansion search. Furthermore, pruning is applied to AES for further acceleration. We achieved significant speedup and a lower word error rate compared with other baselines.
Juntae Kim, Yoonhan Lee, Eesung Kim
IEEE Signal Process. Lett.3
2019 DNN-based Emotion Recognition Based on Bottleneck Acoustic Features and Lexical Features
abstract
In this paper, we propose a novel emotion recognition method to reflect affect salient information using acoustic and lexical features. The acoustic features are extracted from the speech signal by applying statistical functionals of emotionally high-level features derived from Deep Neural Network (DNN). These acoustic features are early fused with two types of lexical features extracted from the text transcription of the speech signal, which are the distributed representation and affective lexicon-based dimensions. The fused features are fed to another DNN for utterance-level emotion classification. Experimental results on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) multimodal dataset showed 75.5% in unweighted accuracy recall, which outperformed the best results reported previously in the multimodal emotion recognition using acoustic and lexical features.
Eesung Kim, Jong Won Shin
ICASSP1