EDBT 2026 Demo / reviewers in the wild / expert
Suyoun Kim
dblp:82/10805
· DBLP profile ↗
19ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0002-6822-337XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 8 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Wonjune Kang, Junteng Jia, Chunyang Wu, Egor Lakomkin, Yashesh Gaur, Leda Sari, Suyoun Kim, Jay Mahadeokar, Ozlem Kalinli |
INTERSPEECH | 8 |
| 2024 | Evaluating Speech Recognition Performance Towards Large Language Model Based Voice Assistants
Zhe Liu 0011, Suyoun Kim, Ozlem Kalinli |
INTERSPEECH | 2 |
| 2023 | Introducing Semantics into Speech EncodersabstractDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim, Zhaojiang Lin, Bing Liu, Akshat Shrivastava, Shang-Wen Li, Liang-Hsuan Tseng, Guan-Ting Lin, Alexei Baevski, Hung-yi Lee, Yizhou Sun, Wei Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Derek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim, Zhaojiang Lin, Akshat Shrivastava, Shang-Wen Li 0001, Liang-Hsuan Tseng, Guan-Ting Lin, Alexei Baevski, Hung-yi Lee, Yizhou Sun, Wei Wang 0010 |
ACL (1) | 4 |
| 2023 | ICASSP 2023 Spoken Language Understanding Grand ChallengeabstractSpoken language understanding (SLU) is a important field between the Speech and NLP community focused on converting a users’ speech utterance into an executable semantic parse. In order to facilitate open research in this space, we introduce the 1st Spoken Language Understanding challenge hosted at ICASSP 2023. We leverage the newly released SLU dataset STOP [1]. In this challenge, participants are asked to compete in 3 tracks of SLU relevant to the field (1) Quality: build the highest performance model (2) On-device: build the highest quality model under 15M parameters and (3) Low-resource: Achieve the highest quality in a low-resource setting. While participants have made significant strides in the challenge, there is still a long way to go in building data and compute efficient SLU models. Akshat Shrivastava, Suyoun Kim, Paden Tomasello, Ali Elkahky, Daniel Lazar, Trang Le, Aleksandr Livshits, Ahmed Aly |
ICASSP | 2 |
| 2023 | Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding
Suyoun Kim, Akshat Shrivastava, Ju Lin, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 1 |
| 2022 | Evaluating User Perception of Speech Recognition System Quality with Semantic Distance MetricabstractMeasuring automatic speech recognition (ASR) system quality is critical for creating user-satisfying voice-driven applications. Word Error Rate (WER) has been traditionally used to evaluate ASR system quality; however, it sometimes correlates poorly with user perception/judgement of transcription quality. This is because WER weighs every word equally and does not consider semantic correctness which has a higher impact on user perception. In this work, we propose evaluating ASR output hypotheses quality with SemDist that can measure semantic correctness by using the distance between the semantic vectors of the reference and hypothesis extracted from a pre-trained language model. Our experimental results of 71K and 36K user annotated ASR output quality show that SemDist achieves higher correlation with user perception than WER. We also show that SemDist has higher correlation with downstream Natural Language Understanding (NLU) tasks than WER. Suyoun Kim, Weiyi Zheng, Tarun Singh, Abhinav Arora, Xiaoyu Zhai, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 1 |
| 2022 | Deliberation Model for On-Device Spoken Language UnderstandingabstractWe propose a novel deliberation-based approach to end-to-end (E2E) spoken language understanding (SLU), where a streaming automatic speech recognition (ASR) model produces the first-pass hypothesis and a second-pass natural language understanding (NLU) component generates the semantic parse by conditioning on both ASR's text and audio embeddings.By formulating E2E SLU as a generalized decoder, our system is able to support complex compositional semantic structures.Furthermore, the sharing of parameters between ASR and NLU makes the system especially suitable for resource-constrained (on-device) environments; our proposed approach consistently outperforms strong pipeline NLU baselines by 0.60% to 0.65% on the spoken version of the TOPv2 dataset (STOP).We demonstrate that the fusion of text and audio features, coupled with the system's ability to rewrite the first-pass hypothesis, makes our approach more robust to ASR errors.Finally, we show that our approach can significantly reduce the degradation when moving from natural speech to synthetic speech training, but more work is required to make text-to-speech (TTS) a viable solution for scaling up E2E SLU. Akshat Shrivastava, Paden Tomasello, Suyoun Kim, Aleksandr Livshits, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 4 |
| 2021 | Improved Neural Language Model Fusion for Streaming Recurrent Neural Network TransducerabstractRecurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily leverage unpaired text data during training. Previous work has proposed various fusion methods to incorporate external NNLMs into end-to-end ASR to address this weakness. In this paper, we propose extensions to these techniques that allow RNN-T to exploit external NNLMs during both training and inference time, resulting in 13-18% relative Word Error Rate improvement on Librispeech compared to strong baselines. Furthermore, our methods do not incur extra algorithmic latency and allow for flexible plug- and-play of different NNLMs without re-training. We also share in-depth analysis to better understand the benefits of the different NNLM fusion methods. Our work provides a reliable technique for leveraging unpaired text data to significantly improve RNN-T while keeping the system streamable, flexible, and lightweight. Suyoun Kim, Yuan Shangguan, Jay Mahadeokar, Antoine Bruguier, Christian Fügen, Michael L. Seltzer |
ICASSP | 1 |
| 2021 | Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language UnderstandingabstractWord Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator for downstream Natural Language Understanding (NLU) tasks, such as intent recognition, slot filling, and semantic parsing in task-oriented dialog systems. This is because WER takes into consideration only literal correctness instead of semantic correctness, the latter of which is typically more important for these downstream tasks. In this study, we propose a novel Semantic Distance (SemDist) measure as an alternative evaluation metric for ASR systems to address this issue. We define SemDist as the distance between a reference and hypothesis pair in a sentence-level embedding space. To represent the reference and hypothesis as a sentence embedding, we exploit RoBERTa, a state-of-the-art pre-trained deep contextualized language model based on the transformer architecture. We demonstrate the effectiveness of our proposed metric on various downstream tasks, including intent recognition, semantic parsing, and named entity recognition. Suyoun Kim, Abhinav Arora, Ching-Feng Yeh, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
Interspeech | 1 |
| 2021 | Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow FusionabstractHow to leverage dynamic contextual information in end-toend speech recognition has remained an active research area.Previous solutions to this problem were either designed for specialized use cases that did not generalize well to open-domain scenarios, did not scale to large biasing lists, or underperformed on rare long-tail words.We address these limitations by proposing a novel solution that combines shallow fusion, trie-based deep biasing, and neural network language model contextualization.These techniques result in significant 19.5% relative Word Error Rate improvement over existing contextual biasing approaches and 5.4%-9.3%improvement compared to a strong hybrid baseline on both open-domain and constrained contextualization tasks, where the targets consist of mostly rare long-tail words.Our final system remains lightweight and modular, allowing for quick modification without model re-training. Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fügen, Ozlem Kalinli, Yatharth Saraf, Michael L. Seltzer |
Interspeech | 4 |
| 2021 | Improving RNN Transducer Based ASR with Auxiliary TasksabstractEnd-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recognizers. Specifically, recurrent neural network transducer (RNN-T) has shown competitive ASR performance on various benchmarks. In this work, we examine ways in which RNN-T can achieve better ASR accuracy via performing auxiliary tasks. We propose (i) using the same auxiliary task as primary RNN-T ASR task, and (ii) performing context-dependent graphemic state prediction as in conventional hybrid modeling. In transcribing social media videos with varying training data size, we first evaluate the streaming ASR performance on three languages: Romanian, Turkish and German. We find that both proposed methods provide consistent improvements. Next, we observe that both auxiliary tasks demonstrate efficacy in learning deep transformer encoders for RNN-T criterion, thus achieving competitive results -2.0%/4.2% WER on LibriSpeech test-clean/other - as compared to prior top performing models. Chunxi Liu, Frank Zhang 0001, Suyoun Kim, Yatharth Saraf, Geoffrey Zweig |
SLT | 4 |
| 2019 | Gated Embeddings in End-to-End Speech Recognition for Conversational-Context FusionabstractWe present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings.Unlike conventional speech recognition models, our model learns longer conversational-context information that spans across sentences and is consequently better at recognizing long conversations.Specifically, we propose to use text-based external word and/or sentence embeddings (i.e., fast-Text, BERT) within an end-to-end framework, yielding significant improvement in word error rate with better conversational-context representation.We evaluated the models on the Switchboard conversational speech corpus and show that our model outperforms standard end-to-end speech recognition models. Suyoun Kim, Siddharth Dalmia, Florian Metze |
ACL (1) | 1 |
| 2019 | Cross-Attention End-to-End ASR for Two-Party ConversationsabstractWe present an end-to-end speech recognition model that learns interaction between two speakers based on the turn-changing information. Unlike conventional speech recognition models, our model exploits two speakers' history of conversational-context information that spans across multiple turns within an end-to-end framework. Specifically, we propose a speaker-specific cross-attention mechanism that can look at the output of the other speaker side as well as the one of the current speaker for better at recognizing long conversations. We evaluated the models on the Switchboard conversational speech corpus and show that our model outperforms standard end-to-end speech recognition models. Suyoun Kim, Siddharth Dalmia, Florian Metze |
INTERSPEECH | 1 |
| 2018 | Towards Language-Universal End-to-End Speech RecognitionabstractBuilding speech recognizers in multiple languages typically involves replicating a monolingual training recipe for each language, or utilizing a multi-task learning approach where models for different languages have separate output labels but share some internal parameters. In this work, we exploit recent progress in end-to-end speech recognition to create a single multilingual speech recognition system capable of recognizing any of the languages seen in training. To do so, we propose the use of a universal character set that is shared among all languages. We also create a language-specific gating mechanism within the network that can modulate the network's internal representations in a language-specific way. We evaluate our proposed approach on the Microsoft Cortana task across three languages and show that our system outperforms both the individual monolingual systems and systems built with a multi-task learning approach. We also show that this model can be used to initialize a monolingual speech recognizer, and can be used to create a bilingual model for use in code-switching scenarios. Suyoun Kim, Michael L. Seltzer |
ICASSP | 1 |
| 2018 | Improved Training for Online End-to-end Speech Recognition SystemsabstractAchieving high accuracy with end-to-end speech recognizers requires careful parameter initialization prior to training.Otherwise, the networks may fail to find a good local optimum.This is particularly true for online networks, such as unidirectional LSTMs.Currently, the best strategy to train such systems is to bootstrap the training from a tied-triphone system.However, this is time consuming, and more importantly, is impossible for languages without a high-quality pronunciation lexicon.In this work, we propose an initialization strategy that uses teacher-student learning to transfer knowledge from a large, well-trained, offline end-to-end speech recognition model to an online end-to-end model, eliminating the need for a lexicon or any other linguistic resources.We also explore curriculum learning and label smoothing and show how they can be combined with the proposed teacher-student learning for further improvements.We evaluate our methods on a Microsoft Cortana personal assistant task and show that the proposed method results in a 19% relative improvement in word error rate compared to a randomly-initialized baseline system. Suyoun Kim, Michael L. Seltzer, Jinyu Li 0001, Rui Zhao 0017 |
INTERSPEECH | 1 |
| 2018 | Dialog-Context Aware end-to-end Speech RecognitionabstractExisting speech recognition systems are typically built at the sentence level, although it is known that dialog context, e.g. higher-level knowledge that spans across sentences or speakers, can help the processing of long conversations. The recent progress in end-to-end speech recognition systems promises to integrate all available information (e.g. acoustic, language resources) into a single model, which is then jointly optimized. It seems natural that such dialog context information should thus also be integrated into the end-to-end models to improve recognition accuracy further. In this work, we present a dialog-context aware speech recognition model, which explicitly uses context information beyond sentence-level information, in an end-to-end fashion. Our dialog-context model captures a history of sentence-level contexts, so that the whole system can be trained with dialog-context information in an end-to-end manner. We evaluate our proposed approach on the Switchboard conversational speech corpus, and show that our system outperforms a comparable sentence-level end-to-end speech recognition system. Suyoun Kim, Florian Metze |
SLT | 1 |
| 2017 | Joint CTC-attention based end-to-end speech recognition using multi-task learningabstractRecently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping between variable-length input and output sequences in one step using a purely data-driven method. The attention model has often been shown to improve the performance over another end-to-end approach, the Connectionist Temporal Classification (CTC), mainly because it explicitly uses the history of the target character without any conditional independence assumptions. However, we observed that the performance of the attention has shown poor results in noisy condition and is hard to learn in the initial training stage with long input sequences. This is because the attention model is too flexible to predict proper alignments in such cases due to the lack of left-to-right constraints as used in CTC. This paper presents a novel method for end-to-end speech recognition to improve robustness and achieve fast convergence by using a joint CTC-attention model within the multi-task learning framework, thereby mitigating the alignment issue. An experiment on the WSJ and CHiME-4 tasks demonstrates its advantages over both the CTC and attention-based encoder-decoder baselines, showing 5.4-14.6% relative improvements in Character Error Rate (CER). Suyoun Kim, Takaaki Hori, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2017 | End-to-End Speech Recognition with Auditory Attention for Multi-Microphone Distance Speech Recognition
Suyoun Kim, Ian Lane |
INTERSPEECH | 1 |
| 2016 | Recurrent Models for Auditory Attention in Multi-Microphone Distant Speech RecognitionabstractIntegration of multiple microphone data is one of the key ways to achieve robust speech recognition in noisy environments or when the speaker is located at some distance from the input device. Signal processing techniques such as beamforming are widely used to extract a speech signal of interest from background noise. These techniques, however, are highly dependent on prior spatial information about the microphones and the environment in which the system is being used. In this work, we present a neural attention network that directly combines multi-channel audio to generate phonetic states without requiring any prior knowledge of the microphone layout or any explicit signal preprocessing for speech enhancement. We embed an attention mechanism within a Recurrent Neural Network (RNN) based acoustic model to automatically tune its attention to a more reliable input source. Unlike traditional multi-channel preprocessing, our system can be optimized towards the desired output in one step. Although attention-based models have recently achieved impressive results on sequence-to-sequence learning, no attention mechanisms have previously been applied to learn potentially asynchronous and non-stationary multiple inputs. We evaluate our neural attention model on the CHiME-3 challenge task, and show that the model achieves comparable performance to beamforming using a purely data-driven method. Suyoun Kim, Ian Lane |
INTERSPEECH | 1 |