EDBT 2026 Demo / reviewers in the wild / expert
Hyung Yong Kim
dblp:173/6571
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-6009-9530ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale InjectionabstractKnowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full access to intrinsic knowledge such as softmax distributions, black-box KD adopts a black-box LLM (e.g., GPT-4) as the teacher, which provides only text-level outputs via API calls. This limited supervision makes black-box KD generally less effective than its white-box counterpart. To bridge the gap between white-box and black-box KD, we propose GrayKD, a novel framework that can effectively distill text-level knowledge from a black-box LLM in a single-stage manner. In particular, rationales generated by the black-box LLM are injected into the student via a lightweight cross-attention module (teacher mode), enabling the model to approximate the black-box teacher’s output distribution without access to internal parameters. The student is then trained with the softmax-level knowledge provided by the teacher mode (student mode). Since both the teacher and student modes share the same backbone, the proposed teacher mode remains highly parameter-efficient, requiring only a small number of additional parameters for rationale injection. Experimental results on instruction-following tasks demonstrate that GrayKD achieves substantial performance improvements over existing KD methods. Hyeongsoo Lim, Hyung Yong Kim, Min Ho Jang, Eun Seo Seo, Youshin Lim, Shukjae Choi, Yunkyu Lim, Hanbin Lee, Byeong-Yeol Kim, Jiwon Yoon 0002 |
AAAI | 2 |
| 2025 | Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence ModelsabstractRecently, Transformer-based encoder-decoder models have demonstrated strong performance in multilingual speech recognition. However, the decoder’s autoregressive nature and large size introduce significant bottlenecks during inference. Additionally, although rare, repetition can occur and negatively affect recognition accuracy. To tackle these challenges, we propose a novel Hybrid Decoding approach that both accelerates inference and alleviates the issue of repetition. Our method extends the transformer encoder-decoder architecture by attaching a lightweight, fast decoder to the pretrained encoder. During inference, the fast decoder rapidly generates an output, which is then verified and, if necessary, selectively corrected by the Transformer decoder. This results in faster decoding and improved robustness against repetitive errors. Experiments on the LibriSpeech and GigaSpeech test sets indicate that, with finetuning limited to the added decoder, our method achieves word error rates comparable to or better than the baseline, while more than doubling the inference speed. Yunkyu Lim, Hyung Yong Kim, Hanbin Lee, Byeong-Yeol Kim |
ASRU | 3 |
| 2024 | Learning Contextualized Representation on Discrete Space Via Hierarchical Product QuantizationabstractSelf-supervised learning has recently demonstrated significant success in various speech processing applications. Recent studies report that pre-training with contextualized continuous targets plays a crucial role in fine-tuning for better speech downstream tasks. However, unlike the continuous targets, it is challenging to produce contextualized targets on discrete space due to unstable training. To address this issue, we introduce a new hierarchical product quantizer that enables the full utilization of multi-layer features by reducing the possible case of quantized targets and preventing mode collapse through diversity loss for all codebooks. Our ablation study confirms the effectiveness of the proposed quantizer and contextualized discrete targets. For supervised ASR, the proposed model outperforms wav2vec2 and showed comparable results with data2vec. In addition, for unsupervised ASR, the proposed method surpasses two baselines. Hyung Yong Kim, Byeong-Yeol Kim, Yunkyu Lim, Youshin Lim, Seung Woo Yu, Hanbin Lee |
ICASSP | 1 |
| 2024 | Self-training ASR Guided by Unsupervised ASR Teacher
Hyung Yong Kim, Byeong-Yeol Kim, Yunkyu Lim, Shukjae Choi, Yooncheol Ju, Youshin Lim, Seung Woo Yu, Hanbin Lee, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2023 | Masked Token Similarity Transfer for Compressing Transformer-Based ASR ModelsabstractRecent self-supervised automatic speech recognition (ASR) models based on transformers are showing best performance, but their footprint is too large to be trained on low-resource environments or deployed to edge devices. Knowledge distillation (KD) can be employed to reduce the model size. However, setting embedding dimension of teacher and student network to different values makes it difficult to transfer token embeddings for better performance. To mitigate this issue, we present a novel KD method in which student mimics the prediction vector of teacher under our proposed masked token similarity transfer (MTST) loss where the temporal relation between a token and the other unmasked ones is encoded into a dimension-agnostic token similarity vector. Under our transfer learning setting with a fine-tuned teacher, our proposed methods reduce the model size of student to 28.3% of teacher’s while word error rate on test-clean subset in LibriSpeech corpus is 4.93%, which surpasses prior works. Our source code will be made available. Euntae Choi, Youshin Lim, Byeong-Yeol Kim, Hyung Yong Kim, Hanbin Lee, Yunkyu Lim, Seung Woo Yu, Sungjoo Yoo |
ICASSP | 4 |
| 2023 | Joint Unsupervised and Supervised Learning for Context-Aware Language IdentificationabstractLanguage identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task only. However, we need additional text labels to train the model to recognize speech, and acquiring the text labels is a cost high. In order to overcome this problem, we propose context-aware language identification using a combination of unsupervised and supervised learning without any text labels. The proposed method learns the context of speech through masked language modeling (MLM) loss and simultaneously trains to determine the language of the utterance with supervised learning loss. The proposed joint learning was found to reduce the error rate by 15.6% compared to the same structure model trained by supervised-only learning on a subset of the VoxLingua107 dataset consisting of sub-three-second utterances in 11 languages. Hyung Yong Kim, Byeong-Yeol Kim, Shukjae Choi, Yunkyu Lim |
ICASSP | 2 |
| 2023 | FACTSpeech: Speaking a Foreign Language Pronunciation Using Only Your Native Characters
Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim, Shukjae Choi, Hyung Yong Kim |
INTERSPEECH | 7 |
| 2023 | Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC ModelsabstractKnowledge distillation (KD), best known as an effective method for model compression, aims at transferring the knowledge of a bigger network (teacher) to a much smaller network (student). Conventional KD methods usually employ the teacher model trained in a supervised manner, where output labels are treated only as targets. Extending this supervised scheme further, we introduce a new type of teacher model for connectionist temporal classification (CTC)-based sequence models, namely Oracle Teacher, that leverages both the source inputs and the output labels as the teacher model's input. Since the Oracle Teacher learns a more accurate CTC alignment by referring to the target information, it can provide the student with more optimal guidance. One potential risk for the proposed approach is a trivial solution that the model's output directly copies the target input. Based on a many-to-one mapping property of the CTC algorithm, we present a training strategy that can effectively prevent the trivial solution and thus enables utilizing both source and target inputs for model training. Extensive experiments are conducted on two sequence learning tasks: speech recognition and scene text recognition. From the experimental results, we empirically show that the proposed model improves the students across these tasks while achieving a considerable speed-up in the teacher model's training time. Jiwon Yoon 0002, Hyung Yong Kim, Hyeon Seung Lee, Sunghwan Ahn, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | ASBERT: ASR-Specific Self-Supervised Learning with Self-TrainingabstractPre-training of self-supervised learning (SSL) generally shows a good performance on various speech processing tasks. However, this pre-training scheme may lead to a sub-optimal solution for fine-tuning a specific task, such as automatic speech recognition (ASR). In order to provide a more optimal pre-trained model for ASR, we introduce an ASR-Specific hidden-unit BERT with self-training, namely ASBERT. Motivated by self-training, we extract linguistic-related pseudo labels from the fine-tuned model, and these labels are used in the next pre-training procedure. Experimental results on LibriSpeech test-clean and test-other datasets show that ASBERT without language model (LM) outperforms the conventional SSL and self-training model, achieving a 6.3/2.0% and 15.4/13.2% relatively word error rate reduction (RERR). Moreover, without using pseudo-transcription, ASBERT yields comparable performance to the conventional self-training method. Hyung Yong Kim, Byeong-Yeol Kim, Seung Woo Yoo, Youshin Lim, Yunkyu Lim, Hanbin Lee |
SLT | 1 |
| 2022 | Neurally Optimized Decoder for Low Bitrate Speech CodecabstractRecently, a conventional neural decoder for speech codec has shown promising performance. However, it typically requires some prior knowledge of decoding such as bit allocation or dequantization scheme, which is not a universal solution for many different kinds of speech codecs. In order to address this limitation, we propose a neurally optimized decoder based on a generative model which can directly reconstruct the speech from the bitstream without a prior knowledge. The proposed decoder mainly consists of two components: 1) a dequantization model to group and dequantize related bits from the bitstream and 2) a generative model to restore the speech conditioned on the output of the dequantization model. Through experiments with mixed excitation linear prediction (MELP), Advanced multi-band excitation (AMBE), and SPEEX at around 2.4 kb/s, it is showed that the proposed model showed better performance in most of the objective and subjective evaluation compared to the conventional speech codecs. Hyung Yong Kim, Jiwon Yoon 0002, Won-Ik Cho, Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2021 | TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech RecognitionabstractIn recent years, there has been a great deal of research in developing end-to-end speech recognition models, which enable simplifying the traditional pipeline and achieving promising results. Despite their remarkable performance improvements, end-to-end models typically require expensive computational cost to show successful performance. To reduce this computational burden, knowledge distillation (KD), which is a popular model compression method, has been used to transfer knowledge from a deep and complex model (teacher) to a shallower and simpler model (student). Previous KD approaches have commonly designed the architecture of the student by reducing the width per layer or the number of layers of the teacher. This structural reduction scheme might limit the flexibility of model selection since the student model structure should be similar to that of the given teacher. To cope with this limitation, we propose a KD method for end-to-end speech recognition, namely TutorNet, that applies KD techniques across different types of neural networks at the hidden representation-level as well as the output-level. For concrete realizations, we firstly apply representation-level knowledge distillation (RKD) during the initialization step, and then apply the softmax-level knowledge distillation (SKD) combined with the original task learning. When the student is trained with RKD, we make use of frame weighting that points out the frames to which the teacher pays more attention. Through a number of experiments, it is verified that TutorNet not only distills the knowledge between networks with different topologies but also significantly contributes to improving the performance of the distilled student. Jiwon Yoon 0002, Hyeon Seung Lee, Hyung Yong Kim, Won-Ik Cho, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation
Hyeongju Kim, Hyeon Seung Lee, Woo Hyun Kang, Hyung Yong Kim, Nam Soo Kim |
IJCAI | 4 |
| 2019 | End-to-End Multi-Channel Speech Enhancement Using Inter-Channel Time-Restricted Attention on Raw Waveform
Hyeon Seung Lee, Hyung Yong Kim, Woo Hyun Kang, Jeunghun Kim, Nam Soo Kim |
INTERSPEECH | 2 |
| 2015 | Discriminative nonnegative matrix factorization using cross-reconstruction error for source separation
Kisoo Kwon, Jong Won Shin, Hyung Yong Kim, Nam Soo Kim |
INTERSPEECH | 3 |