EDBT 2026 Demo / reviewers in the wild / expert
Mun-Hak Lee
dblp:304/7803
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0005-7676-2532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Generalization of End-to-End ASR through Diversity and Independence Regularization
Ye-Eun Ko, Mun-Hak Lee, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2025 | Bayesian Language Model Adaptation for Personalized Speech RecognitionabstractIn deployment environments for speech recognition models, diverse proper nouns such as personal names, song titles, and application names are frequently uttered. These proper nouns are often sparsely distributed within the training dataset, leading to performance degradation and limiting the practical utility of the models. Personalization strategies that leverage userspecific information, such as contact lists or search histories, have proven effective in mitigating performance degradation caused by rare words. In this study, we propose a novel personalization method for combining the scores of a general language model (LM) and a personal LM within a probabilistic framework. The proposed method entails low computational costs, storage requirements, and latency. Through experiments using a realworld dataset collected from the vehicle environment, we demonstrate that the proposed method effectively overcomes the out-ofvocabulary problem and improves recognition performance for rare words Mun-Hak Lee, Ji-Hwan Mo, Ji-Hun Kang, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2024 | Whisper Multilingual Downstream Task Tuning Using Task Vectors
Ji-Hun Kang, Jae-Hong Lee, Mun-Hak Lee, Joon-Hyuk Chang |
INTERSPEECH | 3 |
| 2024 | Balanced-Wav2Vec: Enhancing Stability and Robustness of Representation Learning Through Sample Reweighting Techniques
Mun-Hak Lee, Jae-Hong Lee, Do-Hee Kim, Ye-Eun Ko, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2024 | Proper Error Estimation and Calibration for Attention-Based Encoder-Decoder ModelsabstractAn attention-based automatic speech recognition (ASR) model generates a probability distribution of the tokens set at each time step. Recent studies have shown that calibration errors exist in the output probability distributions of attention-based ASR models trained to minimize the negative log likelihood. This study analyzes the causes of calibration errors in ASR model outputs and their impact on model performance. Based on the analysis, we argue that conventional methods for estimating calibration errors at the token level are unsuitable for ASR tasks. Accordingly, we propose a new calibration measure that estimates the calibration error at the sequence level. Moreover, we present a new post-hoc calibration function and training objective to mitigate the calibration error of the ASR model at the sequence level. Through experiments using the ASR benchmark, we show that the proposed methods effectively alleviate the calibration error of the ASR model and improve the generalization performance. Mun-Hak Lee, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Cross-Modal Learning for CTC-Based ASR: Leveraging CTC-Bertscore and Sequence-Level TrainingabstractDue to the nature of neural networks that easily overfit the training set, neural network-based speech recognition models are vulnerable to prior shifts in data distribution or unseen words. Therefore, studies have been conducted to over-come this problem by using language models trained with a relatively easy-to-obtain unpaired corpus. In this paper, we present a new training method that uses BERT to improve the performance of a connectionist temporal classification (CTC)-based ASR model. The proposed method follows a cross-modal learning scenario and induces the CTC model to better embed contextual information by utilizing an auxiliary objective function operating at the sequence level. We applied the proposed method to fine-tune the pre-trained wav2vec 2.0 model with CTC loss and confirmed that the proposed method improves the generalization performance of the ASR model. Mun-Hak Lee, Sang-Eon Lee, Ji-Eun Choi, Joon-Hyuk Chang |
ASRU | 1 |
| 2023 | Knowledge Distillation From Offline to Streaming Transducer: Towards Accurate and Fast Streaming Model by Matching AlignmentsabstractSequence transducer is a popular end-to-end automatic speech recognition model for streaming scenarios: While, there is a trade-off between accuracy and latency. Latency regularization methods such as FastEmit can reduce latency, but the more they try to reduce latency, the worse accuracy tends to be. Conversely, knowledge distillation (KD) is only used to improve accuracy, and latency is not considered. In this paper, we propose an effective method that combines FastEmit with the KD to reduce latency and improve the accuracy of offline model in scenarios where the latency gap between offline and streaming models gets small. This method reduce the latency gap by applying with FastEmit to both the offline and streaming models. Experimental results on the LibriSpeech dataset show that the model with the best trade-off between accuracy and latency achieves a relative error reduction rate of 7.5% and reduces the latency by $130 \mathrm{~ms}$ compared with the streaming conformer transducer. Ji-Hwan Mo, Jae-Jin Jeon, Mun-Hak Lee, Joon-Hyuk Chang |
ASRU | 3 |
| 2022 | Knowledge Distillation from Language Model to Acoustic Model: A Hierarchical Multi-Task Learning ApproachabstractThe remarkable performance of the pre-trained language model (LM) using self-supervised learning has led to a major paradigm shift in the study of natural language processing. In line with these changes, leveraging the performance of speech recognition systems with massive deep learning-based LMs is a major topic of speech recognition research. Among the various methods of applying LMs to speech recognition systems, in this paper, we focus on a cross-modal knowledge distillation method that transfers knowledge between two types of deep neural networks with different modalities. We propose an acoustic model structure with multiple auxiliary output layers for cross-modal distillation and demonstrate that the proposed method effectively compensates for the shortcomings of the existing label-interpolation-based distillation method. In addition, we extend the proposed method to a hierarchical distillation method using LMs trained in different units (senones, monophones, and subwords) and reveal the effectiveness of the hierarchical distillation method through an ablation study. Mun-Hak Lee, Joon-Hyuk Chang |
ICASSP | 1 |
| 2022 | Regularizing Transformer-based Acoustic Models by Penalizing Attention Weights
Mun-Hak Lee, Joon-Hyuk Chang, Sang-Eon Lee, Ju-Seok Seong, Chanhee Park, Haeyoung Kwon |
INTERSPEECH | 1 |
| 2021 | Deep Neural Network Calibration for E2E Speech Recognition System
Mun-Hak Lee, Joon-Hyuk Chang |
Interspeech | 1 |