EDBT 2026 Demo / reviewers in the wild / expert
Jae-Hong Lee
dblp:62/4284
· DBLP profile ↗
19ranked-venue papers
9as first author
18since 2021 · last 2026
0009-0008-3717-2988ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OnEDIT: Online Editing with Decoupled Implicit Task for Large Language ModelsabstractContinual instruction tuning (CIT) has emerged as a promising strategy for adapting large language models (LLMs) to new tasks while preserving historical knowledge. Most existing CIT methods have focused on offline CIT (offCIT), which assumes clearly defined task boundaries and allows multiple passes over the data. However, such assumptions rarely hold in real-world scenarios, where data arrive in a streaming fashion and task boundaries are unknown. This setting introduces critical challenges: the absence of task identifiers (task IDs), a significant imbalance in task-specific information, and inaccessibility to previously seen data. In this work, we propose Online Editing with Decoupled Implicit Task (OnEDIT), an online CIT(onCIT) approach to tackle these challenges. OnEDIT leverages a fixed-size adapter for the implicit task, balancing current and past knowledge through editing operations every time step without relying on task IDs or backpropagation. Extensive experiments on CIT benchmarks demonstrate that OnEDIT consistently maintains robust and stable performance, whereas existing state-of-the-art baselines often suffer from performance degradation in online settings. It suggests that OnEDIT achieves superior generalization across diverse task orders and model scales, while maintaining high efficiency and low memory overhead. Chae-Won Lee, Jae-Hong Lee, Ji-Hun Kang, Joon-Hyuk Chang |
AAAI | 2 |
| 2026 | Bumper-guided representation interpolation for black-box unsupervised domain adaptation
Jin-Seong Choi, Jae-Hong Lee, Joon-Hyuk Chang |
Comput. Speech Lang. | 2 |
| 2026 | Alignment Regularization for Neural Transducer via Cross-Modal Optimal TransportabstractNeural transducer models learn alignments implicitly by marginalizing over all valid paths, which can lead to unstable or sub-optimal alignments under low-resource and streaming conditions where alignment ambiguity is prevalent. Previous approaches such as knowledge distillation and pruned transducers attempt to mitigate this, but they either rely on large teacher models or impose rigid pruning strategies that limit flexibility. In this work, we propose a self-guided alignment regularization framework based on optimal transport (OT), which computes a soft and globally consistent alignment plan between encoder and prediction embeddings. We further extend this to a dynamic unbalanced OT formulation with relaxed marginal constraints and integrate OT-derived alignment features into the joint network via residual fusion. Experiments on the LibriSpeech 100h and 960h benchmarks showed that our method consistently improved alignment quality and recognition accuracy, yielding relative WER reductions of 14.45%/6.27% ontest-clean/test-otherin the low-resource setting and 25.49%/10.81% in the high-resource setting, without increasing inference complexity. Ji-Hwan Mo, Jae-Hong Lee, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Bayesian Weight Enhancement with Steady-State Adaptation for Test-time Adaptation in Dynamic EnvironmentsabstractTest-time adaptation (TTA) addresses the machine learning challenge of adapting models to unlabeled test data from shifting distributions in dynamic environments. A key issue in this online setting arises from using unsupervised learning techniques, which introduce explicit gradient noise that degrades model weights. To invest in weight degradation, we propose a Bayesian weight enhancement framework, which generalizes existing weight-based TTA methods that effectively mitigate the issue. Our framework enables robust adaptation to distribution shifts by accounting for diverse weights by modeling weight distributions. Building on our framework, we identify a key limitation in existing methods: their neglect of time-varying covariance reflects the influence of the gradient noise. To address this gap, we propose a novel steady-state adaptation (SSA) algorithm that balances covariance dynamics during adaptation. SSA is derived through the solution of a stochastic differential equation for the TTA process and online inference. The resulting algorithm incorporates a covariance-aware learning rate adjustment mechanism. Through extensive experiments, we demonstrate that SSA consistently improves state-of-the-art methods in various TTA scenarios, datasets, and model architectures, establishing its effectiveness in instability and adaptability. Jae-Hong Lee |
ICML | 1 |
| 2025 | Parameter Dynamics of Online Machine Learning and Test-time AdaptationabstractPre-trained models based on deep neural networks hold strong potential for cross-domain adaptability. However, this potential is often impeded in online machine learning (OML) settings, where the breakdown of the independent and identically distributed (i.i.d.) assumption leads to unstable adaptation. While recent advances in test-time adaptation (TTA) have addressed aspects of this challenge under unsupervised learning, most existing methods focus exclusively on unsupervised objectives and overlook the risks posed by non-i.i.d. environments and the resulting dynamics of model parameters. In this work, we present a probabilistic framework that models the adaptation process using stochastic differential equations, enabling a principled analysis of parameter distribution dynamics over time. Within this framework, we find that the log-variance of the parameter transition distribution aligns closely with an inverse-gamma distribution under stable and high-performing adaptation conditions. Motivated by this insight, we propose Structured Inverse-Gamma Model Alignment (SIGMA), a novel algorithm that dynamically regulates parameter evolution to preserve inverse-gamma alignment throughout adaptation. Extensive experiments across diverse models, datasets, and adaptation scenarios show that SIGMA consistently enhances the performance of state-of-the-art TTA methods, highlighting the critical role of parameter dynamics in ensuring robust adaptation. Jae-Hong Lee |
NeurIPS | 1 |
| 2024 | Text-Only Unsupervised Domain Adaptation for Neural Transducer-Based ASR Personalization Using Synthesized DataabstractResearch on personalizing neural transducer-based automatic speech recognition (ASR) systems using the text-only data is currently flourishing. Among various approaches, utilizing synthesized speech offers an advantage of adapting the entire ASR system. In this study, we explore the problem of personalization from a domain adaptation perspective and highlight the potential risk of overfitting associated with synthesized speech. To mitigate this risk, we propose the text-only unsupervised domain adaptation (ToUDA) strategy that robustly finetunes the generic ASR model on synthesized speech by incorporating parameter-averaging over time, model freezing, and filtering out-of-distribution instances. Via various experiments, we not only showcase the effectiveness of our approach but also uncover a noteworthy limitation when it comes to personalizing atypical speech. Jae-Hong Lee, Joon-Hyuk Chang |
ICASSP | 2 |
| 2024 | Continual Momentum Filtering on Parameter Space for Online Test-time AdaptationabstractDeep neural networks (DNNs) have revolutionized tasks such as image classification and speech recognition but often falter when training and test data diverge in distribution. External factors, from weather effects on images to varied speech environments, can cause this discrepancy, compromising DNN performance. Online test-time adaptation (OTTA) methods present a promising solution, recalibrating models in real-time during the test stage without requiring historical data. However, the OTTA paradigm is imperfect, often falling prey to issues such as catastrophic forgetting due to its reliance on noisy, self-trained predictions. Although some contemporary strategies mitigate this by tying adaptations to the static source model, this restricts model flexibility. This paper introduces a continual momentum filtering (CMF) framework, leveraging the Kalman filter (KF) to strike a balance between model adaptability and information retention. The CMF intertwines optimization via stochastic gradient descent with a KF-based inference process. This methodology not only aids in averting catastrophic forgetting but also provides high adaptability to shifting data distributions. We validate our framework on various OTTA scenarios and real-world situations regarding covariate and label shifts, and the CMF consistently shows superior performance compared to state-of-the-art methods. Jae-Hong Lee, Joon-Hyuk Chang |
ICLR | 1 |
| 2024 | Stationary Latent Weight Inference for Unreliable Observations from Online Test-Time AdaptationabstractIn the rapidly evolving field of online test-time adaptation (OTTA), effectively managing distribution shifts is a pivotal concern. State-of-the-art OTTA methodologies often face limitations such as an inadequate target domain information integration, leading to significant issues like catastrophic forgetting and a lack of adaptability in dynamically changing environments. In this paper, we introduce a stationary latent weight inference (SLWI) framework, a novel approach to overcome these challenges. The proposed SLWI uniquely incorporates Bayesian filtering to continually track and update the target model weights along with the source model weight in online settings, thereby ensuring that the adapted model remains responsive to ongoing changes in the target domain. The proposed framework has the peculiar property to identify and backtrack nonlinear weights that exhibit local non-stationarity, thereby mitigating error propagation, a common pitfall of previous approaches. By integrating and refining information from both source and target domains, SLWI presents a robust solution to the persistent issue of domain adaptation in OTTA, significantly improving existing methodologies. The efficacy of SLWI is demonstrated through various experimental setups, showcasing its superior performance in diverse distribution shift scenarios. Jae-Hong Lee, Joon-Hyuk Chang |
ICML | 1 |
| 2024 | Whisper Multilingual Downstream Task Tuning Using Task Vectors
Ji-Hun Kang, Jae-Hong Lee, Mun-Hak Lee, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2024 | Online Subloop Search via Uncertainty Quantization for Efficient Test-Time AdaptationabstractOnline test-time adaptation (OTTA) methods have demonstrated their effectiveness in real-time adapting to the target domain for speech recognition tasks.However, a common thread among these existing methods is their reliance on repetitive learning for each test utterance through a subloop, imposing prohibitive computational costs.This paper highlights the inefficiency inherent in applying a uniform number of subloop iterations to every test sample.To address this issue, we propose the online subloop search (OSS) method, which implicitly adjusts the number of iterations based on the test sample and domain characteristics.The proposed method operates within a framework comprising a chaser model updated via stochastic gradient descent and a leader model updated through the exponential moving average.The OSS method quantifies and quantizes the uncertainty in the chaser model relative to the leader model, using the quantized value to predict the number of iterations for the subloop. Jae-Hong Lee, Sang-Eon Lee, Do-Hee Kim, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2024 | Balanced-Wav2Vec: Enhancing Stability and Robustness of Representation Learning Through Sample Reweighting Techniques
Mun-Hak Lee, Jae-Hong Lee, Do-Hee Kim, Ye-Eun Ko, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2024 | Language Model Personalization for Speech Recognition: A Clustered Federated Learning Approach With Adaptive Weight AverageabstractIn the rapidly evolving field of automatic speech recognition (ASR), the push towards personalization has become a paramount concern. Text-only personalization, while advantageous for data collection and adaptable to text variations, can suffer from overfitting when using personal data and requires extensive data to mitigate this issue. Federated learning (FL) emerges as a solution, facilitating learning from diverse client models while preserving privacy. However, FL addresses the challenges posed by non independent and identically distributed (non-i.i.d) data, potentially leading to poor performance. We propose two approaches for language model personalization in ASR to address these issues. First, adaptive weighted average addresses the limitations of uniform weight average in the existing FL method by combining local language models into a global model. Second, clustered federated learning, based solely on model parameters, improves model stability without relying on information from the local domain. Both strategies aim to enhance personalization and reduce performance degradation, particularly in non-i.i.d scenarios within the FL. Chae-Won Lee, Jae-Hong Lee, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 2 |
| 2024 | Partitioning Attention Weight: Mitigating Adverse Effect of Incorrect Pseudo-Labels for Self-Supervised ASRabstractThe performance of automatic speech recognition (ASR) models has been significantly improved owing to advances in deep learning and end-to-end approaches. However, these require a large amount of labeled data, which are expensive to obtain. Semi-supervised learning techniques, such as pseudo-labeling and self-supervised learning, have emerged as potential solutions to reduce the reliance on labeled data. Recently, some studies have combined self-supervised learning and pseudo-labeling to further enhance ASR performance. However, these methods suffer from incorrect pseudo-labels that propagate errors and reduce ASR performance. In this paper, we propose a novel method calledpartitioning attention weight(PAW) to mitigate the adverse effects of incorrect labels without requiring additional language models. Our proposed method isolates audio segments by partitioning a fully connected attention weight into sub-attention weights to prevent adverse effects that the model learns the wrong context for the entire attention weights from incorrect labels as well as overfitting. The proposed method is simple, requiring few changes to existing learning frameworks, and leverages the alignment information obtained during the pseudo-labeling process. Our experimental results show consistent performance improvements in ASR performance across various semi-supervised learning scenarios. Jae-Hong Lee, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | AWMC: Online Test-Time Adaptation Without Mode Collapse for Continual AdaptationabstractThis paper addresses the critical challenge of distribution shifts in automatic speech recognition (ASR) systems through a novel framework, named adaptation without mode collapse (AWMC). Distribution shifts issue, where the source and target distributions differ, can severely degrade the performance of ASR systems. The proposed AWMC framework, a response to the issue, is designed to facilitate adaptive learning from sequentially streamed utterances and mitigates the effect of mode collapse, which is a common problem with traditional test-time adaptation (TTA) methodologies. Our framework employs three parameter-shared models (anchor, chaser, and leader) in concert to continually adapt from target data, significantly enhancing the performance and reliability of the system. The effectiveness of the AWMC is demonstrated through comprehensive performance comparisons with state-of-the-art TTA methods using widely recognized ASR datasets. Jae-Hong Lee, Do-Hee Kim, Joon-Hyuk Chang |
ASRU | 1 |
| 2023 | M-CTRL: A Continual Representation Learning Framework with Slowly Improving Past Pre-Trained ModelabstractRepresentation models pre-trained on unlabeled data show competitive performance in speech recognition, even when fine-tuned on small amounts of labeled data. The continual representation learning (CTRL) framework combines pre-training and continual learning methods to obtain powerful representation. CTRL relies on two neural networks, online and offline models, where the fixed latter model transfers information to the former model with continual learning loss. In this paper, we present momentum continual representation learning (M-CTRL), a framework that slowly updates the offline model with an exponential moving average of the online model. Our framework aims to capture information from the offline model improved on past and new domains. To evaluate our framework, we continually pre-train wav2vec 2.0 with M-CTRL in the following order: Librispeech, Wall Street Journal, and TED-LIUM V3. Our experiments demonstrate that M-CTRL improves the performance in the new domain and reduces information loss in the past domain compared to CTRL. Jin-Seong Choi, Jae-Hong Lee, Chae-Won Lee, Joon-Hyuk Chang |
ICASSP | 2 |
| 2023 | Repackagingaugment: Overcoming Prediction Error Amplification in Weight-Averaged Speech Recognition Models Subject to Self-TrainingabstractRepresentation-based speech recognition models have demonstrated state-of-the-art performance on downstream tasks. These models are pre-trained on large-scale unlabeled data, fine-tuned on a small amount of labeled data, and subsequently advanced via the self-training procedure by leveraging pseudo-labels. However, a self-trained representation model produces prediction errors caused by training with incorrect labels in the pseudo-labeled data. Weight-averaging methods have been employed to refine the pseudo-labels in a variety of studies; however, these methods amplify the prediction errors of each self-trained model. To alleviate this problem, we propose RepackagingAugment, a data augmentation method that improves the diversity of models while preventing the same incorrect labels from recursively occurring in every epoch. Our data augmentation deconstructs the paired speech–text data into word units and repackages them into a randomly determined number of word sequences. This strategy induces the models to produce different prediction errors by mitigating the problem of incorrect label over-fitting. Through various experiments on representation models, such as wav2vec 2.0 and data2vec, we demonstrate that our approach improves the performance of weight-averaged models. Jae-Hong Lee, Joon-Hyuk Chang |
ICASSP | 1 |
| 2022 | W2V2-Light: A Lightweight Version of Wav2vec 2.0 for Automatic Speech Recognition
Jae-Hong Lee, Ji-Hwan Mo, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | CTRL: Continual Representation Learning to Transfer Information of Pre-trained for WAV2VEC 2.0
Jae-Hong Lee, Chae Won Lee, Jin-Seong Choi, Joon-Hyuk Chang, Woo Kyeong Seong, Jeonghan Lee |
INTERSPEECH | 1 |
| 1997 | A VPN Management Architecture for Supporting CNM Services in ATM Networks
Jong-Tae Park 0001, Jae-Hong Lee, James Won-Ki Hong, Young-Myung Kim, Sung-Bum Kim |
Integrated Network Management | 2 |