EDBT 2026 Demo / reviewers in the wild / expert
Yeachan Kim
dblp:224/6085
· DBLP profile ↗
17ranked-venue papers
10as first author
13since 2021 · last 2026
0009-0004-2069-5265ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey on Memory-Efficient Fine-Tuning for Large Language ModelsabstractAbstract Fine-tuning large language models (LLMs) is a crucial process to align them with human intentions, yet this process remains memoryintensive, varying across tasks and model architectures. These huge and variable memory costs complicate scaling and deployment of LLMs, especially on limited hardware. However, existing surveys on memory efficiency are often either superficial or too narrow in scope, typically focusing on specific subfields. To address this gap, this survey presents the first systematic review of memory-efficient fine-tuning (MEFT) tailored for LLMs. To structure the research landscape, we first categorize existing approaches by their optimization environments (i.e., model itself and systems) and further classify model-based approaches by their specific optimization targets. We also discuss evaluation strategies for assessing MEFT methods and provide empirical analyses. By highlighting challenges and future directions based on current methods, this survey aims to serve as a practical guide for developing MEFT methods. Yeachan Kim, Sangkeun Lee |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language ModelsabstractFine-tuning is widely recognized as a crucial process for aligning large language models (LLMs) with human intentions.However, the substantial memory requirements associated with fine-tuning pose a significant barrier to extending the applicability of LLMs.While parameter-efficient fine-tuning can be a promising approach by reducing trainable parameters, intermediate activations still need to be cached to compute gradients during the backward pass, thereby limiting overall memory efficiency.In this work, we propose Saliency-Guided Gradient Flow (SAGE), a memoryefficient fine-tuning method designed to minimize the memory specifically associated with cached intermediate activations.The key strategy is to selectively cache activations based on their saliency during the forward pass and then use these activations for the backward pass.This process transforms the dense backward pass into a sparse one, thereby enhancing memory efficiency.To verify whether SAGE can serve as an efficient alternative for fine-tuning, we conduct comprehensive experiments across diverse fine-tuning scenarios and setups.The experimental results show that SAGE substantially improves memory efficiency without a significant loss in accuracy, highlighting its broad value in real-world applications 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 1 |
| 2025 | Curriculum Debiasing: Toward Robust Parameter-Efficient Fine-Tuning Against Dataset BiasesabstractParameter-efficient fine-tuning (PEFT) addresses the memory footprint issue of full fine-tuning by modifying only a subset of model parameters. However, on datasets exhibiting spurious correlations, we observed that PEFT slows down the model’s convergence on unbiased examples, while the convergence on biased examples remains fast. This leads to the model’s overfitting on biased examples, causing significant performance degradation in out-of-distribution (OOD) scenarios. Traditional debiasing methods mitigate this issue by emphasizing unbiased examples during training but often come at the cost of in-distribution (ID) performance drops. To address this trade-off issue, we propose a curriculum debiasing framework that presents examples in a biased-to-unbiased order. Our framework initially limits the model’s exposure to unbiased examples, which are harder to learn, allowing it to first establish a foundation on easier-to-converge biased examples. As training progresses, we gradually increase the proportion of unbiased examples in the training set, guiding the model away from reliance on spurious correlations. Compared to the original PEFT methods, our method accelerates convergence on unbiased examples by approximately twofold and improves ID and OOD performance by 1.2% and 8.0%, respectively. Yeachan Kim, Wing-Lam Mok, SangKeun Lee 0001 |
ACL (1) | 2 |
| 2025 | Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware AlignmentabstractMolecule and text representation learning has gained increasing interest due to its potential for enhancing the understanding of chemical information.However, existing models often struggle to capture subtle differences between molecules and their descriptions, as they lack the ability to learn fine-grained alignments between molecular substructures and chemical phrases.To address this limitation, we introduce MolBridge, a novel molecule-text learning framework based on substructure-aware alignments.Specifically, we augment the original molecule-description pairs with additional alignment signals derived from molecular substructures and chemical phrases.To effectively learn from these enriched alignments, Mol-Bridge employs substructure-aware contrastive learning, coupled with a self-refinement mechanism that filters out noisy alignment signals.Experimental results show that MolBridge effectively captures fine-grained correspondences and outperforms state-of-the-art baselines on a wide range of molecular benchmarks, underscoring the importance of substructure-aware alignment in molecule-text learning. Hyuntae Park, Yeachan Kim, SangKeun Lee 0001 |
EMNLP | 2 |
| 2024 | SparseFlow: Accelerating Transformers by Sparsifying Information FlowsabstractTransformers have become the de-facto standard for natural language processing.However, dense information flows within transformers pose significant challenges for real-time and resource-constrained devices, as computational complexity grows quadratically with sequence length.To counteract such dense information flows, we propose SPARSEFLOW, a novel efficient method designed to sparsify the dense pathways of token representations across all transformer blocks.To this end, SPARSEFLOW parameterizes the information flows linking token representations to transformer blocks.These parameterized information flows are optimized to be sparse, allowing only the salient information to pass through into the blocks.To validate the efficacy of SPARSEFLOW, we conduct comprehensive experiments across diverse benchmarks (understanding and generation), scales (ranging from millions to billions), architectures (including encoders, decoders, and seq-to-seq models), and modalities (such as language-only and vision-language).The results convincingly demonstrate that sparsifying the dense information flows leads to substantial speedup gains without compromising task accuracy.For instance, SPARSEFLOW reduces computational costs by half on average, without a significant loss in accuracy 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 1 |
| 2024 | Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label LearningabstractParameter-efficient fine-tuning (PEFT) has enabled the efficient optimization of cumbersome language models in real-world settings.However, as datasets in such environments often contain noisy labels that adversely affect performance, PEFT methods are inevitably exposed to noisy labels.Despite this challenge, the adaptability of PEFT to noisy environments remains underexplored.To bridge this gap, we investigate various PEFT methods under noisy labels.Interestingly, our findings reveal that PEFT has difficulty in memorizing noisy labels due to its inherently limited capacity, resulting in robustness.However, we also find that such limited capacity simultaneously makes PEFT more vulnerable to interference of noisy labels, impeding the learning of clean samples.To address this issue, we propose Clean Routing (CleaR), a novel routing-based PEFT approach that adaptively activates PEFT modules.In CleaR, PEFT modules are preferentially exposed to clean data while bypassing the noisy ones, thereby minimizing the noisy influence.To verify the efficacy of CleaR, we perform extensive experiments on diverse configurations of noisy labels.The results convincingly demonstrate that CleaR leads to substantially improved performance in noisy environments 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 1 |
| 2024 | MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionabstractChemical representation learning has gained increasing interest due to the limited availability of supervised data in fields such as drug and materials design.This interest particularly extends to chemical language representation learning, which involves pre-training Transformers on SMILES sequences -textual descriptors of molecules.Despite its success in molecular property prediction, current practices often lead to overfitting and limited scalability due to early convergence.In this paper, we introduce a novel chemical language representation learning framework, called MolTRES, to address these issues.MolTRES incorporates generator-discriminator training, allowing the model to learn from more challenging examples that require structural understanding.In addition, we enrich molecular representations by transferring knowledge from scientific literature by integrating external materials embedding.Experimental results show that our models outperform existing state-of-the-art models on popular molecular property prediction tasks. github.com/irishev/MolTRES Jun-Hyung Park, Yeachan Kim, Hyuntae Park, SangKeun Lee 0001 |
EMNLP | 2 |
| 2023 | Dynamic Structure Pruning for Compressing CNNsabstractStructure pruning is an effective method to compress and accelerate neural networks. While filter and channel pruning are preferable to other structure pruning methods in terms of realistic acceleration and hardware compatibility, pruning methods with a finer granularity, such as intra-channel pruning, are expected to be capable of yielding more compact and computationally efficient networks. Typical intra-channel pruning methods utilize a static and hand-crafted pruning granularity due to a large search space, which leaves room for improvement in their pruning performance. In this work, we introduce a novel structure pruning method, termed as dynamic structure pruning, to identify optimal pruning granularities for intra-channel pruning. In contrast to existing intra-channel pruning methods, the proposed method automatically optimizes dynamic pruning granularities in each layer while training deep neural networks. To achieve this, we propose a differentiable group learning method designed to efficiently learn a pruning granularity based on gradient-based learning of filter groups. The experimental results show that dynamic structure pruning achieves state-of-the-art pruning performance and better realistic acceleration on a GPU compared with channel pruning. In particular, it reduces the FLOPs of ResNet50 by 71.85% without accuracy degradation on the ImageNet dataset. Our code is available at https://github.com/irishev/DSP. Jun-Hyung Park, Yeachan Kim, Joon-Young Choi, SangKeun Lee 0001 |
AAAI | 2 |
| 2023 | Improving Bias Mitigation through Bias Experts in Natural Language UnderstandingabstractBiases in the dataset often enable the model to achieve high performance on in-distribution data, while poorly performing on out-ofdistribution data.To mitigate the detrimental effect of the bias on the networks, previous works have proposed debiasing methods that down-weight the biased examples identified by an auxiliary model, which is trained with explicit bias labels.However, finding a type of bias in datasets is a costly process.Therefore, recent studies have attempted to make the auxiliary model biased without the guidance (or annotation) of bias labels, by constraining the model's training environment or the capability of the model itself.Despite the promising debiasing results of recent works, the multiclass learning objective, which has been naively used to train the auxiliary model, may harm the bias mitigation effect due to its regularization effect and competitive nature across classes.As an alternative, we propose a new debiasing framework that introduces binary classifiers between the auxiliary model and the main model, coined bias experts.Specifically, each bias expert is trained on a binary classification task derived from the multi-class classification task via the One-vs-Rest approach.Experimental results demonstrate that our proposed strategy improves the bias identification ability of the auxiliary model.Consequently, our debiased model consistently outperforms the state-ofthe-art on various challenge datasets.1 Eojin Jeon, Juhyeong Park, Yeachan Kim, Wing-Lam Mok, SangKeun Lee 0001 |
EMNLP | 4 |
| 2023 | Leap-of-Thought: Accelerating Transformers via Dynamic Token RoutingabstractComputational inefficiency in transformers has been a long-standing challenge, hindering the deployment in resource-constrained or realtime applications.One promising approach to mitigate this limitation is to progressively remove less significant tokens, given that the sequence length strongly contributes to the inefficiency.However, this approach entails a potential risk of losing crucial information due to the irrevocable nature of token removal.In this paper, we introduce Leap-of-Thought (LoT), a novel token reduction approach that dynamically routes tokens within layers.Unlike previous work that irrevocably discards tokens, LoT enables tokens to 'leap' across layers.This ensures that all tokens remain accessible in subsequent layers while reducing the number of tokens processed within layers.We achieve this by pairing the transformer with dynamic token routers, which learn to selectively process tokens essential for the task.Evaluation results clearly show that LoT achieves substantial improvement on computational efficiency.Specifically, LoT attains up to 25× faster inference time without a significant loss in accuracy 1 . Yeachan Kim, Jun-Hyung Park, SangKeun Lee 0001 |
EMNLP | 1 |
| 2023 | Phase-shifted adversarial trainingabstractAdversarial training (AT) has been considered an imperative component for safely deploying neural network-based applications. However, it typically comes with slow convergence and worse performance on clean samples (i.e., non-adversarial samples). In this work, we analyze the behavior of neural networks during learning with adversarial samples through the lens of response frequency. Interestingly, we observe that AT causes neural networks to converge slowly to high-frequency information, resulting in highly oscillatory predictions near each data point. To learn high-frequency content efficiently, we first prove that a universal phenomenon, the frequency principle (i.e., lower frequencies are learned first), still holds in AT. Building upon this theoretical foundation, we present a novel approach to AT, which we call phase-shifted adversarial training (PhaseAT). In PhaseAT, the high-frequency components, which are a contributing factor to slow convergence, are adaptively shifted into the low-frequency range where faster convergence occurs. For evaluation, we conduct extensive experiments on CIFAR-10 and ImageNet, using an adaptive attack that is carefully designed for reliable evaluation. Comprehensive results show that PhaseAT substantially improves convergence for high-frequency information, thereby leading to improved adversarial robustness. Yeachan Kim, Seongyeon Kim, Ihyeok Seo, Bonggun Shin |
UAI | 1 |
| 2022 | In Defense of Core-set: A Density-aware Core-set Selection for Active LearningabstractActive learning enables the efficient construction of a labeled dataset by labeling informative samples from an unlabeled dataset. In a real-world active learning scenario, the use of diversity-based sampling is indispensable because there are many redundant or highly similar samples. Core-set approach is the promising diversity-based method selecting diverse samples by considering the distance between samples. However, the approach poorly performs compared to the uncertainty-based method that selects the most difficult samples where neural models reveal low confidence. In this work, we analyze the feature space through the lens of density and, interestingly, observe that locally sparse regions tend to have more informative samples than dense regions. Motivated by our analysis, we empower the core-set approach with the density-awareness and propose a density-aware core-set (DACS) which estimates the density of the unlabeled samples and selects diverse samples mainly from sparse regions which are treated as the informative regions. To reduce the computational bottlenecks in estimating the density, we introduce a new density approximation based on locality-sensitive hashing. Experimental results demonstrate the efficacy of DACS in both classification and regression tasks and specifically show that DACS can produce state-of-the-art performance in a practical scenario. Since DACS is weakly dependent on architectures, we also present a simple yet effective combination method to show that the existing methods can be beneficially combined with DACS. Yeachan Kim, Bonggun Shin |
KDD | 1 |
| 2022 | Context-based Virtual Adversarial Training for Text Classification with Noisy LabelsabstractDeep neural networks (DNNs) have a high capacity to completely memorize noisy labels given sufficient training time, and its memorization unfortunately leads to performance degradation. Recently, virtual adversarial training (VAT) attracts attention as it could further improve the generalization of DNNs in semi-supervised learning. The driving force behind VAT is to prevent the models from overffiting to data points by enforcing consistency between the inputs and the perturbed inputs. These strategy could be helpful in learning from noisy labels if it prevents neural models from learning noisy samples while encouraging the models to generalize clean samples. In this paper, we propose context-based virtual adversarial training (ConVAT) to prevent a text classifier from overfitting to noisy labels. Unlike the previous works, the proposed method performs the adversarial training in the context level rather than the inputs. It makes the classifier not only learn its label but also its contextual neighbors, which alleviate the learning from noisy labels by preserving contextual semantics on each data point. We conduct extensive experiments on four text classification datasets with two types of label noises. Comprehensive experimental results clearly show that the proposed method works quite well even with extremely noisy settings. Do-Myoung Lee, Yeachan Kim, Chang-gyun Seo |
LREC | 2 |
| 2020 | Adaptive Compression of Word EmbeddingsabstractDistributed representations of words have been an indispensable component for natural language processing (NLP) tasks.However, the large memory footprint of word embeddings makes it challenging to deploy NLP models to memory-constrained devices (e.g., selfdriving cars, mobile devices).In this paper, we propose a novel method to adaptively compress word embeddings.We fundamentally follow a code-book approach that represents words as discrete codes such as (8, 5, 2, 4).However, unlike prior works that assign the same length of codes to all words, we adaptively assign different lengths of codes to each word by learning downstream tasks.The proposed method works in two steps.First, each word directly learns to select its code length in an end-to-end manner by applying the Gumbel-softmax tricks.After selecting the code length, each word learns discrete codes through a neural network with a binary constraint.To showcase the general applicability of the proposed method, we evaluate the performance on four different downstream tasks.Comprehensive evaluation results clearly show that our method is effective and makes the highly compressed word embeddings without hurting the task accuracy.Moreover, we show that our model assigns word to each code-book by considering the significance of tasks. Yeachan Kim, Kang-Min Kim, SangKeun Lee 0001 |
ACL | 1 |
| 2020 | Representation Learning for Unseen Words by Bridging Subwords to Semantic NetworksabstractPre-trained word embeddings are widely used in various fields. However, the coverage of pre-trained word embeddings only includes words that appeared in corpora where pre-trained embeddings are learned. It means that the words which do not appear in training corpus are ignored in tasks, and it could lead to the limited performance of neural models. In this paper, we propose a simple yet effective method to represent out-of-vocabulary (OOV) words. Unlike prior works that solely utilize subword information or knowledge, our method makes use of both information to represent OOV words. To this end, we propose two stages of representation learning. In the first stage, we learn subword embeddings from the pre-trained word embeddings by using an additive composition function of subwords. In the second stage, we map the learned subwords into semantic networks (e.g., WordNet). We then re-train the subword embeddings by using lexical entries on semantic lexicons that could include newly observed subwords. This two-stage learning makes the coverage of words broaden to a great extent. The experimental results clearly show that our method provides consistent performance improvements over strong baselines that use subwords or lexical resources separately. Yeachan Kim, Kang-Min Kim, SangKeun Lee 0001 |
LREC | 1 |
| 2019 | From Small-scale to Large-scale Text ClassificationabstractNeural network models have achieved impressive results in the field of text classification. However, existing approaches often suffer from insufficient training data in a large-scale text classification involving a large number of categories (e.g., several thousands of categories). Several neural network models have utilized multi-task learning to overcome the limited amount of training data. However, these approaches are also limited to small-scale text classification. In this paper, we propose a novel neural network-based multi-task learning framework for large-scale text classification. To this end, we first treat the different scales of text classification (i.e., large and small numbers of categories) as multiple, related tasks. Then, we train the proposed neural network, which learns small- and large-scale text classification tasks simultaneously. In particular, we further enhance this multi-task learning architecture by using a gate mechanism, which controls the flow of features between the small- and large-scale text classification tasks. Experimental results clearly show that our proposed model improves the performance of the large-scale text classification task with the help of the small-scale text classification task. The proposed scheme exhibits significant improvements of as much as 14% and 5% in terms of micro-averaging and macro-averaging F1-score, respectively, over state-of-the-art techniques. Kang-Min Kim, Yeachan Kim, Ji-Min Lee, SangKeun Lee 0001 |
WWW | 2 |
| 2018 | Learning to Generate Word Representations using Subword InformationabstractDistributed representations of words play a major role in the field of natural language processing by encoding semantic and syntactic information of words. However, most existing works on learning word representations typically regard words as individual atomic units and thus are blind to subword information in words. This further gives rise to a difficulty in representing out-of-vocabulary (OOV) words. In this paper, we present a character-based word representation approach to deal with this limitation. The proposed model learns to generate word representations from characters. In our model, we employ a convolutional neural network and a highway network over characters to extract salient features effectively. Unlike previous models that learn word representations from a large corpus, we take a set of pre-trained word embeddings and generalize it to word entries, including OOV words. To demonstrate the efficacy of the proposed model, we perform both an intrinsic and an extrinsic task which are word similarity and language modeling, respectively. Experimental results show clearly that the proposed model significantly outperforms strong baseline models that regard words or their subwords as atomic units. For example, we achieve as much as 18.5% improvement on average in perplexity for morphologically rich languages compared to strong baselines in the language modeling task. Yeachan Kim, Kang-Min Kim, Ji-Min Lee, SangKeun Lee 0001 |
COLING | 1 |