EDBT 2026 Demo / reviewers in the wild / expert
Hyeonseok Moon
dblp:295/3184
· DBLP profile ↗
18ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-0841-4262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and AdaptationabstractCurrent LLM-based services typically require users to submit raw text regardless of its sensitivity.While intuitive, such practice introduces substantial privacy risks, as unauthorized access may expose personal, medical, or legal information.Although prior defenses strived to mitigate these risks, they often incur substantial computational overhead and degrade model performance.To overcome this privacy-efficiency trade-off, we introduce Privacy-Preserving Fine-Tuning (PPFT), a novel training pipeline that eliminates the need for transmitting raw prompt text while maintaining a favorable balance between privacy preservation and model utility for both clients and service providers.Our approach operates in two stages: first, we train a client-side encoder together with a server-side projection module and LLM, enabling the server to condition on k-pooled prompt embeddings instead of raw text; second, we fine-tune the projection module and LLM on private, domain-specific data using noise-injected embeddings, allowing effective adaptation without exposing plain text prompts and requiring access to the decoder's internal parameters.Extensive experiments on domainspecific and general benchmarks demonstrate that PPFT achieves a striking balance between privacy and utility, maintaining competitive performance with minimal degradation compared to noise-free upper bounds. Jeongho Yoon, Chanhee Park, Yongchan Chun, Hyeonseok Moon, Heuiseok Lim |
ACL (1) | 4 |
| 2026 | Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation
Youngjoon Jang 0002, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim |
SIGIR | 3 |
| 2026 | SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
Jaehyung Seo, Hyeonseok Moon, Yoonna Jang, Heuiseok Lim |
Knowl. Based Syst. | 2 |
| 2025 | Cross-Lingual Optimization for Language Transfer in Large Language ModelsabstractAdapting large language models to other languages typically employs supervised finetuning (SFT) as a standard approach.However, it often suffers from an overemphasis on English performance, a phenomenon that is especially pronounced in data-constrained environments.To overcome these challenges, we propose Cross-Lingual Optimization (CLO) that efficiently transfers an English-centric LLM to a target language while preserving its English capabilities.CLO utilizes publicly available English SFT data and a translation model to enable cross-lingual transfer.We conduct experiments using five models on six languages, each possessing varying levels of resource.Our results show that CLO consistently outperforms SFT in both acquiring target language proficiency and maintaining English performance.Remarkably, in low-resource languages, CLO with only 3,200 samples surpasses SFT with 6,400 samples, demonstrating that CLO can achieve better performance with less data.Furthermore, we find that SFT is particularly sensitive to data quantity in medium and lowresource languages, whereas CLO remains robust.Our comprehensive analysis emphasizes the limitations of SFT and incorporates additional training strategies in CLO to enhance efficiency. Jungseob Lee, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim |
ACL (1) | 3 |
| 2025 | MIGRATE: Cross-Lingual Adaptation of Domain-Specific LLMs through Code-Switching and Embedding TransferabstractLarge Language Models (LLMs) have rapidly advanced, with domain-specific expert models emerging to handle specialized tasks across various fields. However, the predominant focus on English-centric models demands extensive data, making it challenging to develop comparable models for middle and low-resource languages. To address this limitation, we introduce Migrate, a novel method that leverages open-source static embedding models and up to 3 million tokens of code-switching data to facilitate the seamless transfer of embeddings to target languages. Migrate enables effective cross-lingual adaptation without requiring large-scale domain-specific corpora in the target language, promoting the accessibility of expert LLMs to a diverse range of linguistic communities. Our experimental results demonstrate that Migrate significantly enhances model performance in target languages, outperforming baseline and existing cross-lingual transfer methods. This approach provides a practical and efficient solution for extending the capabilities of domain-specific expert models. Seongtae Hong, Seungyoon Lee, Hyeonseok Moon, Heuiseok Lim |
COLING | 3 |
| 2025 | Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language ModelsabstractRecent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation.This progress highlights the need for challenging benchmarks that provide objective verification.In this paper, we introduce MCBench, a benchmark designed to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions.Unlike prior benchmarks that depend on subjective judgments or general reasoning, MCBench offers an objective, deterministic and codeverifiable evaluation.This setup allows us to systematically test whether LLMs can maintain accurate step-by-step execution, including instruction adherence, numerical computation, and long-range consistency in handling intermediate results.To ensure objective evaluation of these abilities, we provide a parallel reference code that can evaluate the accuracy of LLM output.We provide three evaluative metrics and three benchmark variants designed to measure the detailed instruction understanding capability of LLMs.Our analyses show that MCBench serves as an effective and objective tool for evaluating the capabilities of cuttingedge LLMs. Hyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok Lim |
EMNLP | 1 |
| 2025 | The Impact of Negated Text on Hallucination with Large Language ModelsabstractRecent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing.However, the impact of negated text on hallucination with LLMs remains largely unexplored.In this paper, we set three important yet unanswered research questions and aim to address them.To derive the answers, we investigate whether LLMs can recognize contextual shifts caused by negation and still reliably distinguish hallucinations comparable to affirmative cases.We also design the NegHalu dataset by reconstructing existing hallucination detection datasets with negated expressions.Our experiments demonstrate that LLMs struggle to detect hallucinations in negated text effectively, often producing logically inconsistent or unfaithful judgments.Moreover, we trace the internal state of LLMs as they process negated inputs at the token level and reveal the challenges of mitigating their unintended effects. Jaehyung Seo, Hyeonseok Moon, Heuiseok Lim |
EMNLP | 2 |
| 2024 | Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean TranslationabstractRecent machine translation (MT) systems have overcome language barriers for a wide range of users, yet they still carry the risk of critical meaning deviation. Critical error detection (CED) is a task that identifies an inherent risk of catastrophic meaning distortions in the machine translation output. With the importance of reflecting cultural elements in detecting critical errors, we introduce the culture-aware “Politeness” type in detecting English-Korean critical translation errors. Besides, we facilitate two tasks by providing multiclass labels: critical error detection and critical error type classification (CETC). Empirical evaluations reveal that our introduced data augmentation approach using a newly presented perturber significantly outperforms existing baselines in both tasks. Further analysis highlights the significance of multiclass labeling by demonstrating its superior effectiveness compared to binary labels. Sugyeong Eo, Jungwoo Lim, Chanjun Park, Dahyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim |
LREC/COLING | 6 |
| 2024 | Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in KoreanabstractCounter-narrative generation, i.e., the generation of fact-based responses to hate speech with the aim of correcting discriminatory beliefs, has been demonstrated to be an effective method to combat hate speech. However, its effectiveness is limited by the resource-intensive nature of dataset construction processes and only focuses on the primary language. To alleviate this problem, we propose a Korean Hate Speech Counter Punch (KHSCP), a cost-effective counter-narrative generation method in the Korean language. To this end, we release the first counter-narrative generation dataset in Korean and pose two research questions. Under the questions, we propose an effective augmentation method and investigate the reasonability of a large language model to overcome data scarcity in low-resource environments by leveraging existing resources. In this regard, we conduct several experiments to verify the effectiveness of the proposed method. Our results reveal that applying pre-existing resources can improve the generation performance by a significant margin. Through deep analysis on these experiments, this work proposes the possibility of overcoming the challenges of generating counter-narratives in low-resource environments. Seungyoon Lee, Chanjun Park, Dahyun Jung, Hyeonseok Moon, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim |
LREC/COLING | 4 |
| 2023 | Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded ConversationsabstractYoonna Jang, Suhyune Son, Jeongwoo Lee, Junyoung Son, Yuna Hur, Jungwoo Lim, Hyeonseok Moon, Kisu Yang, Heuiseok Lim. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yoonna Jang, Suhyune Son, Junyoung Son, Yuna Hur, Jungwoo Lim, Hyeonseok Moon, Kisu Yang, Heuiseok Lim |
EMNLP | 7 |
| 2023 | KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processingabstractAutomatic Speech Recognition (ASR) systems are instrumental across various applications, with their performance being critically tied to user satisfaction.Conventional evaluation metrics for ASR systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.Therefore, we aim to address the limitations of the previous ASR evaluation methods by introducing the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP).KE-BAP enables comprehensive analysis of ASR systems at both speech-and text levels, thereby facilitating a more balanced assessment encompassing speech recognition accuracy and user readability.KEBAP provides 37 newly defined speech-level resources incorporating diverse noise environments and speaker characteristics categories, also presenting 13 distinct textlevel error types.This paper demonstrates detailed statistical analyses of colloquial noise categories and textual error types.Furthermore, we conduct extensive validation and analysis on commercially deployed ASR systems, providing valuable insights into their performance.As a more fine-grained and real-world-centric evaluation method, KEBAP contributes to identifying and mitigating potential weaknesses in ASR systems.* Equally contributed, ‡ Corresponding author 1 Recognition accuracy is the measure of accurately perceiving phonemes as they are externally expressed, regardless of user input quality (Liao et al., 2022).Conventional (WER, CER) 0.45 KEBAP Error types Explainability Noise Type Description Washer/dryer machine Home appliances Vacuum cleaner Difficulty in recognition due to ambient electrical appliance noise.Motorcycle Siren Individual transportation Honk Difficulty in recognition due to surrounding individual transportation noise.Road side Street Crowd Difficulty in recognition due to the surrounding street noise.Conversation Cafe/restaurant Non-conversation Challenges in perception due to the noise in cafes/restaurants.Traditional market Market/shopping mall Shopping mall Difficulties in perception caused by the noise in markets/shopping malls.Subway platform Inside the subway Inside the train (STR/KTX) Public transportation Inside the bus Difficulty in recognition due to surrounding public transportation noise.Train terminal waiting room Terminal Bus terminal waiting room Challenges in perception due to the noise at terminals.Outdoor construction site Construction site Indoor construction site Difficulties in perception caused by the noise at construction sites.processing process Factory Assembly process Difficulties in perception caused by the noise in factories.Sound of rain Nature ambient Sound of the waves Challenges in perception due to natural ambient noise.Noisy environment Etc.Artificial mechanical sound In cases where external noise is present, although not falling into the aforementioned categories. Seonmin Koo, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Hyeonseok Moon, Heuiseok Lim |
EMNLP | 6 |
| 2023 | CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsabstractKorean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction.The complexity of morphological variations allows for diverse sentence forms based on the syntactic-semantic integration of functional morphemes (i.e., affixes) to lexical morphemes (i.e., roots).With this in mind, we propose a method -CHEF, replicating the morphological transformations inherent in sentences based on lexical and functional morpheme combinations through generative data augmentation.CHEF operates using a morpheme blender and a label discriminator, thereby enhancing the diversity of Korean sentence forms by capturing the properties of agglutination while maintaining label consistency.We conduct experiments on Korean multiple classification datasets, improving model performance in full-and few-shot settings.Our proposed method boosts performance beyond the preceding data augmentation methods without incurring external data usage.We demonstrate that our approach achieves comparable results yielded by augmentation techniques that use large language models (LLMs). Jaehyung Seo, Hyeonseok Moon, Jaewook Lee 0008, Sugyeong Eo, Chanjun Park, Heuiseok Lim |
EMNLP | 2 |
| 2023 | Informative Evidence-guided Prompt-based Fine-tuning for English-Korean Critical Error DetectionabstractDaHyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Dahyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim |
IJCNLP (1) | 4 |
| 2023 | Doubts on the reliability of parallel corpus filtering
Hyeonseok Moon, Chanjun Park, Seonmin Koo, Jungseob Lee, Jaehyung Seo, Sugyeong Eo, Yoonna Jang, Hyunjoong Kim, Hyoung-gyu Lee, Heuiseok Lim |
Expert Syst. Appl. | 1 |
| 2022 | QUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine TranslationabstractWith the recent advance in neural machine translation demonstrating its importance, research on quality estimation (QE) has been steadily progressing. QE aims to automatically predict the quality of machine translation (MT) output without reference sentences. Despite its high utility in the real world, there remain several limitations concerning manual QE data creation: inevitably incurred non-trivial costs due to the need for translation experts, and issues with data scaling and language expansion. To tackle these limitations, we present QUAK, a Korean-English synthetic QE dataset generated in a fully automatic manner. This consists of three sub-QUAK datasets QUAK-M, QUAK-P, and QUAK-H, produced through three strategies that are relatively free from language constraints. Since each strategy requires no human effort, which facilitates scalability, we scale our data up to 1.58M for QUAK-P, H and 6.58M for QUAK-M. As an experiment, we quantitatively analyze word-level QE results in various ways while performing statistical analysis. Moreover, we show that datasets scaled in an efficient way also contribute to performance improvements by observing meaningful performance gains in QUAK-M, P when adding data up to 1.58M. Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Gyeongmin Kim, Jungseob Lee, Heuiseok Lim |
COLING | 3 |
| 2022 | Empirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editingabstractAutomatic post-editing (APE) refers to a research field that aims to automatically correct errors included in the translation sentences derived by the machine translation system. This study has several limitations, considering the data acquisition, because there is no official dataset for most language pairs. Moreover, the amount of data is restricted even for language pairs in which official data has been released, such as WMT. To solve this problem and promote universal APE research regardless of APE data existence, this study proposes a method for automatically generating APE data based on a noising scheme from a parallel corpus. Particularly, we propose a human mimicking errors-based noising scheme that considers a practical correction process at the human level. We propose a precise inspection to attain high performance, and we derived the optimal noising schemes that show substantial effectiveness. Through these, we also demonstrate that depending on the type of noise, the noising scheme-based APE data generation may lead to inferior performance. In addition, we propose a dynamic noise injection strategy that enables the acquisition of a robust error correction capability and demonstrated its effectiveness by comparative analysis. This study enables obtaining a high performance APE model without human-generated data and can promote universal APE research for all language pairs targeting English. Hyeonseok Moon, Chanjun Park, Seolhwa Lee, Jaehyung Seo, Jungseob Lee, Sugyeong Eo, Heuiseok Lim |
LREC | 1 |
| 2022 | Priming Ancient Korean Neural Machine TranslationabstractIn recent years, there has been an increasing need for the restoration and translation of historical languages. In this study, we attempt to translate historical records in ancient Korean language based on neural machine translation (NMT). Inspired by priming, a cognitive science theory that two different stimuli influence each other, we propose novel priming ancient-Korean NMT (AKNMT) using bilingual subword embedding initialization with structural property awareness in the ancient documents. Finally, we obtain state-of-the-art results in the AKNMT task. To the best of our knowledge, we confirm the possibility of developing a human-centric model that incorporates the concepts of cognitive science and analyzes the result from the perspective of interference and cognitive dissonance theory for the first time. Chanjun Park, Seolhwa Lee, Jaehyung Seo, Hyeonseok Moon, Sugyeong Eo, Heuiseok Lim |
LREC | 4 |
| 2022 | PU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, Heuiseok Lim |
Knowl. Based Syst. | 6 |