Jaehyung Seo

dblp:298/7721 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-4761-9818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 20 since 2021
YearPublicationVenuePosition
2026 No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
abstract
The Plain Writing Act in the United States requires government documents to be written in clear and simple language.However, existing summarization systems struggle to address diverse linguistic and cognitive barriers among general readers.We propose NRLB (No Reader Left Behind), a unified multi-agent framework for plain language summarization that simulates three representative reader groups: elementary school students, non-native speakers, and readers with attention deficits.NRLB integrates template-based planning with an iterative feedback loop guided by simulated readers and domain expert revision to address comprehension barriers such as unknown terms, missing contexts, and confusing sentences.Evaluations across multiple datasets demonstrate consistent improvements in both readability and factuality.Human evaluation further supports these findings, with annotator preference rates ranging from 55% to 76%, highlighting NRLB's ability to generate summaries that are both faithful to the source and accessible to a wide range of readers.
Jimin Jung, MyoungJin Kim, Jaehyung Seo, Heuiseok Lim
ACL (1)3
2026 HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
abstract
Retrieval-augmented generation (RAG) for document-based open-domain question answering (ODQA) over large industrial corpora faces two core bottlenecks: routing to the correct document and combining scattered evidence.Flat text chunks and page-level images often fail to (i) identify the right document among thousands of candidates and (ii) connect multimodal evidence, such as tables and figures, within a fixed token budget.We propose HiKEY, a hierarchical tree-based multimodal retrieval framework that treats document hierarchy as a first-class retrieval signal.Rather than simply chunking text, HiKEY uses Document Hierarchical Parsing (DHP) to reconstruct a logical heterogeneous graph with explicit parent-child relations.At query time, HiKEY follows a hierarchical coarse-to-fine process: it first performs global routing with hierarchical indexes to prune the corpus, and then ranks sections with a multimodal fusion strategy that selects the most discriminative evidence.It finally builds a token-efficient evidence subgraph through hybrid structural-semantic packing.Experiments on ODQA benchmarks show that HiKEY outperforms page-and chunk-based baselines, improving retrieval recall by up to 12.9 points and end-to-end QA by up to 6.8 points.
Joongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
ACL (1)4
2026 MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
abstract
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
ACL (1)25
2026 SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
Jaehyung Seo, Hyeonseok Moon, Yoonna Jang, Heuiseok Lim
Knowl. Based Syst.1
2025 Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
abstract
Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation.This progress highlights the need for challenging benchmarks that provide objective verification.In this paper, we introduce MCBench, a benchmark designed to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions.Unlike prior benchmarks that depend on subjective judgments or general reasoning, MCBench offers an objective, deterministic and codeverifiable evaluation.This setup allows us to systematically test whether LLMs can maintain accurate step-by-step execution, including instruction adherence, numerical computation, and long-range consistency in handling intermediate results.To ensure objective evaluation of these abilities, we provide a parallel reference code that can evaluate the accuracy of LLM output.We provide three evaluative metrics and three benchmark variants designed to measure the detailed instruction understanding capability of LLMs.Our analyses show that MCBench serves as an effective and objective tool for evaluating the capabilities of cuttingedge LLMs.
Hyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok Lim
EMNLP3
2025 The Impact of Negated Text on Hallucination with Large Language Models
abstract
Recent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing.However, the impact of negated text on hallucination with LLMs remains largely unexplored.In this paper, we set three important yet unanswered research questions and aim to address them.To derive the answers, we investigate whether LLMs can recognize contextual shifts caused by negation and still reliably distinguish hallucinations comparable to affirmative cases.We also design the NegHalu dataset by reconstructing existing hallucination detection datasets with negated expressions.Our experiments demonstrate that LLMs struggle to detect hallucinations in negated text effectively, often producing logically inconsistent or unfaithful judgments.Moreover, we trace the internal state of LLMs as they process negated inputs at the token level and reveal the challenges of mitigating their unintended effects.
Jaehyung Seo, Hyeonseok Moon, Heuiseok Lim
EMNLP1
2025 MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
abstract
RAG-based QA has emerged as a powerful method for processing long industrial documents.However, conventional text chunking approaches often neglect the complex structures of long industrial documents, causing information loss and reduced answer quality.To address this, we introduce MultiDocFusion, a multimodal chunking pipeline that integrates: (i) detection of document regions using visionbased document parsing, (ii) text extraction from these regions via OCR, (iii) reconstruction of document structure into a hierarchical tree using large language model (LLM)based document section hierarchical parsing (DSHP-LLM), and (iv) construction of hierarchical chunks through DFS-based Grouping.Extensive experiments across industrial benchmarks demonstrate that MultiDocFusion improves retrieval precision by 8-15% and ANLS QA scores by 2-3% compared to baselines, emphasizing the critical role of explicitly leveraging document hierarchy for multimodal document-based QA.These significant performance gains underscore the necessity of structure-aware chunking in enhancing the fidelity of RAG-based QA systems.
Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
EMNLP4
2025 K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language Models
abstract
Recent researchers and companies have been developing large language models (LLMs) specifically designed for particular purposes and have achieved significant advancements in various natural language processing tasks. However, LLMs are still prone to generating hallucinations—results that are unfaithful or inconsistent with the given input. As a result, the need for datasets to evaluate and demonstrate the hallucination detection capabilities of LLMs is increasingly recognized. Nonetheless, the Korean NLP community lacks publicly available benchmark datasets demonstrating the faithfulness of knowledge-based information. Furthermore, the few existing datasets that evaluate hallucination are limited in their access to the entire dataset, restricting detailed analysis beyond simple scoring, and are based on translated English knowledge. To address these challenges, we introduce K-HALU, a Korean benchmark designed to evaluate LLMs' hallucination detection in Korean. This benchmark contains seven domains, considering the faithfulness of statements based on knowledge documents compiled from Korean news, magazines, and books. For more strict evaluation, 40% of the dataset is structured as multiple-answer questions, requiring models to select all possible correct answers from the given options. Our empirical results show that open-source LLMs still struggle with hallucination detection in Korean knowledge, emphasizing the need for a more detailed analysis of their limitations.
Jaehyung Seo, Heuiseok Lim
ICLR1
2025 CoME: An Unlearning-based Approach to Conflict-free Model Editing
abstract
Dahyun Jung, Jaehyung Seo, Jaewook Lee, Chanjun Park, Heuiseok Lim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dahyun Jung, Jaehyung Seo, Jaewook Lee 0008, Chanjun Park, Heuiseok Lim
NAACL (Long Papers)2
2025 An analysis on language transfer of pre-trained language model with cross-lingual post-training
Suhyune Son, Chanjun Park, Jungseob Lee, Midan Shim, Chanhee Lee 0004, Yoonna Jang, Jaehyung Seo, Jungwoo Lim, Heuiseok Lim
Expert Syst. Appl.7
2024 Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean Translation
abstract
Recent machine translation (MT) systems have overcome language barriers for a wide range of users, yet they still carry the risk of critical meaning deviation. Critical error detection (CED) is a task that identifies an inherent risk of catastrophic meaning distortions in the machine translation output. With the importance of reflecting cultural elements in detecting critical errors, we introduce the culture-aware “Politeness” type in detecting English-Korean critical translation errors. Besides, we facilitate two tasks by providing multiclass labels: critical error detection and critical error type classification (CETC). Empirical evaluations reveal that our introduced data augmentation approach using a newly presented perturber significantly outperforms existing baselines in both tasks. Further analysis highlights the significance of multiclass labeling by demonstrating its superior effectiveness compared to binary labels.
Sugyeong Eo, Jungwoo Lim, Chanjun Park, Dahyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
LREC/COLING7
2024 Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in Korean
abstract
Counter-narrative generation, i.e., the generation of fact-based responses to hate speech with the aim of correcting discriminatory beliefs, has been demonstrated to be an effective method to combat hate speech. However, its effectiveness is limited by the resource-intensive nature of dataset construction processes and only focuses on the primary language. To alleviate this problem, we propose a Korean Hate Speech Counter Punch (KHSCP), a cost-effective counter-narrative generation method in the Korean language. To this end, we release the first counter-narrative generation dataset in Korean and pose two research questions. Under the questions, we propose an effective augmentation method and investigate the reasonability of a large language model to overcome data scarcity in low-resource environments by leveraging existing resources. In this regard, we conduct several experiments to verify the effectiveness of the proposed method. Our results reveal that applying pre-existing resources can improve the generation performance by a significant margin. Through deep analysis on these experiments, this work proposes the possibility of overcoming the challenges of generating counter-narratives in low-resource environments.
Seungyoon Lee, Chanjun Park, Dahyun Jung, Hyeonseok Moon, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
LREC/COLING5
2023 KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processing
abstract
Automatic Speech Recognition (ASR) systems are instrumental across various applications, with their performance being critically tied to user satisfaction.Conventional evaluation metrics for ASR systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.Therefore, we aim to address the limitations of the previous ASR evaluation methods by introducing the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP).KE-BAP enables comprehensive analysis of ASR systems at both speech-and text levels, thereby facilitating a more balanced assessment encompassing speech recognition accuracy and user readability.KEBAP provides 37 newly defined speech-level resources incorporating diverse noise environments and speaker characteristics categories, also presenting 13 distinct textlevel error types.This paper demonstrates detailed statistical analyses of colloquial noise categories and textual error types.Furthermore, we conduct extensive validation and analysis on commercially deployed ASR systems, providing valuable insights into their performance.As a more fine-grained and real-world-centric evaluation method, KEBAP contributes to identifying and mitigating potential weaknesses in ASR systems.* Equally contributed, ‡ Corresponding author 1 Recognition accuracy is the measure of accurately perceiving phonemes as they are externally expressed, regardless of user input quality (Liao et al., 2022).Conventional (WER, CER) 0.45 KEBAP Error types Explainability Noise Type Description Washer/dryer machine Home appliances Vacuum cleaner Difficulty in recognition due to ambient electrical appliance noise.Motorcycle Siren Individual transportation Honk Difficulty in recognition due to surrounding individual transportation noise.Road side Street Crowd Difficulty in recognition due to the surrounding street noise.Conversation Cafe/restaurant Non-conversation Challenges in perception due to the noise in cafes/restaurants.Traditional market Market/shopping mall Shopping mall Difficulties in perception caused by the noise in markets/shopping malls.Subway platform Inside the subway Inside the train (STR/KTX) Public transportation Inside the bus Difficulty in recognition due to surrounding public transportation noise.Train terminal waiting room Terminal Bus terminal waiting room Challenges in perception due to the noise at terminals.Outdoor construction site Construction site Indoor construction site Difficulties in perception caused by the noise at construction sites.processing process Factory Assembly process Difficulties in perception caused by the noise in factories.Sound of rain Nature ambient Sound of the waves Challenges in perception due to natural ambient noise.Noisy environment Etc.Artificial mechanical sound In cases where external noise is present, although not falling into the aforementioned categories.
Seonmin Koo, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Hyeonseok Moon, Heuiseok Lim
EMNLP4
2023 CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients
abstract
Korean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction.The complexity of morphological variations allows for diverse sentence forms based on the syntactic-semantic integration of functional morphemes (i.e., affixes) to lexical morphemes (i.e., roots).With this in mind, we propose a method -CHEF, replicating the morphological transformations inherent in sentences based on lexical and functional morpheme combinations through generative data augmentation.CHEF operates using a morpheme blender and a label discriminator, thereby enhancing the diversity of Korean sentence forms by capturing the properties of agglutination while maintaining label consistency.We conduct experiments on Korean multiple classification datasets, improving model performance in full-and few-shot settings.Our proposed method boosts performance beyond the preceding data augmentation methods without incurring external data usage.We demonstrate that our approach achieves comparable results yielded by augmentation techniques that use large language models (LLMs).
Jaehyung Seo, Hyeonseok Moon, Jaewook Lee 0008, Sugyeong Eo, Chanjun Park, Heuiseok Lim
EMNLP1
2023 Informative Evidence-guided Prompt-based Fine-tuning for English-Korean Critical Error Detection
abstract
DaHyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Dahyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
IJCNLP (1)5
2023 Doubts on the reliability of parallel corpus filtering
Hyeonseok Moon, Chanjun Park, Seonmin Koo, Jungseob Lee, Jaehyung Seo, Sugyeong Eo, Yoonna Jang, Hyunjoong Kim, Hyoung-gyu Lee, Heuiseok Lim
Expert Syst. Appl.6
2022 QUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine Translation
abstract
With the recent advance in neural machine translation demonstrating its importance, research on quality estimation (QE) has been steadily progressing. QE aims to automatically predict the quality of machine translation (MT) output without reference sentences. Despite its high utility in the real world, there remain several limitations concerning manual QE data creation: inevitably incurred non-trivial costs due to the need for translation experts, and issues with data scaling and language expansion. To tackle these limitations, we present QUAK, a Korean-English synthetic QE dataset generated in a fully automatic manner. This consists of three sub-QUAK datasets QUAK-M, QUAK-P, and QUAK-H, produced through three strategies that are relatively free from language constraints. Since each strategy requires no human effort, which facilitates scalability, we scale our data up to 1.58M for QUAK-P, H and 6.58M for QUAK-M. As an experiment, we quantitatively analyze word-level QE results in various ways while performing statistical analysis. Moreover, we show that datasets scaled in an efficient way also contribute to performance improvements by observing meaningful performance gains in QUAK-M, P when adding data up to 1.58M.
Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Gyeongmin Kim, Jungseob Lee, Heuiseok Lim
COLING4
2022 Empirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editing
abstract
Automatic post-editing (APE) refers to a research field that aims to automatically correct errors included in the translation sentences derived by the machine translation system. This study has several limitations, considering the data acquisition, because there is no official dataset for most language pairs. Moreover, the amount of data is restricted even for language pairs in which official data has been released, such as WMT. To solve this problem and promote universal APE research regardless of APE data existence, this study proposes a method for automatically generating APE data based on a noising scheme from a parallel corpus. Particularly, we propose a human mimicking errors-based noising scheme that considers a practical correction process at the human level. We propose a precise inspection to attain high performance, and we derived the optimal noising schemes that show substantial effectiveness. Through these, we also demonstrate that depending on the type of noise, the noising scheme-based APE data generation may lead to inferior performance. In addition, we propose a dynamic noise injection strategy that enables the acquisition of a robust error correction capability and demonstrated its effectiveness by comparative analysis. This study enables obtaining a high performance APE model without human-generated data and can promote universal APE research for all language pairs targeting English.
Hyeonseok Moon, Chanjun Park, Seolhwa Lee, Jaehyung Seo, Jungseob Lee, Sugyeong Eo, Heuiseok Lim
LREC4
2022 Priming Ancient Korean Neural Machine Translation
abstract
In recent years, there has been an increasing need for the restoration and translation of historical languages. In this study, we attempt to translate historical records in ancient Korean language based on neural machine translation (NMT). Inspired by priming, a cognitive science theory that two different stimuli influence each other, we propose novel priming ancient-Korean NMT (AKNMT) using bilingual subword embedding initialization with structural property awareness in the ancient documents. Finally, we obtain state-of-the-art results in the AKNMT task. To the best of our knowledge, we confirm the possibility of developing a human-centric model that incorporates the concepts of cognitive science and analyzes the result from the perspective of interference and cognitive dissonance theory for the first time.
Chanjun Park, Seolhwa Lee, Jaehyung Seo, Hyeonseok Moon, Sugyeong Eo, Heuiseok Lim
LREC3
2022 PU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, Heuiseok Lim
Knowl. Based Syst.1