Seonmin Koo

dblp:324/3476 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0007-8575-2306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EASE: Entity-Aware Sub-table Generation for Real-world Multi-table QA
abstract
Myunghoon Kang, Dahyun Jung, Suhyune Son, Seonmin Koo, Changwoo Chun, Daniel Rim, Haeyoung Kwon, Yuna Hur, Heuiseok Lim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Myunghoon Kang, Dahyun Jung, Suhyune Son, Seonmin Koo, Changwoo Chun, Daniel Rim, Haeyoung Kwon, Yuna Hur, Heuiseok Lim
ACL (1)4
2026 Evaluating over-empathizing in emotional support conversations: A user-centered framework
Suhyune Son, Seonmin Koo, Evelyn Hayoon Zi, Jungsun Jang, Heuiseok Lim
Expert Syst. Appl.2
2025 Semantic Inversion, Identical Replies: Revisiting Negation Blindness in Large Language Models
abstract
Large language models (LLMs) often fail to capture semantic changes in queries due to negation, and generate incorrect responses.Negation frequently exists in the real world and is useful for understanding the opposite or absence of a statement, so it is an essential element in logical reasoning.Previous studies have explored LLMs' ability to capture negations 'separately' from their ability to properly ground knowledge for positive queries.However, this perspective is limited in that it cannot clearly distinguish whether the cause of incorrect responses is the logical incoherence caused by negations or the lack of grounding ability for the given context.To address this issue, we focus on the phenomenon of the model failing to capture semantic contradictions in negated queries despite its accurate understanding of knowledge about positive queries.We term this phenomenon negation blindness on the query.We propose a verification framework that includes task design and measurement methods to verify this issue.In detail, we establish two criteria for systematic task design-i) 'complexity' and ii) 'constrainedness'-and devise four verification tasks accordingly.Moreover, we analyze the results extensively and provide insights into problem alleviation feasibility through experiments on various approaches 1 .* Equally contributed.† Corresponding author. 1 Our code and resources can be found at https://www. github.com/jin62304/NegationBlindness.(...) in his Seattle home, listening to a thunderstorm raging outside.For a moment, he thought he heard a woman's name being blown in the wind-Ten years later, James changed his name to Jimi Hendrix and formed the band, The Experience.When they debuted at the Monterey Pop Festival in 1967, Hendrix set his guitar on fire and began a new chapter in the history of rock.He died three years later of an accidental drug overdose.Excerpt: 'Jimi: Sounds Like A Rainbow' The guitarist's story is known to many adult fans.But now, the story of young Jimi Hendrix is now told in a new children's book by author Gary Golio and illustrator Javaka Steptoe, called "Jimi Sounds Like a Rainbow: A Story of the Young Jimi Hendrix."In addition to writing children's books Gary Golio is a children's therapist.(...) Q.Who did set fire to his guitar at the Monterey Pop festival in 1967?Q.Who did not set
Seonmin Koo, Heuiseok Lim
EMNLP2
2024 Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean Translation
abstract
Recent machine translation (MT) systems have overcome language barriers for a wide range of users, yet they still carry the risk of critical meaning deviation. Critical error detection (CED) is a task that identifies an inherent risk of catastrophic meaning distortions in the machine translation output. With the importance of reflecting cultural elements in detecting critical errors, we introduce the culture-aware “Politeness” type in detecting English-Korean critical translation errors. Besides, we facilitate two tasks by providing multiclass labels: critical error detection and critical error type classification (CETC). Empirical evaluations reveal that our introduced data augmentation approach using a newly presented perturber significantly outperforms existing baselines in both tasks. Further analysis highlights the significance of multiclass labeling by demonstrating its superior effectiveness compared to binary labels.
Sugyeong Eo, Jungwoo Lim, Chanjun Park, Dahyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
LREC/COLING5
2024 Revisiting Under-Represented Knowledge of Latin American Literature in Large Language Models
abstract
With the advent of large language models (LLMs), concerns about knowledge bias have recently increased. Previously, prevalent research has focused on detecting the bias of model knowledge by providing explicit social terms, such as race, gender, and age, into inputs. However, revealing the subtle and implicit bias of the model knowledge requires verification utilizing language expressed in a more implied form, such as literary works. This is because literature implicitly contains subjective filters of individuals and their living regional culture. Accordingly, this study aims to probe a research question of whether LLMs have a knowledge under-representation problem between two different regions using the same language, Spain and Spanish-speaking countries in Latin America. To this end, we design an under-representation verification task, REGion and Literary Author prediction (REGLA) and dataset based on Spanish-written literary works. Inspired by the knowledge shortcut concept from a previous study, REGLA consists of two tasks to figure out meta-information of poems, i.e., region and author. Moreover, we explore various prompting methods that can unleash the knowledge observed to be under-represented within the verification process. According to the verification and prompt engineering results, knowledge about the literary works of Latin American countries appears to be more under-represented compared to those of Spain in LLMs. It is also observed that the task decomposition prompting method effectively lets under-represented knowledge be generated.
Seonmin Koo, Heuiseok Lim
ECAI2
2024 PANDA: Persona Attributes Navigation for Detecting and Alleviating Overuse Problem in Large Language Models
abstract
In the persona-grounded dialogue (PGD) task, it is required not only to respond fluently, but also to ground the attributes according to the current conversation topic properly.However, due to their tendency to overly ground given attributes, LLMs often generate unnatural responses provoked by using attributes that deviate from the flow of the conversation or by exploiting too many attributes at once.We term this phenomenon the overuse problem of LLMs.Unfortunately, research devising precise criteria and frameworks to quantitatively verify LLMs' overuse problem is obviously insufficient.To address this issue, we propose Persona Attributes Navigation for Detecting and Alleviating the overuse problem (PANDA) framework.PANDA is the first study to quantify the persona overuse problem of LLMs by establishing clear standards of the problem and verifying various LLMs based on them.Moreover, this framework navigates us into understanding persona attributes by introducing diverse and detailed dialogue topics that consider practical conversation situations.We provide insights related to LLMs' persona attribute overuse problem through comprehensive verification and analysis with PANDA in the PGD task.Our code and resources can be found at http://github.com/jin62304/PANDA.
Seonmin Koo, Heuiseok Lim
EMNLP2
2024 Where am I? Large Language Models Wandering between Semantics and Structures in Long Contexts
abstract
As the utilization of Large Language Models (LLMs) becomes more widespread, there is a growing demand for their ability to handle more complex and longer external knowledge across various use cases.Most existing evaluations of the open-ended question answering (ODQA) task, which necessitates the use of external knowledge, focus solely on whether the model provides the correct answer.However, even when LLMs answer correctly, they often fail to provide an obvious source for their responses.Therefore, it is necessary to jointly evaluate and verify the correctness of the answers and the appropriateness of grounded evidence in complex external contexts.To address this issue, we examine the phenomenon of discrepancies in abilities across two distinct tasks-QA and evidence selection-when performed simultaneously, from the perspective of task alignment.To verify LLMs' task alignment, we introduce a verification framework and resources considering both semantic relevancy and structural diversity of the given long context knowledge.Through extensive experiments and detailed analysis, we provide insights into the task misalignment between QA and evidence selection.Our code and resources can be found at https://github.com/seonminkoo/WAI.
Seonmin Koo, Youngjoon Jang 0002, Chanjun Park, Heuiseok Lim
EMNLP1
2024 A large-scale dataset for korean document-level relation extraction from encyclopedia texts
abstract
Abstract Document-level relation extraction (RE) aims to predict the relational facts between two given entities from a document. Unlike widespread research on document-level RE in English, Korean document-level RE research is still at the very beginning due to the absence of a dataset. To accelerate the studies, we present (Toward Document-Level Relation Extraction in Korean) dataset constructed from Korean encyclopedia documents written by the domain experts. We provide detailed statistical analyses for our large-scale dataset and human evaluation results suggest the assured quality of . Also, we introduce the document-level RE model that considers the named entity-type while considering the Korean language’s properties. In the experiments, we demonstrate that our proposed model outperforms the baselines and conduct qualitative analysis.
Suhyune Son, Jungwoo Lim, Seonmin Koo, Youngsik Lim, Dongseok Hyun, Heuiseok Lim
Appl. Intell.3
2023 KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processing
abstract
Automatic Speech Recognition (ASR) systems are instrumental across various applications, with their performance being critically tied to user satisfaction.Conventional evaluation metrics for ASR systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.Therefore, we aim to address the limitations of the previous ASR evaluation methods by introducing the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP).KE-BAP enables comprehensive analysis of ASR systems at both speech-and text levels, thereby facilitating a more balanced assessment encompassing speech recognition accuracy and user readability.KEBAP provides 37 newly defined speech-level resources incorporating diverse noise environments and speaker characteristics categories, also presenting 13 distinct textlevel error types.This paper demonstrates detailed statistical analyses of colloquial noise categories and textual error types.Furthermore, we conduct extensive validation and analysis on commercially deployed ASR systems, providing valuable insights into their performance.As a more fine-grained and real-world-centric evaluation method, KEBAP contributes to identifying and mitigating potential weaknesses in ASR systems.* Equally contributed, ‡ Corresponding author 1 Recognition accuracy is the measure of accurately perceiving phonemes as they are externally expressed, regardless of user input quality (Liao et al., 2022).Conventional (WER, CER) 0.45 KEBAP Error types Explainability Noise Type Description Washer/dryer machine Home appliances Vacuum cleaner Difficulty in recognition due to ambient electrical appliance noise.Motorcycle Siren Individual transportation Honk Difficulty in recognition due to surrounding individual transportation noise.Road side Street Crowd Difficulty in recognition due to the surrounding street noise.Conversation Cafe/restaurant Non-conversation Challenges in perception due to the noise in cafes/restaurants.Traditional market Market/shopping mall Shopping mall Difficulties in perception caused by the noise in markets/shopping malls.Subway platform Inside the subway Inside the train (STR/KTX) Public transportation Inside the bus Difficulty in recognition due to surrounding public transportation noise.Train terminal waiting room Terminal Bus terminal waiting room Challenges in perception due to the noise at terminals.Outdoor construction site Construction site Indoor construction site Difficulties in perception caused by the noise at construction sites.processing process Factory Assembly process Difficulties in perception caused by the noise in factories.Sound of rain Nature ambient Sound of the waves Challenges in perception due to natural ambient noise.Noisy environment Etc.Artificial mechanical sound In cases where external noise is present, although not falling into the aforementioned categories.
Seonmin Koo, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Hyeonseok Moon, Heuiseok Lim
EMNLP1
2023 Doubts on the reliability of parallel corpus filtering
Hyeonseok Moon, Chanjun Park, Seonmin Koo, Jungseob Lee, Jaehyung Seo, Sugyeong Eo, Yoonna Jang, Hyunjoong Kim, Hyoung-gyu Lee, Heuiseok Lim
Expert Syst. Appl.3