Heuiseok Lim

dblp:127/4881 · DBLP profile ↗
← Back
69ranked-venue papers
3as first author
47since 2021 · last 2026
0000-0002-9269-1157ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 1 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
abstract
Although large language models (LLMs) have shown considerable progress in pragmatic language understanding, prior research has focused mainly on their comprehension of verbal behavior.Nonetheless, non-verbal behavior remains a fundamental component of human communication, especially when deliberately utilized in isolation to convey indirect meanings.In this work, we present the first systematic evaluation of LLMs' ability to infer pragmatic meaning in dialogue consisting solely of non-verbal responses.We explore three research questions: (1) Can LLMs recognize indirect intent conveyed through non-verbal responses?(2) When and how do LLMs fail to capture non-verbal intent?(3) How can we improve LLMs' ability to interpret non-verbal intent?.Through the evaluation, we observe that LLMs struggle to infer underlying meaning from non-verbal responses, with accuracy dropping by up to 60% points compared to verbal ones.Further extensive analysis reveals a behavioral pattern in LLMs' interpretations of non-verbal behavior and demonstrates that incontext learning facilitates pragmatic inference.
Sugyeong Eo, Heuiseok Lim
ACL (1)2
2026 No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
abstract
The Plain Writing Act in the United States requires government documents to be written in clear and simple language.However, existing summarization systems struggle to address diverse linguistic and cognitive barriers among general readers.We propose NRLB (No Reader Left Behind), a unified multi-agent framework for plain language summarization that simulates three representative reader groups: elementary school students, non-native speakers, and readers with attention deficits.NRLB integrates template-based planning with an iterative feedback loop guided by simulated readers and domain expert revision to address comprehension barriers such as unknown terms, missing contexts, and confusing sentences.Evaluations across multiple datasets demonstrate consistent improvements in both readability and factuality.Human evaluation further supports these findings, with annotator preference rates ranging from 55% to 76%, highlighting NRLB's ability to generate summaries that are both faithful to the source and accessible to a wide range of readers.
Jimin Jung, MyoungJin Kim, Jaehyung Seo, Heuiseok Lim
ACL (1)4
2026 Towards Scalable Lifelong Knowledge Editing with Selective Knowledge Suppression
abstract
Large language models (LLMs) require frequent knowledge updates to reflect changing facts and mitigate hallucinations.To meet this demand, lifelong knowledge editing has emerged as a continual approach to modify specific pieces of knowledge without retraining the entire model.Existing parameter-editing methods struggle with stability during sequential edits due to catastrophic forgetting.While retrieval-based approaches are proposed to alleviate this issue, their applicability remains limited across various datasets because of high training costs.To address these limitations and enhance scalability in lifelong settings, we propose LightEdit.Our framework first selects relevant knowledge from retrieved information to modify the query effectively.It then incorporates a decoding strategy to suppress the model's original knowledge probabilities, thereby enabling efficient edits based on the selected information.Extensive experiments on ZSRE, Counterfact, and RIPE benchmarks demonstrate that LightEdit outperforms existing lifelong knowledge editing methods.Furthermore, by minimizing training costs, LightEdit achieves cost-effective scalability, enabling easy adaptation to various datasets.1
Dahyun Jung, Jaewook Lee 0008, Heuiseok Lim
ACL (1)3
2026 EASE: Entity-Aware Sub-table Generation for Real-world Multi-table QA
abstract
Myunghoon Kang, Dahyun Jung, Suhyune Son, Seonmin Koo, Changwoo Chun, Daniel Rim, Haeyoung Kwon, Yuna Hur, Heuiseok Lim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Myunghoon Kang, Dahyun Jung, Suhyune Son, Seonmin Koo, Changwoo Chun, Daniel Rim, Haeyoung Kwon, Yuna Hur, Heuiseok Lim
ACL (1)9
2026 LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal
abstract
Dense retrieval in multilingual settings often searches over mixed-language collections, yet multilingual embeddings encode language identity alongside semantics.This language signal can inflate similarity for same-language pairs and crowd out relevant evidence written in other languages.We propose LANGSAE EDIT-ING, a post-hoc sparse autoencoder trained on pooled embeddings that enables controllable removal of language-identity signal directly in vector space.The method identifies languageassociated latent units using cross-language activation statistics, suppresses these units at inference time, and reconstructs embeddings in the original dimensionality, making it compatible with existing vector databases without retraining the base encoder or re-encoding raw text.Experiments across multiple languages show consistent improvements in ranking quality and cross-language coverage, with especially strong gains for script-distinct languages.The LANGSAE model and training code are publicly available.1
Jeongho Yoon, Chanjun Park, Heuiseok Lim
ACL (1)4
2026 CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training
abstract
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang, Dongsuk Oh, Heuiseok Lim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang 0002, Dongsuk Oh, Heuiseok Lim
ACL (1)6
2026 HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
abstract
Retrieval-augmented generation (RAG) for document-based open-domain question answering (ODQA) over large industrial corpora faces two core bottlenecks: routing to the correct document and combining scattered evidence.Flat text chunks and page-level images often fail to (i) identify the right document among thousands of candidates and (ii) connect multimodal evidence, such as tables and figures, within a fixed token budget.We propose HiKEY, a hierarchical tree-based multimodal retrieval framework that treats document hierarchy as a first-class retrieval signal.Rather than simply chunking text, HiKEY uses Document Hierarchical Parsing (DHP) to reconstruct a logical heterogeneous graph with explicit parent-child relations.At query time, HiKEY follows a hierarchical coarse-to-fine process: it first performs global routing with hierarchical indexes to prune the corpus, and then ranks sections with a multimodal fusion strategy that selects the most discriminative evidence.It finally builds a token-efficient evidence subgraph through hybrid structural-semantic packing.Experiments on ODQA benchmarks show that HiKEY outperforms page-and chunk-based baselines, improving retrieval recall by up to 12.9 points and end-to-end QA by up to 6.8 points.
Joongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
ACL (1)5
2026 Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation
abstract
Current LLM-based services typically require users to submit raw text regardless of its sensitivity.While intuitive, such practice introduces substantial privacy risks, as unauthorized access may expose personal, medical, or legal information.Although prior defenses strived to mitigate these risks, they often incur substantial computational overhead and degrade model performance.To overcome this privacy-efficiency trade-off, we introduce Privacy-Preserving Fine-Tuning (PPFT), a novel training pipeline that eliminates the need for transmitting raw prompt text while maintaining a favorable balance between privacy preservation and model utility for both clients and service providers.Our approach operates in two stages: first, we train a client-side encoder together with a server-side projection module and LLM, enabling the server to condition on k-pooled prompt embeddings instead of raw text; second, we fine-tune the projection module and LLM on private, domain-specific data using noise-injected embeddings, allowing effective adaptation without exposing plain text prompts and requiring access to the decoder's internal parameters.Extensive experiments on domainspecific and general benchmarks demonstrate that PPFT achieves a striking balance between privacy and utility, maintaining competitive performance with minimal degradation compared to noise-free upper bounds.
Jeongho Yoon, Chanhee Park, Yongchan Chun, Hyeonseok Moon, Heuiseok Lim
ACL (1)5
2026 Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation
Youngjoon Jang 0002, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim
SIGIR4
2026 Evaluating over-empathizing in emotional support conversations: A user-centered framework
Suhyune Son, Seonmin Koo, Evelyn Hayoon Zi, Jungsun Jang, Heuiseok Lim
Expert Syst. Appl.5
2026 SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
Jaehyung Seo, Hyeonseok Moon, Yoonna Jang, Heuiseok Lim
Knowl. Based Syst.4
2025 Cross-Lingual Optimization for Language Transfer in Large Language Models
abstract
Adapting large language models to other languages typically employs supervised finetuning (SFT) as a standard approach.However, it often suffers from an overemphasis on English performance, a phenomenon that is especially pronounced in data-constrained environments.To overcome these challenges, we propose Cross-Lingual Optimization (CLO) that efficiently transfers an English-centric LLM to a target language while preserving its English capabilities.CLO utilizes publicly available English SFT data and a translation model to enable cross-lingual transfer.We conduct experiments using five models on six languages, each possessing varying levels of resource.Our results show that CLO consistently outperforms SFT in both acquiring target language proficiency and maintaining English performance.Remarkably, in low-resource languages, CLO with only 3,200 samples surpasses SFT with 6,400 samples, demonstrating that CLO can achieve better performance with less data.Furthermore, we find that SFT is particularly sensitive to data quantity in medium and lowresource languages, whereas CLO remains robust.Our comprehensive analysis emphasizes the limitations of SFT and incorporates additional training strategies in CLO to enhance efficiency.
Jungseob Lee, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim
ACL (1)4
2025 MIGRATE: Cross-Lingual Adaptation of Domain-Specific LLMs through Code-Switching and Embedding Transfer
abstract
Large Language Models (LLMs) have rapidly advanced, with domain-specific expert models emerging to handle specialized tasks across various fields. However, the predominant focus on English-centric models demands extensive data, making it challenging to develop comparable models for middle and low-resource languages. To address this limitation, we introduce Migrate, a novel method that leverages open-source static embedding models and up to 3 million tokens of code-switching data to facilitate the seamless transfer of embeddings to target languages. Migrate enables effective cross-lingual adaptation without requiring large-scale domain-specific corpora in the target language, promoting the accessibility of expert LLMs to a diverse range of linguistic communities. Our experimental results demonstrate that Migrate significantly enhances model performance in target languages, outperforming baseline and existing cross-lingual transfer methods. This approach provides a practical and efficient solution for extending the capabilities of domain-specific expert models.
Seongtae Hong, Seungyoon Lee, Hyeonseok Moon, Heuiseok Lim
COLING4
2025 Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
abstract
A sparse Mixture-of-Experts (MoE) architecture has emerged as a highly scalable solution by conditionally activating sub-modules without a proportional increase in computational costs.However, improving expert specialization to enhance performance and generalization remains a challenge for MoE, especially in instruction tuning scenarios characterized by significant input heterogeneity.In this work, we propose the Mixture-of-Clustered-Experts (MoCE) to address this limitation through a dual-stage routing mechanism.The first stage in the mechanism performs expert group routing based on sequence-level features, while the second stage activates the top-k experts within the group at the token level.This approach enables the effective partitioning of heterogeneous inputs based on their knowledge requirements, encouraging expert group specialization while maintaining the advantages of token-level routing.We evaluate MoCE across a comprehensive set of benchmarks, demonstrating its consistent superiority over strong baselines and its enhanced generalization capabilities.Detailed analysis further highlights the robustness and effectiveness of MoCE.
Sugyeong Eo, Jung Jun Lee, Chanjun Park, Heuiseok Lim
EMNLP4
2025 Semantic Inversion, Identical Replies: Revisiting Negation Blindness in Large Language Models
abstract
Large language models (LLMs) often fail to capture semantic changes in queries due to negation, and generate incorrect responses.Negation frequently exists in the real world and is useful for understanding the opposite or absence of a statement, so it is an essential element in logical reasoning.Previous studies have explored LLMs' ability to capture negations 'separately' from their ability to properly ground knowledge for positive queries.However, this perspective is limited in that it cannot clearly distinguish whether the cause of incorrect responses is the logical incoherence caused by negations or the lack of grounding ability for the given context.To address this issue, we focus on the phenomenon of the model failing to capture semantic contradictions in negated queries despite its accurate understanding of knowledge about positive queries.We term this phenomenon negation blindness on the query.We propose a verification framework that includes task design and measurement methods to verify this issue.In detail, we establish two criteria for systematic task design-i) 'complexity' and ii) 'constrainedness'-and devise four verification tasks accordingly.Moreover, we analyze the results extensively and provide insights into problem alleviation feasibility through experiments on various approaches 1 .* Equally contributed.† Corresponding author. 1 Our code and resources can be found at https://www. github.com/jin62304/NegationBlindness.(...) in his Seattle home, listening to a thunderstorm raging outside.For a moment, he thought he heard a woman's name being blown in the wind-Ten years later, James changed his name to Jimi Hendrix and formed the band, The Experience.When they debuted at the Monterey Pop Festival in 1967, Hendrix set his guitar on fire and began a new chapter in the history of rock.He died three years later of an accidental drug overdose.Excerpt: 'Jimi: Sounds Like A Rainbow' The guitarist's story is known to many adult fans.But now, the story of young Jimi Hendrix is now told in a new children's book by author Gary Golio and illustrator Javaka Steptoe, called "Jimi Sounds Like a Rainbow: A Story of the Young Jimi Hendrix."In addition to writing children's books Gary Golio is a children's therapist.(...) Q.Who did set fire to his guitar at the Monterey Pop festival in 1967?Q.Who did not set
Seonmin Koo, Heuiseok Lim
EMNLP3
2025 TORSO: Template-Oriented Reasoning Towards General Tasks
abstract
The approaches that guide Large Language Models (LLMs) to emulate human reasoning during response generation have emerged as an effective method for enabling them to solve complex problems in a step-by-step manner, thereby achieving superior performance.However, most existing approaches using few-shot prompts to generate responses heavily depend on the provided examples, limiting the utilization of the model's inherent reasoning capabilities.Moreover, constructing task-specific few-shot prompts is often costly and may lead to inconsistencies across different tasks.In this work, we introduce Template-Oriented Reasoning (TORSO), which elicits the model to utilize internal reasoning abilities to generate proper responses across various tasks without the need for manually crafted few-shot examples.Our experimental results demonstrate that TORSO achieves strong performance on diverse LLMs benchmarks with reasonable rationales.
Minhyuk Kim, Seungyoon Lee, Heuiseok Lim
EMNLP3
2025 Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
abstract
Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands.For example, ARC is assumed to test reasoning, while HellaSwag is designed to evaluate commonsense.However, we lack a systematic way to verify if these benchmarks actually measure these labels.We introduce BENCHMARK PROFILING, a diagnostic framework that decomposes benchmark performance into ten cognitively grounded abilities.The method combines gradient-based importance scoring with targeted parameter ablation to compute an Ability Impact Score (AIS) that quantifies how much each ability contributes to a model's success on a given benchmark.Profiling three instruction-tuned models across ten widely used benchmarks yields four key findings: (i) most benchmarks draw on several abilities rather than one, (ii) datasets with similar labels rely on distinct ability mixtures, (iii) code-generation benchmarks reward broad, multi-skill improvement and thus show only modest gains from narrow domain-specific fine-tuning, and (iv) abilities irrelevant to the task could negatively affect performance.BENCHMARK PROFILING therefore explains why performance gains do not always translate into user-perceived competence and offers a transparent tool for benchmark audit and model interpretability.The code is available on https://github.com/ junkim100/Benchmark-Profiling Leonard Bereska and Efstratios Gavves.
Gyuho Shim, Yongchan Chun, Minhyuk Kim, Chanjun Park, Heuiseok Lim
EMNLP6
2025 Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
abstract
Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation.This progress highlights the need for challenging benchmarks that provide objective verification.In this paper, we introduce MCBench, a benchmark designed to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions.Unlike prior benchmarks that depend on subjective judgments or general reasoning, MCBench offers an objective, deterministic and codeverifiable evaluation.This setup allows us to systematically test whether LLMs can maintain accurate step-by-step execution, including instruction adherence, numerical computation, and long-range consistency in handling intermediate results.To ensure objective evaluation of these abilities, we provide a parallel reference code that can evaluate the accuracy of LLM output.We provide three evaluative metrics and three benchmark variants designed to measure the detailed instruction understanding capability of LLMs.Our analyses show that MCBench serves as an effective and objective tool for evaluating the capabilities of cuttingedge LLMs.
Hyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok Lim
EMNLP4
2025 The Impact of Negated Text on Hallucination with Large Language Models
abstract
Recent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing.However, the impact of negated text on hallucination with LLMs remains largely unexplored.In this paper, we set three important yet unanswered research questions and aim to address them.To derive the answers, we investigate whether LLMs can recognize contextual shifts caused by negation and still reliably distinguish hallucinations comparable to affirmative cases.We also design the NegHalu dataset by reconstructing existing hallucination detection datasets with negated expressions.Our experiments demonstrate that LLMs struggle to detect hallucinations in negated text effectively, often producing logically inconsistent or unfaithful judgments.Moreover, we trace the internal state of LLMs as they process negated inputs at the token level and reveal the challenges of mitigating their unintended effects.
Jaehyung Seo, Hyeonseok Moon, Heuiseok Lim
EMNLP3
2025 MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
abstract
RAG-based QA has emerged as a powerful method for processing long industrial documents.However, conventional text chunking approaches often neglect the complex structures of long industrial documents, causing information loss and reduced answer quality.To address this, we introduce MultiDocFusion, a multimodal chunking pipeline that integrates: (i) detection of document regions using visionbased document parsing, (ii) text extraction from these regions via OCR, (iii) reconstruction of document structure into a hierarchical tree using large language model (LLM)based document section hierarchical parsing (DSHP-LLM), and (iv) construction of hierarchical chunks through DFS-based Grouping.Extensive experiments across industrial benchmarks demonstrate that MultiDocFusion improves retrieval precision by 8-15% and ANLS QA scores by 2-3% compared to baselines, emphasizing the critical role of explicitly leveraging document hierarchy for multimodal document-based QA.These significant performance gains underscore the necessity of structure-aware chunking in enhancing the fidelity of RAG-based QA systems.
Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
EMNLP5
2025 K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language Models
abstract
Recent researchers and companies have been developing large language models (LLMs) specifically designed for particular purposes and have achieved significant advancements in various natural language processing tasks. However, LLMs are still prone to generating hallucinations—results that are unfaithful or inconsistent with the given input. As a result, the need for datasets to evaluate and demonstrate the hallucination detection capabilities of LLMs is increasingly recognized. Nonetheless, the Korean NLP community lacks publicly available benchmark datasets demonstrating the faithfulness of knowledge-based information. Furthermore, the few existing datasets that evaluate hallucination are limited in their access to the entire dataset, restricting detailed analysis beyond simple scoring, and are based on translated English knowledge. To address these challenges, we introduce K-HALU, a Korean benchmark designed to evaluate LLMs' hallucination detection in Korean. This benchmark contains seven domains, considering the faithfulness of statements based on knowledge documents compiled from Korean news, magazines, and books. For more strict evaluation, 40% of the dataset is structured as multiple-answer questions, requiring models to select all possible correct answers from the given options. Our empirical results show that open-source LLMs still struggle with hallucination detection in Korean knowledge, emphasizing the need for a more detailed analysis of their limitations.
Jaehyung Seo, Heuiseok Lim
ICLR2
2025 CoME: An Unlearning-based Approach to Conflict-free Model Editing
abstract
Dahyun Jung, Jaehyung Seo, Jaewook Lee, Chanjun Park, Heuiseok Lim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dahyun Jung, Jaehyung Seo, Jaewook Lee 0008, Chanjun Park, Heuiseok Lim
NAACL (Long Papers)5
2025 An analysis on language transfer of pre-trained language model with cross-lingual post-training
Suhyune Son, Chanjun Park, Jungseob Lee, Midan Shim, Chanhee Lee 0004, Yoonna Jang, Jaehyung Seo, Jungwoo Lim, Heuiseok Lim
Expert Syst. Appl.9
2024 Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean Translation
abstract
Recent machine translation (MT) systems have overcome language barriers for a wide range of users, yet they still carry the risk of critical meaning deviation. Critical error detection (CED) is a task that identifies an inherent risk of catastrophic meaning distortions in the machine translation output. With the importance of reflecting cultural elements in detecting critical errors, we introduce the culture-aware “Politeness” type in detecting English-Korean critical translation errors. Besides, we facilitate two tasks by providing multiclass labels: critical error detection and critical error type classification (CETC). Empirical evaluations reveal that our introduced data augmentation approach using a newly presented perturber significantly outperforms existing baselines in both tasks. Further analysis highlights the significance of multiclass labeling by demonstrating its superior effectiveness compared to binary labels.
Sugyeong Eo, Jungwoo Lim, Chanjun Park, Dahyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
LREC/COLING8
2024 Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in Korean
abstract
Counter-narrative generation, i.e., the generation of fact-based responses to hate speech with the aim of correcting discriminatory beliefs, has been demonstrated to be an effective method to combat hate speech. However, its effectiveness is limited by the resource-intensive nature of dataset construction processes and only focuses on the primary language. To alleviate this problem, we propose a Korean Hate Speech Counter Punch (KHSCP), a cost-effective counter-narrative generation method in the Korean language. To this end, we release the first counter-narrative generation dataset in Korean and pose two research questions. Under the questions, we propose an effective augmentation method and investigate the reasonability of a large language model to overcome data scarcity in low-resource environments by leveraging existing resources. In this regard, we conduct several experiments to verify the effectiveness of the proposed method. Our results reveal that applying pre-existing resources can improve the generation performance by a significant margin. Through deep analysis on these experiments, this work proposes the possibility of overcoming the challenges of generating counter-narratives in low-resource environments.
Seungyoon Lee, Chanjun Park, Dahyun Jung, Hyeonseok Moon, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
LREC/COLING7
2024 Ask, Assess, and Refine: Rectifying Factual Consistency and Hallucination in LLMs with Metric-Guided Feedback Learning
abstract
Recent advancements in Large Language Models (LLMs) have heralded unprecedented capabilities in information-seeking and text generation, as evidenced by applications like Bing Chat and perplexity.ai.Despite these strides, challenges on hallucination and factual inconsistency continue to impede their wider real-world adoption.Contemporary methods, including retrieval-augmented LLMs and feedback-based learning, serve as alternatives to mitigate these challenges.However, challenges remain, particularly regarding referencing erroneous evidence (citation errors) and generating information not present in the evidence (hallucination).In this paper, we introduce the A 2 R framework: Ask, Assess, and Refine.Our approach utilizes an explicit evaluation paradigm, incorporating metrics specifically tailored to assess citation errors and hallucination, aiming to address these prevalent challenges robustly.Capitalizing on these evaluations, we devise a strategy to formulate actionable natural language feedback, enabling iterative refinements that yield improved factual consistency and reduced hallucinations in responses.Our experiments on ASQA, ELI5, and QAMPARI datasets demonstrate our method's superiority in enhancing correctness, fluency, and citation quality.
Dongyub Lee, Eunhwan Park, Hodong Lee, Heuiseok Lim
EACL (1)4
2024 Revisiting Under-Represented Knowledge of Latin American Literature in Large Language Models
abstract
With the advent of large language models (LLMs), concerns about knowledge bias have recently increased. Previously, prevalent research has focused on detecting the bias of model knowledge by providing explicit social terms, such as race, gender, and age, into inputs. However, revealing the subtle and implicit bias of the model knowledge requires verification utilizing language expressed in a more implied form, such as literary works. This is because literature implicitly contains subjective filters of individuals and their living regional culture. Accordingly, this study aims to probe a research question of whether LLMs have a knowledge under-representation problem between two different regions using the same language, Spain and Spanish-speaking countries in Latin America. To this end, we design an under-representation verification task, REGion and Literary Author prediction (REGLA) and dataset based on Spanish-written literary works. Inspired by the knowledge shortcut concept from a previous study, REGLA consists of two tasks to figure out meta-information of poems, i.e., region and author. Moreover, we explore various prompting methods that can unleash the knowledge observed to be under-represented within the verification process. According to the verification and prompt engineering results, knowledge about the literary works of Latin American countries appears to be more under-represented compared to those of Spain in LLMs. It is also observed that the task decomposition prompting method effectively lets under-represented knowledge be generated.
Seonmin Koo, Heuiseok Lim
ECAI3
2024 PANDA: Persona Attributes Navigation for Detecting and Alleviating Overuse Problem in Large Language Models
abstract
In the persona-grounded dialogue (PGD) task, it is required not only to respond fluently, but also to ground the attributes according to the current conversation topic properly.However, due to their tendency to overly ground given attributes, LLMs often generate unnatural responses provoked by using attributes that deviate from the flow of the conversation or by exploiting too many attributes at once.We term this phenomenon the overuse problem of LLMs.Unfortunately, research devising precise criteria and frameworks to quantitatively verify LLMs' overuse problem is obviously insufficient.To address this issue, we propose Persona Attributes Navigation for Detecting and Alleviating the overuse problem (PANDA) framework.PANDA is the first study to quantify the persona overuse problem of LLMs by establishing clear standards of the problem and verifying various LLMs based on them.Moreover, this framework navigates us into understanding persona attributes by introducing diverse and detailed dialogue topics that consider practical conversation situations.We provide insights related to LLMs' persona attribute overuse problem through comprehensive verification and analysis with PANDA in the PGD task.Our code and resources can be found at http://github.com/jin62304/PANDA.
Seonmin Koo, Heuiseok Lim
EMNLP3
2024 Where am I? Large Language Models Wandering between Semantics and Structures in Long Contexts
abstract
As the utilization of Large Language Models (LLMs) becomes more widespread, there is a growing demand for their ability to handle more complex and longer external knowledge across various use cases.Most existing evaluations of the open-ended question answering (ODQA) task, which necessitates the use of external knowledge, focus solely on whether the model provides the correct answer.However, even when LLMs answer correctly, they often fail to provide an obvious source for their responses.Therefore, it is necessary to jointly evaluate and verify the correctness of the answers and the appropriateness of grounded evidence in complex external contexts.To address this issue, we examine the phenomenon of discrepancies in abilities across two distinct tasks-QA and evidence selection-when performed simultaneously, from the perspective of task alignment.To verify LLMs' task alignment, we introduce a verification framework and resources considering both semantic relevancy and structural diversity of the given long context knowledge.Through extensive experiments and detailed analysis, we provide insights into the task misalignment between QA and evidence selection.Our code and resources can be found at https://github.com/seonminkoo/WAI.
Seonmin Koo, Youngjoon Jang 0002, Chanjun Park, Heuiseok Lim
EMNLP5
2024 A large-scale dataset for korean document-level relation extraction from encyclopedia texts
abstract
Abstract Document-level relation extraction (RE) aims to predict the relational facts between two given entities from a document. Unlike widespread research on document-level RE in English, Korean document-level RE research is still at the very beginning due to the absence of a dataset. To accelerate the studies, we present (Toward Document-Level Relation Extraction in Korean) dataset constructed from Korean encyclopedia documents written by the domain experts. We provide detailed statistical analyses for our large-scale dataset and human evaluation results suggest the assured quality of . Also, we introduce the document-level RE model that considers the named entity-type while considering the Korean language’s properties. In the experiments, we demonstrate that our proposed model outperforms the baselines and conduct qualitative analysis.
Suhyune Son, Jungwoo Lim, Seonmin Koo, Youngsik Lim, Dongseok Hyun, Heuiseok Lim
Appl. Intell.8
2024 Adaptive Multi-Domain Dialogue State Tracking on Spoken Conversations
abstract
The main objective of the task-oriented dialogue system is to identify the intent and needs of human dialogue. Many existing studies are conducted under the setting of written dialogue, but there always exists a difficulty in coping with real-world spoken dialogues. To this end, DSTC10 challenge organizers propose the task of building robust dialogue state tracking (DST) models on spoken dialogues. With the powerful existing DST model (i.e., MinTL), this article suggests integral components for building a dialogue state tracker; 1) Data augmentation effectively enhances the capability of the model to catch the entities that exist in the evaluation dataset. 2) Levenshtein post-processing aims to prevent the distortion in model prediction caused by automatic speech recognition errors. To validate the effectiveness of our methods, we evaluate our model on DSTC10 datasets and conduct qualitative analysis by ablating each component of the model. Experimental results show that our model significantly outperforms baselines in all evaluation metrics and took 3rd place in the challenge.
Jungwoo Lim, Taesun Whang, Dongyub Lee, Heuiseok Lim
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations
abstract
Yoonna Jang, Suhyune Son, Jeongwoo Lee, Junyoung Son, Yuna Hur, Jungwoo Lim, Hyeonseok Moon, Kisu Yang, Heuiseok Lim. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yoonna Jang, Suhyune Son, Junyoung Son, Yuna Hur, Jungwoo Lim, Hyeonseok Moon, Kisu Yang, Heuiseok Lim
EMNLP9
2023 KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processing
abstract
Automatic Speech Recognition (ASR) systems are instrumental across various applications, with their performance being critically tied to user satisfaction.Conventional evaluation metrics for ASR systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.Therefore, we aim to address the limitations of the previous ASR evaluation methods by introducing the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP).KE-BAP enables comprehensive analysis of ASR systems at both speech-and text levels, thereby facilitating a more balanced assessment encompassing speech recognition accuracy and user readability.KEBAP provides 37 newly defined speech-level resources incorporating diverse noise environments and speaker characteristics categories, also presenting 13 distinct textlevel error types.This paper demonstrates detailed statistical analyses of colloquial noise categories and textual error types.Furthermore, we conduct extensive validation and analysis on commercially deployed ASR systems, providing valuable insights into their performance.As a more fine-grained and real-world-centric evaluation method, KEBAP contributes to identifying and mitigating potential weaknesses in ASR systems.* Equally contributed, ‡ Corresponding author 1 Recognition accuracy is the measure of accurately perceiving phonemes as they are externally expressed, regardless of user input quality (Liao et al., 2022).Conventional (WER, CER) 0.45 KEBAP Error types Explainability Noise Type Description Washer/dryer machine Home appliances Vacuum cleaner Difficulty in recognition due to ambient electrical appliance noise.Motorcycle Siren Individual transportation Honk Difficulty in recognition due to surrounding individual transportation noise.Road side Street Crowd Difficulty in recognition due to the surrounding street noise.Conversation Cafe/restaurant Non-conversation Challenges in perception due to the noise in cafes/restaurants.Traditional market Market/shopping mall Shopping mall Difficulties in perception caused by the noise in markets/shopping malls.Subway platform Inside the subway Inside the train (STR/KTX) Public transportation Inside the bus Difficulty in recognition due to surrounding public transportation noise.Train terminal waiting room Terminal Bus terminal waiting room Challenges in perception due to the noise at terminals.Outdoor construction site Construction site Indoor construction site Difficulties in perception caused by the noise at construction sites.processing process Factory Assembly process Difficulties in perception caused by the noise in factories.Sound of rain Nature ambient Sound of the waves Challenges in perception due to natural ambient noise.Noisy environment Etc.Artificial mechanical sound In cases where external noise is present, although not falling into the aforementioned categories.
Seonmin Koo, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Hyeonseok Moon, Heuiseok Lim
EMNLP7
2023 CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients
abstract
Korean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction.The complexity of morphological variations allows for diverse sentence forms based on the syntactic-semantic integration of functional morphemes (i.e., affixes) to lexical morphemes (i.e., roots).With this in mind, we propose a method -CHEF, replicating the morphological transformations inherent in sentences based on lexical and functional morpheme combinations through generative data augmentation.CHEF operates using a morpheme blender and a label discriminator, thereby enhancing the diversity of Korean sentence forms by capturing the properties of agglutination while maintaining label consistency.We conduct experiments on Korean multiple classification datasets, improving model performance in full-and few-shot settings.Our proposed method boosts performance beyond the preceding data augmentation methods without incurring external data usage.We demonstrate that our approach achieves comparable results yielded by augmentation techniques that use large language models (LLMs).
Jaehyung Seo, Hyeonseok Moon, Jaewook Lee 0008, Sugyeong Eo, Chanjun Park, Heuiseok Lim
EMNLP6
2023 Informative Evidence-guided Prompt-based Fine-tuning for English-Korean Critical Error Detection
abstract
DaHyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Dahyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
IJCNLP (1)6
2023 Doubts on the reliability of parallel corpus filtering
Hyeonseok Moon, Chanjun Park, Seonmin Koo, Jungseob Lee, Jaehyung Seo, Sugyeong Eo, Yoonna Jang, Hyunjoong Kim, Hyoung-gyu Lee, Heuiseok Lim
Expert Syst. Appl.11
2022 Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge
abstract
Humans usually have conversations by making use of prior knowledge about a topic and background information of the people whom they are talking to. However, existing conversational agents and datasets do not consider such comprehensive information, and thus they have a limitation in generating the utterances where the knowledge and persona are fused properly. To address this issue, we introduce a call For Customized conversation (FoCus) dataset where the customized answers are built with the user's persona and Wikipedia knowledge. To evaluate the abilities to make informative and customized utterances of pre-trained language models, we utilize BART and GPT-2 as well as transformer-based models. We assess their generation abilities with automatic scores and conduct human evaluations for qualitative results. We examine whether the model reflects adequate persona and knowledge with our proposed two sub-tasks, persona grounding (PG) and knowledge grounding (KG). Moreover, we show that the utterances of our data are constructed with the proper knowledge and persona through grounding quality assessment.
Yoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh, Suhyune Son, Yeonsoo Lee, Dong-Hoon Shin, Seungryong Kim, Heuiseok Lim
AAAI9
2022 QUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine Translation
abstract
With the recent advance in neural machine translation demonstrating its importance, research on quality estimation (QE) has been steadily progressing. QE aims to automatically predict the quality of machine translation (MT) output without reference sentences. Despite its high utility in the real world, there remain several limitations concerning manual QE data creation: inevitably incurred non-trivial costs due to the need for translation experts, and issues with data scaling and language expansion. To tackle these limitations, we present QUAK, a Korean-English synthetic QE dataset generated in a fully automatic manner. This consists of three sub-QUAK datasets QUAK-M, QUAK-P, and QUAK-H, produced through three strategies that are relatively free from language constraints. Since each strategy requires no human effort, which facilitates scalability, we scale our data up to 1.58M for QUAK-P, H and 6.58M for QUAK-M. As an experiment, we quantitatively analyze word-level QE results in various ways while performing statistical analysis. Moreover, we show that datasets scaled in an efficient way also contribute to performance improvements by observing meaningful performance gains in QUAK-M, P when adding data up to 1.58M.
Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Gyeongmin Kim, Jungseob Lee, Heuiseok Lim
COLING7
2022 KoCHET: A Korean Cultural Heritage Corpus for Entity-related Tasks
abstract
As digitized traditional cultural heritage documents have rapidly increased, resulting in an increased need for preservation and management, practical recognition of entities and typification of their classes has become essential. To achieve this, we propose KoCHET - a Korean cultural heritage corpus for the typical entity-related tasks, i.e., named entity recognition (NER), relation extraction (RE), and entity typing (ET). Advised by cultural heritage experts based on the data construction guidelines of government-affiliated organizations, KoCHET consists of respectively 112,362, 38,765, 113,198 examples for NER, RE, and ET tasks, covering all entity types related to Korean cultural heritage. Moreover, unlike the existing public corpora, modified redistribution can be allowed both domestic and foreign researchers. Our experimental results make the practical usability of KoCHET more valuable in terms of cultural heritage. We also provide practical insights of KoCHET in terms of statistical and linguistic analysis. Our corpus is freely available at https://github.com/Gyeongmin47/KoCHET.
Gyeongmin Kim, Junyoung Son, Heuiseok Lim
COLING4
2022 Don't Judge a Language Model by Its Last Layer: Contrastive Learning with Layer-Wise Attention Pooling
abstract
Recent pre-trained language models (PLMs) achieved great success on many natural language processing tasks through learning linguistic features and contextualized sentence representation. Since attributes captured in stacked layers of PLMs are not clearly identified, straightforward approaches such as embedding the last layer are commonly preferred to derive sentence representations from PLMs. This paper introduces the attention-based pooling strategy, which enables the model to preserve layer-wise signals captured in each layer and learn digested linguistic features for downstream tasks. The contrastive learning objective can adapt the layer-wise attention pooling to both unsupervised and supervised manners. It results in regularizing the anisotropic space of pre-trained embeddings and being more uniform. We evaluate our model on standard semantic textual similarity (STS) and semantic search tasks. As a result, our method improved the performance of the base contrastive learned BERT_{base} and variants.
Dongsuk Oh, Yejin Kim 0005, Hodong Lee, H. Howie Huang, Heuiseok Lim
COLING5
2022 GRASP: Guiding Model with RelAtional Semantics Using Prompt for Dialogue Relation Extraction
abstract
The dialogue-based relation extraction (DialogRE) task aims to predict the relations between argument pairs that appear in dialogue. Most previous studies utilize fine-tuning pre-trained language models (PLMs) only with extensive features to supplement the low information density of the dialogue by multiple speakers. To effectively exploit inherent knowledge of PLMs without extra layers and consider scattered semantic cues on the relation between the arguments, we propose a Guiding model with RelAtional Semantics using Prompt (GRASP). We adopt a prompt-based fine-tuning approach and capture relational semantic clues of a given dialogue with 1) an argument-aware prompt marker strategy and 2) the relational clue detection task. In the experiments, GRASP achieves state-of-the-art performance in terms of both F1 and F1c scores on a DialogRE dataset even though our method only leverages PLMs without adding any extra layers.
Junyoung Son, Jungwoo Lim, Heuiseok Lim
COLING4
2022 Empirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editing
abstract
Automatic post-editing (APE) refers to a research field that aims to automatically correct errors included in the translation sentences derived by the machine translation system. This study has several limitations, considering the data acquisition, because there is no official dataset for most language pairs. Moreover, the amount of data is restricted even for language pairs in which official data has been released, such as WMT. To solve this problem and promote universal APE research regardless of APE data existence, this study proposes a method for automatically generating APE data based on a noising scheme from a parallel corpus. Particularly, we propose a human mimicking errors-based noising scheme that considers a practical correction process at the human level. We propose a precise inspection to attain high performance, and we derived the optimal noising schemes that show substantial effectiveness. Through these, we also demonstrate that depending on the type of noise, the noising scheme-based APE data generation may lead to inferior performance. In addition, we propose a dynamic noise injection strategy that enables the acquisition of a robust error correction capability and demonstrated its effectiveness by comparative analysis. This study enables obtaining a high performance APE model without human-generated data and can promote universal APE research for all language pairs targeting English.
Hyeonseok Moon, Chanjun Park, Seolhwa Lee, Jaehyung Seo, Jungseob Lee, Sugyeong Eo, Heuiseok Lim
LREC7
2022 FreeTalky: Don't Be Afraid! Conversations Made Easier by a Humanoid Robot using Persona-based Dialogue
abstract
We propose a deep learning-based foreign language learning platform, named FreeTalky, for people who experience anxiety dealing with foreign languages, by employing a humanoid robot NAO and various deep learning models. A persona-based dialogue system that is embedded in NAO provides an interesting and consistent multi-turn dialogue for users. Also, an grammar error correction system promotes improvement in grammar skills of the users. Thus, our system enables personalized learning based on persona dialogue and facilitates grammar learning of a user using grammar error feedback. Furthermore, we verified whether FreeTalky provides practical help in alleviating xenoglossophobia by replacing the real human in the conversation with a NAO robot, through human evaluation.
Chanjun Park, Yoonna Jang, Seolhwa Lee, Heuiseok Lim
LREC5
2022 Priming Ancient Korean Neural Machine Translation
abstract
In recent years, there has been an increasing need for the restoration and translation of historical languages. In this study, we attempt to translate historical records in ancient Korean language based on neural machine translation (NMT). Inspired by priming, a cognitive science theory that two different stimuli influence each other, we propose novel priming ancient-Korean NMT (AKNMT) using bilingual subword embedding initialization with structural property awareness in the ancient documents. Finally, we obtain state-of-the-art results in the AKNMT task. To the best of our knowledge, we confirm the possibility of developing a human-centric model that incorporates the concepts of cognitive science and analyzes the result from the perspective of interference and cognitive dissonance theory for the first time.
Chanjun Park, Seolhwa Lee, Jaehyung Seo, Hyeonseok Moon, Sugyeong Eo, Heuiseok Lim
LREC6
2022 PU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, Heuiseok Lim
Knowl. Based Syst.8
2022 Unifying user preference and item knowledge-based similarity models for top-N recommendation
YeongWook Yang, Jaechoon Jo, Heuiseok Lim
Pers. Ubiquitous Comput.3
2021 Neural spelling correction: translating incorrect sentences to correct sentences for multimedia
Chanjun Park, Kuekyeng Kim, YeongWook Yang, Minho Kang, Heuiseok Lim
Multim. Tools Appl.5
2020 I Know What You Asked: Graph Path Learning using AMR for Commonsense Reasoning
abstract
CommonsenseQA is a task in which a correct answer is predicted through commonsense reasoning with pre-defined knowledge.Most previous works have aimed to improve the performance with distributed representation without considering the process of predicting the answer from the semantic representation of the question.To shed light upon the semantic interpretation of the question, we propose an AMR-ConceptNet-Pruned (ACP) graph.The ACP graph is pruned from a full integrated graph encompassing Abstract Meaning Representation (AMR) graph generated from input questions and an external commonsense knowledge graph, ConceptNet (CN).Then the ACP graph is exploited to interpret the reasoning path as well as to predict the correct answer on the CommonsenseQA task.This paper presents the manner in which the commonsense reasoning process can be interpreted with the relations and concepts provided by the ACP graph.Moreover, ACP-based models are shown to outperform the baselines.
Jungwoo Lim, Dongsuk Oh, Yoonna Jang, Kisu Yang, Heuiseok Lim
COLING5
2020 An Effective Domain Adaptive Post-Training Method for BERT in Response Selection
abstract
We focus on multi-turn response selection in a retrieval-based dialog system. In this paper, we utilize the powerful pre-trained language model Bi-directional Encoder Representations from Transformer (BERT) for a multi-turn dialog system and propose a highly effective post-training method on domain-specific corpus. Although BERT is easily adopted to various NLP tasks and outperforms previous baselines of each task, it still has limitations if a task corpus is too focused on a certain domain. Post-training on domain-specific corpus (e.g., Ubuntu Corpus) helps the model to train contextualized representations and words that do not appear in general corpus (e.g., English Wikipedia). Experimental results show that our approach achieves new state-of-the-art on two response selection benchmarks (i.e., Ubuntu Corpus V1, Advising Corpus) performance improvement by 5.9% and 6% on R@1.
Taesun Whang, Dongyub Lee, Chanhee Lee 0004, Kisu Yang, Dongsuk Oh, Heuiseok Lim
INTERSPEECH6
2020 Predicting course achievement of university students based on their procrastination behaviour on Moodle
YeongWook Yang, Danial Hooshyar, Margus Pedaste, Minhong Wang 0001, Yueh-Min Huang, Heuiseok Lim
Soft Comput.6
2019 A Comparative Analysis of Emotional Words for Learning Effectiveness in Online Education
Jaechoon Jo, YeongWook Yang, Gyeongmin Kim, Heuiseok Lim
EDM4
2019 GPS: Factorized group preference-based similarity models for sparse sequential recommendation
YeongWook Yang, Danial Hooshyar, Heuiseok Lim
Inf. Sci.3
2018 Character-Level Feature Extraction with Densely Connected Networks
abstract
Generating character-level features is an important step for achieving good results in various natural language processing tasks. To alleviate the need for human labor in generating hand-crafted features, methods that utilize neural architectures such as Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) to automatically extract such features have been proposed and have shown great results. However, CNN generates position-independent features, and RNN is slow since it needs to process the characters sequentially. In this paper, we propose a novel method of using a densely connected network to automatically extract character-level features. The proposed method does not require any language or task specific assumptions, and shows robustness and effectiveness while being faster than CNN- or RNN-based methods. Evaluating this method on three sequence labeling tasks - slot tagging, Part-of-Speech (POS) tagging, and Named-Entity Recognition (NER) - we obtain state-of-the-art performance with a 96.62 F1-score and 97.73% accuracy on slot tagging and POS tagging, respectively, and comparable performance to the state-of-the-art 91.13 F1-score on NER.
Chanhee Lee 0004, Young-Bum Kim, Dongyub Lee, Heuiseok Lim
COLING4
2018 Rich Character-Level Information for Korean Morphological Analysis and Part-of-Speech Tagging
abstract
Due to the fact that Korean is a highly agglutinative, character-rich language, previous work on Korean morphological analysis typically employs the use of sub-character features known as graphemes or otherwise utilizes comprehensive prior linguistic knowledge (i.e., a dictionary of known morphological transformation forms, or actions). These models have been created with the assumption that character-level, dictionary-less morphological analysis was intractable due to the number of actions required. We present, in this study, a multi-stage action-based model that can perform morphological transformation and part-of-speech tagging using arbitrary units of input and apply it to the case of character-level Korean morphological analysis. Among models that do not employ prior linguistic knowledge, we achieve state-of-the-art word and sentence-level tagging accuracy with the Sejong Korean corpus using our proposed data-driven Bi-LSTM model.
Andrew Matteson, Chanhee Lee 0004, Young-Bum Kim, Heuiseok Lim
COLING4
2017 A Study of Keywords Based on the Word Frequency Effect Theory in Video Lectures of Software Engineering Education for Detecting Mind
abstract
The increased popularity of Massive Open Online Courses (MOOC) and e-learning has constantly increased video-based online education platforms. There are also many video lectures for software engineering education in online education platforms. Although online lectures have many advantages, there are also limitations. We performed a verification research to see if high frequency words can detect mind wandering to resolve existing limitations. In this verification study, experiments to identify whether high frequency words can represent the software engineering video lecture, the minimum number of words needed to detect mind wandering, and whether mind wandering detection standards should be changed according to the length of the video lecture. The results of this study confirmed that mind wandering can be detected through high frequency words and they can be used as an important feature in various learning analysis investigations to resolve existing limitations of online education.
Jaechoon Jo, Heuiseok Lim
CSEE&T2
2016 Comparing Programming Language Comprehension between Novice and Expert Programmers Using EEG Analysis
abstract
For programming language comprehension, high cognitive skills (e.g., reading, writing, working memory, etc.) and information processing are required. However, there are few papers that approach this from a neuroscientific perspective. In this paper, we examine program comprehension neuroscientifically and also observe the differences between novice and expert programmers. We designed an EEG (electroencephalogram) experiment and observed 18 participants during a series of program comprehension tasks. We found clear differences in program comprehension ability between novice and expert programmers. Experts exhibited higher brainwave activation than novices in electrodes F3 and P8. These results indicate that experts have outstanding program comprehension-associated abilities such as digit encoding, coarse coding, short-term memory, and subsequent memory effect. Our findings can serve as a foundation for future research in this pioneering field.
Seolhwa Lee, Andrew Matteson, Danial Hooshyar, SongHyun Kim, JaeBum Jung, GiChun Nam, Heuiseok Lim
BIBE7
2016 How to Judge Learning on Online Learning: Minimum Learning Judgment System
Jaechoon Jo, Heuiseok Lim
EDM2
2016 Mining students activities from a computer supported collaborative learning system based on peer to peer network
Hyesung Ji, Kinam Park, Jaechoon Jo, Heuiseok Lim
Peer-to-Peer Netw. Appl.4
2015 The biometric signature delegation scheme to balance the load of digital signing in hybrid P2P networks
Sung-Hyun Yun, Heuiseok Lim, Kyung-Yong Chung
Peer-to-Peer Netw. Appl.2
2013 Real-time vehicle tracking mechanism with license plate recognition from road images
Jae-Khun Chang, Seungtaek Ryoo, Heuiseok Lim
J. Supercomput.3
2012 Automatic extraction of user's search intention from web search logs
Kinam Park, Hyesung Jee, Taemin Lee, Soon Young Jung, Heuiseok Lim
Multim. Tools Appl.5
2010 Phonological Recoding in the Second Language Processing
Chang H. Lee, Kyungill Kim, Heuiseok Lim
ICCSA (4)3
2010 A Personalized CALL System Considering Users Cognitive Abilities
Saebyeok Lee, WonGye Lee, Hyeoncheol Kim, Soon Young Jung, Heuiseok Lim
ICCSA (4)5
2010 An MLP-based feature subset selection for HIV-1 protease cleavage site analysis
Gilhan Kim, Yeonjoo Kim, Heuiseok Lim, Hyeoncheol Kim
Artif. Intell. Medicine3
2006 A Computational Korean Lexical Access Model Using Artificial Neural Network
Heuiseok Lim, Kichun Nam, Kinam Park
ICIC (3)1
2006 Word Frequency Effect and Word Similarity Effect in Korean Lexical Decision Task and Their Computational Model
YouAn Kwon, Kinam Park, Heuiseok Lim, Kichun Nam, Soon Young Jung
ICONIP (3)3
2006 Mental Representation and Processing Involved in Comprehending Korean Regular and Irregular Verb Eojeols: An fMRI and Reaction Time Study
Hyungwook Yim, Changsu Park, Heuiseok Lim, Kichun Nam
ICONIP (1)3
2005 A Computational Model of Korean Mental Lexicon
Heuiseok Lim, Kichun Nam, Yumi Hwang
ICCSA (1)1
2004 Automated Classification of Industry and Occupation Codes Using Document Classification Method
Heuiseok Lim, Hyeoncheol Kim
ICONIP1