EDBT 2026 Demo / reviewers in the wild / expert
Hwanhee Lee
dblp:218/5402
· DBLP profile ↗
19ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-9367-9811ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian LanguagesabstractSocial media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is more informal, spontaneous, and adheres to many different sociolects, the performance of NLP models often deteriorates. One solution to this problem is to transform data to a standard variant before processing it, which is also called lexical normalization. There has been a wide variety of benchmarks and models proposed for this task. The MultiLexNorm benchmark proposed to unify these efforts, but it consists almost solely of languages from the Indo-European language family in the Latin script. Hence, we propose an extension to MultiLexNorm, which covers five Asian languages from different language families in four different scripts. We show that the previous state-of-the-art model performs worse on the new languages and propose a new architecture based on Large Language Models (LLMs), which shows more robust performance. Finally, we analyze remaining errors, revealing future directions for this task. Weerayut Buaphet, Thanh-Nhi Nguyen, Risa Kondo, Tomoyuki Kajiwara, Yumin Kim, Jimin Lee 0001, Hwanhee Lee, Holy Lovenia, Peerat Limkonchotiwat, Sarana Nutanong, Rob van der Goot |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2025 | Exploring Persona Sentiment Sensitivity in Personalized Dialogue GenerationabstractPersonalized dialogue systems have advanced considerably with the integration of user-specific personas into large language models (LLMs). However, while LLMs can effectively generate personalized responses, the influence of persona sentiment on dialogue quality remains underexplored. In this work, we conduct a large-scale analysis of dialogues generated using a range of polarized user profiles. Our experiments reveal that dialogues involving negatively polarized users tend to overemphasize persona attributes. In contrast, positively polarized profiles yield dialogues that selectively incorporate persona information, resulting in smoother interactions. Furthermore, we find that personas with weak or neutral sentiment generally produce lower-quality dialogues. Motivated by these findings, we propose a dialogue generation approach that explicitly accounts for persona polarity by combining a turn-based generation strategy with a profile ordering mechanism and sentiment-aware prompting. Our study provides new insights into the sensitivity of LLMs to persona sentiment and offers guidance for developing more robust and nuanced personalized dialogue systems. Yonghyun Jun, Hwanhee Lee |
ACL (1) | 2 |
| 2025 | Intraoperative Absolute Depth Estimation in MVD SurgeryabstractMicrovascular decompression (MVD) is a neurosurgical procedure that relieves nerve compression by repositioning or separating offending blood vessels, effectively reducing pain or spasms. Accurate localization of the compression site is crucial for optimal surgical outcomes, as it enables precise identification and decompression of the offending vessel. While horizontal anatomical relationships are easily identified in the surgical view, compressions occurring along the depth axis are more challenging to discern. In this study, we propose a method to measure accurate intraoperative distances during MVD surgery using Depth-Anything-V2. By leveraging the optical properties of standard imaging equipment in conjunction with the depth estimation model, our method computes precise, absolute distances rather than relying solely on relative measurements, achieving distance estimation errors of less than 2 mm compared to intraoperative and preoperative reference measurements. Hwanhee Lee, Jay J. Park, Ethan Htun, Bruce Changlong Xu, Sang-Hoon Cho, Vivek P. Buch |
CBMS | 2 |
| 2025 | How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR HeadsabstractDespite significant advancements in Large Vision Language Models (LVLMs), a gap remains, particularly regarding their interpretability and how they locate and interpret textual information within images.In this paper, we explore various LVLMs to identify the specific heads responsible for recognizing text from images, which we term the Optical Character Recognition Head (OCR Head).Our findings regarding these heads are as follows:(1) Less Sparse: Unlike previous retrieval heads, a large number of heads are activated to extract textual information from images.(2) Qualitatively Distinct: OCR heads possess properties that differ significantly from general retrieval heads, exhibiting low similarity in their characteristics.(3) Statically Activated: The frequency of activation for these heads closely aligns with their OCR scores.We validate our findings in downstream tasks by applying Chain-of-Thought (CoT) to both OCR and conventional retrieval heads and by masking these heads.We also demonstrate that redistributing sink-token values within the OCR heads improves performance.These insights provide a deeper understanding of the internal mechanisms LVLMs employ in processing embedded textual information in images. Ingeol Baek, Hwan Chang, Sunghyun Ryu, Hwanhee Lee |
EMNLP | 4 |
| 2025 | Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question AnsweringabstractAs Large Language Models (LLMs) are increasingly deployed in sensitive domains such as enterprise and government, ensuring that they adhere to user-defined security policies within context is critical-especially with respect to information non-disclosure.While prior LLM studies have focused on general safety and socially sensitive data, large-scale benchmarks for contextual security preservation against attacks remain lacking.To address this, we introduce a novel large-scale benchmark dataset, CoPriva, evaluating LLM adherence to contextual non-disclosure policies in question answering.Derived from realistic contexts, our dataset includes explicit policies and queries designed as direct and challenging indirect attacks seeking prohibited information.We evaluate 10 LLMs on our benchmark and reveal a significant vulnerability: many models violate user-defined policies and leak sensitive information.This failure is particularly severe against indirect attacks, highlighting a critical gap in current LLM safety alignment for sensitive applications.Our analysis reveals that while models can often identify the correct answer to a query, they struggle to incorporate policy constraints during generation.In contrast, they exhibit a partial ability to revise outputs when explicitly prompted.Our findings underscore the urgent need for more robust methods to guarantee contextual security.1 Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee |
EMNLP | 4 |
| 2025 | ToDi: Token-wise Distillation via Fine-Grained Divergence ControlabstractLarge language models (LLMs) offer impressive performance but are impractical for resource-constrained deployment due to high latency and energy consumption.Knowledge distillation (KD) addresses this by transferring knowledge from a large teacher to a smaller student model.However, conventional KD, notably approaches like Forward KL (FKL) and Reverse KL (RKL), apply uniform divergence loss across the entire vocabulary, neglecting token-level prediction discrepancies.By investigating these representative divergences via gradient analysis, we reveal that FKL boosts underestimated tokens, while RKL suppresses overestimated ones, showing their complementary roles.Based on this observation, we propose Token-wise Distillation (ToDi), a novel method that adaptively combines FKL and RKL per token using a sigmoid-based weighting function derived from the teacherstudent probability log-ratio.ToDi dynamically emphasizes the appropriate divergence for each token, enabling precise distribution alignment.We demonstrate that ToDi consistently outperforms recent distillation baselines using uniform or less granular strategies across instruction-following benchmarks.Extensive ablation studies and efficiency analysis further validate ToDi's effectiveness and practicality.1 Seongryong Jung, Suwan Yoon, DongGeon Kim, Hwanhee Lee |
EMNLP | 4 |
| 2025 | SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQLabstractText-to-SQL aims to convert natural language questions into executable SQL queries.While previous approaches, such as skeleton-masked selection, have demonstrated strong performance by retrieving similar training examples to guide large language models (LLMs), they struggle in real-world scenarios where such examples are unavailable.To overcome this limitation, we propose Self-Augmentation incontext learning with Fine-grained Example selection for Text-to-SQL (SAFE-SQL), a novel unsupervised framework that enhances SQL generation by generating and intelligently filtering self-augmented examples.SAFE-SQL leverages an LLM to generate diverse Textto-SQL examples, which are then filtered by a novel fine-grained mechanism using criteria for semantic similarity, structural alignment, and reasoning path quality to curate highquality in-context learning examples.Leveraging these carefully selected self-generated examples, SAFE-SQL significantly surpasses previous zero-shot and few-shot Text-to-SQL frameworks, achieving superior execution accuracy.Notably, our approach demonstrates substantial performance gains in challenging extra hard and unseen scenarios, where conventional methods often struggle. Jimin Lee 0001, Ingeol Baek, Byeongjeong Kim, Hyunkyung Bae, Hwanhee Lee |
EMNLP | 5 |
| 2025 | Event-Driven Storytelling with Multiple Lifelike Humans in a 3D SceneabstractIn this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human and human-scene interactions. We adapt the power of a large language model (LLM) to digest the contextual complexity within textual input and convert the task into tangible subproblems such that we can generate multi-agent behavior beyond the scale that was not considered before. Specifically, our event generator formulates the temporal progression of a dynamic scene into a sequence of small events. Each event calls for a well-defined motion involving relevant characters and objects. Next, we synthesize the motions of characters at positions sampled based on spatial guidance. We employ a high-level module to deliver scalable yet comprehensive context, translating events into relative descriptions that enable the retrieval of precise coordinates. As the first to address this problem at scale and with diversity, we offer a benchmark to assess diverse aspects of contextual reasoning. Benchmark results and user studies show that our framework effectively captures scene context with high scalability. The code and benchmark, along with result videos, are available at our project page: https://rms0329.github.io/Event-Driven-Storytelling/. Donggeun Lim, Jinseok Bae, Inwoo Hwang, Hwanhee Lee, Young Min Kim 0001 |
ICCV | 5 |
| 2025 | AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective IntelligenceabstractMinbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Minbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung |
NAACL (Long Papers) | 2 |
| 2024 | Kosmic: Korean Text Similarity Metric Reflecting Honorific DistinctionsabstractExisting English-based text similarity measurements primarily focus on the semantic dimension, neglecting the unique linguistic attributes found in languages like Korean, where honorific expressions are explicitly integrated. To address this limitation, this study proposes Kosmic, a novel Korean text-similarity metric that encompasses the semantic and tonal facets of a given text pair. For the evaluation, we introduce a novel benchmark annotated by human experts, empirically showing that Kosmic outperforms the existing method. Moreover, by leveraging Kosmic, we assess various Korean paraphrasing methods to determine which techniques are most effective in preserving semantics and tone. Yerin Hwang, Yongil Kim, Hyunkyung Bae, Jeesoo Bang, Hwanhee Lee, Kyomin Jung |
LREC/COLING | 5 |
| 2024 | KoCoSa: Korean Context-aware Sarcasm Detection DatasetabstractSarcasm is a way of verbal irony where someone says the opposite of what they mean, often to ridicule a person, situation, or idea. It is often difficult to detect sarcasm in the dialogue since detecting sarcasm should reflect the context (i.e., dialogue history). In this paper, we introduce a new dataset for the Korean dialogue sarcasm detection task, KoCoSa (Korean Context-aware Sarcasm Detection Dataset), which consists of 12.8K daily Korean dialogues and the labels for this task on the last response. To build the dataset, we propose an efficient sarcasm detection dataset generation pipeline: 1) generating new sarcastic dialogues from source dialogues with large language models, 2) automatic and manual filtering of abnormal and toxic dialogues, and 3) human annotation for the sarcasm detection task. We also provide a simple but effective baseline for the Korean sarcasm detection task trained on our dataset. Experimental results on the dataset show that our baseline system outperforms strong baselines like large language models, such as GPT-3.5, in the Korean sarcasm detection task. We show that the sarcasm detection task relies deeply on the existence of sufficient context. We will release the dataset at https://github.com/Yu-billie/KoCoSa_sarcasm_detection. Yumin Kim, Heejae Suh, Dongyeon Won, Hwanhee Lee |
LREC/COLING | 5 |
| 2024 | FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out DocumentabstractThrough the advent of pre-trained language models, there have been notable advancements in abstractive summarization systems.Simultaneously, a considerable number of novel methods for evaluating factual consistency in abstractive summarization systems has been developed.But these evaluation approaches incorporate substantial limitations, especially on refinement and interpretability.In this work, we propose highly effective and interpretable factual inconsistency detection method FIZZ (Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document) for abstractive summarization systems that is based on fine-grained atomic facts decomposition.Moreover, we align atomic facts decomposed from the summary with the source document through adaptive granularity expansion.These atomic facts represent a more fine-grained unit of information, facilitating detailed understanding and interpretability of the summary's factual inconsistency.Experimental results demonstrate that our proposed factual consistency checking system significantly outperforms existing systems.We release the code at https://github.com/plm3332/FIZZ. Joonho Yang, Seunghyun Yoon 0002, Byeongjeong Kim, Hwanhee Lee |
EMNLP | 4 |
| 2024 | IterCQR: Iterative Conversational Query Reformulation with Retrieval GuidanceabstractYunah Jang, Kang-il Lee, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yunah Jang, Kangil Lee 0001, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung |
NAACL-HLT | 4 |
| 2023 | Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesabstractTo address the data scarcity issue in Conversational question answering (ConvQA), a dialog inpainting method, which utilizes documents to generate ConvQA datasets, has been proposed.However, the original dialog inpainting model is trained solely on the dialog reconstruction task, resulting in the generation of questions with low contextual relevance due to insufficient learning of question-answer alignment.To overcome this limitation, we propose a novel framework called Dialogizer, which has the capability to automatically generate ConvQA datasets with high contextual relevance from textual sources.The framework incorporates two training tasks: question-answer matching (QAM) and topic-aware dialog generation (TDG).Moreover, re-ranking is conducted during the inference phase based on the contextual relevance of the generated questions.Using our framework, we produce four Con-vQA datasets by utilizing documents from multiple domains as the primary source.Through automatic evaluation using diverse metrics, as well as human evaluation, we validate that our proposed framework exhibits the ability to generate datasets of higher quality compared to the baseline dialog inpainting model. Yerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee, Jeesoo Bang, Kyomin Jung |
EMNLP | 4 |
| 2021 | CrossAug: A Contrastive Data Augmentation Method for Debiasing Fact Verification ModelsabstractFact verification datasets are typically constructed using crowdsourcing techniques due to the lack of text sources with veracity labels. However, the crowdsourcing process often produces undesired biases in data that cause models to learn spurious patterns. In this paper, we propose CrossAug, a contrastive data augmentation method for debiasing fact verification models. Specifically, we employ a two-stage augmentation pipeline to generate new claims and evidences from existing samples. The generated samples are then paired cross-wise with the original pair, forming contrastive samples that facilitate the model to rely less on spurious patterns and learn more robust representations. Experimental results show that our method outperforms the previous state-of-the-art debiasing technique by 3.6% on the debiased extension of the FEVER dataset, with a total performance boost of 10.13% from the baseline. Furthermore, we evaluate our approach in data-scarce settings, where models can be more susceptible to biases due to the lack of training data. Experimental results demonstrate that our approach is also effective at debiasing in these low-resource conditions, exceeding the baseline performance on the Symmetric dataset with just 1% of the original data. Minwoo Lee 0003, Seungpil Won, Juae Kim, Hwanhee Lee, Cheon-Eum Park, Kyomin Jung |
CIKM | 4 |
| 2021 | KPQA: A Metric for Generative Question Answering Using Keyphrase WeightsabstractHwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Joongbo Shin, Kyomin Jung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hwanhee Lee, Seunghyun Yoon 0002, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Joongbo Shin, Kyomin Jung |
NAACL-HLT | 1 |
| 2021 | Entangled Bidirectional Encoder to Autoregressive Decoder for Sequential RecommendationabstractRecently, BERT has shown overwhelming performance in sequential recommendation by using a bidirectional attention mechanism. Although the bidirectional model effectively captures dynamics from user interaction, its training strategy does not fit well to the inference stage in sequential recommendation which generally proceeds in a left-to-right way. To address this problem, we introduce a new recommendation system built upon BART, which is widely used in NLP tasks. BART uses a left-to-right decoder and injects noise into its bidirectional encoder, which can reduce the gap between training and inference. However, direct usage of BART for recommendation system is challenging due to its model property and domain difference. BART is an auto-regressive generative model, and its noising transformation techniques are originally developed for text sequence. In this paper, we present a novel sequential recommendation model, Entangled BART for Recommendation (E-BART4Rec) that entangles bidirectional encoder and auto-regressive decoder with noisy transformations for user interaction. Unlike BART, where the final output only depends on its output of the decoder, E-BART4Rec dynamically integrates the output of the bidirectional encoder and auto-regressive decoder based on a gating mechanism that calculates the importance of each output. We also employ noisy transformation that imitates the real users' behaviors, such as item deletion, item cropping, item reverse, and item infilling, to the input of the encoder. Extensive experiments on widely used real-world datasets demonstrate that our models significantly outperform the baselines. Taegwan Kang, Hwanhee Lee, Byeongjin Choe, Kyomin Jung |
SIGIR | 2 |
| 2020 | Attentive Modality Hopping Mechanism for Speech Emotion RecognitionabstractIn this work, we explore the impact of visual modality in addition to speech and text for improving the accuracy of the emotion detection system. The traditional approaches tackle this task by independently fusing the knowledge from the various modalities for performing emotion classification. In contrast to these approaches, we tackle the problem by introducing an attention mechanism to combine the information. In this regard, we first apply a neural network to obtain hidden representations of the modalities. Then, the attention mechanism is defined to select and aggregate important parts of the video data by conditioning on the audio and text data. Furthermore, the attention mechanism is again applied to attend the essential parts of speech and textual data by considering other modalities. Experiments are performed on the standard IEMOCAP dataset using all three modalities (audio, text, and video). The achieved results show a significant improvement of 3.65% in terms of weighted accuracy compared to the baseline system. Seunghyun Yoon 0002, Subhadeep Dey, Hwanhee Lee, Kyomin Jung |
ICASSP | 3 |
| 2019 | Improving Neural Question Generation Using Answer SeparationabstractNeural question generation (NQG) is the task of generating a question from a given passage with deep neural networks. Previous NQG models suffer from a problem that a significant proportion of the generated questions include words in the question target, resulting in the generation of unintended questions. In this paper, we propose answer-separated seq2seq, which better utilizes the information from both the passage and the target answer. By replacing the target answer in the original passage with a special token, our model learns to identify which interrogative word should be used. We also propose a new module termed keyword-net, which helps the model better capture the key information in the target answer and generate an appropriate question. Experimental results demonstrate that our answer separation method significantly reduces the number of improper questions which include answers. Consequently, our model significantly outperforms previous state-of-the-art NQG models. Yanghoon Kim, Hwanhee Lee, Joongbo Shin, Kyomin Jung |
AAAI | 2 |