Hyunkyung Bae

dblp:321/0513 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 45% Question answering and dialogue systems · 24% Information extraction and text analysis · 15%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
1.422024
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs · EMNLP 2024
Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources · EMNLP 2023
Natural language and speech › Language models and text generation › in-context learning
demonstration selection
0.912025
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL · EMNLP 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL · EMNLP 2025
Natural language and speech › Language models and text generation
instruction following
0.912025
LLMs can be easily Confused by Instructional Distractions · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.912025
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL · EMNLP 2025
Machine learning › Generative modeling › synthetic data generation
dataset generation
0.712023
Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources · EMNLP 2023
Machine learning › Trustworthy machine learning › large language model trustworthiness
large language model robustness
0.312025
LLMs can be easily Confused by Instructional Distractions · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

large language model · 2.4passage-to-dialogue generation · 1.5self-augmentation · 0.9fine-grained example selection · 0.9benchmark construction · 0.9topic-aware dialog generation · 0.7reranking · 0.7question-answer matching · 0.7
YearPublicationVenuePosition
2025 LLMs can be easily Confused by Instructional Distractions
abstract
Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions.Instruction following tasks typically involve a clear task description and input text containing the target data to be processed.However, when the input itself resembles an instruction, confusion may arise, even if there is explicit prompting to distinguish between the task instruction and the input.We refer to this phenomenon as instructional distraction.In this paper, we introduce a novel benchmark, named DIM-Bench, specifically designed to assess LLMs' performance under instructional distraction.The benchmark categorizes real-world instances of instructional distraction and evaluates LLMs across four instruction tasks: rewriting, proofreading, translation, and style transfer-alongside five input tasks: reasoning, code generation, mathematical reasoning, bias detection, and question answering.Our experimental results reveal that even the most advanced LLMs are susceptible to instructional distraction, often failing to accurately follow user intent in such cases. Instruction Input ExampleRewrite Reasoning Instruction: Paraphrase the following text.Input: Laundry detergents were once manufactured to contain high ... which would a lake become as a result of the phosphorous in the detergent?Options : A. canyon B. desert C. swamp D. river
Yerin Hwang, Yongil Kim, Jahyun Koo 0004, Taegwan Kang, Hyunkyung Bae, Kyomin Jung
ACL (1)5
2025 SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL
abstract
Text-to-SQL aims to convert natural language questions into executable SQL queries.While previous approaches, such as skeleton-masked selection, have demonstrated strong performance by retrieving similar training examples to guide large language models (LLMs), they struggle in real-world scenarios where such examples are unavailable.To overcome this limitation, we propose Self-Augmentation incontext learning with Fine-grained Example selection for Text-to-SQL (SAFE-SQL), a novel unsupervised framework that enhances SQL generation by generating and intelligently filtering self-augmented examples.SAFE-SQL leverages an LLM to generate diverse Textto-SQL examples, which are then filtered by a novel fine-grained mechanism using criteria for semantic similarity, structural alignment, and reasoning path quality to curate highquality in-context learning examples.Leveraging these carefully selected self-generated examples, SAFE-SQL significantly surpasses previous zero-shot and few-shot Text-to-SQL frameworks, achieving superior execution accuracy.Notably, our approach demonstrates substantial performance gains in challenging extra hard and unseen scenarios, where conventional methods often struggle.
Jimin Lee 0001, Ingeol Baek, Byeongjeong Kim, Hyunkyung Bae, Hwanhee Lee
EMNLP4
2024 Kosmic: Korean Text Similarity Metric Reflecting Honorific Distinctions
abstract
Existing English-based text similarity measurements primarily focus on the semantic dimension, neglecting the unique linguistic attributes found in languages like Korean, where honorific expressions are explicitly integrated. To address this limitation, this study proposes Kosmic, a novel Korean text-similarity metric that encompasses the semantic and tonal facets of a given text pair. For the evaluation, we introduce a novel benchmark annotated by human experts, empirically showing that Kosmic outperforms the existing method. Moreover, by leveraging Kosmic, we assess various Korean paraphrasing methods to determine which techniques are most effective in preserving semantics and tone.
Yerin Hwang, Yongil Kim, Hyunkyung Bae, Jeesoo Bang, Hwanhee Lee, Kyomin Jung
LREC/COLING3
2024 MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
abstract
Despite advancements in on-topic dialogue systems, effectively managing topic shifts within dialogues remains a persistent challenge, largely attributed to the limited availability of training datasets.To address this issue, we propose Multi-Passage to Dialogue (MP2D), a data generation framework that automatically creates conversational questionanswering datasets with natural topic transitions.By leveraging the relationships between entities in a knowledge graph, MP2D maps the flow of topics within a dialogue, effectively mirroring the dynamics of human conversation.It retrieves relevant passages corresponding to the topics and transforms them into dialogues through the passage-to-dialogue method.Through quantitative and qualitative experiments, we demonstrate MP2D's efficacy in generating dialogue with natural topic shifts.Furthermore, this study introduces a novel benchmark for topic shift dialogues, TS-WikiDialog.Utilizing the dataset, we demonstrate that even Large Language Models (LLMs) struggle to handle topic shifts in dialogue effectively, and we showcase the performance improvements of models trained on datasets generated by MP2D across diverse topic shift dialogue tasks.
Yerin Hwang, Yongil Kim, Yunah Jang, Jeesoo Bang, Hyunkyung Bae, Kyomin Jung
EMNLP5
2024 IterCQR: Iterative Conversational Query Reformulation with Retrieval Guidance
abstract
Yunah Jang, Kang-il Lee, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yunah Jang, Kangil Lee 0001, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung
NAACL-HLT3
2023 Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources
abstract
To address the data scarcity issue in Conversational question answering (ConvQA), a dialog inpainting method, which utilizes documents to generate ConvQA datasets, has been proposed.However, the original dialog inpainting model is trained solely on the dialog reconstruction task, resulting in the generation of questions with low contextual relevance due to insufficient learning of question-answer alignment.To overcome this limitation, we propose a novel framework called Dialogizer, which has the capability to automatically generate ConvQA datasets with high contextual relevance from textual sources.The framework incorporates two training tasks: question-answer matching (QAM) and topic-aware dialog generation (TDG).Moreover, re-ranking is conducted during the inference phase based on the contextual relevance of the generated questions.Using our framework, we produce four Con-vQA datasets by utilizing documents from multiple domains as the primary source.Through automatic evaluation using diverse metrics, as well as human evaluation, we validate that our proposed framework exhibits the ability to generate datasets of higher quality compared to the baseline dialog inpainting model.
Yerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee, Jeesoo Bang, Kyomin Jung
EMNLP3
2022 Subgraph Representation Learning with Hard Negative Samples for Inductive Link Prediction
abstract
The inductive link prediction in knowledge graphs (KGs) is often addressed to induce logical rules that capture entity-independent relational semantics. Recent studies suggest graph representation learning to encode these logical rules within the local subgraph structures. With this approach, the model can have the inductive ability to cope with unseen entities, which is practical for evolving nature of real-world KGs. However, despite the importance of a high-quality negative sample in link prediction, there is currently no method for selecting hard negatives for inductive link prediction. To overcome this limitation, we propose a new sampling method for selecting hard negative samples given a positive triplet. We also propose Subgraph Infomax (SGI), a novel inductive link prediction model, with a newly-proposed training objective that maximizes the mutual information (MI) between the target relation and the enclosing subgraph. We select hard negative samples by using the pretrained MI estimator of SGI. The model is then fine-tuned using the selected hard negative samples. Empirically, we demonstrate superior performances of our model on multiple datasets of the inductive KGC benchmark, showing the enhanced connectivity between the target relation embedding and the subgraph representation.
Heeyoung Kwak, Hyunkyung Bae, Kyomin Jung
ICASSP2