Haeun Yu

dblp:305/0149 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-1287-703XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 50% Trustworthy machine learning · 35% Information extraction and text analysis · 14%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization
1.922026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
A Reality Check on Context Utilisation for Retrieval-Augmented Generation · ACL (1) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.922026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
A Reality Check on Context Utilisation for Retrieval-Augmented Generation · ACL (1) 2025
Machine learning › Trustworthy machine learning
interpretability
1.022024
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods · ACL (1) 2024
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution · WSDM 2023
Natural language and speech › Information extraction and text analysis
fact-checking
0.912025
A Reality Check on Context Utilisation for Retrieval-Augmented Generation · ACL (1) 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
neuron attribution
0.812024
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods · ACL (1) 2024
Machine learning › Trustworthy machine learning › interpretability
training data attribution
0.812024
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
coreference resolution
0.712023
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution · WSDM 2023
Machine learning › Trustworthy machine learning
fairness
0.712023
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution · WSDM 2023
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation
0.712023
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution · WSDM 2023
Natural language and speech › Language models and text generation
natural language understanding
0.712023
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution · WSDM 2023

Methods — techniques the papers use, named apart from their topics

faithfulness tests · 0.8attribution methods · 0.8regularization · 0.7embedding visualization · 0.7elastic weight consolidation · 0.7
YearPublicationVenuePosition
2026 CUB: Benchmarking Context Utilisation Techniques for Language Models
abstract
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples.
Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein
ACL (1)3
2025 A Reality Check on Context Utilisation for Retrieval-Augmented Generation
abstract
Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings.
Lovisa Hagström, Sara Marjanovic, Haeun Yu, Arnav Arora, Christina Lioma, Maria Maistro, Pepa Atanasova, Isabelle Augenstein
ACL (1)3
2024 Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
abstract
Language Models (LMs) acquire parametric knowledge from their training process, embedding it within their weights.The increasing scalability of LMs, however, poses significant challenges for understanding a model's inner workings and further for updating or correcting this embedded knowledge without the significant cost of retraining.This underscores the importance of unveiling exactly what knowledge is stored and its association with specific model components.Instance Attribution (IA) and Neuron Attribution (NA) offer insights into this training-acquired knowledge, though they have not been compared systematically.Our study introduces a novel evaluation framework to quantify and compare the knowledge revealed by IA and NA.To align the results of the methods we introduce the attribution method NA-Instances to apply NA for retrieving influential training instances, and IA-Neurons to discover important neurons of influential instances discovered by IA.We further propose a comprehensive list of faithfulness tests to evaluate the comprehensiveness and sufficiency of the explanations provided by both methods.Through extensive experiments and analysis, we demonstrate that NA generally reveals more diverse and comprehensive information regarding the LM's parametric knowledge compared to IA.Nevertheless, IA provides unique and valuable insights into the LM's parametric knowledge, which are not revealed by NA.Our findings further suggest the potential of a synergistic approach of combining the diverse findings of IA and NA for a more holistic understanding of an LM's parametric knowledge.
Haeun Yu, Pepa Atanasova, Isabelle Augenstein
ACL (1)1
2023 Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution
abstract
Leveraging pre-trained language models (PLMs) as initializers for efficient transfer learning has become a universal approach for text-related tasks. However, the models not only learn the language understanding abilities but also reproduce prejudices for certain groups in the datasets used for pre-training. Recent studies show that the biased knowledge acquired from the datasets affects the model predictions on downstream tasks. In this paper, we mitigate and analyze the gender biases in PLMs with coreference resolution, which is one of the natural language understanding (NLU) tasks. PLMs exhibit two types of gender biases: stereotype and skew. The primary causes for the biases are the imbalanced datasets with more male examples and the stereotypical examples on gender roles. While previous studies mainly focused on the skew problem, we aim to mitigate both gender biases in PLMs while maintaining the model's original linguistic capabilities. Our method employs two regularization terms, Stereotype Neutralization (SN) and Elastic Weight Consolidation (EWC). The models trained with the methods show to be neutralized and reduce the biases significantly on the WinoBias dataset compared to the public BERT. We also invented a new gender bias quantification metric called the Stereotype Quantification (SQ) score. In addition to the metrics, embedding visualizations were used to interpret how our methods have successfully debiased the models.
Kyuri Choi, Haeun Yu, Youngjoong Ko
WSDM3
2023 Knowledge-grounded dialogue modelling with dialogue-state tracking, domain tracking, and entity extraction
Taesuk Hong, Haeun Yu, Youngjoong Ko, Jungyun Seo
Comput. Speech Lang.3
2023 Graph-based query reformulation system for descriptive queries of jargon words using definitions
Hyewon Choi, Haeun Yu, Youngjoong Ko
Expert Syst. Appl.3
2023 Enriching the dialogue state tracking model with a asyntactic discourse graph
Haeun Yu, Youngjoong Ko
Pattern Recognit. Lett.1
2021 Query Reformulation for Descriptive Queries of Jargon Words Using a Knowledge Graph based on a Dictionary
abstract
Query reformulation (QR) is a key factor in overcoming the problems faced by the lexical chasm in information retrieval (IR) systems. In particular, when searching for jargon, people tend to use descriptive queries, such as "a medical examination of the colon" rather than "colonoscopy," or they often use them interchangeably. Thus, transforming users' descriptive queries into appropriate jargon queries helps to retrieve more relevant documents. In this paper, we propose a new graph-based QR system that uses a dictionary, where the model does not require human-labeled data. Given a descriptive query, our system predicts the corresponding jargon word over a graph consisting of pairs of a headword and its description in the dictionary. First, we train a graph neural network to represent the relational properties between words and to infer a jargon word using compositional information of the descriptive query's words. Moreover, we propose a graph search model that finds the target node in real time using the relevance scores of neighborhood nodes. By adding this fast graph search model to the front of the proposed system, we reduce the reformulating time significantly. Experimental results on two datasets show that the proposed method can effectively reformulate descriptive queries to corresponding jargon words as well as improve retrieval performance under several search frameworks.
Hyewon Choi, Haeun Yu, Youngjoong Ko
CIKM3