Youna Kim

dblp:340/0327 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 78% Information extraction and text analysis · 19% Trustworthy machine learning · 4%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization
1.012026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.912025
When to Speak, When to Abstain: Contrastive Decoding with Abstention · ACL (1) 2025
Natural language and speech › Language models and text generation › natural language understanding
ambiguity handling
0.812024
Aligning Language Models to Explicitly Handle Ambiguity · EMNLP 2024
Natural language and speech › Information extraction and text analysis
text classification
0.712023
CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.712023
CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023
Natural language and speech › Language models and text generation
instruction tuning
0.212024
Aligning Language Models to Explicitly Handle Ambiguity · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

knowledge relevance estimation · 0.9contrastive decoding · 0.9supervised fine-tuning · 0.8preference optimization · 0.8prompting · 0.7linear discriminative analysis · 0.7clustering · 0.7
YearPublicationVenuePosition
2026 CUB: Benchmarking Context Utilisation Techniques for Language Models
abstract
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples.
Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein
ACL (1)2
2025 When to Speak, When to Abstain: Contrastive Decoding with Abstention
abstract
Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks by leveraging pre-trained (i.e., parametric) and external (i.e., contextual) knowledge. While substantial efforts have been made to enhance the utilization of both forms of knowledge, situations in which models lack relevant information remain underexplored. To investigate this challenge, we first present a controlled testbed featuring four distinct knowledge access scenarios, including the aforementioned edge case, revealing that conventional LLM usage exhibits insufficient robustness in handling all instances. Addressing this limitation, we propose Contrastive Decoding with Abstention (CDA), a novel training-free decoding method that allows LLMs to generate responses when relevant knowledge is available and to abstain otherwise. CDA estimates the relevance of both knowledge sources for a given input, adaptively deciding which type of information to prioritize and which to exclude. Through extensive experiments, we demonstrate that CDA can effectively perform accurate generation and abstention simultaneously, enhancing reliability and preserving user trust.
Hyuhng Joon Kim, Youna Kim, Sang-goo Lee, Taeuk Kim
ACL (1)2
2024 Aligning Language Models to Explicitly Handle Ambiguity
abstract
Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim
EMNLP2
2023 CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels
abstract
Utilizing language models (LMs) without internal access is becoming an attractive paradigm in the field of NLP as many cutting-edge LMs are released through APIs and boast a massive scale.The de-facto method in this type of black-box scenario is known as prompting, which has shown progressive performance enhancements in situations where data labels are scarce or unavailable.Despite their efficacy, they still fall short in comparison to fully supervised counterparts and are generally brittle to slight modifications.In this paper, we propose Clustering-Enhanced Linear Discriminative Analysis (CELDA), a novel approach that improves the text classification accuracy with a very weak-supervision signal (i.e., name of the labels).Our framework draws a precise decision boundary without accessing weights or gradients of the LM model or data labels.The core ideas of CELDA are twofold: (1) extracting a refined pseudo-labeled dataset from an unlabeled dataset, and (2) training a lightweight and robust model on the top of LM, which learns an accurate decision boundary from an extracted noisy dataset.Throughout in-depth investigations on various datasets, we demonstrated that CELDA reaches new state-of-theart in weakly-supervised text classification and narrows the gap with a fully-supervised model.Additionally, our proposed methodology can be applied universally to any LM and has the potential to scale to larger models, making it a more viable option for utilizing large LMs.
Hyunsoo Cho, Youna Kim, Sang-goo Lee
ACL (1)2
2023 Multi-scale Contrastive Learning for Complex Scene Generation
abstract
Recent advances in Generative Adversarial Networks (GANs) have enabled photo-realistic synthesis of single object images. Yet, modeling more complex distributions, such as scenes with multiple objects, remains challenging. The difficulty stems from the incalculable variety of scene configurations which contain multiple objects of different categories placed at various locations. In this paper, we aim to alleviate the difficulty by enhancing the discriminative ability of the discriminator through a locally defined self-supervised pretext task. To this end, we design a discriminator to leverage multi-scale local feedback that guides the generator to better model local semantic structures in the scene. Then, we require the discriminator to carry out pixel-level contrastive learning at multiple scales to enhance discriminative capability on local regions. Experimental results on several challenging scene datasets show that our method improves the synthesis quality by a substantial margin compared to state-of-the-art baselines.
Hanbit Lee, Youna Kim, Sang-goo Lee
WACV2