YunSeok Choi

dblp:197/9031 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-9971-1501ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 77% Information extraction and text analysis · 23%
Software engineering, system software, and programming languages
3 papers
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 65% Recommender systems · 35%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
adversarial attack
1.222023
DIP: Dead code Insertion based Black-box Attack for Programming Language Model · ACL (1) 2023
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Program synthesis and code generation › code model
code model robustness
1.222023
DIP: Dead code Insertion based Black-box Attack for Programming Language Model · ACL (1) 2023
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Natural language and speech › Language models and text generation › large language model evaluation › truthfulness evaluation
factual consistency evaluation
1.012026
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026
Program synthesis and code generation
code summarization
1.012026
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026
Program synthesis and code generation › code summarization
code summarization evaluation
1.012026
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026
Natural language and speech › Language models and text generation
in-context learning
0.912025
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.912025
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025
Information retrieval › retrieval augmentation
demonstration retrieval
0.912025
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025
Recommender systems
session-based recommendation
0.712023
LOAM: Improving Long-tail Session-based Recommendation via Niche Walk Augmentation and Tail Session Mixup · SIGIR 2023
Security and privacy of machine learning › adversarial attack
textual adversarial attack
0.612022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Information retrieval › document retrieval › domain-specific retrieval
code search
0.212022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Information retrieval › document retrieval › domain-specific retrieval › code search
semantic code search
0.212022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

large language model · 2.0in-context learning · 1.7graph-based retrieval · 1.7contextual semantic filtering · 1.7beam search · 1.7dead code insertion · 1.3black-box attack · 1.3tail session mixup · 0.7self-supervised learning · 0.7niche walk augmentation · 0.7
YearPublicationVenuePosition
2026 ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
abstract
As Large Language Models (LLMs) have become capable of generating long and descriptive code summaries, accurate and reliable evaluation of factual consistency has become a critical challenge.However, previous evaluation methods are primarily designed for short summaries of isolated code snippets.Consequently, they struggle to provide fine-grained evaluation of multi-sentence functionalities and fail to accurately assess dependency context commonly found in real-world code summaries.To address this, we propose ReFEree, a referencefree and fine-grained method for evaluating factual consistency in real-world code summaries.We define factual inconsistency criteria specific to code summaries and evaluate them at the segment level using these criteria along with dependency information.These segment-level results are then aggregated into a fine-grained score.We construct a code summarization benchmark with human-annotated factual consistency labels.The evaluation results demonstrate that ReFEree achieves the highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art.Our code and data are available at https: //github.com/bsy99615/ReFEree.git.
Suyoung Bae, CheolWon Na, Yumin Lee, YunSeok Choi, Jee-Hyong Lee 0001
ACL (1)5
2025 DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
abstract
Text-to-SQL, which translates a natural language question into an SQL query, has advanced with in-context learning of Large Language Models (LLMs). However, existing methods show little improvement in performance compared to randomly chosen demonstrations, and significant performance drops when smaller LLMs (e.g., Llama 3.1-8B) are used. This indicates that these methods heavily rely on the intrinsic capabilities of hyper-scaled LLMs, rather than effectively retrieving useful demonstrations. In this paper, we propose a novel approach for effectively retrieving demonstrations and generating SQL queries. We construct a Deep Contextual Schema Link Graph, which contains key information and semantic relationship between a question and its database schema items. This graph-based structure enables effective representation of Text-to-SQL samples and retrieval of useful demonstrations for in-context learning. Experimental results on the Spider benchmark demonstrate the effectiveness of our approach, showing consistent improvements in SQL generation performance and efficiency across both hyper-scaled LLMs and small LLMs. The code is available at https://github.com/jjklle/DCG-SQL.
Jihyung Lee, Jin-Seop Lee, YunSeok Choi, Jee-Hyong Lee 0001
ACL (1)4
2025 SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
abstract
Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001
NAACL (Long Papers)2
2025 DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
abstract
Suyoung Bae, YunSeok Choi, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Suyoung Bae, YunSeok Choi, Jee-Hyong Lee 0001
NAACL (Long Papers)2
2025 CoRAC: Integrating Selective API Document Retrieval with Question Semantic Intent for Code Question Answering
abstract
YunSeok Choi, CheolWon Na, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
YunSeok Choi, CheolWon Na, Jee-Hyong Lee 0001
NAACL (Long Papers)1
2024 Code Defect Detection Using Pre-trained Language Models with Encoder-Decoder via Line-Level Defect Localization
abstract
Recently, code Pre-trained Language Models (PLMs) trained on large amounts of code and comment, have shown great success in code defect detection tasks. However, most PLMs simply treated the code as a single sequence and only used the encoder of PLMs to determine if there exist defects in the entire code. For a more analyzable and explainable approach, it is crucial to identify which lines contain defects. In this paper, we propose a novel method for code defect detection that integrates line-level defect localization into a unified training process. To identify code defects at the line-level, we convert the code into a sequence separated by lines using a special token. Then, to utilize the characteristic that both the encoder and decoder of PLMs process information differently, we leverage both the encoder and decoder for line-level defect localization. By learning code defect detection and line-level defect localization tasks in a unified manner, our proposed method promotes knowledge sharing between the two tasks. We demonstrate that our proposed method significantly improves performance on four benchmark datasets for code defect detection. Additionally, we show that our method can be easily integrated with ChatGPT.
Jimin An, YunSeok Choi, Jee-Hyong Lee 0001
LREC/COLING2
2023 DIP: Dead code Insertion based Black-box Attack for Programming Language Model
abstract
Automatic processing of source code, such as code clone detection and software vulnerability detection, is very helpful to software engineers.Large pre-trained Programming Language (PL) models (such as CodeBERT, Graph-CodeBERT, CodeT5, etc.), show very powerful performance on these tasks.However, these PL models are vulnerable to adversarial examples that are generated with slight perturbation.Unlike natural language, an adversarial example of code must be semantic-preserving and compilable.Due to the requirements, it is hard to directly apply the existing attack methods for natural language models.In this paper, we propose DIP (Dead code Insertion based Blackbox Attack for Programming Language Model), a high-performance and efficient black-box attack method to generate adversarial examples using dead code insertion.We evaluate our proposed method on 9 victim downstream-task large code models.Our method outperforms the state-of-the-art black-box attack in both attack efficiency and attack quality, while generated adversarial examples are compiled preserving semantic functionality.
CheolWon Na, YunSeok Choi, Jee-Hyong Lee 0001
ACL (1)2
2023 LOAM: Improving Long-tail Session-based Recommendation via Niche Walk Augmentation and Tail Session Mixup
abstract
Session-based recommendation aims to predict the user's next action based on anonymous sessions without using side information. Most of the real-world session datasets are sparse and have long-tail item distribution. Although long-tail item recommendation plays a crucial role in improving user satisfaction, only a few methods have been proposed to take the long-tail session recommendation into consideration. Previous works in handling data sparsity problems are mostly limited to self-supervised learning techniques with heuristic augmentation which can ruin the original characteristic of session datasets, sequential and co-occurrences, and make noisier short sessions by dropping items and cropping sequences. We propose a novel method, LOAM, improving LOng-tail session-based recommendation via niche walk Augmentation and tail session Mixup, that alleviates popularity bias and enhances long-tail recommendation performance. LOAM consists of two modules, Niche Walk Augmentation (NWA) and Tail Session Mixup (TSM). NWA can generate synthetic sessions considering long-tail distribution which are likely to be found in original datasets, unlike previous heuristic methods, and expose a recommender model to various item transitions with global information. This improves the item coverage of recommendations. TSM makes the model more generalized and robust by interpolating sessions at the representation level. It encourages the recommender system to predict niche items with more diversity and relevance. We conduct extensive experiments with four real-world datasets and verify that our methods greatly improve tail performance while balancing overall performance.
Heeyoon Yang, YunSeok Choi, Gahyung Kim, Jee-Hyong Lee 0001
SIGIR2
2022 TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search
abstract
As pre-trained models have shown successful performance in program language processing as well as natural language processing, adversarial attacks on these models also attract attention.However, previous works on blackbox adversarial attacks generated adversarial examples in a very inefficient way with simple greedy search.They also failed to find out better adversarial examples because it was hard to reduce the search space without performance loss.In this paper, we propose TABS, an efficient beam search black-box adversarial attack method.We adopt beam search to find out better adversarial examples, and contextual semantic filtering to effectively reduce the search space.Contextual semantic filtering reduces the number of candidate adversarial words considering the surrounding context and the semantic similarity.Our proposed method shows good performance in terms of attack success rate, the number of queries, and semantic similarity in attacking models for two tasks: NL code search classification and retrieval tasks.
YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001
EMNLP1
2020 Neural attention model with keyword memory for abstractive document summarization
abstract
Summary Abstractive summarization is the task of creating summaries by generating a set of novel sentences based on the information extracted from the original document, while most of summarization researches are based on extractive or compressive approaches. These approaches extract phrases from the original document and concatenate them by post‐processing and cannot truly encapsulate the contents of summaries, because they only reuse the phrases in the given document. Moreover, there are limits for paraphrasing and re‐organizing of the original contents with the current Natural Language Processing (NLP) techniques. With these reasons, we propose a novel abstractive summarization method. The main goal of our paper is to generate a long sequence of words with coherent sentences by reflecting the key concepts of the original document and the contents of summaries. To achieve this goal, we propose an attention mechanism that uses Document Content Memory for learning the language model effectively. To evaluate its effectiveness, the proposed methods are compared with other language models and an extractive summarization method. We demonstrate that our proposed methods improve summarization results in ROUGE score using ACL dataset. The experimental results show that our proposed methods using keyword memory are effective to generate long sequence summary.
YunSeok Choi, Dahae Kim, Jee-Hyong Lee 0001
Concurr. Comput. Pract. Exp.1