EDBT 2026 Demo / reviewers in the wild / expert
YunSeok Choi
dblp:197/9031
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-9971-1501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 77% Information extraction and text analysis · 23% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 100% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 65% Recommender systems · 35% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning
adversarial attack |
1.2 | 2 | 2023 | DIP: Dead code Insertion based Black-box Attack for Programming Language Model · ACL (1) 2023 TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Program synthesis and code generation › code model
code model robustness |
1.2 | 2 | 2023 | DIP: Dead code Insertion based Black-box Attack for Programming Language Model · ACL (1) 2023 TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Natural language and speech › Language models and text generation › large language model evaluation › truthfulness evaluation
factual consistency evaluation |
1.0 | 1 | 2026 | ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.0 | 1 | 2026 | ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026 |
Program synthesis and code generation
code summarization |
1.0 | 1 | 2026 | ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026 |
Program synthesis and code generation › code summarization
code summarization evaluation |
1.0 | 1 | 2026 | ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization · ACL (1) 2026 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.9 | 1 | 2025 | DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025 |
Information retrieval › retrieval augmentation
demonstration retrieval |
0.9 | 1 | 2025 | DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph · ACL (1) 2025 |
Recommender systems
session-based recommendation |
0.7 | 1 | 2023 | LOAM: Improving Long-tail Session-based Recommendation via Niche Walk Augmentation and Tail Session Mixup · SIGIR 2023 |
Security and privacy of machine learning › adversarial attack
textual adversarial attack |
0.6 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.2 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Information retrieval › document retrieval › domain-specific retrieval › code search
semantic code search |
0.2 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.0in-context learning · 1.7graph-based retrieval · 1.7contextual semantic filtering · 1.7beam search · 1.7dead code insertion · 1.3black-box attack · 1.3tail session mixup · 0.7self-supervised learning · 0.7niche walk augmentation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationabstractAs Large Language Models (LLMs) have become capable of generating long and descriptive code summaries, accurate and reliable evaluation of factual consistency has become a critical challenge.However, previous evaluation methods are primarily designed for short summaries of isolated code snippets.Consequently, they struggle to provide fine-grained evaluation of multi-sentence functionalities and fail to accurately assess dependency context commonly found in real-world code summaries.To address this, we propose ReFEree, a referencefree and fine-grained method for evaluating factual consistency in real-world code summaries.We define factual inconsistency criteria specific to code summaries and evaluate them at the segment level using these criteria along with dependency information.These segment-level results are then aggregated into a fine-grained score.We construct a code summarization benchmark with human-annotated factual consistency labels.The evaluation results demonstrate that ReFEree achieves the highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art.Our code and data are available at https: //github.com/bsy99615/ReFEree.git. Suyoung Bae, CheolWon Na, Yumin Lee, YunSeok Choi, Jee-Hyong Lee 0001 |
ACL (1) | 5 |
| 2025 | DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link GraphabstractText-to-SQL, which translates a natural language question into an SQL query, has advanced with in-context learning of Large Language Models (LLMs). However, existing methods show little improvement in performance compared to randomly chosen demonstrations, and significant performance drops when smaller LLMs (e.g., Llama 3.1-8B) are used. This indicates that these methods heavily rely on the intrinsic capabilities of hyper-scaled LLMs, rather than effectively retrieving useful demonstrations. In this paper, we propose a novel approach for effectively retrieving demonstrations and generating SQL queries. We construct a Deep Contextual Schema Link Graph, which contains key information and semantic relationship between a question and its database schema items. This graph-based structure enables effective representation of Text-to-SQL samples and retrieval of useful demonstrations for in-context learning. Experimental results on the Spider benchmark demonstrate the effectiveness of our approach, showing consistent improvements in SQL generation performance and efficiency across both hyper-scaled LLMs and small LLMs. The code is available at https://github.com/jjklle/DCG-SQL. Jihyung Lee, Jin-Seop Lee, YunSeok Choi, Jee-Hyong Lee 0001 |
ACL (1) | 4 |
| 2025 | SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented DataabstractSuyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001 |
NAACL (Long Papers) | 2 |
| 2025 | DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language ModelsabstractSuyoung Bae, YunSeok Choi, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Suyoung Bae, YunSeok Choi, Jee-Hyong Lee 0001 |
NAACL (Long Papers) | 2 |
| 2025 | CoRAC: Integrating Selective API Document Retrieval with Question Semantic Intent for Code Question AnsweringabstractYunSeok Choi, CheolWon Na, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. YunSeok Choi, CheolWon Na, Jee-Hyong Lee 0001 |
NAACL (Long Papers) | 1 |
| 2024 | Code Defect Detection Using Pre-trained Language Models with Encoder-Decoder via Line-Level Defect LocalizationabstractRecently, code Pre-trained Language Models (PLMs) trained on large amounts of code and comment, have shown great success in code defect detection tasks. However, most PLMs simply treated the code as a single sequence and only used the encoder of PLMs to determine if there exist defects in the entire code. For a more analyzable and explainable approach, it is crucial to identify which lines contain defects. In this paper, we propose a novel method for code defect detection that integrates line-level defect localization into a unified training process. To identify code defects at the line-level, we convert the code into a sequence separated by lines using a special token. Then, to utilize the characteristic that both the encoder and decoder of PLMs process information differently, we leverage both the encoder and decoder for line-level defect localization. By learning code defect detection and line-level defect localization tasks in a unified manner, our proposed method promotes knowledge sharing between the two tasks. We demonstrate that our proposed method significantly improves performance on four benchmark datasets for code defect detection. Additionally, we show that our method can be easily integrated with ChatGPT. Jimin An, YunSeok Choi, Jee-Hyong Lee 0001 |
LREC/COLING | 2 |
| 2023 | DIP: Dead code Insertion based Black-box Attack for Programming Language ModelabstractAutomatic processing of source code, such as code clone detection and software vulnerability detection, is very helpful to software engineers.Large pre-trained Programming Language (PL) models (such as CodeBERT, Graph-CodeBERT, CodeT5, etc.), show very powerful performance on these tasks.However, these PL models are vulnerable to adversarial examples that are generated with slight perturbation.Unlike natural language, an adversarial example of code must be semantic-preserving and compilable.Due to the requirements, it is hard to directly apply the existing attack methods for natural language models.In this paper, we propose DIP (Dead code Insertion based Blackbox Attack for Programming Language Model), a high-performance and efficient black-box attack method to generate adversarial examples using dead code insertion.We evaluate our proposed method on 9 victim downstream-task large code models.Our method outperforms the state-of-the-art black-box attack in both attack efficiency and attack quality, while generated adversarial examples are compiled preserving semantic functionality. CheolWon Na, YunSeok Choi, Jee-Hyong Lee 0001 |
ACL (1) | 2 |
| 2023 | LOAM: Improving Long-tail Session-based Recommendation via Niche Walk Augmentation and Tail Session MixupabstractSession-based recommendation aims to predict the user's next action based on anonymous sessions without using side information. Most of the real-world session datasets are sparse and have long-tail item distribution. Although long-tail item recommendation plays a crucial role in improving user satisfaction, only a few methods have been proposed to take the long-tail session recommendation into consideration. Previous works in handling data sparsity problems are mostly limited to self-supervised learning techniques with heuristic augmentation which can ruin the original characteristic of session datasets, sequential and co-occurrences, and make noisier short sessions by dropping items and cropping sequences. We propose a novel method, LOAM, improving LOng-tail session-based recommendation via niche walk Augmentation and tail session Mixup, that alleviates popularity bias and enhances long-tail recommendation performance. LOAM consists of two modules, Niche Walk Augmentation (NWA) and Tail Session Mixup (TSM). NWA can generate synthetic sessions considering long-tail distribution which are likely to be found in original datasets, unlike previous heuristic methods, and expose a recommender model to various item transitions with global information. This improves the item coverage of recommendations. TSM makes the model more generalized and robust by interpolating sessions at the representation level. It encourages the recommender system to predict niche items with more diversity and relevance. We conduct extensive experiments with four real-world datasets and verify that our methods greatly improve tail performance while balancing overall performance. Heeyoon Yang, YunSeok Choi, Gahyung Kim, Jee-Hyong Lee 0001 |
SIGIR | 2 |
| 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam SearchabstractAs pre-trained models have shown successful performance in program language processing as well as natural language processing, adversarial attacks on these models also attract attention.However, previous works on blackbox adversarial attacks generated adversarial examples in a very inefficient way with simple greedy search.They also failed to find out better adversarial examples because it was hard to reduce the search space without performance loss.In this paper, we propose TABS, an efficient beam search black-box adversarial attack method.We adopt beam search to find out better adversarial examples, and contextual semantic filtering to effectively reduce the search space.Contextual semantic filtering reduces the number of candidate adversarial words considering the surrounding context and the semantic similarity.Our proposed method shows good performance in terms of attack success rate, the number of queries, and semantic similarity in attacking models for two tasks: NL code search classification and retrieval tasks. YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001 |
EMNLP | 1 |
| 2020 | Neural attention model with keyword memory for abstractive document summarizationabstractSummary Abstractive summarization is the task of creating summaries by generating a set of novel sentences based on the information extracted from the original document, while most of summarization researches are based on extractive or compressive approaches. These approaches extract phrases from the original document and concatenate them by post‐processing and cannot truly encapsulate the contents of summaries, because they only reuse the phrases in the given document. Moreover, there are limits for paraphrasing and re‐organizing of the original contents with the current Natural Language Processing (NLP) techniques. With these reasons, we propose a novel abstractive summarization method. The main goal of our paper is to generate a long sequence of words with coherent sentences by reflecting the key concepts of the original document and the contents of summaries. To achieve this goal, we propose an attention mechanism that uses Document Content Memory for learning the language model effectively. To evaluate its effectiveness, the proposed methods are compared with other language models and an extractive summarization method. We demonstrate that our proposed methods improve summarization results in ROUGE score using ACL dataset. The experimental results show that our proposed methods using keyword memory are effective to generate long sequence summary. YunSeok Choi, Dahae Kim, Jee-Hyong Lee 0001 |
Concurr. Comput. Pract. Exp. | 1 |