VLDB 2026 Research / reviewers in the wild / expert
Youngkyoung Kim
dblp:219/4400
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-5457-7997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Production and test bug report classification based on transfer learning
Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001 |
Inf. Softw. Technol. | 2 |
| 2023 | Preliminary Study on the Reproducibility of Fix Templates in Static Analysis ToolabstractAutomated Program Repair (APR) automatically generates patches for identified defects. As a result, APR can encourage novice students to learn coding by providing appropriate patches. Since students have their coding conventions, we should be able to deal with many types of defects. Existing studies have used predefined RuleId-driven templates to automatically fix defects in Static Analysis Tools(SATs). However, the community periodically adds, deletes, or changes SAT’s RuleIds. This is difficult for existing RuleId-based templates to reflect those changes immediately. Existing studies only cover about 10 RuleIds, making it difficult to address all defects faced by all students. Therefore, it is necessary to establish appropriate criteria for classifying templates. The SAT has a predefined format for how Error Messages are written, and since Error Messages contain fixing actions, defects with similar Error Messages tend to have similar fixing actions. These characteristics of the Error Message are suitable for reproducible template classification criteria. Our preliminary study demonstrated that by classifying patterns based on Error Messages, we could effectively address various defects, including those in different programming languages, using a single template. This means that if a newly added RuleId corresponds to an Error Message format already in the predefined Error Message-based template, it can be modified without additional effort. We plan to construct reproducible templates for each Error Message and provide ongoing patching of defects to students. Youngkyoung Kim, Eunseok Lee 0001 |
CSEE&T | 2 |
| 2023 | Improving Transformer-based Program Repair Model through False Behavior DiagnosisabstractResearch on automated program repairs using transformer-based models has recently gained considerable attention.The comprehension of the erroneous behavior of a model enables the identification of its inherent capacity and provides insights for improvement.However, the current landscape of research on program repair models lacks an investigation of their false behavior.Thus, we propose a methodology for diagnosing and treating the false behaviors of transformer-based program repair models.Specifically, we propose 1) a behavior vector that quantifies the behavior of the model when it generates an output, 2) a behavior discriminator (BeDisc) that identifies false behaviors, and 3) two methods for false behavior treatment.Through a large-scale experiment on 55,562 instances employing four datasets and three models, the BeDisc exhibited a balanced accuracy of 86.6% for false behavior classification.The first treatment, namely, early abortion, successfully eliminated 60.4% of false behavior while preserving 97.4% repair accuracy.Furthermore, the second treatment, namely, masked bypassing, resulted in an average improvement of 40.5% in the top-1 repair accuracy.These experimental results demonstrated the importance of investigating false behaviors in program repair models.* corresponding author 1 We refer to these as bugs in a broad sense. Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001 |
EMNLP | 1 |
| 2022 | Tracking Down Misguiding Terms for Locating Bugs in Deep Learning-Based Software (Student Abstract)abstractBugs in source files (SFs) may cause software malfunction, inconveniencing users and even leading to catastrophic accidents. Therefore, the bugs in SFs should be found and fixed quickly. However, from hundreds of candidate SFs, finding buggy SFs is tedious and time consuming. To lessen the burden on developers, deep learning-based bug localization (DLBL) tools can be utilized. Text terms in bug reports and SFs play an important role. However, some terms provide incorrect information and degrade bug localization performance. Therefore, those terms are defined here as "misguiding terms," and an explainable-artificial-intelligence-based identification method is proposed. The effectiveness of the proposed method for DLBL was investigated. When misguiding terms were removed, the mean average precision of the bug localization model improved by 33% on average. Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001 |
AAAI | 1 |
| 2022 | Impact of Defect Instances for Successful Deep Learning-based Automatic Program RepairabstractDeep learning-based automatic program repair (DL-APR) returns a patch code when given a defect code. Recent studies on DL-APR techniques have focused on the training phase to generate more accurate patches; however, a trained model cannot always generate an accurate patch for every new defect code, as the training dataset does not completely represent the new defects to be input in the future. DL-APR researchers should study a method to elicit the best performance on new inputs from the trained and deployed model. A new defect instance (i.e., defect codes and their context codes) is one of the crucial input data that determine the accuracy of the DL-APR, which can be changed and improved. We improve the quality of new input defect instances by focusing on the presence of noise tokens which compromise the defect instances’ quality, thus impairing the accuracy of generated patches. This paper shows that 1) there are noise tokens which prevent correct patch generation (inference) in a new defect instance, and 2) it is necessary to mask these noise tokens to avoid their usage in inferencing patch codes. In order to validate these two assertions, we use a state-of-the-art DL-APR technique and a genetic algorithm to generate near-optimal defect instances which maximize the patch generation accuracy (i.e., the BLEU score) of 4,573 defect instances. Based on optimization results, we found that 1) noise tokens impair patch generation accuracy in approximately 49% of instances, and 2) if these tokens are precluded from inference by masking them, we can improve patch generation accuracy by 88%. The results suggest that future work is required to automatically remove noise tokens from new defect instances so that the trained patch generator generates better patches. Misoo Kim, Youngkyoung Kim, Jinseok Heo, Hohyeon Jeong, Sungoh Kim, Eunseok Lee 0001 |
ICSME | 2 |
| 2022 | An Empirical Study of IR-based Bug Localization for Deep Learning-based SoftwareabstractAs the impact of deep-learning-based software (DLSW) increases, automatic debugging techniques for guaranteeing DLSW quality are becoming increasingly important. Information-retrieval-based bug localization (IRBL) techniques can aid in debugging by automatically localizing buggy entities (files and functions). The low-cost advantage of IRBL can alleviate the difficulty of identifying bug locations due to the complexity of DLSW. However, there are significant differences between DLSW and traditional software, and these differences lead to differences in search space and query quality for IRBL. That is, IRBL performance must be validated in DLSW. We empirically validated IRBL performance for DLSW from the following four perspectives: 1) similarity model, 2) query generation, 3) ranking model for buggy file localization, and 4) ranking model for buggy function localization. Based on four research questions and a large-scale experiment using 2,365 bug reports from 136 DLSW projects, we confirmed the salient char-acteristics of DLSW from the perspective of IRBL and derived four recommendations for practical IRBL usage in DLSW from the empirical results. Regarding IRBL performance, we validated that IRBL performance with the combination of bug-related features outperformed that of using only file similarity by 15 % and IRBL ranked buggy files and functions on average of 1.6th and 2.9th, respectively. Our study is valuable as a baseline for IRBL researchers and as a guideline for DLSW developers who wish to apply IRBL to ensure DLSW quality. Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001 |
ICST | 2 |
| 2022 | Multi-objective Optimization-based Bug-fixing Template Mining for Automated Program RepairabstractTemplate-based automatic program repair (T-APR) techniques depend on the quality of bug-fixing templates. For such templates to be of sufficient quality for T-APR techniques to succeed, they must satisfy three criteria: applicability, fixability, and efficiency. Existing template mining approaches select templates based only on the first criteria, and are thus suboptimal in their performance. This study proposes a multi-objective optimization-based bug-fixing template mining method for T-APR in which we estimate template quality based on nine code abstraction tasks and three objective functions. Our method determines the optimal code abstraction strategy (i.e., the optimal combination of abstraction tasks) which maximizes the values of three objective functions and generates a final set of bug-fixing templates by clustering template candidates to which the optimal abstraction strategy is applied. Our preliminary experiment demonstrated that our optimized strategy can improve templates’ applicability and efficiency by 7% and 146% over the existing mining technique, respectively. We therefore conclude that the multi-objective optimization-based template mining technique effectively finds high-quality bug-fixing templates. Misoo Kim, Youngkyoung Kim, Kicheol Kim, Eunseok Lee 0001 |
ASE | 2 |
| 2022 | An empirical study of deep transfer learning-based program repair for Kotlin projectsabstractDeep learning-based automated program repair (DL-APR) can automatically fix software bugs and has received significant attention in the industry because of its potential to significantly reduce software development and maintenance costs. The Samsung mobile experience (MX) team is currently switching from Java to Kotlin projects. This study reviews the application of DL-APR, which automatically fixes defects that arise during this switching process; however, the shortage of Kotlin defect-fixing datasets in Samsung MX team precludes us from fully utilizing the power of deep learning. Therefore, strategies are needed to effectively reuse the pretrained DL-APR model. This demand can be met using the Kotlin defect-fixing datasets constructed from industrial and open-source repositories, and transfer learning. This study aims to validate the performance of the pretrained DL-APR model in fixing defects in the Samsung Kotlin projects, then improve its performance by applying transfer learning. We show that transfer learning with open source and industrial Kotlin defect-fixing datasets can improve the defect-fixing performance of the existing DL-APR by 307%. Furthermore, we confirmed that the performance was improved by 532% compared with the baseline DL-APR model as a result of transferring the knowledge of an industrial (non-defect) bug-fixing dataset. We also discovered that the embedded vectors and overlapping code tokens of the code-change pairs are valuable features for selecting useful knowledge transfer instances by improving the performance of APR models by up to 696%. Our study demonstrates the possibility of applying transfer learning to practitioners who review the application of DL-APR to industrial software. Misoo Kim, Youngkyoung Kim, Hohyeon Jeong, Jinseok Heo, Sungoh Kim, Hyunhee Chung, Eunseok Lee 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2021 | A Novel Automatic Query Expansion with Word Embedding for IR-based Bug LocalizationabstractInformation retrieval-based bug localization (IRBL) aims at finding buggy files using a bug report as a query. IRBL performance is highly dependent on the query quality. To improve the query quality for IRBL, automatic query expansion (AQE) method has been proposed for identifying query-related terms from the first-retrieved source files. This approach inevitably depends on two determinant of post- retrieval results, the retrieval model and the initial query quality. We propose a novel word embedding-based AQE technique, WEQE, to avoid the heavy dependency of the current AQE approach. Word embedding model enables to fetch terms semantically related to a query by representing words in a vector space. Our method embeds the words from both the global corpus and project-specific-corpus. The initial query is extended by adding words semantically similar to it based on vector representations from our embedding model. We validated the effectiveness of WEQE by using 4,583 bug reports from seven projects, four IRBL models, and two em-bedding models. Our large-scale experimental results show that WEQE can improve the average precision for bug localization for at least 42% of all queries. Our expanded queries on the best IRBL model achieve a 6% higher mean average precision for bug localization than the initial query. Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001 |
ISSRE | 2 |
| 2021 | Denchmark: A Bug Benchmark of Deep Learning-related SoftwareabstractA growing interest in deep learning (DL) has instigated a concomitant rise in DL-related software (DLSW). Therefore, the importance of DLSW quality has emerged as a vital issue. Simultaneously, researchers have found DLSW more complicated than traditional SW and more difficult to debug owing to the black-box nature of DL. These studies indicate the necessity of automatic debugging techniques for DLSW. Although several validated debugging techniques exist for general SW, no such techniques exist for DLSW. There is no standard bug benchmark to validate these automatic debugging techniques. In this study, we introduce a novel bug benchmark for DLSW, Denchmark, consisting of 4,577 bug reports from 193 popular DLSW projects, collected through a systematic dataset construction process. These DLSW projects are further classified into eight categories: framework, platform, engine, compiler, tool, library, DL-based application, and others. All bug reports in Denchmark contain rich textual information and links with bug-fixing commits, as well as three levels of buggy entities, such as files, methods, and lines. Our dataset aims to provide an invaluable starting point for the automatic debugging techniques of DLSW. Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001 |
MSR | 2 |
| 2020 | Feature Combination to Alleviate Hubness Problem of Source Code Representation for Bug LocalizationabstractDeep learning-based bug localization (DLBL) can effectively reduce software maintenance costs. However, the inherent hub ness problem of the high-dimensional vector of the source code file used in DLBL leads to inaccurate bug localization. To solve this problem, we analyzed 10,359 defects and found that the call graph and flow of the program can distinguish buggy files from non-buggy files, and provide functional semantic information for bug localization. Based on our observations, we propose a feature combination to alleviate the hubness problem of the source file representation by using functional semantic information. Our proposed method models the functional semantics with the call graph and program flow based on the raw abstract syntax tree. We evaluated the effectiveness of the proposed approach on 19 widely used projects and conducted an ablation study. The experimental results show that the proposed method can improve the current approaches by 12 % to 45 %, with differentiating buggy files and non-buggy files. In our ablation study, functional information shows its significance as the absence of functional semantics deteriorates performance by 8.5 %. Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001 |
APSEC | 1 |