Misoo Kim

dblp:221/1657 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-8274-5457ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Production and test bug report classification based on transfer learning
Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001
Inf. Softw. Technol.1
2023 Improving Transformer-based Program Repair Model through False Behavior Diagnosis
abstract
Research on automated program repairs using transformer-based models has recently gained considerable attention.The comprehension of the erroneous behavior of a model enables the identification of its inherent capacity and provides insights for improvement.However, the current landscape of research on program repair models lacks an investigation of their false behavior.Thus, we propose a methodology for diagnosing and treating the false behaviors of transformer-based program repair models.Specifically, we propose 1) a behavior vector that quantifies the behavior of the model when it generates an output, 2) a behavior discriminator (BeDisc) that identifies false behaviors, and 3) two methods for false behavior treatment.Through a large-scale experiment on 55,562 instances employing four datasets and three models, the BeDisc exhibited a balanced accuracy of 86.6% for false behavior classification.The first treatment, namely, early abortion, successfully eliminated 60.4% of false behavior while preserving 97.4% repair accuracy.Furthermore, the second treatment, namely, masked bypassing, resulted in an average improvement of 40.5% in the top-1 repair accuracy.These experimental results demonstrated the importance of investigating false behaviors in program repair models.* corresponding author 1 We refer to these as bugs in a broad sense.
Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001
EMNLP2
2022 Tracking Down Misguiding Terms for Locating Bugs in Deep Learning-Based Software (Student Abstract)
abstract
Bugs in source files (SFs) may cause software malfunction, inconveniencing users and even leading to catastrophic accidents. Therefore, the bugs in SFs should be found and fixed quickly. However, from hundreds of candidate SFs, finding buggy SFs is tedious and time consuming. To lessen the burden on developers, deep learning-based bug localization (DLBL) tools can be utilized. Text terms in bug reports and SFs play an important role. However, some terms provide incorrect information and degrade bug localization performance. Therefore, those terms are defined here as "misguiding terms," and an explainable-artificial-intelligence-based identification method is proposed. The effectiveness of the proposed method for DLBL was investigated. When misguiding terms were removed, the mean average precision of the bug localization model improved by 33% on average.
Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001
AAAI2
2022 Systematic Analysis of Defect-Specific Code Abstraction for Neural Program Repair
abstract
Automated program repair(APR) is in the spotlight in academia and the field to reduce the time and cost of maintenance for developers. Recently, APR has continued to study based on deep-learning models to understand and learn how to fix software bugs. Text-to-Text Transfer Transformer(T5), which scored state-of-the-art in natural language processing benchmarks, also showed promising results on program repair in recent studies. In deep-learning-based program repair studies, studies commonly propose code abstraction techniques to avoid vocabulary problems and learn fine code transformation to generate bug-fixing patches. However, there is not enough systematic analysis of code abstraction according to each bug type in deep-learning-based program repair. Therefore, We leverage TFix, T5-based program repair, to evaluate how code abstraction techniques affect neural program repair. Our experimental results showed that defect-specific code abstraction achives a higher average BLEU score than the existing code abstraction technique in both T5 and multilingual-T5(mT5) model-based TFix results. Also, mT5 model-based TFix, which is applied defect-specific code abstraction, gets a higher BLEU score in 37 error types of 52 ESLint error types than TFix.
Kicheol Kim, Misoo Kim, Eunseok Lee 0001
APSEC2
2022 Impact of Defect Instances for Successful Deep Learning-based Automatic Program Repair
abstract
Deep learning-based automatic program repair (DL-APR) returns a patch code when given a defect code. Recent studies on DL-APR techniques have focused on the training phase to generate more accurate patches; however, a trained model cannot always generate an accurate patch for every new defect code, as the training dataset does not completely represent the new defects to be input in the future. DL-APR researchers should study a method to elicit the best performance on new inputs from the trained and deployed model. A new defect instance (i.e., defect codes and their context codes) is one of the crucial input data that determine the accuracy of the DL-APR, which can be changed and improved. We improve the quality of new input defect instances by focusing on the presence of noise tokens which compromise the defect instances’ quality, thus impairing the accuracy of generated patches. This paper shows that 1) there are noise tokens which prevent correct patch generation (inference) in a new defect instance, and 2) it is necessary to mask these noise tokens to avoid their usage in inferencing patch codes. In order to validate these two assertions, we use a state-of-the-art DL-APR technique and a genetic algorithm to generate near-optimal defect instances which maximize the patch generation accuracy (i.e., the BLEU score) of 4,573 defect instances. Based on optimization results, we found that 1) noise tokens impair patch generation accuracy in approximately 49% of instances, and 2) if these tokens are precluded from inference by masking them, we can improve patch generation accuracy by 88%. The results suggest that future work is required to automatically remove noise tokens from new defect instances so that the trained patch generator generates better patches.
Misoo Kim, Youngkyoung Kim, Jinseok Heo, Hohyeon Jeong, Sungoh Kim, Eunseok Lee 0001
ICSME1
2022 An Empirical Study of IR-based Bug Localization for Deep Learning-based Software
abstract
As the impact of deep-learning-based software (DLSW) increases, automatic debugging techniques for guaranteeing DLSW quality are becoming increasingly important. Information-retrieval-based bug localization (IRBL) techniques can aid in debugging by automatically localizing buggy entities (files and functions). The low-cost advantage of IRBL can alleviate the difficulty of identifying bug locations due to the complexity of DLSW. However, there are significant differences between DLSW and traditional software, and these differences lead to differences in search space and query quality for IRBL. That is, IRBL performance must be validated in DLSW. We empirically validated IRBL performance for DLSW from the following four perspectives: 1) similarity model, 2) query generation, 3) ranking model for buggy file localization, and 4) ranking model for buggy function localization. Based on four research questions and a large-scale experiment using 2,365 bug reports from 136 DLSW projects, we confirmed the salient char-acteristics of DLSW from the perspective of IRBL and derived four recommendations for practical IRBL usage in DLSW from the empirical results. Regarding IRBL performance, we validated that IRBL performance with the combination of bug-related features outperformed that of using only file similarity by 15 % and IRBL ranked buggy files and functions on average of 1.6th and 2.9th, respectively. Our study is valuable as a baseline for IRBL researchers and as a guideline for DLSW developers who wish to apply IRBL to ensure DLSW quality.
Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001
ICST1
2022 Multi-objective Optimization-based Bug-fixing Template Mining for Automated Program Repair
abstract
Template-based automatic program repair (T-APR) techniques depend on the quality of bug-fixing templates. For such templates to be of sufficient quality for T-APR techniques to succeed, they must satisfy three criteria: applicability, fixability, and efficiency. Existing template mining approaches select templates based only on the first criteria, and are thus suboptimal in their performance. This study proposes a multi-objective optimization-based bug-fixing template mining method for T-APR in which we estimate template quality based on nine code abstraction tasks and three objective functions. Our method determines the optimal code abstraction strategy (i.e., the optimal combination of abstraction tasks) which maximizes the values of three objective functions and generates a final set of bug-fixing templates by clustering template candidates to which the optimal abstraction strategy is applied. Our preliminary experiment demonstrated that our optimized strategy can improve templates’ applicability and efficiency by 7% and 146% over the existing mining technique, respectively. We therefore conclude that the multi-objective optimization-based template mining technique effectively finds high-quality bug-fixing templates.
Misoo Kim, Youngkyoung Kim, Kicheol Kim, Eunseok Lee 0001
ASE1
2022 ECench: An Energy Bug Benchmark of Ethereum Client Software
abstract
With the introduction of smart contacts, Ethereum has become one of the most popular blockchain networks. In the wake of its popularity, an increasing number of Ethereum-based software have been developed. However, the carbon emissions resulting from these software has been pointed out as a global issue. It is necessary to reduce the energy consumed by these software to reduce carbon emissions. Recently, most studies have focused on smart contracts and proposed energy-efficient methods for the development of carbon friendly Ethereum networks. However, in addition to smart contracts, the energy used by client software in Ethereum networks should also be reviewed. This is because the client software performs all functions occurring in the Ethereum network, including smart contracts. Therefore, energy bugs that waste energy in Ethereum client software should be investigated and solved. The first task to enable this is to build an energy bug benchmark of Ethereum client software. This study introduces ECench, an energy bug benchmark of Ethereum client software. ECench includes 507 energy buggy commits from 7 series of client software that are officially operated in the Ethereum network. We carefully collected and manually reviewed them for cleaner commits. A key strength of our benchmark is that it provides eight energy wastage categories, which can serve as a cornerstone for researchers to identify energy waste codes. ECench can provide a valuable starting point for studies on energy reduction and carbon reduction in Ethereum.
Misoo Kim, Eunseok Lee 0001
MSR2
2022 An empirical study of deep transfer learning-based program repair for Kotlin projects
abstract
Deep learning-based automated program repair (DL-APR) can automatically fix software bugs and has received significant attention in the industry because of its potential to significantly reduce software development and maintenance costs. The Samsung mobile experience (MX) team is currently switching from Java to Kotlin projects. This study reviews the application of DL-APR, which automatically fixes defects that arise during this switching process; however, the shortage of Kotlin defect-fixing datasets in Samsung MX team precludes us from fully utilizing the power of deep learning. Therefore, strategies are needed to effectively reuse the pretrained DL-APR model. This demand can be met using the Kotlin defect-fixing datasets constructed from industrial and open-source repositories, and transfer learning. This study aims to validate the performance of the pretrained DL-APR model in fixing defects in the Samsung Kotlin projects, then improve its performance by applying transfer learning. We show that transfer learning with open source and industrial Kotlin defect-fixing datasets can improve the defect-fixing performance of the existing DL-APR by 307%. Furthermore, we confirmed that the performance was improved by 532% compared with the baseline DL-APR model as a result of transferring the knowledge of an industrial (non-defect) bug-fixing dataset. We also discovered that the embedded vectors and overlapping code tokens of the code-change pairs are valuable features for selecting useful knowledge transfer instances by improving the performance of APR models by up to 696%. Our study demonstrates the possibility of applying transfer learning to practitioners who review the application of DL-APR to industrial software.
Misoo Kim, Youngkyoung Kim, Hohyeon Jeong, Jinseok Heo, Sungoh Kim, Hyunhee Chung, Eunseok Lee 0001
ESEC/SIGSOFT FSE1
2021 A Novel Automatic Query Expansion with Word Embedding for IR-based Bug Localization
abstract
Information retrieval-based bug localization (IRBL) aims at finding buggy files using a bug report as a query. IRBL performance is highly dependent on the query quality. To improve the query quality for IRBL, automatic query expansion (AQE) method has been proposed for identifying query-related terms from the first-retrieved source files. This approach inevitably depends on two determinant of post- retrieval results, the retrieval model and the initial query quality. We propose a novel word embedding-based AQE technique, WEQE, to avoid the heavy dependency of the current AQE approach. Word embedding model enables to fetch terms semantically related to a query by representing words in a vector space. Our method embeds the words from both the global corpus and project-specific-corpus. The initial query is extended by adding words semantically similar to it based on vector representations from our embedding model. We validated the effectiveness of WEQE by using 4,583 bug reports from seven projects, four IRBL models, and two em-bedding models. Our large-scale experimental results show that WEQE can improve the average precision for bug localization for at least 42% of all queries. Our expanded queries on the best IRBL model achieve a 6% higher mean average precision for bug localization than the initial query.
Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001
ISSRE1
2021 Denchmark: A Bug Benchmark of Deep Learning-related Software
abstract
A growing interest in deep learning (DL) has instigated a concomitant rise in DL-related software (DLSW). Therefore, the importance of DLSW quality has emerged as a vital issue. Simultaneously, researchers have found DLSW more complicated than traditional SW and more difficult to debug owing to the black-box nature of DL. These studies indicate the necessity of automatic debugging techniques for DLSW. Although several validated debugging techniques exist for general SW, no such techniques exist for DLSW. There is no standard bug benchmark to validate these automatic debugging techniques. In this study, we introduce a novel bug benchmark for DLSW, Denchmark, consisting of 4,577 bug reports from 193 popular DLSW projects, collected through a systematic dataset construction process. These DLSW projects are further classified into eight categories: framework, platform, engine, compiler, tool, library, DL-based application, and others. All bug reports in Denchmark contain rich textual information and links with bug-fixing commits, as well as three levels of buggy entities, such as files, methods, and lines. Our dataset aims to provide an invaluable starting point for the automatic debugging techniques of DLSW.
Misoo Kim, Youngkyoung Kim, Eunseok Lee 0001
MSR1
2021 Are datasets for information retrieval-based bug localization techniques trustworthy?
Misoo Kim, Eunseok Lee 0001
Empir. Softw. Eng.1
2020 Feature Combination to Alleviate Hubness Problem of Source Code Representation for Bug Localization
abstract
Deep learning-based bug localization (DLBL) can effectively reduce software maintenance costs. However, the inherent hub ness problem of the high-dimensional vector of the source code file used in DLBL leads to inaccurate bug localization. To solve this problem, we analyzed 10,359 defects and found that the call graph and flow of the program can distinguish buggy files from non-buggy files, and provide functional semantic information for bug localization. Based on our observations, we propose a feature combination to alleviate the hubness problem of the source file representation by using functional semantic information. Our proposed method models the functional semantics with the call graph and program flow based on the raw abstract syntax tree. We evaluated the effectiveness of the proposed approach on 19 widely used projects and conducted an ablation study. The experimental results show that the proposed method can improve the current approaches by 12 % to 45 %, with differentiating buggy files and non-buggy files. In our ablation study, functional information shows its significance as the absence of functional semantics deteriorates performance by 8.5 %.
Youngkyoung Kim, Misoo Kim, Eunseok Lee 0001
APSEC2
2020 ManQ: Many-objective optimization-based automatic query reduction for IR-based bug localization
Misoo Kim, Eunseok Lee 0001
Inf. Softw. Technol.1