VLDB 2026 Research / reviewers in the wild / expert
Jaechang Nam
dblp:86/982
· DBLP profile ↗
17ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-1678-2185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EnCus: Customizing Search Space for Automated Program RepairabstractThe primary challenge faced by Automated Program Repair (APR) techniques in fixing buggy programs is the search space problem. To generate a patch, APR techniques must address three critical decisions: where to fix (location), how to fix (operation), and what to fix with (ingredient). In this study, we propose EnCus, a novel approach that customizes the search space of ingredients and mutation operators during patch generation. EnCus acts as an APR wingman, using an ensemble-based strategy to customize the search space. The search space is customized by extracting edit operations that are used to fix similar bug-introducing changes from existing patches. EnCus applies an ensemble of edit operations extracted from three open source project pools and three Abstract Syntax Tree (AST)-level code differencing tools. This ensemble provides complementary perspectives on the buggy context. To evaluate this approach, we integrate EnCus to an existing context-based APR tool, ConFix. Using EnCus, the extensive search space of ConFix is reduced to ten recommended patches. EnCus was evaluated on single-line Defects4J bugs, successfully generating 20 correct patches which performs comparably to state-of-the-art context-based APR techniques. Seongbin Kim, Sechang Jang, Jindae Kim 0001, Jaechang Nam |
ICST | 4 |
| 2025 | Pre-trained Models for Bytecode InstructionsabstractRecent advancements in pre-trained models have rapidly expanded their applicability to various software engineering challenges. Despite this progress, current research predominantly focuses on source code and natural language processing, largely overlooking Java bytecode. Java bytecode, with its well-defined structure and high availability, presents a promising yet under-explored domain for leveraging pre-trained models. Its inherent properties, such as platform independence and optimized performance, make Java bytecode an ideal candidate for developing robust and efficient software engineering solutions. Addressing this gap could unlock new opportunities for enhancing automated program analysis, bug detection, and code generation tasks. In this study, we propose byteT5 and byteBERT, which are pre-trained models with hexadecimal bytecode. To build our models, we developed a bytecode tokenizer, ByteTok, to generate hexadecimal input representations for our pre-trained models. We conduct an empirical study comparing our models and GPT-40. The results indicate that byteT5 and byteBERT outperform GPT-40 in the span masking task. We anticipate these findings will pave the way for novel approaches to addressing various software engineering challenges, particularly live patching. Heeyoul Choi, Jaechang Nam |
ICST | 6 |
| 2023 | An Empirical Study on the Stability of Explainable Software Defect PredictionabstractExplaining the results of software defect prediction (SDP) models is practical but challenging. Jiarpakdee et al. proposed using two model-agnostic techniques (i.e., LIME and BreakDown) to explain prediction results. They showed that model-agnostic techniques can achieve remarkable performance and that the generated explanations can assist developers in understanding the prediction results. However, the fact that they examined these model-agnostic techniques only under a specific SDP setting calls into question their reliability on SDP models under various settings. In this paper, we set out to investigate the reliability and stability of model-agnostic-based explanation generation approaches on SDP models under different settings, e.g., different data sampling techniques, machine learning classifiers, and prediction scenarios used when building SDP models. We use model-agnostic techniques to generate explanations for the same instance under various SDP models with different settings and then check the stability of the generated explanations for the instance. We reused the same defect data and experiment configurations from Jiarpakdee et al. in our experiments. The results show that the examined model-agnostic techniques generate inconsistent explanations under different SDP settings for the same test instances. Our user case study further confirms that inconsistent explanations can significantly affect developers' understanding of the prediction results, which implies that the model-agnostic techniques can be unreliable for practical explanation generation under different SDP settings. To conclude, we urge a revisit of existing model-agnostic-based studies in software engineering and call for more research in explainable SDP toward achieving stable explanation generation. Reem Aleithan, Jaechang Nam, Junjie Wang 0001, Nima Shiri Harzevili, Song Wang 0009 |
APSEC | 3 |
| 2023 | WINE: Warning miner for improving bug finders
Yoon-Ho Choi, Jaechang Nam |
Inf. Softw. Technol. | 2 |
| 2022 | On the Naturalness of Bytecode InstructionsabstractBytecode is used in software analysis and other approaches due to its advantages such as high availability and simple specification. Therefore, to leverage these advantages in training language models with bytecode, it is important to clearly recognize the characteristics of the naturalness of bytecode. However, the naturalness of bytecode has not been actively explored. Yoon-Ho Choi, Jaechang Nam |
ASE | 2 |
| 2021 | Continuous Software Bug PredictionabstractBackground: Many software bug prediction models have been proposed and evaluated on a set of well-known benchmark datasets. We conducted pilot studies on the widely used benchmark datasets and observed common issues among them. Specifically, most of existing benchmark datasets consist of randomly selected historical versions of software projects, which poses non-trivial threats to the validity of existing bug prediction studies since the real-world software projects often evolve continuously. Yet how to conduct software bug prediction in the real-world continuous software development scenarios is not well studied. Song Wang 0009, Junjie Wang 0001, Jaechang Nam, Nachiappan Nagappan |
ESEM | 3 |
| 2020 | Deep Semantic Feature Learning for Software Defect PredictionabstractSoftware defect prediction, which predicts defective code regions, can assist developers in finding bugs and prioritizing their testing efforts. Traditional defect prediction features often fail to capture the semantic differences between different programs. This degrades the performance of the prediction models built on these traditional features. Thus, the capability to capture the semantics in programs is required to build accurate prediction models. To bridge the gap between semantics and defect prediction features, we propose leveraging a powerful representation-learning algorithm, deep learning, to learn the semantic representations of programs automatically from source code files and code changes. Specifically, we leverage a deep belief network (DBN) to automatically learn semantic features using token vectors extracted from the programs' abstract syntax trees (AST) (for file-level defect prediction models) and source code changes (for change-level defect prediction models). We examine the effectiveness of our approach on two file-level defect prediction tasks (i.e., file-level within-project defect prediction and file-level cross-project defect prediction) and two change-level defect prediction tasks (i.e., change-level within-project defect prediction and change-level cross-project defect prediction). Our experimental results indicate that the DBN-based semantic features can significantly improve the examined defect prediction tasks. Specifically, the improvements of semantic features against existing traditional features (in F1) range from 2.1 to 41.9 percentage points for file-level within-project defect prediction, from 1.5 to 13.4 percentage points for file-level cross-project defect prediction, from 1.0 to 8.6 percentage points for change-level within-project defect prediction, and from 0.6 to 9.9 percentage points for change-level cross-project defect prediction. Song Wang 0009, Taiyue Liu, Jaechang Nam, Lin Tan 0001 |
IEEE Trans. Software Eng. | 3 |
| 2019 | A bug finder refined by a large set of open-source projects
Jaechang Nam, Song Wang 0009, Yuan Xi, Lin Tan 0001 |
Inf. Softw. Technol. | 1 |
| 2018 | Heterogeneous Defect PredictionabstractMany recent studies have documented the success of cross-project defect prediction (CPDP) to predict defects for new projects lacking in defect data by using prediction models built by other projects. However, most studies share the same limitations: it requires homogeneous data; i.e., different projects must describe themselves using the same metrics. This paper presents methods for heterogeneous defect prediction (HDP) that matches up different metrics in different projects. Metric matching for HDP requires a “large enough” sample of distributions in the source and target projects-which raises the question on how large is “large enough” for effective heterogeneous defect prediction. This paper shows that empirically and theoretically, “large enough” may be very small indeed. For example, using a mathematical model of defect prediction, we identify categories of data sets were as few as 50 instances are enough to build a defect prediction model. Our conclusion for this work is that, even when projects use different metric sets, it is possible to quickly transfer lessons learned about defect prediction. Jaechang Nam, Wei Fu 0002, Sunghun Kim 0001, Tim Menzies, Lin Tan 0001 |
IEEE Trans. Software Eng. | 1 |
| 2017 | QTEP: quality-aware test case prioritizationabstractTest case prioritization (TCP) is a practical activity in software testing for exposing faults earlier. Researchers have proposed many TCP techniques to reorder test cases. Among them, coverage-based TCPs have been widely investigated. Specifically, coverage-based TCP approaches leverage coverage information between source code and test cases, i.e., static code coverage and dynamic code coverage, to schedule test cases. Existing coverage-based TCP techniques mainly focus on maximizing coverage while often do not consider the likely distribution of faults in source code. However, software faults are not often equally distributed in source code, e.g., around 80% faults are located in about 20% source code. Intuitively, test cases that cover the faulty source code should have higher priorities, since they are more likely to find faults. Song Wang 0009, Jaechang Nam, Lin Tan 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2016 | Developer Micro Interaction Metrics for Software Defect PredictionabstractTo facilitate software quality assurance, defect prediction metrics, such as source code metrics, change churns, and the number of previous defects, have been actively studied. Despite the common understanding that developer behavioral interaction patterns can affect software quality, these widely used defect prediction metrics do not consider developer behavior. We therefore propose micro interaction metrics (MIMs), which are metrics that leverage developer interaction information. The developer interactions, such as file editing and browsing events in task sessions, are captured and stored as information by Mylyn, an Eclipse plug-in. Our experimental evaluation demonstrates that MIMs significantly improve overall defect prediction accuracy when combined with existing software measures, perform well in a cost-effective manner, and provide intuitive feedback that enables developers to recognize their own inefficient behaviors during software development. Taek Lee, Jaechang Nam, DongGyun Han, Sunghun Kim 0001, Hoh Peter In |
IEEE Trans. Software Eng. | 2 |
| 2015 | CLAMI: Defect Prediction on Unlabeled Datasets (T)abstractDefect prediction on new projects or projects with limited historical data is an interesting problem in software engineering. This is largely because it is difficult to collect defect information to label a dataset for training a prediction model. Cross-project defect prediction (CPDP) has tried to address this problem by reusing prediction models built by other projects that have enough historical data. However, CPDP does not always build a strong prediction model because of the different distributions among datasets. Approaches for defect prediction on unlabeled datasets have also tried to address the problem by adopting unsupervised learning but it has one major limitation, the necessity for manual effort. In this study, we propose novel approaches, CLA and CLAMI, that show the potential for defect prediction on unlabeled datasets in an automated manner without need for manual effort. The key idea of the CLA and CLAMI approaches is to label an unlabeled dataset by using the magnitude of metric values. In our empirical study on seven open-source projects, the CLAMI approach led to the promising prediction performances, 0.636 and 0.723 in average f-measure and AUC, that are comparable to those of defect prediction based on supervised learning. Jaechang Nam, Sunghun Kim 0001 |
ASE | 1 |
| 2015 | REMI: defect prediction for efficient API testingabstractQuality assurance for common APIs is important since the the reliability of APIs affects the quality of other systems using the APIs. Testing is a common practice to ensure the quality of APIs, but it is a challenging and laborious task especially for industrial projects. Due to a large number of APIs with tight time constraints and limited resources, it is hard to write enough test cases for all APIs. To address these challenges, we present a novel technique, REMI that predicts high risk APIs in terms of producing potential bugs. REMI allows developers to write more test cases for the high risk APIs. We evaluate REMI on a real-world industrial project, Tizen-wearable, and apply REMI to the API development process at Samsung Electronics. Our evaluation results show that REMI predicts the bug-prone APIs with reasonable accuracy (0.681 f-measure on average). The results also show that applying REMI to the Tizen-wearable development process increases the number of bugs detected, and reduces the resources required for executing test cases. Mijung Kim, Jaechang Nam, Jaehyuk Yeon, Soonhwang Choi, Sunghun Kim 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2015 | Heterogeneous defect predictionabstractSoftware defect prediction is one of the most active research areas in software engineering. We can build a prediction model with defect data collected from a software project and predict defects in the same project, i.e. within-project defect prediction (WPDP). Researchers also proposed cross-project defect prediction (CPDP) to predict defects for new projects lacking in defect data by using prediction models built by other projects. In recent studies, CPDP is proved to be feasible. However, CPDP requires projects that have the same metric set, meaning the metric sets should be identical between projects. As a result, current techniques for CPDP are difficult to apply across projects with heterogeneous metric sets. To address the limitation, we propose heterogeneous defect prediction (HDP) to predict defects across projects with heterogeneous metric sets. Our HDP approach conducts metric selection and metric matching to build a prediction model between projects with heterogeneous metric sets. Our empirical study on 28 subjects shows that about 68% of predictions using our approach outperform or are comparable to WPDP with statistical significance. Jaechang Nam, Sunghun Kim 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2013 | Automatic patch generation learned from human-written patchesabstractPatch generation is an essential software maintenance task because most software systems inevitably have bugs that need to be fixed. Unfortunately, human resources are often insufficient to fix all reported and known bugs. To address this issue, several automated patch generation techniques have been proposed. In particular, a genetic-programming-based patch generation technique, GenProg, proposed by Weimer et al., has shown promising results. However, these techniques can generate nonsensical patches due to the randomness of their mutation operations. To address this limitation, we propose a novel patch generation approach, Pattern-based Automatic program Repair (Par), using fix patterns learned from existing human-written patches. We manually inspected more than 60,000 human-written patches and found there are several common fix patterns. Our approach leverages these fix patterns to generate program patches automatically. We experimentally evaluated Par on 119 real bugs. In addition, a user study involving 89 students and 164 developers confirmed that patches generated by our approach are more acceptable than those generated by GenProg. Par successfully generated patches for 27 out of 119 bugs, while GenProg was successful for only 16 bugs. Dongsun Kim 0001, Jaechang Nam, Jaewoo Song, Sunghun Kim 0001 |
ICSE | 2 |
| 2013 | Transfer defect learningabstractMany software defect prediction approaches have been proposed and most are effective in within-project prediction settings. However, for new projects or projects with limited training data, it is desirable to learn a prediction model by using sufficient training data from existing source projects and then apply the model to some target projects (cross-project defect prediction). Unfortunately, the performance of cross-project defect prediction is generally poor, largely because of feature distribution differences between the source and target projects. In this paper, we apply a state-of-the-art transfer learning approach, TCA, to make feature distributions in source and target projects similar. In addition, we propose a novel transfer defect learning approach, TCA+, by extending TCA. Our experimental results for eight open-source projects show that TCA+ significantly improves cross-project prediction performance. Jaechang Nam, Sinno Jialin Pan, Sunghun Kim 0001 |
ICSE | 1 |
| 2011 | Micro interaction metrics for defect predictionabstractThere is a common belief that developers' behavioral interaction patterns may affect software quality. However, widely used defect prediction metrics such as source code metrics, change churns, and the number of previous defects do not capture developers' direct interactions. We propose 56 novel micro interaction metrics (MIMs) that leverage developers' interaction information stored in the Mylyn data. Mylyn is an Eclipse plug-in, which captures developers' interactions such as file editing and selection events with time spent. To evaluate the performance of MIMs in defect prediction, we build defect prediction (classification and regression) models using MIMs, traditional metrics, and their combinations. Our experimental results show that MIMs significantly improve defect classification and regression accuracy. Taek Lee, Jaechang Nam, DongGyun Han, Sunghun Kim 0001, Hoh Peter In |
SIGSOFT FSE | 2 |