VLDB 2026 Research / reviewers in the wild / expert
Scott Uk-Jin Lee
dblp:71/2905
· DBLP profile ↗
16ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-8457-3097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 2 first-author · 6 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Diffusion-Driven Test Input Generation using Large Language ModelabstractDeep learning-based image classifiers have achieved remarkable accuracy on standard benchmarks, yet they can fail on unusual or unseen inputs with potentially serious consequences in real-world applications. In this study, we present a test input generation pipeline that leverages a large language model (LLM) and a diffusion-based image generator to produce realistic, diverse test images for classifier evaluation. In our approach, an LLM autonomously produces semantically rich textual prompts describing plausible yet challenging instances of a target class, and a diffusion model then renders each prompt into a high-fidelity image. Unlike static approaches with fixed human-written prompts, this dynamic prompting strategy automatically explores a broader range of edge case scenarios without manual intervention. Experiments show that LLM-driven prompts significantly increase input diversity and uncover more classifier misclassifications compared to static prompting. Human evaluators also judged the images generated by our approach to be more realistic and representative of the target class. These results demonstrate that the proposed approach effectively exposes model blind spots and can enhance the robustness of image classifiers. Joonwoo Lee, Seungho Kim, Scott Uk-Jin Lee |
APSEC | 3 |
| 2025 | Can We Trust the Actionable Guidance from Explainable AI Techniques in Defect Prediction?abstractDespite advances in high-performance Software Defect Prediction (SDP) models, practitioners remain hesitant to adopt them due to opaque decision-making and a lack of actionable insights. Recent research has applied various explainable AI (XAI) techniques to provide explainable and actionable guidance for SDP results to address these limitations, but the trustworthiness of such guidance for practitioners has not been sufficiently investigated. Practitioners may question the feasibility of implementing the proposed changes, and if these changes fail to resolve predicted defects or prove inaccurate, their trust in the guidance may diminish. In this study, we empirically evaluate the effectiveness of current XAI approaches for SDP across 32 releases of 9 large-scale projects, focusing on whether the guidance meets practitioners' expectations. Our findings reveal that their actionable guidance (i) does not guarantee that predicted defects are resolved; (ii) fails to pinpoint modifications required to resolve predicted defects; and (iii) deviates from the typical code changes practitioners make in their projects. These limitations indicate that the guidance is not yet reliable enough for developers to justify investing their limited debugging resources. We suggest that future XAI research for SDP incorporate feedback loops that offer clear rewards for practitioners' efforts, and propose a potential alternative approach utilizing counterfactual explanations. Gichan Lee, Hansae Ju, Scott Uk-Jin Lee |
SANER | 3 |
| 2024 | Less is More: An Empirical Study of Undersampling Techniques for Technical Debt Prediction
Gichan Lee, Scott Uk-Jin Lee |
ICECCS | 2 |
| 2024 | NeuroJIT: Improving Just-In-Time Defect Prediction Using Neurophysiological and Empirical Perceptions of Modern DevelopersabstractModern developers make new changes based on their understanding of the existing code context and review these changes by analyzing the modified code and its context (i.e., commits). If commits are difficult to comprehend, the likelihood of human errors increases, making it harder for practitioners to identify commits that might introduce unintended defects. Nevertheless, research on predicting defect-inducing commits based on the difficulty of understanding them has been limited. In this study, we present a novel approach NeuroJIT, that leverages the correlation between modern developers' neurophysiological and empirical reactions to different code segments and their code characteristics to find the features that can capture the understandability of each commit. We investigate the understandability features of NeuroJIT in three key aspects: (i) their correlation with defect-inducing risks; (ii) their differences from widely adopted features used to predict these risks; and (iii) whether they can improve the performance of just-in-time defect prediction models. Based on our findings, we conclude that neurophysiological and empirical understandability of commits can be a competitive predictor and provide more actionable guidance from a unique perspective on defect-inducing commits. Gichan Lee, Hansae Ju, Scott Uk-Jin Lee |
ASE | 3 |
| 2023 | An Empirical Comparison of Model-Agnostic Techniques for Defect Prediction ModelsabstractRecently, software defect prediction studies have attempted to make black-box defect prediction models explainable and actionable. State-of-the-art defect prediction studies have utilized various model-agnostic techniques derived from the explainable AI domain as key tools to make the predictions easier for practitioners to understand. However, it has not been sufficiently investigated whether there is inconsistent information within local explanations generated by different model-agnostic techniques when interpreting a defect prediction. If local explanations generated by heterogeneous model-agnostic techniques consist of different information, the derivable insights to understand and act upon defect predictions becomes less viable and it may cause ineffective or even incorrect defect corrections. In this research, we empirically analyzed 323,844 local explanations generated by three different model-agnostic techniques: (1) LIME, (2) SHAP, and (3) BreakDown. These local explanations were analyzed in terms of how the contributions of features were distributed, how the contributions were ranked, and whether the contributions were contradictory. We concluded that (i) different model-agnostic techniques provide practitioners with local explanations where average contributions of the top-ranked features are different; (ii) different model-agnostic techniques provide practitioners with local explanations consisting of different contribution rankings and inconsistent contribution directions of top-ranked features. Therefore, we recommend that practitioners should avoid using model-agnostic techniques interchangeably and must perform a multi-faceted manual validation when planning actions based on the local explanations. Gichan Lee, Scott Uk-Jin Lee |
SANER | 2 |
| 2023 | A systematic literature review on Android-specific smells
Zhiqiang Wu 0004, Xin Chen 0064, Scott Uk-Jin Lee |
J. Syst. Softw. | 3 |
| 2022 | Fast Automated Abstract Machine Repair Using Simultaneous Modifications and RefactoringabstractAutomated model repair techniques enable machines to synthesise patches that ensure models meet given requirements. B-repair, which is an existing model repair approach, assists users in repairing erroneous models in the B formal method, but repairing large models is inefficient due to successive applications of repair. In this work, we improve the performance of B-repair using simultaneous modifications, repair refactoring, and better classifiers. The simultaneous modifications can eliminate multiple invariant violations at a time so the average time to repair each fault can be reduced. Further, the modifications can be refactored to reduce the length of repair. The purpose of using better classifiers is to perform more accurate and general repairs and avoid inefficient brute-force searches. We conducted an empirical study to demonstrate that the improved implementation leads to the entire model process achieving higher accuracy, generality, and efficiency. Jing Sun 0002, Gillian Dobbie, Hadrien Bride, Jin Song Dong 0001, Scott Uk-Jin Lee |
Formal Aspects Comput. | 7 |
| 2020 | Integrated Formal Tools for Software Architecture Smell DetectionabstractThe architecture smells are the poor design practices applied to the software architecture design. The smells in software architecture design can be cascaded to cause the issues in the system implementation and significantly affect the maintainability and reliability attribute of the software system. The prevention of architecture smells at the design phase can therefore improve the overall quality of the software system. This paper presents a framework that supports the detection of architecture smells based on the formalization of architecture design. Our modeling specification supports representing both structural and behavioral aspect of software architecture design; it allows the smells to be analyzed and detected with the provided tools. Our framework has been applied to seven architecture smells that violate different design principles. The evaluation has been conducted and the result shows that our detection approach gives accurate results and performs well on different size of models. With the proposed framework, other architecture smells can be defined and detected using the process and tools presented in this paper. Nacha Chondamrongkul, Jing Sun 0002, Ian Warren, Scott Uk-Jin Lee |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2019 | Achieving Abstract Machine Reachability with Learning-Based Model FulfilmentabstractThis paper proposes a probabilistic reachability repair solution that enables abstract machines to automatically evolve and satisfy desired requirements. The solution is a combination of the B-method, machine learning and program synthesis. The B-method is used to formally specify an abstract machine and analyse the reachability of the abstract machine. Machine learning models are used to approximate features hidden in the semantics of the abstract machine. When the abstract machine fails to reach a desired state, the machine learning models are used to discover missing transitions to the state. Inserting the discovered transitions into the original abstract machine will lead to a repaired abstract machine that is capable of achieving the state. To obtain the repaired abstract machine, a set of insertion repairs are synthesised from the discovered transitions and are simplified using context-free grammars. Experimental results reveal that the reachability repair solution is applicable to a wide range of abstract machines and can accurately discover transitions that satisfy the requirements of reachability. Moreover, the results demonstrate that random forests are efficient machine learning models on transition discovery tasks. Additionally, we argue that the automated reachability repair process can improve the efficiency of software development. Jing Sun 0002, Gillian Dobbie, Scott Uk-Jin Lee |
APSEC | 4 |
| 2014 | Predicting Student Blood Pressure by Support Vector Machine Using FacebookabstractPredicting human blood pressure (B.P) is an important aspect of primary emotion using Facebook has not yet been investigated. Primary emotions help a person to express her/his feelings, thoughts and understanding the importance of social connections using Facebook. Facebook contribute rich environment of primary emotion and famous social site having collection of information that concerned with primary emotions. The well-known machine learning approaches have known as novel methods for doing prediction using SNS. Support Vector Machine (SVM) has recently been a strong machine learning and data mining tool. Our article, it is used to predict human BP. The dataset contain primary emotion and blood pressure that are collected using Facebook post that consists of formal text from come forward of hanyang university student. Current human B.P and those belonging up to six previous primary emotions and B.P values with respect to human emotion are given as input variables, while the blood pressure used as output parameter. The outcome shows that SVM can be prosperously applied for prediction of B.P through primary emotion. On the contrary, validations signify that the error statistics of SVM model marginally outperforms. Shazada Muhammad Umair Khan, Javeria Shaikh Manzoor, Scott Uk-Jin Lee |
SERVICES | 3 |
| 2011 | MELO 2011 - 1st Workshop on Model-Driven Engineering, Logic and Optimization
Jordi Cabot, Patrick Albert, Grégoire Dupé, Marcos Didonet Del Fabro, Scott Uk-Jin Lee |
ECMFA | 5 |
| 2011 | An MDE-Based Approach for Solving Configuration Problems: An Application to the Eclipse Platform
Guillaume Doux, Patrick Albert, Gabriel Barbier, Jordi Cabot, Marcos Didonet Del Fabro, Scott Uk-Jin Lee |
ECMFA | 6 |
| 2010 | Theorem prover approach to semistructured data design
Scott Uk-Jin Lee, Gillian Dobbie, Jing Sun 0002, Lindsay Groves |
Formal Methods Syst. Des. | 1 |
| 2009 | Verifying Semistructured Data Normalization Using SWRLabstractSemistructured data has become more and more prominent in the fast growing areas of web information technology. XML has been used as a standard format for semistructured data in representing and exchanging information in various applications. However, the lack of formality and verification support in the design of a good semistructured data model may hinder its development. For example, redundant data in XML must be removed or minimized to avoid inconsistent and inefficient information processing. Normalization algorithms have been developed to overcome these problems by transforming the schema of a semistructured document into a better form. Therefore, it is essential to ensure that a transformed schema model preserves the same information that its original form holds. In this paper, we present an approach to investigate and verify the no-data-loss property of semistructured data normalization. We encode the verification criteria in the Semantic Web Rule Language (SWRL) and make use of its ontology reasoning engine to provide automated support for the checking process. In summary, our approach not only investigates the information preserving aspect of semistructured data normalization, but also provides a scalable and automated solution towards the problem. Yuan-Fang Li, Jing Sun 0002, Gillian Dobbie, Scott Uk-Jin Lee, Hai H. Wang |
TASE | 4 |
| 2008 | Verifying Semistructured Data Normalization Using PVSabstractThe dramatic expansion of semistructured data has led to the development of database systems for manipulating the data. Despite its huge potential, there is still a lack of formality and verification support in the design of good semistructured databases. Like traditional database systems, developed semistructured database systems should contain minimal redundancies and update anomalies, in order to store and manage the data effectively. Several normalization algorithms have been proposed to satisfy these needs, by transforming the schema of the semistructured data into a better form. It is essential to ensure that the normalized schema remains semantically equivalent to its original form. In this paper, we present tool support for reasoning about the correctness of semistructured data normalization. The proposed approach uses the ORA-SS data modeling notation and defines its correctness criteria and rules in the PVS formal language. It further utilizes the PVS theorem prover to perform automated checking on the normalized schema, checking that functional dependencies are preserved, no data is lost and no spurious data is created. In summary, our approach not only investigates the characteristics of semistructured data normalization, but also provides a scalable and automated first step towards reasoning about the correctness of normalization algorithms on semistructured data. Scott Uk-Jin Lee, Jing Sun 0002, Gillian Dobbie, Lindsay Groves |
ICECCS | 1 |
| 2006 | A PVS Approach to Verifying ORA-SS Data Models
Scott Uk-Jin Lee, Gillian Dobbie, Jing Sun 0002, Lindsay Groves |
SEKE | 1 |