Myoungkyu Song

dblp:85/2806 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0003-4477-8933ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 9 since 2021
YearPublicationVenuePosition
2025 SYNC: SYnergistic aNnotation Collaboration between Humans and LLMs for Enhanced Model Training
abstract
Large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing tasks, highlighting their potential as effective data annotators. While LLM-generated annotations tend to be costeffective, they are often error-prone and may inadvertently introduce bias. It is advantageous to harness the strengths of both LLMs and humans to ensure higher accuracy and reliability in the annotation process. In this paper, we present a multi-step, human-LLM collaborative approach to optimize data annotation for Stack Overflow datasets. We begin by applying TF-IDF to rank and prioritize relevant elements. Next, we utilize NLP Transformer and UniXcoder to leverage their deep contextual understanding for handling the code-related queries and discussions typical of Stack Overflow, resulting in more consistent automated labels. Finally, human annotators re-annotate to correct potential errors and mitigate bias introduced during earlier stages. To support human-LLM collaboration, we developed SYNC a research prototype that implements SYnergistic aNnotation Collaboration through an intuitive graphical user interface, enabling real-time interaction between human annotators and the LLM, allowing for iterative refinements. Overall, our approach is designed to integrate automated efficiency with human oversight to improve annotation outcomes, particularly for complex or domain-specific tasks such as those found in Stack Overflow datasets.
Tammy Le, Will Taylor, Shradha Maharjan, Myoungkyu Song
SERA5
2025 Intelligent Code Completion by a Unified Multi-task Learning with a Large Language Model
abstract
Code completion has become an essential tool in modern software development. It helps developers by predicting the next token (e.g., an API function call) based on the current coding context. Its widespread use highlights the need for efficient and context-aware solutions that streamline the development process. Despite ongoing efforts to enhance code completion performance, many existing approaches remain limited, offering ranked lists primarily based on alphabetical order or usage frequency from partially typed code fragments. While these studies have seen incremental improvements, the level of meaningful assistance provided to developers has not advanced in parallel. To address this limitation, we propose CODECOM, a deep learning-based code completion technique that leverages a large language model (LLM) with multi-task learning. Our approach processes sequences of source code tokens, their corresponding abstract syntax trees (ASTs), and program dependencies, enabling more context-aware and accurate code predictions. In our case study, CODECOM demonstrates state-of-the-art performance in the code completion downstream task. The evaluation results highlight significant improvements over the baseline, achieving 29.49% in BLEU, 71.16% in Acc@1, and 67.19% in Acc@5. These advancements are further validated by the Wilcoxon signed-rank test, which confirms strong statistical significance across all metrics. These findings indicate that Code Com can accelerate software development and assist developers in reducing potential errors effectively. Index Terms-software maintenance, automated code completion, deep learning, large language model.
Shradha Maharjan, Tae-Hyuk Ahn, Myoungkyu Song
SERA4
2025 Automated Code Summarization by Training Large Language Models with Crowdsourced Knowledge
abstract
In modern software development, efficient program comprehension is essential for maintaining and evolving software systems. Developers often dedicate over $50 \%$ of their time to understanding code due to the complexity and time demands involved. Code summarization-generating concise natural language descriptions of source code-has emerged as a potential solution. However, existing automated summarization techniques frequently produce summaries that are incomplete or lack accuracy. Moreover, as software evolves, documentation often becomes outdated, leading to discrepancies between code and comments. To address these challenges, we present DeepKnowCode, an automated approach that utilizes a DEEP learning technique based on a large language model and crowdsourced KNOWledge for CODE summarization. This approach aids developers by generating summaries that elucidate (1) the internal behavior of the code, (2) the rationale behind its implementation, and (3) practical guidelines for its use. We implemented a research prototype to assess real-world applicability and rigorously evaluated DeepKnowCode against state-of-the-art approaches. In our evaluation, DeepKnowCode demonstrates performance improvements in BLEU scores of 38.8% and 39.2% over two baselines, with statistical validation underscoring its effectiveness. By incorporating crowdsourced knowledge, DeepKnowCode captures essential elements of code semantics, context, and patterns, enhancing its ability to produce accurate and contextually relevant code summaries that facilitate program comprehension.
Shradha Maharjan, Myoungkyu Song
SERA3
2023 Programming Model for Information Sharing among IoT Devices: Software Engineering Perspective
abstract
The Internet of Things (IoT) has become an essential part of our daily lives and society as a whole, but it is still hard to develop and deliver IoT applications because of the complexity of the inherent heterogeneity nature of IoT devices. In this paper, we present a programming model for IoT, especially focusing on effective sharing of information based on software engineering ideas such as conceptual integrity and managing complexity. We analyze four programming models, Lisp, Fortran, Smalltalk, and Haskell, to understand what factors or ideas made them successful in managing complexity to solve problems in various domains. Then, based on the analysis, we propose an IoT programming model with the conceptual integrity that ‘every information is represented as a map.’ When we share information only in the form of a map data structure—a set of (key, value) pairs, we can simplify the process that manages the lifetime of the information shared. Even more, when we share information only in the form of a map, we can use both probabilistic data structures to reduce the footprint of the information when we need size efficiency and a JSON (JavaScript Object Notation) type information when we need to share high-quality information. We propose an architecture to accomplish this goal and implement a virtual machine to explain how information is generated, processed, and stored in the form of a map data structure.
Samuel Sungmin Cho, Myoungkyu Song
SERA2
2023 A Statistical Method for API Usage Learning and API Misuse Violation Finding
abstract
A large corpus of software repositories enables an opportunity for using machine learning (ML) approaches to create new software engineering tools. In this paper, we propose a novel technique which leverages ML approaches for automating software engineering tasks and thus improves software quality. Our concrete goal is to (1) explore the abundance of predictable repetitive regularities of such a massive codebase, (2) develop an ML approach for training a statistical model to identify common patterns in software corpora, and then (3) use these patterns to statistically detect anomalous, likely buggy, program behavior that significantly deviates from these typical patterns. These internal regularities and repetitive properties of software can be captured as patterns to detect violations of these common patterns. Such violations have a critical impact on program behavior such as bugs, security vulnerabilities, or even program crashes. Our approach focuses on usage patterns of application programming interfaces (APIs). API usage patterns are commonly recurring, representative examples of how real-world applications use APIs in software corpora. These desirable patterns of API usage are learnable to validate or improve developers' implementations. This paper shows preliminary results that we use standard cross-entropy and perplexity to measure how surprising a test subject application is to a statistical model estimated from a software corpus. We continue to develop our approach and evaluate the effectiveness to focus on the following research questions. Are our ML models effectively trainable on large code corpora to learn desirable API usage patterns? How does the performance of our ML-based approach compare to state-of-the-art language models for software when learning API usage for detecting API misuse violations?
Deepak Panda, Piyush Basia, Kushal Nallavolu, Xin Zhong 0001, Harvey P. Siy, Myoungkyu Song
SERA6
2023 SSDTutor: A feedback-driven intelligent tutoring system for secure software development
Dip Kiran Pradhan Newar, Rui Zhao 0005, Harvey P. Siy, Leen-Kiat Soh, Myoungkyu Song
Sci. Comput. Program.5
2022 An IDE Support for Validating Machine Learning Applications in Bioengineering Text Corpora
abstract
Modeling in machine learning (ML) is critical for software systems in practice. ML applications are required to validate their models and implementations but quality validation is a challenging and time-consuming process for developers. To address this limitation, we present a novel validation technique for ML applications to help developers or researchers (e.g., bioengineering domain) inspect (1) software code (ML API usages) and (2) ML model (extracted features).
Piyush Basia, Tae-Hyuk Ahn, Myoungkyu Song
BIBM3
2022 BioMDSE: A Multimodal Deep Learning-Based Search Engine Framework for Biofilm Documents Classifications
abstract
As biofilms research grows rapidly, a corpus of bibliographic literature (i.e., documents) is increasing at an incredible rate. Many researchers often need to inspect these large document collections, including (1) text, (2) images, and (3) captions, to understand underlying biological mechanisms and make a critical decision. However, researchers have great difficulty in exploring such ever-growing large datasets in labor-intensive processes. Thus, automation of such tasks is urgently required for the automatic identification or classification of a large volume of document collections. To address this problem, we present a multimodal deep learning-based approach to automatically classify documents for a specialized information retrieval technique based on biofilm images, captions, and texts, which is a major source of information for the classification of documents. Images, captions, and texts from biofilm documents are represented in a large vector space. Then, they are fed into convolutional neural networks (CNNs), to improve similarity matching and relevance. Our extensive experiments and analysis will take captions, texts, or images as unimodal models as inputs and concatenate them all into multimodal models. The trained models for this classification approach in turn help a search engine to precisely identify relevant and domain-specific documents from a large volume of document collections for further research direction in biofilm development.
Pei-Chi Huang, Ejan Shakya, Myoungkyu Song, Mahadevan Subramaniam
BIBM3
2022 Debugging Support for Machine Learning Applications in Bioengineering Text Corpora
abstract
Modeling in machine learning (ML) is becoming an essential part of software systems in practice. Validating ML applications is a challenging and time-consuming process for developers since the accuracy of prediction heavily relies on generated models. ML applications are written by relatively more data-driven programming based on the blackbox of ML frameworks. If all of the datasets and the ML application need to be individually investigated, the ML debugging tasks would take a lot of time and effort. To address this limitation, we present a novel debugging technique for machine learning applications, called MLDBUG that helps ML application developers inspect the training data and the generated features for the ML model. Inspired by software debugging for reproducing the potential reported bugs, MLDBUG takes as input an ML application and its training datasets to build the ML models, helping ML application developers easily reproduce and understand anomalies on the ML application. We have implemented an Eclipse plugin for MLDBUG which allows developers to validate the prediction behavior of their ML applications, the ML model, and the training data on the Eclipse IDE. In our evaluation, we used 23,500 documents in the bioengineering research domain. We assessed the MLDBUG's capability of how effectively our debugging technique can help ML application developers investi-gate the connection between the produced features and the labels in the training model and the relationship between the training instances and the instances the model predicts.
Kwok Sun Cheng, Tae-Hyuk Ahn, Myoungkyu Song
COMPSAC3
2022 A Feasibility Study of Using Code Clone Detection for Secure Programming Education
abstract
Secure library reuse is critical for modern ap-plications to protect private information in software security engineering. Teaching secure programming is also more critical to tackle the challenges of new and evolving threats. However, novice students often make mistakes by API misuses due to a lack of understanding of secure libraries or a false sense of security. In this paper, we study the feasibility of applying code clone detection (CCD) for finding relevant examples to effectively teach secure programming to computer science students. CCD is an emerging new technology that extracts syntactically or semantically similar code fragments to support many software engineering tasks, such as program understanding, code quality analysis, software evolution analysis, and bug detection. We have developed a prototype implementation ExTUTOR that allows students to search for relevant examples as feedback when they want to fix their programming issues or vulnerabilities. In our evaluation, we applied ExTUTOR to open source subject applications in the security domain. Our approach should help novice students gain benefits from feedback and identify how to effectively make use of APIs, encouraging students to fix their own security violations in their own applications.
Michael Menard, Tommy Nelson, Milan Shahi, Hugh Morton, Adam DeTavernier, Harvey P. Siy, Rui Zhao 0005, Myoungkyu Song
COMPSAC8
2022 RepChaBug: Automatically Repairing Incorrect Change Bugs in Software Evolution
abstract
Software systems become inevitable to have bugs due to the change complexity in a large portion of code. Program repair is a critical maintenance task in software evolution. To reduce maintenance effort and time, the technique of copying and pasting code snippets (code clone) is generally practiced by both the novice and large parts of the developer community. Debugging errors on such replication changes is an effort-prone and time-consuming task. To address this problem, we present an automated program repair approach, called REPCHABUG that helps developers find and repair code change anomalies when they mistakenly apply contradictory updates (contradiction bugs) or miss required edits (exclusion bugs) on code cloning activity. Given the code change portion, REPCHABUG generalizes the code modification into a dependency-aware change template by computing data and control dependency relationships. In our evaluation, a case study with open source applications and a user study with computer science students showed that REPCHABUG should improve developer productivity in finding and fixing replication change bugs.
Samuel Sungmin Cho, Myoungkyu Song
COMPSAC3
2022 An Intelligent Tutoring System for API Misuse Correction by Instant Quality Feedback
abstract
Computer science students have difficulty understanding correct usages of an Application Programming Interface (API) and programming violations that cause compilation or runtime errors. Despite high-quality documentation for programming, the students typically need an instructor's feedback when their programs cause bugs, crashes, and vulnerabilities. This paper presents a pedagogical approach that is based on an Intelligent Tutoring System called INTTuToR. Briefly, INTTUTORprovides novice students with instant feedback to fix their programming issues or vulnerabilities. We have implemented our approach as a plug-in application in the Integrated Development Environment (IDE) for an interactive educational environment. In our proposed evaluation, we plan to perform empirical studies with CS students to assess how effectively INTTUTORimproves their ability to identify and fix potential bugs or vulnerabilities in the cryptography-related programming assignments.
Rui Zhao 0005, Harvey P. Siy, Chulwoo Pack, Leen-Kiat Soh, Myoungkyu Song
COMPSAC5
2021 Learning To Rank Relevant Documents for Information Retrieval in Bioengineering Text Corpora
abstract
In this paper, we present a Learning To Rank-based approach that helps EXPLORE and understand Relevant documents for bioengineering text corpora, called LTREXPLORER. Based on the likelihood of being the most relevance to a search query, the ranking model sorts documents according to their degrees of relevance, preference, or importance with various domain-specific features. The evaluation results demonstrated that our approach has the potential to effectively provide the retrieval scoring functions to researchers, who focus on the most relevant documents in bioengineering information retrieval.
Kwok Sun Cheng, Myoungkyu Song
COMPSAC2
2021 Analyzing Bug Reports by Topic Mining in Software Evolution
abstract
Reporting bugs is one of the vital activities for evolving software systems. Given such reports, developers cope with unanticipated behaviors during software development, maintenance, and operations. The description of bug reports typically includes (1) what errors occurred previously and (2) how a failure can be reproduced through specific steps, test inputs, and original configurations when a failure was created. However, analyzing bug reports is a tedious and error-prone process due to overflowing, complex terminologies. For example, diverse terms are used to represent similar or divergent elucidations by surrounding contexts during software development and maintenance. To address this problem, we present an approach that applies a topic mining technique to bug reports for finding an adequate code reviewer, who can potentially cope with reported failures, by inferring some hidden topics of a textual document.
Uy Nguyen, Kwok Sun Cheng, Samuel Sungmin Cho, Myoungkyu Song
COMPSAC4
2021 FireBugs: Finding and Repairing Cryptography API Misuses in Mobile Applications
abstract
In this paper, we present FireBugs for Finding and Repairing Bugs based on security patterns. For the common misuse patterns of cryptography APIs (crypto APIs), we encode common cryptography rules into the pattern representations for bug detection and program repair regarding cryptography rule violations. In the evaluation, we conducted a case study to assess the bug detection capability by applying FireBugs to datasets mined from both open source and commercial projects. Also, we conducted a user study with professional software engineers at Mutual of Omaha Insurance Company to estimate the program repair capability. This evaluation showed that FireBugs can help professional engineers develop various cryptographic requirements in a resilient application.
Larry Singleton, Rui Zhao 0005, Harvey P. Siy, Myoungkyu Song
COMPSAC4
2021 Cocoa: Towards a Scalable Compute Cost-aware Data Analytics System
abstract
Recently, many data analytics systems have focused on adopting a newly emerging compute resource, serverless, which offers scalability and agility to deal with peak workloads in a timely and cost-efficient manner, i.e., serverless data analytics (SDA). Unfortunately, these systems may encounter a cost bottleneck ($) because they have ignored the per unit time cost ($) of serverless, which is more expensive by up to 5.8 times for the same compute capacity than a traditional compute resource such as a virtual machine (VM). In addition, SDA may also encounter a performance bottleneck due to serverless' worse performance than VM. In this paper, we first study and report when serverless is beneficial for data analytics. Then, we present a scalable compute cost-aware data analytics system, Cocoa, that exploits serverless and VM together to achieve composite benefits. A Cocoa prototype was implemented on Spark. Evaluation results show a richer cost-performance tradeoff space opened by exploiting heterogeneous compute resources together, and identify substantial opportunities for future serverless-enabled systems research.
Kwangsung Oh, Myoungkyu Song
IC2E2
2020 Code Inspection Support for Recurring Changes with Deep Learning in Evolving Software
abstract
Developers often make recurring changes, similar but different changes across multiple locations. They inspect such code changes per source file (i.e., a diff patch) during code reviews; however, diff patches represent low-level code modification without summarizing recurring changes, leading to tedious and error-prone code inspection. To address this problem, we propose a novel code review approach, Recurring Code Changes Inspection with Deep Learning (RIDL) that leverages change patterns of an edit script by learning code clones, identical or nearly similar code fragments. To train a classifier, RIDL learns 13,940 clones with four different clone types (e.g., Type-1, Type-2, Type-3, and Type-4 clones) from a clone database mined from 25,000 subject programs. Our approach then leverages the classifier to (1) interactively summarize recurring changes and (2) detect change mistakes, potential anomalies in a given codebase. In the evaluation, after 2 hours of training, RIDL analyzes code changes in four open source projects. It summarizes recurring changes with 95.1% accuracy and detects change anomalies with 93.1% accuracy. Our results show that RIDL should help developers effectively inspect recurring changes during code reviews.
Krishna Teja Ayinala, Kwok Sun Cheng, Kwangsung Oh, Teukseob Song, Myoungkyu Song
COMPSAC5
2019 LAC: Locating and Applying Consistent and Repetitive Changes
abstract
As a software system evolves, constant changes are usually made in the code. Similar changes to code fragments often occur in multiple locations. It is very tedious and error prone for developers to manually update similar changes in multiple locations. To address this problem, we propose an approach for Locating and Applying Consistent and Repetitive Changes (LAC) as an Eclipse plug-in. LAC analyzes the code fragments to detect change anomalies, such as omissions or inconsistent edits, and automatically repairs these anomalies by applying required changes. Our static analysis technique to detect anomalies is carried out by inferring change patterns. The evaluation results in a user study show that automatically applying repetitive changes in identified locations using LAC correctly matches with that of manually updated ones, which significantly decreases manual efforts and error-prone edits.
Sushma Sakala, Vamshi Krishna Epuri, Samuel Sungmin Cho, Myoungkyu Song
COMPSAC (1)4
2019 Tool support for managing repetitive program changes in evolving software
abstract
Software modification often requires consistent program changes , a group of similar, related changes, at multiple locations in a program. Developers are typically uneasy to (i) detect potential change anomalies such as omission errors or incorrect edits and (ii) determine related locations to apply consistent changes, which is a tedious and error‐prone process. To address this problem, this study presents a technique for managing consistent program changes, checking and applying repetitive program transformation (CARP). Given program differencing results between original and edited program versions, CARP (i) infers change patterns to detect change anomalies , (ii) identifies required edit locations , and (iii) automatically applies adequate edits . It has been implemented in the context of the integrated development environment as an Eclipse plug‐in. The authors evaluated CARP on three open‐source projects and found that CARP detects seeded anomalies with 99.1% accuracy on average. Furthermore, it identifies change locations and transforms them with 98% accuracy. Their results show that CARP should help developers detect potential change anomalies in repetitive program changes and perform consistent changes automatically.
Vamshi Krishna Epuri, Sushma Sakala, Tae-Hyuk Ahn, Myoungkyu Song
IET Softw.4
2019 Decomposing Composite Changes for Code Review and Regression Test Selection in Evolving Software
Bo Guo 0004, Young-Woo Kwon 0001, Myoungkyu Song
J. Comput. Sci. Technol.3
2019 A declarative enhancement of JavaScript programs by leveraging the Java metadata infrastructure
Kwok Sun Cheng, Myoungkyu Song, Eli Tilevich
Sci. Comput. Program.3
2018 SORA: Scalable Overlap-graph Reduction Algorithms for Genome Assembly using Apache Spark in the Cloud
Alexander J. Paul, Dylan Lawrence, Myoungkyu Song, Seung-Hwan Lim, Chongle Pan, Tae-Hyuk Ahn
BIBM3
2018 Improving regression test efficiency with an awareness of refactoring changes
Zhiyuan Chen 0006, Myoungkyu Song
Inf. Softw. Technol.3
2018 Clone refactoring inspection by summarizing clone refactorings and detecting inconsistent changes during software evolution
abstract
It has been broadly assumed that removing code clones by refactorings would solve the problems of code duplication. Despite recent empirical studies on the benefit of refactorings, contradicting evidence shows that it is often difficult or impossible to remove clones by using standard refactoring techniques. Developers cannot easily determine which clones can be refactored or how they should be maintained scattered throughout a large code base in evolving systems. We propose pattern‐based clone refactoring inspection (PRI), a technique for managing clone refactorings. PRI summarizes refactorings of clones and detects clones that are not consistently refactored. To help developers refactor these anomalies, PRI also visualizes clone evolution and refactorings and fixes refactoring anomalies to prevent the clone group from being left in an inconsistent state. We evaluated PRI on 6 open‐source projects and showed that it identifies clone refactorings with 94.1% accuracy and detects inconsistent refactorings with 98.4% accuracy, tracking clone change histories. In a study with 10 student developers, the participants reported that flexible PRI's summarization and detection features can be valuable for novice developers to learn about refactorings to clones. These results show that PRI should improve developer productivity in inspecting clone refactorings distributed across multiple files in evolving systems.
Zhiyuan Chen 0006, Young-Woo Kwon 0001, Myoungkyu Song
J. Softw. Evol. Process.3
2018 Refactoring Inspection Support for Manual Refactoring Edits
abstract
Refactoring is commonly performed manually, supported by regression testing, which serves as a safety net to provide confidence on the edits performed. However, inadequate test suites may prevent developers from initiating or performing refactorings. We propose RefDistiller, a static analysis approach to support the inspection of manual refactorings. It combines two techniques. First, it applies predefined templates to identify potential missed edits during manual refactoring. Second, it leverages an automated refactoring engine to identify extra edits that might be incorrect. RefDistiller also helps determine the root cause of detected anomalies. In our evaluation, RefDistiller identifies 97 percent of seeded anomalies, of which 24 percent are not detected by generated test suites. Compared to running existing regression test suites, it detects 22 times more anomalies, with 94 percent precision on average. In a study with 15 professional developers, the participants inspected problematic refactorings with RefDistiller versus testing only. With RefDistiller, participants located 90 percent of the seeded anomalies, while they located only 13 percent with testing. The results show RefDistiller can help check the correctness of manual refactorings.
Everton L. G. Alves, Myoungkyu Song, Tiago Massoni, Patrícia Duarte de Lima Machado, Miryung Kim
IEEE Trans. Software Eng.2
2017 Tool Support for Managing Clone Refactorings to Facilitate Code Review in Evolving Software
abstract
Developers often perform copy-and-paste activities. This practice causes the similar code fragment (aka code clones) to be scattered throughout a code base. Refactoring for clone removal is beneficial, preventing clones from having negative effects on software quality, such as hidden bug propagation and unintentional inconsistent changes. However, recent research has provided evidence that factoring out clones is not always to reduce the risk of introducing defects, and it is often difficult or impossible to remove clones using standard refactoring techniques. To investigate which or how clones can be refactored, developers typically spend a significant amount of their time managing individual clone instances or clone groups scattered across a large code base. To address the problem, this paper presents a technique for managing clone refactorings, Pattern-based clone Refactoring Inspection (PRI), using refactoring pattern templates. By matching the refactoring pattern templates against a code base, it summarizes refactoring changes of clones, and detects the clone instances not consistently factored out as potential anomalies. PRI also provides novel visualization user interfaces specifically designed for inspecting clone refactorings. In the evaluation, PRI analyzes clone instances in six open source projects. It identifies clone refactorings with 94.1% accuracy and detects inconsistent refactorings with 98.4% accuracy. Our results show that PRI should help developers effectively inspect evolving clones and correctly apply refactorings to clone groups.
Zhiyuan Chen 0006, Maneesh Mohanavilasam, Young-Woo Kwon 0001, Myoungkyu Song
COMPSAC (1)4
2017 Interactively Decomposing Composite Changes to Support Code Review and Regression Testing
abstract
Developers often address multiple development issues to make composite code changes, as opposed to atomic changes that address one single issue. Investigating and testing such code changes is a tedious and error-prone process for developers. To address the problem, this paper presents a technique, called CHGCUTTER, for (1) interactively decomposing composite changes into atomic changes, (2) building related change subsets using program dependence relationships without syntactic violation, and (3) safely selecting only related test cases from the test suite to reduce the time to conduct regression testing. In the evaluation, CHGCUTTER analyzes 28 composite changes in four open source projects. It identifies related change subsets with 95.7% accuracy, and it selects test cases affected by these changes with 89.0% accuracy. Our results show that CHGCUTTER should help developers effectively inspect changes and validate modified applications during development.
Bo Guo 0004, Myoungkyu Song
COMPSAC (1)2
2017 Enabling Flexible and Efficient Remote Execution in Opportunistic Networks through Message-Oriented Middleware
abstract
Computation offloading has received much attention to improve the performance or energy efficiency of mobile systems that have usually limited and constrained resource capacities. Yet, applying the computation offloading technique in opportunistic networks that are highly dynamic and often become volatile still remains a challenge due to the following reasons: (1) technical difficulties in constructing efficient and reliable execution environment using commodity devices using WiFi, (2) a lack of runtime support for multiple clients that request diverse computational tasks and execute them concurrently with minimum performance impacts. In this paper, we introduce a new middleware system that provides an offloading framework operated in opportunistic networks. In particular, our middleware employs a publish-subscribe communication mechanism to provide multiple different communication models (e.g., one-to-one, one-to-many, many-to-one, and many-to-many) for different use cases. Furthermore, when distributing computational tasks to nearby nodes, our middleware takes their resource capabilities into consideration for efficient execution. Finally, since partial failure is an unavoidable artifact in highly dynamic and volatile opportunistic networks, we provide a simple, but effective failure handling mechanism. Our benchmarks and experimental results indicate that our approach enables programmers to easily apply computation offloading techniques in opportunistic networks when compared with the local execution.
Minh Le, Myoungkyu Song, Young-Woo Kwon 0001
COMPSAC (1)2
2015 Interactive Code Review for Systematic Changes
abstract
Developers often inspect a diff patch during peer code reviews. Diff patches show low-level program differences per file without summarizing systematic changes -- similar, related changes to multiple contexts. We present Critics, an interactive approach for inspecting systematic changes. When a developer specifies code change within a diff patch, Critics allows developers to customize the change template by iteratively generalizing change content and context. By matching a generalized template against the codebase, it summarizes similar changes and detects potential mistakes. We evaluated Critics using two methods. First, we conducted a user study at Salesforce.com, where professional engineers used Critics to investigate diff patches authored by their own team. After using Critics, all six participants indicated that they would like Critics to be integrated into their current code review environment. This also attests to the fact that Critics scales to an industry-scale project and can be easily adopted by professional engineers. Second, we conducted a user study where twelve participants reviewed diff patches using Critics and Eclipse diff. The results show that human subjects using Critics answer questions about systematic changes 47.3% more correctly with 31.9% saving in time during code review tasks, in comparison to the baseline use of Eclipse diff. These results show that Critics should improve developer productivity in inspecting systematic changes during peer code reviews.
Tianyi Zhang 0001, Myoungkyu Song, Joseph Pinedo, Miryung Kim
ICSE (1)2
2015 Reusing metadata across components, applications, and languages
Myoungkyu Song, Eli Tilevich
Sci. Comput. Program.1
2014 RefDistiller: a refactoring aware code review tool for inspecting manual refactoring edits
abstract
Manual refactoring edits are error prone, as refactoring requires developers to coordinate related transformations and understand the complex inter-relationship between affected types, methods, and variables. We present RefDistiller, a refactoring-aware code review tool that can help developers detect potential behavioral changes in manual refactoring edits. It first detects the types and locations of refactoring edits by comparing two program versions. Based on the reconstructed refactoring information, it then detects potential anomalies in refactoring edits using two techniques: (1) a template-based checker for detecting missing edits and (2) a refactoring separator for detecting extra edits that may change a program's behavior. By helping developers be aware of deviations from pure refactoring edits, RefDistiller can help developers have high confidence about the correctness of manual refactoring edits. RefDistiller is available as an Eclipse plug-in at https://sites.google.com/site/refdistiller/ and its demonstration video is available at http://youtu.be/0Iseoc5HRpU.
Everton L. G. Alves, Myoungkyu Song, Miryung Kim
SIGSOFT FSE2
2014 Critics: an interactive code review tool for searching and inspecting systematic changes
abstract
During peer code reviews, developers often examine program differences. When using existing program differencing tools, it is difficult for developers to inspect systematic changes—similar, related changes that are scattered across multiple files. Developers cannot easily answer questions such as "what other code locations changed similar to this change?" and "are there any other locations that are similar to this code but are not updated?" In this paper, we demonstrate Critics, an Eclipse plug-in that assists developers in inspecting systematic changes. It (1) allows developers to customize a context-aware change template, (2) searches for systematic changes using the template, and (3) detects missing or inconsistent edits. Developers can interactively refine the customized change template to see corresponding search results. Critics has potential to improve developer productivity in inspecting large, scattered edits during code reviews. The tool's demonstration video is available at https://www.youtube.com/watch?v=F2D7t_Z5rhk
Tianyi Zhang 0001, Myoungkyu Song, Miryung Kim
SIGSOFT FSE2
2012 Metadata invariants: Checking and inferring metadata coding conventions
abstract
As the prevailing programming model of enterprise applications is becoming more declarative, programmers are spending an increasing amount of their time and efforts writing and maintaining metadata, such as XML or annotations. Although metadata is a cornerstone of modern software, automatic bug finding tools cannot ensure that metadata maintains its correctness during refactoring and enhancement. To address this shortcoming, this paper presents metadata invariants, a new abstraction that codifies various naming and typing relationships between metadata and the main source code of a program. We reify this abstraction as a domain-specific language. We also introduce algorithms to infer likely metadata invariants and to apply them to check metadata correctness in the presence of program evolution. We demonstrate how metadata invariant checking can help ensure that metadata remains consistent and correct during program evolution; it finds metadata-related inconsistencies and recommends how they should be corrected. Similar to static bug finding tools, a metadata invariant checker identifies metadata-related bugs as a program is being refactored and enhanced. Because metadata is omnipresent in modern software applications, our approach can help ensure the overall consistency and correctness of software as it evolves.
Myoungkyu Song, Eli Tilevich
ICSE1
2012 Detecting metadata bugs on the fly
abstract
Programmers are spending a large and increasing amount of their time writing and modifying metadata, such as Java annotations and XML deployment descriptors. And yet, automatic bug finding tools cannot find metadata-related bugs introduced during program refactoring and enhancement. To address this shortcoming, we have created metadata invariants, a new programming abstraction that expresses naming and typing relationships between metadata and the main source code of a program. A paper that appears in the main technical program of ICSE 2012 describes the idea, concept, and prototype of metadata invariants [4]. The goal of this demo is to supplement that paper with a demonstration of our Eclipse plugin, Metadata Bug Finder (MBF). MBF takes as input a script written in our domain-specific language that describes a set of metadata coding conventions the programmer wishes to enforce. Then after each file save operation, MBF checks the edited codebase for the presence of any violations of the given metadata programming conventions. These violations are immediately reported to the programmer as potential metadata-related bugs. By making the programmer aware of these potential bugs, MBF prevents them from seeping into production, thereby improving the overall correctness of the edited codebase.
Myoungkyu Song, Eli Tilevich
ICSE1
2009 Enhancing source-level programming tools with an awareness of transparent program transformations
abstract
Programs written in managed languages are compiled to a platform-independent intermediate representation, such as Java bytecode. The relative high level of Java bytecode has engendered a widespread practice of changing the bytecode directly, without modifying the maintained version of the source code. This practice, called bytecode engineering or enhancement, has become indispensable in introducing various concerns, including persistence, distribution, and security, transparently. For example, transparent persistence architectures help avoid the entanglement of business and persistence logic in the source code by changing the bytecode directly to synchronize objects with stable storage. With functionality added directly at the bytecode level, the source code reflects only partial semantics of the program. Specifically, the programmer can neither ascertain the program's runtime behavior by browsing its source code, nor map the runtime behavior back to the original source code.
Myoungkyu Song, Eli Tilevich
OOPSLA1