Eunjong Choi

dblp:21/10488 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-2196-033XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Development and benchmarking of multilingual code clone detector
Norihiro Yoshida, Toshihiro Kamiya, Eunjong Choi, Hiroaki Takada
J. Syst. Softw.4
2024 Global Alignment Learning for Code Search
abstract
Code search plays a role in bridging code and query. However, recent code search studies mainly rely on affinity-matrix-based cross-modal attention to learn the word alignments between code and query, which may lead to incorrect alignments. In this paper, we propose a Global Alignment Learning Model (GALM) to learn global alignments and demonstrate that better-learned correct alignments can significantly improve code search performance. Specifically, GALM characterizes the query and code embedding into an alignment graph to enhance the feature representation and further learns global alignments by a dense graph convolutional network. To evaluate the performance of GALM, we compared it with several baseline models on two pop-ular datasets. The results demonstrate that GALM outperforms the best baseline models by 9.8% and 6.8% with the Top@1 accuracy of 0.601 and 0.671 on two datasets, respectively.
Juntong Hong, Eunjong Choi, Kinari Nishiura, Osamu Mizuno
SERA2
2024 Analyzing the Inpact of Formal Methods on Isuue Trends Using BERTopic
abstract
Software specification quality significantly influ-ences overall software quality. Ambiguity in natural language specifications has been identified as a major obstacle to achieving high quality. Formal methods have been proposed as a solution to mitigate this issue, but their adoption in software development remains limited. While previous studies have primarily focused on the challenges of introducing formal methods, few reports exist on the effects of applying formal methods in software development. To tackle this issue, this study analyzes the trends in OSS (open source software) with and without formal methods, based on issues collected from GitHub repositories. We employed BERTopic, a topic modeling technique, to extract topics associated with issues. Subsequently, we labeled the issues and statistically compared topics of issues in OSS with and without formal methods. The results indicate a significant difference in the tendency of issues between OSSs with formal methods and without formal methods, particularly in the frequency of reported errors. These findings suggest the existence of issues specific to OSSs with formal methods and highlight the potential benefits of formal methods in reducing errors during development.
Soshi Inoue, Kinari Nishiura, Eunjong Choi, Osamu Mizuno
SERA3
2024 Benefits and Pitfalls of Token-Level SZZ: An Empirical Study on OSS Projects
abstract
SZZ is the de facto standard method for identifying bug-inducing commits. The accuracy of this method heavily relies on source code management systems, such as Git, as it requires tracing the history of source code changes (i.e., commit histories) to bug-inducing commits. However, it has been reported that these systems introduce biases in commit histories because they only store line-level changes. It is known that such coarse-grained line-level changes can result in the failure to accurately track the commit history and reduce the performance of SZZ. To relieve this challenge, we explore the accuracy of SZZ in token-level changes, which provide finer-grained information to trace commit histories compared to line-level ones, and we discuss the potential benefits and pitfalls of utilizing token-level changes for SZZ. As a result of experiments on 68 OSS projects, we found that SZZ, which uses token-level histories, identifies two new bug-inducing commits that are missed when using line-level histories. Furthermore, our manual analysis of the identified commits indicates that they reduce false-positive bug-inducing commits caused by source code formatting and whitespace changes. However, this improvement in detecting bug-inducing commits comes with a trade-off of 0.081 decrease in overall accuracy, as measured by the F1 score. Consequently, we summarized three potential benefits and five pitfalls of using token-level and line-level tracking for SZZ.
Hiroya Watanabe, Masanari Kondo, Eunjong Choi, Osamu Mizuno
SANER3
2024 Two improving approaches for faulty interaction localization using logistic regression analysis
abstract
Abstract Faulty Interaction Localization (FIL) is a process to identify which combination of input parameter values induced test failures in combinatorial testing. An accurate and fast FIL provides helpful information to fix defects causing the test failure. One type of conventional FIL approach, which analyzes test results of whole test cases and estimates the suspiciousness of each combination, has two main concerns; (1) the accuracy is not enough, (2) the huge time cost is sometimes needed. In this paper, we propose two novel approaches to improve those concerns. attempts to estimate suspiciousness more accurately using logistic regression analysis. attempts to estimate failure-inducing combinations at high speed by estimating the subsets of them using logistic regression analysis and exploring just their supersets. Through evaluation experiments using a large number of artificial test results based on several real software systems, we observed that has very high accuracy, and can drastically reduce time cost for targets that have been difficult to complete by the conventional method.
Kinari Nishiura, Eun-Hye Choi, Eunjong Choi, Osamu Mizuno
Softw. Qual. J.3
2023 Cost-Benefit Analysis for Modernizing a Large-Scale Industrial System
abstract
Legacy systems pose significant challenges to companies. Software modernization approaches have been proposed to address this issue. However, a lack of standardization and reliance on ad hoc processes often lead to software modernization failures. Incremental modernization, a strategy that improves software systems in a step-by-step manner rather than attempting to simultaneously overhaul the entire system, aims to mitigate the risk of failure. However, this approach can increase costs owing to the complexity of integrating legacy and modernized products. In this paper, we present a case study that employs a cost-benefit estimation analysis in a large-scale industrial project that underwent incremental modernization in the past. We compare the actual and estimated cost-benefit values in the context of incremental modernization. As a result, we confirmed that the cost estimates were valid, but we could not judge whether the benefit estimates were valid.
Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Joji Okada, Yoshiki Higo
APSEC2
2023 Towards Better Online Communication for Future Software Development in Industry
abstract
COVID-19 has transformed face-to-face software development into distributed development (e. g., remote work). While the company authors belong to studies microtask programming, an open source software (OSS) -like development, as a solution to employ distributed development, a prior study reports a challenge: online communication in microtask programming takes longer; such lengthy communication discourages developers and affects their completion of assigned tasks. OSS, however, is successfully developed using online communication, such as issues. Hence, we have a question: how does OSS address the online communication challenge? In this experience report, we answer this question based on an empirical study on OSS communication. We found that (1) OSS prefers burst communication similar to face-to-face development, and (2) attracting developers’ attention may be a possible solution. Based on the findings, we discuss the direction of future studies to achieve better online communication in microtask programming in the company. The main contributions of this report are (1) to empirically reveal the actual communication times in OSS and (2) to show how an empirical approach helps industrial collaborators.
Masanari Kondo, Shinobu Saito, Yukako Iimura, Eunjong Choi, Osamu Mizuno, Yasutaka Kamei, Naoyasu Ubayashi
COMPSAC4
2023 Investigating the Generalizability of Deep Learning-based Clone Detectors
abstract
The generalizability of Deep Learning (DL) models is a significant challenge, as poor generalizability indicates that the model has overfitted to the training data and is not able to generalize to new data. Despite numerous DL-based clone detectors emerging in recent years, their generalizability has not been thoroughly assessed. This study investigates the generalizability of three DL-based clone detectors (CCLearner, ASTNN, and CodeBERT) by comparing their detection accuracy on different training and testing clone benchmarks. The results show that all three clone detectors do not generalize well to new data and there is a strong relationship between clone types and generalizability for CCLearner and ASTNN.
Eunjong Choi, Norihiro Fuke, Yuji Fujiwara, Norihiro Yoshida, Katsuro Inoue
ICPC1
2022 MSCCD: grammar pluggable clone detection based on ANTLR parser generation
abstract
For various reasons, programming languages continue to multiply and evolve. It has become necessary to have a multilingual clone detection tool that can easily expand supported programming languages and detect various code clones is needed. However, research on multilingual code clone detection has not received sufficient attention. In this study, we propose MSCCD (Multilingual Syntactic Code Clone Detector), a grammar pluggable code clone detection tool that uses a parser generator to generate a code block extractor for the target language. The extractor then extracts the semantic code blocks from a parse tree. MSCCD can detect Type-3 clones at various granularities. We evaluated MSCCD's language extensibility by applying MSCCD to 20 modern languages. Sixteen languages were perfectly supported, and the remaining four were provided with the same detection capabilities at the expense of execution time. We evaluated MSCCD's recall by using BigCloneEval and conducted a manual experiment to evaluate precision. MSCCD achieved equivalent detection performance equivalent to state-of-the-art tools.
Norihiro Yoshida, Toshihiro Kamiya, Eunjong Choi, Hiroaki Takada
ICPC4
2022 Challenges and Future Research Direction for Microtask Programming in Industry
abstract
Microtask programming [4] is a solution to promote distributed development in industry. The key idea of microtask programming is to reduce face-to-face communication across developers by splitting the development task of software into independent microtasks. Such microtasks can be completed by crowd workers who work remotely and at their preferable time such as early morning. Dedicated developers who have the responsibility for the progress of development split the task into microtasks, and distribute them to crowd workers. Hence, microtask programming has these two actors. Our research team reported that microtask programming has potential benefits such as the fluidity of project assignments in industrial companies [4]. However, we suppose it still has challenges. In addition, it is still unclear what are future research direction to support both actors in microtask programming, though our research team has conducted three studies for microtask programming so far [2--4].
Masanari Kondo, Shinobu Saito, Yukako Iimura, Eunjong Choi, Osamu Mizuno, Yasutaka Kamei, Naoyasu Ubayashi
MSR4
2020 Clone Notifier: Developing and Improving the System to Notify Changes of Code Clones
abstract
A code clone is a code fragment that is identical or similar to it in the source code. It has been identified as one of the main problems in software maintenance. When a developer fixes a defect, they need to find the code clones corresponding to the code fragments. In this paper, we present Clone Notifier, a system that alerts on creations and changes of code clones to software developers. First, Clone Notifier identifies creations and changes of code clones. Subsequently, it groups them into four categories (new, deleted, changed, stable) and assigns labels (e.g., consistent, inconsistent) to them. Finally, it notifies on creations and changes of code clones along with the corresponding categories and labels. Clone Notifier and its video are available at: https://github.com/s-tokui/CloneNotifier.
Shogo Tokui, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
SANER3
2019 CCEvovis: a clone evolution visualization system for software maintenance
abstract
Understanding the evolution of code clones is important in software maintenance. With the information about how code clones evolve, both developers and researchers can understand the impacts of code clones and build a more robust code clone management system. So far, many studies have investigated the evolution of code clones to better understand the effects of code clones. However, only a few systems have been presented to support managing code clones based on the information about how code clone evolves. To mitigate this problem, in this paper, we present CCEvovis, a system that visualizes the evolved code clones across multiple versions of a program. CCEvovis highlights and visualizes the clone change to support software maintenance. CCEvovis is available at: https://github.com/hirotaka0616/CCEvovis.
Hirotaka Honda, Shogo Tokui, Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
ICPC4
2019 Identifying and predicting key features to support bug reporting
abstract
Abstract Bug reports are the primary means through which developers triage and fix bugs. To achieve this effectively, bug reports need to clearly describe those features that are important for the developers. However, previous studies have found that reporters do not always provide such features. Therefore, we first perform an exploratory study to identify the key features that reporters frequently miss in their initial bug report submissions. Then, we propose an approach that predicts whether reporters should provide certain key features to ensure a good bug report. A case study of the bug reports for Camel, Derby, Wicket, Firefox, and Thunderbird projects shows that Steps to Reproduce, Test Case, Code Example, Stack Trace, and Expected Behavior are the additional features that reporters most often omit from their initial bug report submissions. We also find that these features significantly affect the bug‐fixing process. On the basis of our findings, we build and evaluate classification models using four different text‐classification techniques to predict key features by leveraging historical bug‐fixing knowledge. The evaluation results show that our models can effectively predict the key features. Our comparative study of different text‐classification techniques shows that naïve Bayes multinomial (NBM) outperforms other techniques. Our findings can benefit reporters to improve the contents of bug reports.
Md. Rejaul Karim, Akinori Ihara, Eunjong Choi, Hajimu Iida
J. Softw. Evol. Process.3
2018 An Investigation of the Relationship between Extract Method and Change Metrics: A Case Study of JEdit
abstract
Extract Method is one of the most widely used refactoring patterns. So far, low quality of source code has been regarded as an indicator for Extract Method opportunities. However, recent studies showed that there is no clear relationship between source code quality and Extract Method. Change metrics can be indicators for Extract Method because the characteristics of software evolution strongly affect software quality. However, there has been no study that investigated the relationship between change metrics and Extract Method. In this study, we conducted two studies investigating the relationship between Extract Method and change metrics. As a result, we found that (1) change metrics have a clear relationship with Extract Method and (2) both product and change metrics are necessary to recommend candidates for Extract Method with high accuracy.
Eunjong Choi, Daiki Tanaka, Norihiro Yoshida, Kenji Fujiwara, Daniel Port, Hajimu Iida
APSEC1
2018 Multilingual Detection of Code Clones Using ANTLR Grammar Definitions
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose an approach for the multilingual detection of code clones using grammar files for a parser generator ANTLR. We extended a clone detection tool CCFinderSW with the proposed approach and then apply the extended CCFinderSW to ANTLR grammar files for 43 languages. As a result, the files for 39 out of the 43 languages can be analyzed correctly by the extended CCFinderSW.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC3
2018 Investigating Vector-Based Detection of Code Clones Using BigCloneBench
abstract
In a vector-based approach to detecting code clones from source code, all code fragments in the source are mapped to a vector space and then code fragments are detected as code clones if they are neighbors in the vector space. So far, our research group has developed a vector-based approach using TF-IDF and cosine similarity. For the improvement of the vector-based approach, we preliminary investigated what kind of vectorization algorithms and similarity measurements are effective in terms of recall and detection time. In this paper, we present preliminary investigation results using BigCloneBench, a large-scale code clone benchmark.
Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
APSEC2
2018 FaCoY: a code-to-code search engine
abstract
Code search is an unavoidable activity in software development. Various approaches and techniques have been explored in the literature to support code search tasks. Most of these approaches focus on serving user queries provided as natural language free-form input. However, there exists a wide range of use-case scenarios where a code-to-code approach would be most beneficial. For example, research directions in code transplantation, code diversity, patch recommendation can leverage a code-to-code search engine to find essential ingredients for their techniques. In this paper, we propose FaCoY, a novel approach for statically finding code fragments which may be semantically similar to user input code. FaCoY implements a query alternation strategy: instead of directly matching code query tokens with code in the search space, FaCoY first attempts to identify other tokens which may also be relevant in implementing the functional behavior of the input code. With various experiments, we show that (1) FaCoY is more effective than online code-to-code search engines; (2) FaCoY can detect more semantic code clones (i.e., Type-4) in BigCloneBench than the state-of-the-art; (3) FaCoY, while static, can detect code fragments which are indeed similar with respect to runtime execution behavior; and (4) FaCoY can be useful in code/patch recommendation.
Kisub Kim, Dongsun Kim 0001, Tegawendé F. Bissyandé, Eunjong Choi, Li Li 0029, Jacques Klein, Yves Le Traon
ICSE4
2017 CCFinderSW: Clone Detection Tool with Flexible Multilingual Tokenization
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose a clone detection tool CCFinderSW that has extension mechanism to handle addition language on demand from practitioners.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC3
2017 Frame-based behavior preservation in refactoring
abstract
Behavior preservation often bothers programmers in refactoring. This poster paper proposes a new approach that tames the behavior preservation by introducing the concept of a frame. A frame in refactoring defines stakeholder's individual concerns about the refactored code. Frame-based refactoring preserves the observable behavior within a particular frame. Therefore, it helps programmers distinguish the behavioral changes that they should observe from those that they can ignore.
Katsuhisa Maruyama, Shinpei Hayashi, Norihiro Yoshida, Eunjong Choi
SANER4
2016 Revisiting the relationship between code smells and refactoring
abstract
Refactoring is a critical technique in evolving software systems. Martin Fowler presented a catalogue of refactoring patterns that defines a list of code smells and their corresponding refactoring patterns. This list aimed at supporting programmers in finding suitable refactoring patterns that remove code smells from their systems. However, a recent empirical study by Bavota et al. shows that refactoring rarely removes code smells which do not align with Fowler's catalog. To bridge the gap between them, we revisit the relationship between code smells and refactorings. In this study, we investigate whether developers apply appropriate refactoring patterns to fix code smells in three open source software systems.
Norihiro Yoshida, Tsubasa Saika, Eunjong Choi, Ali Ouni 0001, Katsuro Inoue
ICPC3
2016 On the Effectiveness of Vector-Based Approach for Supporting Simultaneous Editing of Software Clones
Seiya Numata, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
PROFES3
2015 Quick Trigger on Stack Overflow: A Study of Gamification-Influenced Member Tendencies
abstract
In recent times, gamification has become a popular technique to aid online communities stimulate active member participation. Gamification promotes a reward-driven approach, usually measured by response-time. Possible concerns of gamification could a trade-off between speedy over quality responses. Conversely, bias toward easier question selection for maximum reward may exist. In this study, we analyze the distribution gamification-influenced tendencies on the Q&A Stack Overflow online community. In addition, we define some gamification-influenced metrics related to response time to a question post. We carried experiments of a four-month period analyzing 101,291 members posts. Over this period, we determined a Rapid Response time of 327 seconds (5.45 minutes). Key findings suggest that around 92% of SO members have fewer rapid responses that non-rapid responses. Accepted answers have no clear relationship with rapid responses. However, we did find that rapid responses significantly contain tags that did not follow their usual tagging tendencies.
Xin Yang 0018, Raula Gaikovina Kula, Eunjong Choi, Katsuro Inoue, Hajimu Iida
MSR4
2013 Applying clone change notification system into an industrial development process
abstract
Programmers tend to write code clones unintentionally even in the case that they can easily avoid them. Clone change management is one of crucial issues in open source software (OSS) development as well as in industrial software development (e.g., development of social infrastructure, financial system, and medical equipment). When an industrial developer fixes a defect, he/she has to find the code clones corresponding to the code fragment including it. So far, several studies performed on the analysis of clone evolution in OSS. However, to our knowledge, a few researches have been reported on an application of a clone change notification system to industrial development process. In this paper, we introduce a system for notifying creation and change of code clones, and then report on the experience with 40-days application of it into a development process in NEC Corporation. In the industrial application, a developer successfully identified ten unintentionally-developed clones that should be refactored.
Yuki Yamanaka, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue, Tateki Sano
ICPC2