Dong Wang 0044

dblp:40/3934-44 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0002-2004-0902ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 7 first-author · 28 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Humans Integrate, Agents Fix: How Agent-Authored Pull Requests Are Referenced in Practice
abstract
Although coding agents have introduced new coordination dynamics in collaborative software development, detailed interactions in practice remain underexplored, especially for the code review process. In this study, we mine agent-authored PR references from the AIDev dataset [15] and introduce a taxonomy to characterize the intent of these references across Human-to-Agent and Agent-to-Agent interactions in the form of Pull Requests (i.e. PRs). Our analysis shows that while humans initiate most references to agent-authored PRs, a substantial portion of these interactions are AI-assisted, indicating the emergence of meta-collaborative workflows, where humans mostly use references to build new features, whereas agents make them to fix errors.
Islem Khemissi, Moataz Chouchen, Dong Wang 0044, Raula Gaikovina Kula
MSR3
2026 Evaluating large language models for multilingual vulnerability detection at dual granularities
Honglin Shu, Junji Yu, Dong Wang 0044, Chakkrit Tantithamthavorn, Junjie Chen 0003, Yasutaka Kamei
Empir. Softw. Eng.4
2026 An Empirical Study on Language Models for Generating Log Statements in Test Code
abstract
Log statements play a critical role in modern software development, capturing essential run-time information necessary for software maintenance. Recently, new techniques have been developed to automate logging activities, allowing log statements to be injected into code by identifying specific code locations, selecting the appropriate log level, and generating meaningful log messages that describe the behavior being logged. Although automated logging in production code has attracted significant attention, little focus has been given to the injection of logs in test code. To fill this gap, we conduct an empirical study on 5,206,759 Java test methods collected from 6,405 GitHub projects to explore and disclose the effectiveness and limitations of Pre-Trained Language Models (PLMs) and Large Language Models (LLMs) for generating and injecting test log statements. Our findings demonstrate that general-purpose LLMs like GPT-3.5-Turbo, when properly instructed to inject logging statements in test methods, performs comparably to the best-performing PLMs on predicting log level. Additionally, GPT-3.5-Turbo substantially outperforms the best in PLMs on predicting log position, with a 33.97% improvement while also achieving superior performance in predicting log messages in terms of BLEU and ROUGE . This work takes the first step toward evaluating the capability of PLMs and LLMs to generate test log statement. To facilitate future research, we have open sourced all data and source code used in this work.
Honglin Shu, Dong Wang 0044, Antonio Mastropaolo, Gabriele Bavota, Yasutaka Kamei
ACM Trans. Softw. Eng. Methodol.2
2026 Cross-Project Flakiness: A Case Study of the OpenStack Ecosystem
abstract
Automated regression testing is a cornerstone of modern software development, often contributing directly to code review and Continuous Integration (CI). Yet some tests suffer from flakiness, where their outcomes vary non-deterministically. Flakiness erodes developer trust in test results, wastes computational resources, and undermines CI reliability. While prior research has examined test flakiness within individual projects, its broader ecosystem-wide impact remains largely unexplored. In this paper, we present an empirical study of test flakiness in the OpenStack ecosystem, which focuses on (1) cross-project flakiness, where flaky tests impact multiple projects, and (2) inconsistent flakiness, where a test exhibits flakiness in some projects but remains stable in others. By analyzing 649 OpenStack projects, we identify 1,535 cross-project flaky tests and 1,105 inconsistently flaky tests. We find that cross-project flakiness affects 55% of OpenStack projects and significantly increases both review time and computational costs. Surprisingly, 70% of unit tests exhibit cross-project flakiness, challenging the assumption that unit tests are inherently insulated from issues that span modules like integration and system-level tests. Through qualitative analysis, we observe that race conditions in CI, inconsistent build configurations, and dependency mismatches are the primary causes of inconsistent flakiness. These findings underline the need for better coordination across complex ecosystems, standardized CI configurations, and improved test isolation strategies.
Tao Xiao 0001, Dong Wang 0044, Shane McIntosh, Hideaki Hata, Yasutaka Kamei
IEEE Trans. Software Eng.2
2025 Selecting Initial Seeds for Better JVM Fuzzing
abstract
JVM fuzzing techniques serve as a cornerstone for guaranteeing the quality of implementations. In typical fuzzing workflows, initial seeds are crucial as they form the basis of the process. Literature in traditional program fuzzing has confirmed that effectiveness is largely impacted by redundancy among initial seeds, thereby proposing a series of seed selection methods. JVM fuzzing, compared to traditional ones, presents unique characteristics, including large-scale and intricate code, and programs with both syntactic and semantic features. However, it remains unclear whether the existing initial seed selection methods are suitable for JVM fuzzing and whether utilizing program features can enhance effectiveness. To address this, we devise a total of 10 initial seed selection methods, comprising coverage-based, prefuzz-based, and program-feature-based methods. We then conduct an empirical study on three JVM implementations to extensively evaluate the performance of the initial seed selection methods within two state-of-the-art fuzzing techniques (JavaTailor and VECT). Specifically, we examine performance from three aspects: (i) effectiveness and efficiency using widely studied initial seeds, (ii) effectiveness using the programs in the wild, and (iii) the ability to detect new bugs. Evaluation results first show that the program-feature-based method that utilizes the control flow graph not only has a significantly lower time overhead (i.e., 30s), but also outperforms other methods, achieving 142% to 269% improvement compared to the full set of initial seeds. Second, results reveal that the initial seed selection greatly improves the quality of wild programs and exhibits complementary effectiveness by detecting new behaviors. Third, results demonstrate that given the same testing period, initial seed selection improves the JVM fuzzing techniques by detecting more unknown bugs. Particularly, 21 out of the 25 detected bugs have been confirmed or fixed by developers. This work takes the first look at initial seed selection in JVM fuzzing, confirming its importance in fuzzing effectiveness and efficiency.
Tianchang Gao, Junjie Chen 0003, Dong Wang 0044, Yile Guo, Yingquan Zhao
ICSE3
2025 Clarifying Semantics of In-Context Examples for Unit Test Generation
abstract
Recent advances in large language models (LLMs) have enabled promising performance in unit test generation through in-context learning (ICL). However, the quality of in-context examples significantly influences the effectiveness of generated tests—poorly structured or semantically unclear test examples often lead to suboptimal outputs. In this paper, we propose CLAST, a novel technique that systematically refines unit tests to improve their semantic clarity, thereby enhancing their utility as in-context examples. The approach decomposes complex tests into logically clearer ones and improves semantic clarity through a combination of program analysis and LLM-based rewriting. We evaluated CLAST on four open-source and three industrial projects. The results demonstrate that CLAST largely outperforms UTgen, the state-of-the-art refinement technique, in both preserving test effectiveness and enhancing semantic clarity. Specifically, CLAST fully retains the original effectiveness of unit tests, while UTgen reduces compilation success rate (CSR), pass rate (PR), test coverage (Cov), and mutation score (MS) by an average of 12.90%, 35.82%, 4.65%, and 5.07%, respectively. Over 85.33% of participants in our user study preferred the semantic clarity of CLAST-refined tests. Notably, incorporating CLAST-refined tests as examples effectively improves ICL-based unit test generation approaches such as RAGGen and TELPA, resulting in an average increase of 25.97% in CSR, 28.22% in PR, and 45.99% in Cov for generated tests, compared to incorporating UTgen-refined tests. The insights from the follow-up user study not only reinforce CLAST’s potential impact in software testing practice but also illuminate avenues for future research.
Lin Yang 0030, Dong Wang 0044, Junjie Chen 0003
ASE4
2025 An empirical study of test case prioritization on the Linux Kernel
Haichi Wang, Dong Wang 0044, Yiheng Du, Yingquan Zhao, Junjie Chen 0003
Autom. Softw. Eng.3
2025 Developer reactions to protestware in open source software: the cases of color.js and es5.ext
abstract
There is growing concern about maintainers self-sabotaging their work in order to take political or economic stances, a practice referred to as “protestware”. Our objective is to understand the discourse around discussions on such an attack, how it is received by the community, and whether developers respond to the attack in a timely manner. We study two notable protestware cases i.e., colors.js and es5-ext. Results indicate that protestware discussions are spread more quickly on the GitHub platform, while security vulnerabilities are faster on social media. By establishing a taxonomy of protestware discussions, we identify posts that express stances and provide technical mitigation instructions. We applied a thematic analysis to 684 protestware related posts to identify five major themes during the discussions: i. disseminate and response, ii. stance, iii. reputation, iv. communicative styles, v. rights and ethics. This work sheds light on the nuanced landscape of protestware discussions, offering insights for both researchers and developers into maintaining a healthy balance between the political or social actions of developers and the collective well-being of the open-source community.
Youmei Fan, Dong Wang 0044, Supatsara Wattanakriengkrai, Hathaichanok Damrongsiri, Christoph Treude, Hideaki Hata, Raula Gaikovina Kula
Empir. Softw. Eng.2
2025 LEAM++: Learning for Selective Mutation Fault Construction
abstract
Mutation faults are the core of mutation testing and have been widely used in many software testing tasks. Hence, efficiently constructing high-quality mutation faults is critical. To address the effectiveness limitations of traditional and deep learning-based mutation techniques, we first proposed LEAM , utilizing a syntax-guided encoder–decoder architecture with extended grammar rules. While LEAM significantly enhances the effectiveness, it does not consider the associated testing cost. To further improve the efficiency of LEAM , we propose LEAM++ , adopting a novel selective mutation fault construction module based on the probability of grammar rule sequences and the similarity of mutation faults. We extensively evaluate LEAM++ using Defects4J. Regarding effectiveness, the results demonstrate that the mutation faults constructed by LEAM++ can better represent real faults than two traditional techniques ( Major and PIT ) and the deep learning-based technique ( DeepMutation ), and substantially boost three downstream applications, i.e., mutation-based test case prioritization, mutation-based fault localization, and mutation-based bug detection. Regarding efficiency, LEAM++ demonstrates superiority over the four selective mutation testing techniques across three scenarios, i.e., mutation testing, mutation-based test case prioritization, and mutation-based fault localization. Our work serves as an important step toward the efficiently automated construction of mutation faults.
Zhao Tian 0002, Junjie Chen 0003, Dong Wang 0044, Qihao Zhu, Xingyu Fan, Lingming Zhang 0001
ACM Trans. Softw. Eng. Methodol.3
2024 Towards Identifying Code Proficiency Through the Analysis of Python Textbooks
abstract
Python, one of the most prevalent programming languages today, is widely utilized in various domains, including web development, data science, machine learning, and DevOps. Recent scholarly efforts have proposed a methodology to assess Python competence levels, similar to how proficiency in natural languages is evaluated. This method involves assigning levels of competence to Python constructs—for instance, placing simple ‘print’ statements at the most basic level and abstract base classes at the most advanced. The aim is to gauge the level of proficiency a developer must have to understand a piece of source code. This is particularly crucial for software maintenance and evolution tasks, such as debugging or adding new features. For example, in a code review process, this method could determine the competence level required for reviewers. However, categorizing Python constructs by proficiency levels poses significant challenges. Prior attempts, which relied heavily on expert opinions and developer surveys, have led to considerable discrepancies. In response, this paper presents a new approach to identifying Python competency levels through the systematic analysis of introductory Python programming textbooks. By comparing the sequence in which Python constructs are introduced in these textbooks with the current state of the art, we have uncovered notable discrepancies in the order of introduction of Python constructs. Our study underscores a misalignment in the sequences, demonstrating that pinpointing proficiency levels is not trivial. Insights from the study serve as pivotal steps toward reinforcing the idea that textbooks serve as a valuable source for evaluating developers' proficiency, and particularly in terms of their ability to undertake maintenance and evolution tasks.
Ruksit Rojpaisarnkit, Gregorio Robles, Raula Gaikovina Kula, Dong Wang 0044, Chaiyong Ragkhitwetsagul, Jesús M. González-Barahona, Ken-ichi Matsumoto
ICSME4
2024 Large Language Models for Equivalent Mutant Detection: How Far Are We?
abstract
Mutation testing is vital for ensuring software quality. However, the presence of equivalent mutants is known to introduce redundant cost and bias issues, hindering the effectiveness of mutation testing in practical use. Although numerous equivalent mutant detection (EMD) techniques have been proposed, they exhibit limitations due to the scarcity of training data and challenges in generalizing to unseen mutants. Recently, large language models (LLMs) have been extensively adopted in various code-related tasks and have shown superior performance by more accurately capturing program semantics. Yet the performance of LLMs in equivalent mutant detection remains largely unclear. In this paper, we conduct an empirical study on 3,302 method-level Java mutant pairs to comprehensively investigate the effectiveness and efficiency of LLMs for equivalent mutant detection. Specifically, we assess the performance of LLMs compared to existing EMD techniques, examine the various strategies of LLMs, evaluate the orthogonality between EMD techniques, and measure the time overhead of training and inference. Our findings demonstrate that LLM-based techniques significantly outperform existing techniques (i.e., the average improvement of 35.69% in terms of F1-score), with the fine-tuned code embedding strategy being the most effective. Moreover, LLM-based techniques offer an excellent balance between cost (relatively low training and inference time) and effectiveness. Based on our findings, we further discuss the impact of model size and embedding quality, and provide several promising directions for future research. This work is the first to examine LLMs in equivalent mutant detection, affirming their effectiveness and efficiency.
Zhao Tian 0002, Honglin Shu, Dong Wang 0044, Xuejie Cao, Yasutaka Kamei, Junjie Chen 0003
ISSTA3
2024 Exploring the Effect of Multiple Natural Languages on Code Suggestion Using GitHub Copilot
abstract
GitHub Copilot is an AI-enabled tool that automates program synthesis. It has gained significant attention since its launch in 2021. Recent studies have extensively examined Copilot's capabilities in various programming tasks, as well as its security issues. However, little is known about the effect of different natural languages on code suggestion. Natural language is considered a social bias in the field of NLP, and this bias could impact the diversity of software engineering. To address this gap, we conducted an empirical study to investigate the effect of three popular natural languages (English, Japanese, and Chinese) on Copilot. We used 756 questions of varying difficulty levels from AtCoder contests for evaluation purposes. The results highlight that the capability varies across natural languages, with Chinese achieving the worst performance. Furthermore, regardless of the type of natural language, the performance decreases significantly as the difficulty of questions increases. Our work represents the initial step in comprehending the significance of natural languages in Copilot's capability and introduces promising opportunities for future endeavors.
Kei Koyanagi, Dong Wang 0044, Kotaro Noguchi, Masanari Kondo, Alexander Serebrenik, Yasutaka Kamei, Naoyasu Ubayashi
MSR2
2024 GitRev: An LLM-Based Gamification Framework for Modern Code Review Activities
abstract
Modern code review (MCR) is recognized as an effective software quality assurance practice that is broadly adopted by open-source and commercial software projects. MCR is most effective when developers follow best practices, as it improves code quality, enhances knowledge transfer, increases team awareness and shares code ownership. However, prior work highlights that poor code review practices are common and often manifest in the form of low review participation and engagement, shallow review, and toxic communications. To address these issues, we introduce GitRev, a novel approach that applies gamification mechanisms to boost developer motivation and engagement. GitRev is built on top of a Large Language Model (LLM), used as a points-based reward system that leverages the code change context, and code review activities. We implement GitRev as a GitHub app with a web browser extension that consists of a client-side web browser extension that gamifies the GitHub user interface, and a server-side composed of a Node.js server for authentication and data management. To evaluate GitRev, we conduct a controlled experiment with 86 graduate and undergraduate students. Results indicate the promising potential of our approach for improving the code review process and developers' engagement. GitRev is publicly available at https://anonymous.40pen.science/r/GitRev-OB74
Jasem Khelifi, Moataz Chouchen, Ali Ouni 0001, Dong Wang 0044, Raula Gaikovina Kula, Salma Hamza, Mohamed Wiem Mkaouer
SCAM4
2024 Understanding the characteristics and the role of visual issue reports
Hiroki Kuramoto, Dong Wang 0044, Masanari Kondo, Yutaro Kashiwa, Yasutaka Kamei, Naoyasu Ubayashi
Empir. Softw. Eng.2
2024 Quantifying and characterizing clones of self-admitted technical debt in build systems
Tao Xiao 0001, Zhili Zeng, Dong Wang 0044, Hideaki Hata, Shane McIntosh, Ken-ichi Matsumoto
Empir. Softw. Eng.3
2024 A Disruptive Research Playbook for Studying Disruptive Innovations
abstract
As researchers today, we are witnessing a fundamental change in our technologically-enabled world due to the advent and diffusion of highly disruptive technologies such as generative Artificial Intelligence (AI), Augmented Reality (AR) and Virtual Reality (VR). In particular, software engineering has been profoundly affected by the transformative power of disruptive innovations for decades, with a significant impact of technical advancements on social dynamics due to its socio-technical nature. In this article, we reflect on the importance of formulating and addressing research problems in software engineering through a socio-technical lens, thus ensuring a holistic understanding of the complex phenomena in this field. We propose a research playbook with the aim of providing a guide to formulate compelling and socially relevant research questions and to identify the appropriate research strategies for empirical investigations, with an eye on the long-term implications of technologies or their use. We showcase how to apply the research playbook. Firstly, we show how it can be used retrospectively to reflect on a prior disruptive technology, Stack Overflow, and its impact on software development. Secondly, we show how it can be used to question the impact of two current disruptive technologies: AI and AR/VR. Finally, we introduce a specialized GPT model to support the researcher in framing future investigations. We conclude by discussing the broader implications of adopting the playbook for both researchers and practitioners in software engineering and beyond.
Margaret-Anne D. Storey, Daniel Russo 0002, Nicole Novielli, Takashi Kobayashi 0001, Dong Wang 0044
ACM Trans. Softw. Eng. Methodol.5
2023 Repeated Builds During Code Review: An Empirical Study of the OpenStack Community
abstract
Code review is a popular practice where developers critique each others' changes. Since automated builds can identify low-level issues (e.g., syntactic errors, regression bugs), it is not uncommon for software organizations to incorporate automated builds in the code review process. In such code review deployment scenarios, submitted change sets must be approved for integration by both peer code reviewers and automated build bots. Since automated builds may produce an unreliable signal of the status of a change set (e.g., due to “flaky” or non-deterministic execution behaviour), code review tools, such as Gerrit, allow developers to request a “recheck”, which repeats the build process without updating the change set. We conjecture that an unconstrained recheck command will waste time and resources if it is not applied judiciously. To explore how the recheck command is applied in a practical setting, in this paper, we conduct an empirical study of 66,932 code reviews from the OpenStack community. We quantitatively analyze (i) how often build failures are rechecked; (ii) the extent to which invoking recheck changes build failure outcomes; and (iii) how much waste is generated by invoking recheck. We observe that (i) 55% of code reviews invoke the recheck command after a failing build is reported; (ii) invoking the recheck command only changes the outcome of a failing build in 42% of the cases; and (iii) invoking the recheck command increases review waiting time by an average of 2,200% and equates to 187.4 compute years of waste-enough compute resources to compete with the oldest land living animal on earth. Our observations indicate that the recheck command is frequently used after the builds fail, but does not achieve a high likelihood of build success. Based on a developer survey and our history-based quantitative findings, we encourage reviewer teams to think twice before rechecking and be considerate of waste. While recheck currently generates plenty of wasted computational resources and bloats waiting times, it also presents exciting future opportunities for researchers and tool builders to propose solutions that can reduce waste.
Rungroj Maipradit, Dong Wang 0044, Patanamon Thongtanunam, Raula Gaikovina Kula, Yasutaka Kamei, Shane McIntosh
ASE2
2023 Understanding the Role of Images on Stack Overflow
abstract
Images are increasingly being shared by software developers in diverse channels including question-and-answer forums like Stack Overflow. Although prior work has pointed out that these images are meaningful and provide complementary information compared to their associated text, how images are used to support questions is empirically unknown. To address this knowledge gap, in this paper we specifically conduct an empirical study to investigate (I) the characteristics of images, (II) the extent to which images are used in different question types, and (III) the role of images on receiving answers. Our results first show that user interface is the most common image content and undesired output is the most frequent purpose for sharing images. Moreover, these images essentially facilitate the understanding of 68% of sampled questions. Second, we find that discrepancy questions are more relatively frequent compared to those without images, but there are no significant differences observed in description length in all types of questions. Third, the quantitative results statistically validate that questions with images are more likely to receive accepted answers, but do not speed up the time to receive answers. Our work demonstrates the crucial role that images play by approaching the topic from a new angle and lays the foundation for future opportunities to use images to assist in tasks like generating questions and identifying question-relatedness.
Dong Wang 0044, Tao Xiao 0001, Christoph Treude, Raula Gaikovina Kula, Hideaki Hata, Yasutaka Kamei
MSR1
2023 Towards Privacy Preserving Cross Project Defect Prediction with Federated Learning
abstract
Defect prediction models can predict defects in software projects, and many researchers study defect prediction models to assist debugging efforts in software development. In recent years, there has been growing interest in Cross Project Defect Prediction (CPDP), which predicts defects in a project using a defect prediction model learned from other projects’ data when there is insufficient data to construct a defect prediction model. Since CPDP uses other projects’ data, data privacy preservation is one of the most significant issues. However, prior CPDP studies still require data sharing among projects to train models, and do not fully consider protecting project confidentiality. To address this, we propose a CPDP model FLR employing federated learning, a distributed machine learning approach that does not require data sharing. We evaluate FLR, using 25 projects, to investigate its effectiveness and feature interpretation. Our key results show that first, FLR outperforms the existing privacy-preserving methods (i.e., LACE2). Meanwhile, the performance is relatively comparable to the conventional methods (e.g., supervised and unsupervised learning). Second, the results of the interpretation analysis show that scale-related features have a common effect on the prediction performance of the FLR. In addition, further insights demonstrate that parameters of federated learning (e.g., learning rates and the number of clients) also play a role in the performance. This study is served as a first step to confirm the feasibility of the employment of federated learning in CPDP to ensure privacy preservation and lays the groundwork for future research on applying other machine learning models to federated learning.
Hiroki Yamamoto, Dong Wang 0044, Gopi Krishnan Rajbahadur, Masanari Kondo, Yasutaka Kamei, Naoyasu Ubayashi
SANER2
2023 When conversations turn into work: a taxonomy of converted discussions and issues in GitHub
Dong Wang 0044, Masanari Kondo, Yasutaka Kamei, Raula Gaikovina Kula, Naoyasu Ubayashi
Empir. Softw. Eng.1
2023 More than React: Investigating the Role of Emoji Reaction in GitHub Pull Requests
Dong Wang 0044, Tao Xiao 0001, Teyon Son, Raula Gaikovina Kula, Takashi Ishio, Yasutaka Kamei, Ken-ichi Matsumoto
Empir. Softw. Eng.1
2023 Giving Back: Contributions Congruent to Library Dependency Changes in a Software Ecosystem
abstract
The widespread adoption of third-party libraries for contemporary software development has led to the creation of large inter-dependency networks, where sustainability issues of a single library can have widespread network effects. Maintainers of these libraries are often overworked, relying on the contributions of volunteers to sustain these libraries. To understand these contributions, in this work, we leverage socio-technical techniques to introduce and formalise dependency-contribution congruence (DC congruence) at both ecosystem and library level, i.e., to understand the degree and origins of contributions congruent to dependency changes, analyze whether they contribute to library dormancy (i.e., a lack of activity), and investigate similarities between these congruent contributions compared to typical contributions. We conduct a large-scale empirical study to measure the DC congruence for the npm ecosystem using 1.7 million issues, 970 thousand pull requests (PRs), and over 5.3 million commits belonging to 107,242 npm libraries. We find that the most congruent contributions originate from contributors who can only submit (not commit) to both a client and a library. At the project level, we find that DC congruence shares an inverse relationship with the likelihood that a library becomes dormant. Specifically, a library is less likely to become dormant if the contributions are congruent with upgrading dependencies. Finally, by comparing the source code of contributions, we find statistical differences in the file path and added lines in the source code of congruent contributions when compared to typical contributions. Our work has implications to encourage dependency contributions, especially to support library maintainers in sustaining their projects.
Supatsara Wattanakriengkrai, Dong Wang 0044, Raula Gaikovina Kula, Christoph Treude, Patanamon Thongtanunam, Takashi Ishio, Ken-ichi Matsumoto
IEEE Trans. Software Eng.2
2022 Newcomer OSS-Candidates: Characterizing Contributions of Novice Developers to GitHub
abstract
Abstract The ability of an Open Source Software (OSS) project to attract, onboard, and retain any newcomer is vital to its livelihood. Although, evidence suggests an upsurge in novice developers joining social coding platforms (such as GitHub), the extent to which their activities result in a OSS contribution is unknown. Henceforth, we execute the protocols of a registered report to study activities of a “Newcomer OSS-Candidate”, who is a novice developer that is new to that social coding platform, and has the intention to later onboard an OSS project. Using GitHub as a case platform, we analyze 171 identified Newcomer OSS-Candidates to characterize their contribution activities. Results show that Newcomer OSS-Candidates are likely to target software based repositories (i.e., 66%), and their first contributions are mainly associated with development (commits) and maintenance (PRs). Newcomer OSS-Candidates are less likely to practice social coding, but eventually end up onboarding (i.e., 30% quantitative, 70% follow-up survey) an OSS project. Furthermore, they cite finding a way to start as the most challenging barrier to contribute. Our work reveals insights on how newcomers to social coding platforms are potential sources of OSS contributions.
Ifraz Rehman, Dong Wang 0044, Raula Gaikovina Kula, Takashi Ishio, Ken-ichi Matsumoto
Empir. Softw. Eng.2
2022 Characterizing and Mitigating Self-Admitted Technical Debt in Build Systems
abstract
Technical Debt is a metaphor used to describe the situation in which long-term software artifact quality is traded for short-term goals in software projects. In recent years, the concept of self-admitted technical debt (SATD) was proposed, which focuses on debt that is intentionally introduced and described by developers. Although prior work has made important observations about admitted technical debt in source code, little is known about SATD in build systems. In this paper, we set out to better understand the characteristics of SATD in build systems. To do so, through a qualitative analysis of 500 SATD comments in the Maven build system of 291 projects, we characterize SATD by location and rationale (reason and purpose). Our results show that limitations in tools and libraries, and complexities of dependency management are the most frequent causes, accounting for 50% and 24% of the comments. We also find that developers often document SATD as issues to be fixed later. As a first step towards the automatic detection of SATD rationale, we train classifiers to detect the two most frequently occurring reasons and the four most frequently occurring purposes of SATD in the content of comments in Maven build systems. The classifier performance is promising, achieving an F1-score of 0.71–0.79. Finally, within 16 identified ‘ready-to-be-addressed’ SATD instances, the three SATD submitted by pull requests and the five SATD submitted by issue reports were resolved after developers were made aware. Our work presents the first step towards understanding technical debt in build systems and opens up avenues for future work, such as tool support to track and manage SATD backlogs.
Tao Xiao 0001, Dong Wang 0044, Shane McIntosh, Hideaki Hata, Raula Gaikovina Kula, Takashi Ishio, Ken-ichi Matsumoto
IEEE Trans. Software Eng.2
2021 An End-to-End Sleep Staging Simulator Based on Mixed Deep Neural Networks
abstract
Sleep screening is not only a major tool in the assessment of pathophysiology, but also a bridge between the central neuronal systems and behaviour/cognition. Automatic sleep staging is an alternative for the time-consuming gold standard manual scoring procedure. Most of the existing works designed such procedure by using deep neural networks without considering the medical criterion of the sleep staging task. We argue that capturing the stage-specific features which meet the criterion is of significant importance for the automatic sleep staging alternative. In this work we propose an end-to-end sleep staging simulator based on mixed neural networks, i.e., CNN, LSTM, and Transformer. The framework consists of two subnetworks: stage dependent feature mapping network which is constructed by the idea of physiological sleep nature, and an attention-based parallel staging network. Moreover, we adopt a mixed precision training strategy to quantize the model for exploring feasible usage in the clinical settings. Through an experiment with a large EEG database (Sleep Heart Health Study), the proposed method has a competitive stage scoring performance, especially in stages Wake, N2, and N3, with higher precision of 0.92, 0.85, and 0.86, respectively. Our study proves that the quantized model has potential capability for further application in the clinical staging task.
Zheng Chen 0012, Ziwei Yang 0002, Dong Wang 0044, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM3
2021 An Integrated Multi-Omics Approach for AMR Phenotype Prediction of Gut Microbiota
abstract
The gut microbiota is crucial for human physiology and susceptibility to diseases. Knowing the AMR phenotype canfacilitate the understanding of the impact of antibiotics administration on the gut microbiota. Nowadays, whole-genome sequencing for antibiotic susceptibility testing (WGS-AST) is widely used in clinical microbiology to predict the AMR phenotype. To release the limitations of the genomic information and improve the WGS-AST prediction, we propose an integrated multi-omics approach, employing a deep generative neural network (VAE: variational auto-encoder). We evaluate the proposed approach by two machine learning techniques (i.e., K-means for clustering and Random Forest for classification). Our evaluation results show that the integrated multi-omics approach achieves relatively better performance than the conventional WGS-AST. Moreover, the integrated multi-omics approach is able to visually reveal AMR phenotype of the gutmicrobiota via antibacterial spectrum. Our work provides evidence that multi-omics information is useful to enhance the WGS-AST prediction.
Pei Gao, Zheng Chen 0012, Dong Wang 0044, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM3
2021 Exploring Feasibility of Truth-Involved Automatic Sleep Staging Combined with Transformer
abstract
Recently, deep learning-based methods have been successfully proposed for electrophysiology signal-based sleep staging with promising results. Most existing methods use convolutional layers and recurrent-based architectures to implement a model structure from feature extraction to sequence signal classification. In this study, we propose a method of segmenting electroencephalogram (EEG) and electrooculogram (EOG) data according to frequency bands and construct a Transformer based automatic sleep classification model on top of it. The results show that the classifications of the stage Wake, N3, and REM outperform the state-of-art works, with the Fl-scores of 0.92, 0.85 and 0.91. Our work is the first attempt to explore the feasibility of a truth-involved Transformer-based model with a large-scale sleep database.
Ziwei Yang 0002, Dong Wang 0044, Zheng Chen 0012, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM2
2021 Anti-patterns in Modern Code Review: Symptoms and Prevalence
abstract
Modern code review (MCR) is now broadly adopted as an established and effective software quality assurance practice, with an increasing number of open-source as well as commercial software projects identifying code review as a crucial practice. During the MCR process, developers review, provide constructive feedback, and/or critique each others’ patches before a code change is merged into the codebase. Nevertheless, code review is basically a human task that involves technical, personal and social aspects. Existing literature hint the existence of poor reviewing practices i.e., anti-patterns, that may contribute to a tense reviewing culture, degradation of software quality, slow down integration, and may affect the overall sustainability of the project. To better understand these practices, we present in this paper the concept of Modern Code Review Anti-patterns (MCRA) and take a first step to define a catalog that enumerates common poor code review practices. In detail we explore and characterize MCRA symptoms, causes, and impacts. We also conduct a series of preliminary experiments to investigate the prevalence and co-occurrences of such anti-patterns on a random sample of 100 code reviews from various OpenStack projects.
Moataz Chouchen, Ali Ouni 0001, Raula Gaikovina Kula, Dong Wang 0044, Patanamon Thongtanunam, Mohamed Wiem Mkaouer, Ken-ichi Matsumoto
SANER4
2021 Understanding shared links and their intentions to meet information needs in modern code review
abstract
Abstract Code reviews serve as a quality assurance activity for software teams. Especially for Modern Code Review, sharing a link during a review discussion serves as an effective awareness mechanism where “Code reviews are good FYIs [for your information].”. Although prior work has explored link sharing and the information needs of a code review, the extent to which links are used to properly conduct a review is unknown. In this study, we performed a mixed-method approach to investigate the practice of link sharing and their intentions. First, through a quantitative study of the OpenStack and Qt projects, we identify 19,268 reviews that have 39,686 links to explore the extent to which the links are shared, and analyze a correlation between link sharing and review time. Then in a qualitative study, we manually analyze 1,378 links to understand the role and usefulness of link sharing. Results indicate that internal links are more widely referred to (93% and 80% for the two projects). Importantly, although the majority of the internal links are referencing to reviews, bug reports and source code are also shared in review discussions. The statistical models show that the number of internal links as an explanatory factor does have an increasing relationship with the review time. Finally, we present seven intentions of link sharing, with providing context being the most common intention for sharing links. Based on the findings and a developer survey, we encourage the patch author to provide clear context and explore both internal and external resources, while the review team should continue link sharing activities. Future research directions include the investigation of causality between sharing links and the review process, as well as the potential for tool support.
Dong Wang 0044, Tao Xiao 0001, Patanamon Thongtanunam, Raula Gaikovina Kula, Ken-ichi Matsumoto
Empir. Softw. Eng.1
2021 Automatic patch linkage detection in code review using textual content and file location features
Dong Wang 0044, Raula Gaikovina Kula, Takashi Ishio, Ken-ichi Matsumoto
Inf. Softw. Technol.1
2021 Can we benchmark Code Review studies? A systematic mapping study of methodology, dataset, and metric
Dong Wang 0044, Yuki Ueda, Raula Gaikovina Kula, Takashi Ishio, Ken-ichi Matsumoto
J. Syst. Softw.1
2020 Newcomer Candidate: Characterizing Contributions of a Novice Developer to GitHub
abstract
To attract, onboard, and retain any newcomer in Open Source Software (OSS) projects is vital to their livelihood. Recent studies conclude that OSS projects risk failure due to abandonment and poor participation of newcomers. Evidence suggests more new users are joining GitHub, however, the extent to which they contribute to OSS projects is unknown. In this study, we coin the term `newcomer candidate' to describe new users to the GitHub platform. Our objective is to track and characterize their initial contributions. As a preliminary survey, we collected 208 newcomer candidate contributions in GitHub. Using this dataset, we then plan to track their contributions to reveal insights. We will use a mixed-methods approach, i.e., quantitative and qualitative, to identify whether or not newcomer candidates practice social coding, the kinds of their contributions, projects they target, and the proportion that they eventually onboard to an OSS project.
Ifraz Rehman, Dong Wang 0044, Raula Gaikovina Kula, Takashi Ishio, Ken-ichi Matsumoto
ICSME2
2018 An Exploratory Study to Identify Similar Patches: A Case Study in Modern Code Review
abstract
Due to the distributed nature of Modern Code Review (MCR) tools, developers risk submitting similar patches (i.e., patches that attempt to achieve similar objectives), which potentially causes extra efforts both for the contributors and reviewers. Although researches on other duplicate software artifact exist, there is no prior work that explores the impact of such similar patches in MCR. In this paper, we conduct an empirical study to understand the impact of similar patches on reviewing efforts in MCR. We extracted over 3,400 similar patches from the OpenStack project. Results of the exploratory study confirm that similar patches take just as much time and patch revisions as merged patches.
Dong Wang 0044, Raula Gaikovina Kula, Ken-ichi Matsumoto
APSEC1