Dong Shao

dblp:29/6082 · DBLP profile ↗
← Back
42ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 32 · 15 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions
Hongyu Kuang, Sebastian Baltes, Xin Zhou 0016, He Zhang 0001, Xiaoxing Ma, Guoping Rong, Dong Shao, Christoph Treude
IEEE Trans. Software Eng.8
2025 Brevity is the Soul of Wit: Condensing Code Changes to Improve Commit Message Generation
abstract
Commit messages are valuable resources for describing why code changes are committed to repositories in version control systems (e.g., Git).They effectively help developers understand code changes and better perform software maintenance tasks.Unfortunately, developers often neglect to write high-quality commit messages in practice.Therefore, a growing body of work is proposed to generate commit messages automatically.These works all demonstrated that how to organize and represent code changes is vital in generating good commit messages, including the use of fine-grained graphs or embeddings to better represent code changes.In this study, we choose an alternative way to condense code changes before generation, i.e., proposing brief yet concise text templates consisting of the following three parts: (1) summarized code changes, (2) elicited comments, and (3) emphasized code identifiers.Specifically, we first condense code changes by using our proposed templates with the help of a heuristic-based tool named ChangeScribe, and then fine-tune CodeLlama-7B on the pairs of our proposed templates and corresponding commit messages.Our proposed templates better utilize pre-trained language models, while being naturally brief and readable to complement generated commit messages for developers.
Hongyu Kuang, Xin Zhou 0016, Wesley K. G. Assunção, Xiaoxing Ma, Dong Shao, Guoping Rong, He Zhang 0001
Internetware7
2025 AUCAD: Automated Construction of Alignment Dataset from Log-Related Issues for Enhancing LLM-based Log Generation
abstract
Log statements have become an integral part of modern software systems.Prior research efforts have focused on supporting the decisions of placing log statements, such as where/what to log.With the increasing adoption of Large Language Models (LLMs) for coderelated tasks such as code completion or generation, automated approaches for generating log statements have gained much momentum.However, the performance of these approaches still has a long way to go.This paper explores enhancing the performance of LLM-based solutions for automated log statement generation by post-training LLMs with a purpose-built dataset.Thus the primary contribution is a novel approach called AUCAD, which automatically constructs such a dataset with information extracting from log-related issues.Researchers have long noticed that a significant portion of the issues in the open-source community are related to log statements.However, distilling this portion of data requires manual efforts, which is labor-intensive and costly, rendering it impractical.Utilizing our approach, we automatically extract logrelated issues from 1,537 entries of log data across 88 projects and identify 808 code snippets (i.e., methods) with retrievable source code both before and after modification of each issue (including log statements) to construct a dataset.Each entry in the dataset consists of a data pair representing high-quality and problematic log statements, respectively.With this dataset, we proceed to post-train multiple LLMs (primarily from the Llama series) for automated * Corresponding author.
Hao Zhang 0210, Dongjun Yu, Lei Zhang 0160, Guoping Rong, Yongda Yu, Haifeng Shen, He Zhang 0001, Dong Shao, Hongyu Kuang
Internetware8
2025 Correction to: A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli
Empir. Softw. Eng.6
2025 Fine-Tuning Large Language Models to Improve Accuracy and Comprehensibility of Automated Code Review
abstract
As code review is a tedious and costly software quality practice, researchers have proposed several machine learning-based methods to automate the process. The primary focus has been on accuracy, that is, how accurately the algorithms are able to detect issues in the code under review. However, human intervention still remains inevitable since results produced by automated code review are not 100% correct. To assist human reviewers in making their final decisions on automatically generated review comments, the comprehensibility of the comments underpinned by accurate localization and relevant explanations for the detected issues with repair suggestions is paramount. However, this has largely been neglected in the existing research. Large language models (LLMs) have the potential to generate code review comments that are more readable and comprehensible by humans, thanks to their remarkable processing and reasoning capabilities. However, even mainstream LLMs perform poorly in detecting the presence of code issues because they have not been specifically trained for this binary classification task required in code review. In this article, we contribute Comprehensibility of Automated Code Review using Large Language Models ( Carllm ), a novel fine-tuned LLM that has the ability to improve not only the accuracy but, more importantly, the comprehensibility of automated code review, as compared to state-of-the-art pre-trained models and general LLMs.
Yongda Yu, Guoping Rong, Haifeng Shen, He Zhang 0001, Dong Shao, Zhao Wei, Juhong Wang
ACM Trans. Softw. Eng. Methodol.5
2024 An Explainable Automated Model for Measuring Software Engineer Contribution
abstract
Software engineers play an important role throughout the software development life-cycle, particularly in industry emphasizing quality assurance and timely delivery. Contribution measurement provides proper incentives to software engineers that motivate them to continuously improve the quality and efficiency of their work. However, existing research tends to ignore contribution measurement for software engineers in practice, relying heavily on peer review and lacking objectivity and transparency. Specifically, these studies still have two weaknesses. First, a few studies explore which metrics can be useful for contribution measurement in practice. Second, managers measure the contribution of software engineers based on their experience and lack of explainable automated tools to assist them.
Yue Li 0047, He Zhang 0001, Yuzhe Jin, Liming Dong 0001, Lanxin Yang, David Lo 0001, Dong Shao
ASE9
2024 AVIATE: Exploiting Translation Variants of Artifacts to Improve IR-based Traceability Recovery in Bilingual Software Projects
abstract
Traceability plays a vital role in facilitating various software development activities by establishing the traces between different types of artifacts (e.g., issues and commits in software repositories). Among the explorations for automated traceability recovery, the IR (Information Retrieval)-based approaches leverage textual similarity to measure the likelihood of traces between artifacts and show advantages in many scenarios. However, the globalization of software development has introduced new challenges, such as the possible multilingualism on the same concept (e.g., "[SEE PDF]" vs. "attribute") in the artifact texts, thus significantly hampering the performance of IR-based approaches. Existing research has shown that machine translation can help address the term inconsistency in bilingual projects. However, the translation can also bring in synonymous terms that are not consistent with those in the bilingual projects (e.g., another translation of "[SEE PDF]" as "property"). Therefore, we propose an enhancement strategy called AVIATE that exploits translation variants from different translators by utilizing the word pairs that appear simultaneously across the translation variants from different kinds artifacts (a.k.a. consensual biterms). We use these biterms to first enrich the artifact texts, and then to enhance the calculated IR values for improving IR-based trace-ability recovery for bilingual software projects. The experiments on 17 bilingual projects (involving English and 4 other languages) demonstrate that AVIATE significantly outperformed the IR-based approach with machine translation (the state-of-the-art in this field) with an average increase of 16.67 in Average Precision (31.43%) and 8.38 (11.22%) in Mean Average Precision, indicating its effectiveness in addressing the challenges of multilingual traceability recovery.
Yiding Ren, Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001
ASE7
2024 A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli
Empir. Softw. Eng.6
2024 Towards a security-optimized approach for the microservice-oriented decomposition
abstract
Abstract Microservice architecture (MSA) is a mainstream architectural style due to its high maintainability and scalability. In practice, an appropriate microservice‐oriented decomposition is the foundation to make a system enjoy the benefits of MSA. In terms of decomposing monolithic systems into microservices, researchers have been exploring many optimization objectives, of which modularity is a predominantly focused quality attribute. Security is also a critical quality attribute, that measures the extent to which a system protects data from malicious access or use by attackers. Considering security in microservices‐oriented decomposition can help avoid the risk of leaking critical data and other unexpected software security issues. However, few researchers consider the security objective during microservice‐oriented decomposition, because the measurement of security and the trade‐off with other objectives are challenging in reality. To bridge this research gap, we propose a security‐optimized approach for microservice‐oriented decomposition (So4MoD). In this approach, we adapt five metrics from previous studies for the measurement of the data security of candidate microservices. A multi‐objective optimization algorithm based on NSGA‐II is designed to search for microservices with optimized security and modularity. To validate the effectiveness of the proposed So4MoD, we perform several experiments on eight open‐source projects and compare the decomposition results to other three state‐of‐the‐art approaches, that is, FoSCI, CO‐GCN, and MSExtractor. The experiment results show that our approach can achieve at least an 11.5% improvement in terms of security metrics. Moreover, the decomposition results of So4MoD outperform other approaches in four modularity metrics, demonstrating that So4MoD can optimize data security while pursuing a well‐modularized MSA.
Chenxing Zhong, Shanshan Li 0002, Dong Shao
J. Softw. Evol. Process.7
2024 Distilling Quality Enhancing Comments From Code Reviews to Underpin Reviewer Recommendation
abstract
Code review is an important practice in software development. One of its main objectives is for the assurance of code quality. For this purpose, the efficacy of code review is subject to the credibility of reviewers, i.e., reviewers who have demonstrated strong evidence of previously making quality-enhancing comments are more credible than those who have not. Code reviewer recommendation (CRR) is designed to assist in recommending suitable reviewers for a specific objective and, in this context, assurance of code quality. Its performance is susceptible to the relevance of its training dataset to this objective, composed of all reviewers’ historical review comments, which, however, often contains a plethora of comments that are irrelevant to the enhancement of code quality. Furthermore, recommendation accuracy has been adopted as the sole metric to evaluate a recommender's performance, which is inadequate as it does not take reviewers’ relevant credibility into consideration. These two issues form the ground truth problem in CRR as they both originate from the relevance of dataset used to train and evaluate CRR algorithms. To tackle this problem, we first propose the concept of Quality-Enhancing Review Comments (QERC), which includes three types of comments - change-triggering inline comments, informative general comments, and approve-to-merge comments. We then devise a set of algorithms and procedures to obtain a distilled dataset by applyingQERCto the original dataset. We finally introduce a new metric – reviewer's credibility for quality enhancement (RCQE) – as a complementary metric to recommendation accuracy for evaluating the performance of recommenders. To validate the proposed QERC-based approach to CRR, we conduct empirical studies using real data from seven projects containing over 82K pull requests and 346K review comments. Results show that: (a)QERCcan effectively address the ground truth problem by distilling quality-enhancing comments from the dataset containing original code reviews, (b)QERCcan assist recommenders in finding highly credible reviewers at a slight cost of recommendation accuracy, and (c) even “wrong” recommendations using the distilled dataset are likely to be more credible than those using the original dataset.
Guoping Rong, Yongda Yu, He Zhang 0001, Haifeng Shen, Dong Shao, Hongyu Kuang, Zhao Wei, Juhong Wang
IEEE Trans. Software Eng.6
2023 Locating Anomaly Clues for Atypical Anomalous Services: An Industrial Exploration
abstract
Continuity and steadiness are vital for services with massive users, which requires the anomalies of services should be detected and resolved in a timely manner. Our previous work proposed a tool, namelyImpAPTr (Impact Analysis based on Pruning Tree), to identify the combination of multiple dimensional attributes as the clues leading to the root cause of service anomalies. However,ImpAPTrapplies a threshold driven strategy, i.e., it needs to be triggered by a$\geq 0.05\%$drop of the success rate of the service calls (abbr.SRSC), which may face problems in an atypical yet pervasive situation in field application. For example, the combination of trivial anomalies (i.e., each causes a drop less than 0.05% toSRSC) can lead to a far more than 0.05% drop onSRSC. Besides, a suitable threshold is usually hard to be determined, etc. To address these problems, we propose a new method, namelyImpAPTr+in this paper to free the constraint of the 0.05% threshold. The basic idea is to involve time dimension and identify clues across multiple time intervals of data. We performed evaluation on three typical methods (i.e.,ImpAPTr+,R-AdtributorandSqueeze) with both production environment dataset and simulation dataset. The former dataset is directly retrieved from the service monitoring data inMeituan, one of the largest on-line service providers worldwide. The latter dataset is fabricated also using the monitoring data from the same company. The results indicate: (1)ImpAPTr+outperforms previous approaches to a large degree in terms of accuracy. (2) BothImpAPTr+andR-Adtributorare able to find proper clues within seconds. (3)ImpAPTr+tends to find proper clues with shorter time intervals (i.e., less data), which implies that the method is more suitable for near real-time monitoring scenarios.
Guoping Rong, Shenghui Gu, Yangchen Xu, Dong Shao, He Zhang 0001
IEEE Trans. Dependable Secur. Comput.6
2022 Incorporating Pre-trained Transformer Models into TextCNN for Sentiment Analysis on Software Engineering Texts
abstract
Software information sites (e.g., Jira, Stack Overflow) are now wide-ly used in software development. These online platforms for collaborative development preserve a large amount of Software Engineering (SE) texts. These texts enable researchers to detect developers’ attitudes toward their daily development by analyzing the sentiments expressed in the texts. Unfortunately, recent works reported that neither off-the-shelf tools nor SE-specified tools for sentiment analysis on SE texts can provide satisfying and reliable results. In this paper, we propose to incorporate pre-trained transformer models into the sentence-classification oriented deep learning framework named TextCNN to better capture the unique expression of sentiments in SE texts. Specifically, we introduce an optimized BERT model named RoBERTa as the word embedding layer of TextCNN, along with additional residual connections between RoBERTa and TextCNN for better cooperation in our training framework. An empirical evaluation based on four datasets from different software information sites shows that our training framework can achieve overall better accuracy and generalizability than the four baselines.
Xiaobo Shi, Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001
Internetware7
2022 Using Consensual Biterms from Text Structures of Requirements and Code to Improve IR-Based Traceability Recovery
abstract
Traceability approves trace links among software artifacts based on whether two artifacts are related by system functionalities. The traces are valuable for software development, but are difficult to obtain manually. To cope with the costly and fallible manual recovery, automated approaches are proposed to recover traces through textual similarities among software artifacts, such as those based on Information Retrieval (IR). However, the low quality & quantity of artifact texts negatively impact the calculated IR values, thus greatly hindering the performance of IR-based approaches. In this study, we propose to extract co-occurred word pairs from the text structures of both requirements and code (i.e., consensual biterms) to improve IR-based traceability recovery. We first collect a set of biterms based on the part-of-speech of requirement texts, and then filter them through the code texts. We then use these consensual biterms to both enrich the input corpus for IR techniques and enhance the calculations of IR values. A nine-system-based evaluation shows that in general, when solely used to enhance IR techniques, our approach can outperform pure IR-based approaches and another baseline by 21.9% & 21.8% in AP, and 9.3% & 7.2% in MAP, respectively. Moreover, when used to collaborate with another enhancing strategy from different perspectives, it can outperform this baseline by 5.9% in AP and 4.8% in MAP.
Hongyu Kuang, Xiaoxing Ma, Alexander Egyed, Patrick Mäder, Guoping Rong, Dong Shao, He Zhang 0001
ASE8
2021 A Research Landscape of Software Engineering Education
abstract
Nowadays, software permeates almost every aspect of our lives. To produce complex and large-scale software products, a large number of software engineers are required. Accordingly, researchers and educators recognize the importance of Software Engineering Education (SEE), and many studies related to SEE have been published in recent years. To synthesize the large amount of research in SEE, some Systematic Literature Reviews (SLRs) focusing on different areas of SEE have been conducted and reported. However, due to their limited focuses, none of these SLRs is able to depict an overall state-of-the-art for SEE. To remedy this, we conducted a tertiary study on SEE, which identifies 26 relevant SLRs published between 2004 and 2019. By classifying and positioning these SLRs in two dimensions, i.e. the education methods/tools applied for SEE and the research topics related to SEE, we present a landscape of SEE, which locates the SLRs on SEE and their research dimensions. Further, we collected the issues studied in the published research and those that need to be addressed for instructors. This paper also discusses the challenges of the current SEE research landscape.
Xin Huang 0019, He Zhang 0001, Xin Zhou 0016, Dong Shao, Letizia Jaccheri
APSEC4
2021 Exploiting the Unique Expression for Improved Sentiment Analysis in Software Engineering Text
abstract
Sentiment analysis on software engineering (SE) texts has been widely used in the SE research, such as evaluating app reviews or analyzing developers' sentiments in commit messages. To better support the use of automated sentiment analysis for SE tasks, researchers built an SE-domain-specified sentiment dictionary to further improve the accuracy of the results. Unfortunately, recent work reported that current mainstream tools for sentiment analysis still cannot provide reliable results when analyzing the sentiments in SE texts. We suggest that the reason for this situation is because the way of expressing sentiments in SE texts is largely different from the way in social network or movie comments. In this paper, we propose to improve sentiment analysis in SE texts by using sentence structures, a different perspective from building a domain dictionary. Specifically, we use sentence structures to first identify whether the author is expressing her sentiment in a given clause of an SE text, and to further adjust the calculation of sentiments which are confirmed in the clause. An empirical evaluation based on four different datasets shows that our approach can outperform two dictionary-based baseline approaches, and is more generalizable compared to a learning-based baseline approach.
Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001
ICPC6
2021 Quality Assessment in Systematic Literature Reviews: A Software Engineering Perspective
Lanxin Yang, He Zhang 0001, Haifeng Shen, Xin Huang 0019, Xin Zhou 0016, Guoping Rong, Dong Shao
Inf. Softw. Technol.7
2020 Can You Capture Information As You Intend To? A Case Study on Logging Practice in Industry
abstract
Background: Logs provide crucial information to understand the dynamic behavior of software systems in modern software development and maintenance. Usually, logs are produced by log statements which will be triggered and executed under certain conditions. However, current studies paid very limited attention to developers' Intentions and Concerns (I&C) on logging practice, leading uncertainty that whether the developers' I&C are properly reflected by log statements and questionable capability to capture the expected information of system behaviors in logs. Objective: This study aims to reveal the status of developers' I&C on logging practice and more importantly, how the I&C are properly reflected in software source code in real-world software development. Method: We collected evidence from two sources of a series of interviews and source code analysis which are conducted in a big-data company, followed by consolidation and analysis of the evidence. Results: Major gaps and inconsistencies have been identified between the developers' I&C and real log statements in source code. Many code snippets contained no log statements that the interviewees claimed to have inserted. Conclusion: Developers' original I&C towards logging practice are usually poorly realized, which inevitably impacted the motivation and purpose to conduct this practice.
Guoping Rong, Yangchen Xu, Shenghui Gu, He Zhang 0001, Dong Shao
ICSME5
2020 Locating the Clues of Declining Success Rate of Service Calls
abstract
For many on-line systems with massive users, to provide services continuously and steadily is vital for business, which requires the anomalies of services should be located and resolved in a timely manner. As a common IT infrastructure, various APM (Application Performance Management) systems/frameworks have been adopted to monitor each call request to a service. Nevertheless, the call request may contain multidimensional attributes (e.g., City, ISP, Platform, etc.), which may further contain multiple values (e.g., ISP could be T-Mobile, CMCC, etc.). As a result, an anomaly such as DSR (Declining Success Rate) to service typically occurs with a combination of such attribute values, which creates major challenges to locate the root cause of the anomaly due to potentially huge numbers of the combinations. In this paper, we propose a novel method, ImpAPTr (Impact Analysis based on Pruning Tree), to identify the combination of dimensional attributes as the clues leading to the root cause of anomalies regarding DSR timely. In the evaluation with the simulated dataset, ImpAPTr detects valid clues in milliseconds with an accuracy of 99.37% (within the top 10 candidate results), 97.72% (top 5), and 94.51% (top 3), respectively, which outperforms previous approaches to a large degree. A field test with a production environment dataset indicates that ImpAPTr is able to detect valid clues in a few seconds.
Guoping Rong, Yong You, He Zhang 0001, Dong Shao, Yangchen Xu
ISSRE6
2020 Fireteam: a small-team development practice in industry
abstract
Software development is a collective undertaking, and the team’s efficiency is critical in development. In order to reduce project management overheads and improve productivity, a global information and communication technology enterprise institutionalizes an organization wide small-team practice, called fireteams, to tackle the problems arising from human and social aspects, such as amicability, talent, skill, and communications. This paper reports a mixed-method research, which combines archive analysis, interviews and survey, to empirically investigate the characteristics and impacts of fireteam in this industrial setting. We identify three categories of fireteam in terms of its demonstrated characteristics: ordinary agile team with extensions, single-function team, and entire life-cycle team; elaborate four key activities of fireteam, i.e. team formation, maintenance, communication, and meeting. Less communication and management overheads, higher agility & concurrency, and improved personal ability are the three important contributors that increase the productivity of fireteams. Whereas management & leadership effort, divergent understanding of fireteam, and self-organized team are discovered as the three major problems associated with fireteams. Although the benefits of fireteam can be observed from its adoption, this practice does not achieve the enterprise’s anticipations very well. Some considerations and recommendations are also discussed to improve this small-team practice.
He Zhang 0001, Dong Shao, Xin Huang 0019
ESEC/SIGSOFT FSE3
2020 DevDocOps: Enabling continuous documentation in alignment with DevOps
abstract
Summary The proliferation of DevOps enables significant acceleration and automation of the delivery and deployment of massive software products. Unfortunately, the development of supporting documents that is vital for large‐scale software systems in many cases does not keep pace with the rhythm of feature delivery using DevOps in practice, which becomes the bottleneck for many software organizations to deliver full value to the customers as claimed by the DevOps. This paper proposes, implements, and evaluates an integrated approach, DevDocOps, for continuous automated documentation, in particular for DevOps. With DevDocOps, supporting documents are created along with the development process simultaneously by various roles within a DevOps project, which largely guarantees the accuracy and integrity of documents as well as significantly increases their delivery speed. Within an established delivery chain, a set of templates are created to collect and transform the required information from its origin to the target documents for delivery. A real system, iDoc, is implemented to map, collect, and synthesize the information from document templates and automate the documentation process. DevDocOps has been successfully adopted in a top‐tier global telecommunication enterprise to support more than 5000 users with different roles related to documentation. The lag time between the releases of the product version and its supporting document has been shortened from 1 to 2 months on average to less than 2 days. DevDocOps extends the scope of DevOps and enhances the value delivery by supporting continuous documentation and bridges the gap between feature delivery and document delivery with automation.
Guoping Rong, Zefeng Jin, He Zhang 0001, Wenhua Ye, Dong Shao
Softw. Pract. Exp.6
2019 JLLAR: A Logging Recommendation Plug-in Tool for Java
abstract
Logs are the execution results of logging statements in software systems after being triggered by various events, which is able to capture the dynamic behavior of software systems during runtime and provide important information for software analysis, e.g., issue tracking, performance monitoring, etc. Obviously, to meet this purpose, the quality of the logs is critical, which requires appropriately placement of logging statements. Existing research on this topic reveals that where to log? and what to log? are two most concerns when conducting logging practice in software development, which mainly relies on developers' personal skills, expertise and preference, rendering several problems impacting the quality of the logs inevitably. One of the reasons leading to this phenomenon might be that several recognized best practices(strategies as well) are easily neglected by software developers. Especially in those software projects with relatively large number of participants. To address this issue, we designed and implemented a plug-in tool (i.e., JLLAR) based on the Intellij IDEA, which applied machine learning technology to identify and create a set of rules reflecting commonly recognized logging practices. Based on this rule set, JLLAR can be used to scan existing source code to identify issues regarding the placement of logging statements. Moreover, JLLAR also provides automatic code completion and semi code completion (i.e., to provide recommendations) regarding logging practice to support software developers during coding.
Guoping Rong, Guocheng Huang, Shenghui Gu, He Zhang 0001, Dong Shao
Internetware6
2018 A replicated experiment for evaluating the effectiveness of pairing practice in PSP education
Guoping Rong, He Zhang 0001, Bohan Liu 0003, Qi Shan, Dong Shao
J. Syst. Softw.5
2017 A Simulation Model of Kanban Software Process
abstract
Kanban has been confirmed as an effective and promising software development method. However, using of Kanban still relies on the experiences of project managers. The balance between development and test is one of the critical problems in Kanban development process. In this work, we develop a System Dynamics based Kanban development process model for the analysis of the adjustment of the number of developers and testers.
Haojie Gong, Dong Shao
APSEC3
2017 A Goal-Driven Framework in Support of Knowledge Management
abstract
Knowledge management nowadays usually focuses on the choice among some models or methodologies as a whole, but not on some specific, quantitative contributions of particular goals of the organization. Such a simplification misses some important chances for knowledge integration and transformation. What's worse, this simplification depresses the motivation of team members to accumulate and use the knowledge. In this paper, we propose a knowledge management framework which features in its goal-driven philosophy to manage project development, organize the knowledge and effectively integrate the knowledge management process into the development process. This method helps software project teams comprehensively and systematically identify and track knowledge management goals as far as possible. With a common framework, an organization is able to exchange knowledge and expertise within itself, which helps to glue the company together; while at the same time ensures that knowledge is shared over time so that the company benefits from past experience. Team members come to a common understanding on how to accumulate knowledge by establishing goals and corresponding solutions to meet the goals, and this consensus and clear vision on knowledge management motivates members to create knowledge and reduce the "gulf" between knowledge creation and application. It was successfully applied in several projects of different companies. The framework helps them establish an initial knowledge and experience repository. Software engineers are able to have more information available than they could understand and apply.
Guoping Rong, Xinbei Liu, Shenghui Gu, Dong Shao
APSEC4
2017 DevOpsEnvy: An Education Support System for DevOps
abstract
As an emerging approach to support fast delivery of software features with reliable quality, DevOps attracts more and more practitioners and shows the potential to become one of the mainstream approach for software development and operation. Many universities begin to offer DevOps related courses to the students majored in software engineering and computer science. However, as a critical part of a DevOps course, the project practicing using DevOps might cast big challenges for teachers, compared to traditional project practicing. For example, the more frequent than ever delivery in DevOps practicing will inevitably increase the workload vastly for teachers to conduct effective evaluation. In this paper, we introduce a web based system (DevOpsEnvy) to support the management and monitoring of student teams practicing DevOps. By integrating several popular open source tools, this system provides students with features such as group management, project status monitoring and student performance data analysis, etc. Meanwhile, DevOpsEnvy system also provides teachers with sufficient evidence to perform evaluation. Our preliminary trial in Nanjing University revealed several advantages of DevOpsEnvy system.
Guoping Rong, Shenghui Gu, He Zhang 0001, Dong Shao
CSEE&T4
2017 Towards Confidence with Capture-recapture Estimation: An Exploratory Study of Dependence within Inspections
abstract
Background: Capture-ReCapture (CRC), as a technique for post-inspection defect estimation, has been studied in Software Engineering (SE) community since 1990s. While most studies focused on the performance evaluation of various CRC models and estimators, few have been done on the assessment of the credibility of estimation results, rendering the difficulty of decision-making for quality management when applying CRC for defect estimation. Objective: This research aims to explore and investigate a reliable and practical approach to assess the credibility of CRC based defect estimation. Method: One fundamental assumption of applying CRC method is the statistical independence of samples that can be measured by 'Coefficient of CoVariation' (CCV). We applied CCV as an indicator of the statistical dependence between the observations (i.e., the defects detected by inspectors), and assessed the estimation results of CRC with the published datasets in SE literature by examining the correlation between Relative Error (RE) and CCV. Based on the observed correlation, we further propose CĈV, which replaces the unknown N (the actual number of defects) with the estimated number (N), to assess the credibility of CRC estimates. Results: We found that most datasets are with non-zero CCVs and the R2 (Coefficient of Determination) of non-linear curve-fitting for their CCVs and REs is higher than 0.8. Conclusions: Our study shows the evidence that the statistical dependence among inspectors is ubiquitous in the existing CRC-related studies. Besides, the significant correlation between CCV (by CĈV in practice) and RE may enable the possibility of the assessment of CRC-based estimation in support of quality management.
Guoping Rong, Bohan Liu 0003, He Zhang 0001, Qiuping Zhang, Dong Shao
EASE5
2016 CMMI guided process improvement for DevOps projects: an exploratory case study
abstract
Very recently, an increasing number of software companies adopted DevOps to adapt themselves to the ever-changing business environment. While it is important to mature adoption of the DevOps for these companies, no dedicated maturity models for DevOps exist. Meanwhile, maturity models such as CMMI models have demonstrated their effects in the traditional paradigm of software industry, however, it is not clear whether the CMMI models could guide the improvements with the context of DevOps. This paper reports a case study aiming at evaluating the feasibility to apply the CMMI models to guide process improvement for DevOps projects and identifying possible gaps. Using a structured method(i.e., SCAMPI C), we conducted a case study by interviewing four employees from one DevOps project. Based on evidence we collected in the case study, we managed to characterize the maturity/capability of the DevOps project, which implies the possibility to use the CMMI models to appraise the current processes in this DevOps project and guide future improvements. Meanwhile, several gaps also are identified between the CMMI models and the DevOps mode. In this sense, the CMMI models could be taken as a good foundation to design suitable maturity models so as to guide process improvement for projects adopting the DevOps.
Guoping Rong, He Zhang 0001, Dong Shao
ICSSP3
2015 The Impacts of Supporting Materials on Code Reading: A Controlled Experiment
abstract
Background: Code inspection has been accepted as an effective method to detect and remove defects and code reading is a critical step in code inspection. However, there are very limited empirical studies on the content and appropriate forms of the suitable software artifacts as the supporting materials, hence inspectors may not be well-supported with necessary knowledge to carry out code reading. Objective: This research aims to investigate the impact of different common supporting materials (i.e., comments vs. design documents) on code reading. Method: A relatively large-scale controlled experiment with 135 senior students was designed and executed to compare the impacts of different supporting materials on code reading. The subjects were randomly separated into three groups with different treatments, i.e, the comments, the design documents and the comments+design documents, respectively. Two metrics regarding the code reading performance (i.e., Effectiveness and Defect Detection Rate) were used to compare the different impacts derived from the two different types of supporting materials. Qualitative feedbacks were also collected using questionnaires for the final analysis. Results: The results indicate that students performed better when being provided with comments than comments+design documents. Also, the removal of design documents shows little impact on inspection effectiveness and may lead to an increase in defect detection rate. Conclusion: Comments may provide more help and value than design documents as supporting material in small to median sized code reading.
Guoping Rong, He Zhang 0001, Qi Shan, Gaoxuan Liu, Dong Shao
APSEC5
2015 Process simulation for software engineering education
abstract
Training and learning are one important purpose of Software Process Simulation (SPS). Some previous reviews showed a noticeable number of studies that combine SPS and Soft- ware Engineering Education (SEE). The objective of this research is to present the latest state-of-the-art of this area, and more importantly provide practical support for the effective adoption of SPS in educational context. We conducted an extended Systematic Literature Review (SLR) based on our previous reviews. The review identified 42 primary studies from 1992 to 2013. This paper presents the preliminary results by answering the research questions. The overall findings confirmed the positive impact of SPS on education. The detailed discussions and recommendations may offer reference value to the community.
He Zhang 0001, Dong Shao, Guoping Rong
ICSSP4
2014 Where does experience matter in software process education? An experience report
abstract
In order to enhance the understanding of important concepts and strengthen the awareness of software process, we designed a special project-practicing course in Nanjing University as an attempt to solve typical issues in these courses (e.g., focusing on aspects of software process, participation, limited time in a regular semester, etc.). The course is composed of 6-hour lecture and 32-hour bidding game. Preliminary results indicated several advantages with this new education approach on process-specific practicing course, which we already reported on CSEE&T2013. Since this course has been delivered to students from school (less experiences) and industry (more experiences), we noticed students' different performances on this course. In this paper, we collected course results from six classes, based on a comprehensive analysis from 8 different aspects; we try to understand where “EXPERIENCE” impacts students' difference performance and benefit from the understanding to improve our education on software engineering.
Guoping Rong, He Zhang 0001, Dong Shao
CSEE&T3
2014 Investigating code reading techniques for novice inspectors: an industrial case study
abstract
Code inspection is believed to be an effective technique to remove defects and improve software quality. However, the adoption of code inspection in industry is far less than it should be, which may lead to many novice inspectors in industry. For these novice inspectors, a suitable reading technique should be of the first step to begin this quality journey. While reports indicated that Checklist-Based Reading (CBR) and Ad Hoc Reading (AHR) had been the most adopted inspection techniques in industry, we deem it is necessary to investigate these two techniques first. In this paper, we present a case study of the adoption of code reading techniques in one small-sized software company. In this study, five engineers used different techniques (i.e., CBR vs. AHR) to read source code in 20 modules. Both quantitative data and qualitative data are collected during the case study. Initial analysis of these data indicates that industrial novice inspectors using CBR tended to have a lower reading speed than those using AHR. Both techniques could help these novice inspectors to remove a certain portion of defects during code review, and compared to AHR approach, CBR may help them find larger percentage of defects. However, there still exist several issues, for example, missing large portion of review-removable defects could not be avoided for novice inspectors. What's more, CBR may limit reviewers' ability to find defects outside the checklist, and to establish effective checklist remains a big challenge for novice inspectors. Besides, both internal factors (e.g., faith in inspection to achieve high quality) as well as external factors (e.g., schedule pressure) may also impact novice inspectors to adopt code reading.
Guoping Rong, He Zhang 0001, Dong Shao
EASE3
2014 Processes for embedded systems development: preliminary results from a systematic review
abstract
With the proliferation of embedded ubiquitous systems in all aspects of human life, the development of embedded systems has been facing more and more challenges (e.g., quality, time to market, etc.). Meanwhile, lots of software processes have been reported to be applied in Embedded Systems Development (ESD) with various advantages and disadvantages. Therefore, it’s important to portrait a big picture of the state-of-the-practice of the adoption of the software processes in ESD, which may benefit both practitioners and researchers in this area. This paper presents our investigation on this topic using systematic review that is intended to: 1) identify typical challenging factors and how software processes and practices address them; and 2) discover improvement opportunities from both academic and industrial perspectives.
Guoping Rong, Mingjuan Xie, Jieyu Chen, Dong Shao
ICSSP6
2014 Analyses and Modeling of Power Line Channel Attenuation Characteristics for Low Voltage Access Network in China
abstract
This paper presents the measurement results of channel attenuation characteristics of low voltage access network in China. The measurement campaign was performed in typical urban and rural residential areas, which represents the underground cable and the overhead line topologies respectively. Both narrow-band (30-500 kHz) and broad-band (500 kHz-20 MHz) attenuations are investigated. Based on the extensive measurement results, statistical methods are used in the comparison of the average signal attenuation obtained in different areas, the attenuation profile with coupling mode match/mismatch, the attenuation dynamic range at different frequencies. These analyses may provide a comprehensive understanding of the representative channel attenuation characteristics for the access domain. Besides, the classical multipath model was used to model the broad-band (0.5-20 MHz) PLC channel after simplified. Results indicated that the simplified model covers the practical channel quite well.
Dong Shao, Qiang Wang 0007, Yuquan Shu, Conglin Lai, Kangle Zhang
VTC Fall1
2014 Interference Neutralization and Alignment in Cognitive Relay Assisted 3-User Interference Channels
abstract
It is well known that relay is able to neutralize some interferences in destinations. In this paper we use not only relay to neutralize interferences but also interference alignment to reduce the number of antennas needed in destinations. To this end, a scheme of cognitive relay-aided interference neutralization and alignment is proposed. To neutralize and align interferences, the relay needs to retransmit the signals using proper transmitting vectors. We demonstrate that the 3-user MIMO interference channels with cognitive relay can achieve 2M degrees of freedom (DoF) when each node has M antennas. It is a big improvement compared to k-user system using interference alignment which is able to obtain 3M/2 DoF. The transmitting vectors in sources and relay are carefully designed to not only satisfy the neutralization and alignment constraints but also achieve higher sum-rate, which can be proved by simulation results.
Yuquan Shu, Qiang Wang 0007, Dong Shao, Jianhua Zhang 0001
VTC Fall3
2013 Applying competitive bidding games in software process education
abstract
In order to enhance the understanding of important concepts and strengthen the awareness of software process, students need to learn from their experiences in process-specific project practices. However, it's often difficult to design and carry out such practices in tertiary education environment. Typical challenges may include: 1) the difficulty to separate process-specific project practices from other (e.g., technical) practices in a software project, which may result in students paying more attention on technical aspects than process-specific aspects. 2) The limitation of a habitual technical-alone perspective may neglect concerns of other project stakeholders (e.g., project owner). We designed a special project-practicing course in Nanjing University as an attempt to solve these issues. The course is composed of 6-hour lecture and 32-hour bidding game. We found several positive results with this new education approach on process-specific practicing course. For example, it was short and flexible, which is easy to be placed in a regular semester. Besides, students were also forced to pay close attention only to process-specific aspects of the practice project. What's more, students were able to think from different perspectives, e.g., the senior management and customers.
Guoping Rong, He Zhang 0001, Dong Shao
CSEE&T3
2012 Delivering Software Process-Specific Project Courses in Tertiary Education Environment: Challenges and Solution
abstract
The importance of delivering software process courses to software engineering students has been more and more recognized in China in recent years. However, students usually cannot fully appreciate the value of software process courses by only learning methodology and principle in the classroom. Therefore, a process-specific project course was designed to fill the gap between the software process theoretical and experiential knowledge. But to design the course also has many challenges, such as: to provide enough guideline for students; to monitor every process task; to gather and use the process data, especially considering the large class size. We designed a summer school 6-weeks project course in Nanjing University based on TSP (Team Software Process) methodology. To support the course, we developed a supporting tool, the Advance Process Improvement Solution (APIS), which can record and use historical data, support teamwork, and provide process data to both students and teachers in real time. This paper describes the methodology, course organization, supporting tool, and evaluation in details. Based on our two years' experience, this course plays a key role for SE students to better understand software process.
Guoping Rong, Dong Shao
CSEE&T2
2012 Improving PSP education by pairing: An empirical study
abstract
Handling large-sized classes and maintaining students' involvement are two of the major challenges in Personal Software Process (PSP) education in universities. In order to tackle these two challenges, we adapted and incorporated some typical practices of Pair Programming (PP) into the PSP class at summer school in Software Institute of Nanjing University in 2010, and received positive results, such as higher students' involvement and conformity of process discipline, as well as (half) workload reduction in evaluating assignments. However, the experiment did not confirm the improved performance of the paired students as expected. Based on the experience and feedbacks, we improved this approach in our PSP course in 2011. Accordingly, by analyzing the previous experiment results, we redesigned the experiment with a number of improvements, such as lab environment, evaluation methods and student selection, to further investigate the effects of this approach in PSP education, in particular students' performance. We also introduced several new metrics to enable the comparison analysis of the data collected from both paired students and solo students. The new experiment confirms the value of pairing practices in PSP education. The results show that in PSP class, compared to solo students, paired students can achieve better performance in terms of program quality and exam scores.
Guoping Rong, He Zhang 0001, Mingjuan Xie, Dong Shao
ICSE4
2011 Research and practice on software engineering undergraduate curriculum NJU-SEC2006
abstract
Training a large number of qualified software engineers is a great challenge for universities, and curriculum design is an important issue. Based on IEEE-CS/ACM SE2004, Nanjing University in China designs the software engineering undergraduate curriculum NJU-SEC2006. There are three main concerns about the curriculum design. Firstly, the knowledge delivering sequence is designed to match the different scales (small/medium/large) software development. Secondly, the knowledge of professional practices is integrated into courses throughout the whole undergraduate program. Thirdly, traditional computer science courses are reformed according to the situation of China. NJU-SEC2006 has been executed for years, and received positive feedback from students, instructors and employers.
Eryu Ding, Bin Luo 0003, Daliang Zhang, Jidong Ge, Dong Shao
CSEE&T5
2011 Delivering PSP course in tertiary education environment: Challenges and solution
abstract
Nowadays, many universities include Personal Software Process (PSP) into their software engineering curriculum. However, delivering PSP course in tertiary education environment always faces at least two challenges. Firstly, in a typical PSP course in education environment, one teacher may teach much more students than a typical PSP class in industry, hence it is extremely difficult to provide evaluation of students' assignments in time. Secondly, participation of students in university often has significantly different characteristics compared to those trainees who had industry experiences. Based on education practice in Software Institute of Nanjing University, this paper proposed an approach to teaching PSP in tertiary education environment with higher efficiency and effectiveness. In this approach, a complete PSP course is delivered and cooperative learning (in pair) is encouraged. Besides, an evaluation team is established to provide timely evaluation on students' submissions and to help students correct their development behaviors. To validate this teaching approach, we conducted an experiment which involved all the freshman students enrolled in software engineering. We compared some process data collected from the submissions of both groups (individual and pair) of students. The results of the experiment show that the load of students' submissions reduced by half while students' interest of learning increased.
Guoping Rong, He Zhang 0001, Zhenyu Chen 0001, Dong Shao
CSEE&T4
2011 An introductory software engineering course for software engineering program
abstract
One important issue in undergraduate software engineering curriculum is how to help students establish the concept of software engineering at the beginning of software engineering undergraduate program and to provide a reasonable basis of knowledge and skills for subsequent courses. The "Computing and Software Engineering (CSE)", a three-semester course, is designed as the introductory course for undergraduate software engineering program at NJU in China; it tries to help students learn the comprehensive knowledge and skills in constructing small-to-medium size software. The course includes not only technical topics, such as programming and software development technology, but also professionalism and teamwork through constructing different scales of software. The knowledge is organized with the complete software example development demonstration, which makes it easier for students to synthesize all knowledge related in software development. CSE has been executed from 2009, and it has been refined according to feedback from students, lecturers and TAs. This paper describes the design and teaching practice of CSE.
Dong Shao, Bin Luo 0003, Eryu Ding
CSEE&T1
2011 Goal-Driven Development Method for Managing Embedded System Projects: An Industrial Experience Report
abstract
Technologies and methods for the development of embedded system projects are highly constrained by predefined hardware and software platforms. In this sense, embedded system projects may have more goals (derived from constraints) to achieve than regular software projects. Without pragmatic support, engineers from different disciplines are likely to neglect some project goals in the real-world embedded system projects. As a consequence, the success of embedded system projects may be more difficult to achieve than regular software projects. In this paper we report experiences gained during applying a goal driven project management methodology on several embedded system projects in a software company. We evaluated the effectiveness and efficiency of our Goal-Driven Development (GDD) methodology in practice by both projects results and feedbacks from relevant stakeholders. The results of our study show that GDD enables embedded system project teams to systematically and effectively identify, understand, track, and ultimately realize the project goals to meet relevant stakeholders' expectations. Being supported by GDD, explicit linkages and assignments are established between goals and solutions with project team's commitments.
Guoping Rong, Dong Shao, He Zhang 0001
ESEM2
2010 SCRUM-PSP: Embracing Process Agility and Discipline
abstract
With the research and debates on software process, the mainstream software processes can be grouped into two categories, the plan-driven (disciplined) processes and the agile processes. In terms of the classification, personal software process (PSP) is a typical plan-driven process while SCRUM is an agile-style instance. Although they are distinct from each other per se, our research found that PSP and SCRUM may also complement each other when SCRUM provides an agile process management framework, and PSP provides the skills and disciplines that a qualified team member needs to estimate, plan and manage his/her job. This paper proposes an integrated process model, SCRUM-PSP, which combines the strengths of each. We also verified that this integrated process by adopting it into a real project environment where typical agile processes are favored, i.e. change-prone requirements, rapid development, fast delivery, etc. As a result, manageability and predictability which traditional plan-driven processes usually benefit can also be achieved. The work described in this paper is a worthy attempt to embrace both process agility and discipline.
Guoping Rong, Dong Shao, He Zhang 0001
APSEC2