LiGuo Huang

dblp:28/3956 · also Liguo Huang · DBLP profile ↗
← Back
57ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0001-7790-0195ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 46 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Refactoring ≠ Bug-Inducing: Improving Defect Prediction with Code Change Tactics Analysis
abstract
Just-in-time defect prediction (JIT-DP) aims to predict the likelihood of code changes resulting in software defects at an early stage. Although code change metrics and semantic features have enhanced prediction accuracy, prior research has largely ignored code refactoring during both the evaluation and methodology phases, despite its prevalence. Refactoring and its propagation often tangle with bug-fixing and bug-inducing changes within the same commit and statement. Neglecting refactoring can introduce bias into the learning and evaluation of JIT-DP models. To address this gap, we investigate the impact of refactoring and its propagation on six state-of-the-art JIT-DP approaches. We propose Code chAnge Tactics (CAT) analysis to categorize code refactoring and its propagation, which improves labeling accuracy in the JIT-Defects4J dataset by $\mathbf{1 3. 7}$. Our experiments reveal that failing to consider refactoring information in the dataset can diminish the performance of models, particularly semantic-based models, by $\mathbf{1 8. 6}$ % and $\mathbf{3 7. 3} \%$ in F1-score. Additionally, we propose integrating refactoring information to enhance six baseline approaches, resulting in overall improvements in recall and $\mathbf{F 1}$-score, with increases of up to $43.2 \%$ and $32.5 \%$, respectively. Our research underscores the importance of incorporating refactoring information in the methodology and evaluation of JIT-DP. Furthermore, our CAT has broad applicability in analyzing refactoring and its propagation for software maintenance.
Feifei Niu, Junqian Shao, Christoph Mayr-Dorn, LiGuo Huang, Wesley K. G. Assunção, Chuanyi Li, Jidong Ge, Alexander Egyed
ISSRE4
2025 Effective Hard Negative Mining for Contrastive Learning-Based Code Search
abstract
Background . Code search aims to find the most relevant code snippet in a large codebase based on a given natural language query. An accurate code search engine can increase code reuse and improve programming efficiency. The focus of code search is how to represent the semantic similarity of code and query. With the development of code pre-trained models, the pattern of using numeric feature vectors (embeddings) to represent code semantics and using vector distance to represent semantic similarity has replaced traditional string matching methods. The quality of semantic representations is critical to the effectiveness of downstream tasks such as code search. Currently, the state-of-the-art (SOTA) learning method uses the contrastive learning paradigm. The objective of contrastive learning is to maximize the similarity between matching code and query (positive samples) and minimize the similarity between mismatched pairs (negative samples). To increase the reusing of negative samples, prior contrastive learning approaches use a large queue (memory bank) to store embeddings. Problem . However, there is still a lot of room for improvement in using negative examples for code search: ① Due to the random selection of negative samples, semantic representations learned by existing models cannot distinguish similar codes well. ② Since semantic vectors in the memory bank are reused from previous inference results and then directly used for loss function calculation without gradient descent, the model cannot effectively learn the negative sample semantic information. Method . To solve the above problems, we propose a contrastive learning code search model with hard negative mining called CoCoHaNeRe: ❶ To enable the model to distinguish similar codes, we introduce hard negative examples into contrastive training, which are negative examples in the codebase that are most similar to positive examples. As a result, hard negative examples are most likely to make the model make mistakes. ❷ To improve the learning efficiency of negative samples during training, we add all hard negative examples to the model's gradient descent process. Result . To verify the effectiveness of CoCoHaNeRe, we conducted experiments on large code search datasets with six programming languages, as well as similar retrieval tasks code clone detection and code question answering. Experimental results show that our model achieves SOTA performance. In the code search task, the average MRR score of CoCoHaNeRe exceeds CodeBERT, GraphCodeBERT, and UniXcoder by 11.25%, 8.13%, and 7.38%, respectively. It has also made great progress in code clone detection and code question answering. In addition, our method performs well in different programming languages and code pre-training models. Furthermore, qualitative analysis shows that our model effectively distinguishes high-order semantic differences between similar codes.
Chuanyi Li, Jidong Ge, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.4
2025 Patch Correctness Assessment: A Survey
abstract
Most automated program repair methods rely on test cases to determine the correctness of the generated patches. However, due to the incompleteness of available test suites, some patches that pass all the test cases may still be incorrect. This issue is known as the patch overfitting problem. Overfitting problem is a longstanding problem in automated program repair. Due to overfitting patches, the patches obtained by automated program repair tools require further validation to determine their correctness. Researchers have proposed many methods to automatically assess the correctness of patches, but no systematic review provides a detailed introduction to this problem, the existing solutions, and the challenges. To address this deficiency, we systematically review the existing approaches to patch correctness assessment. We first offer a few examples of overfitting patches to acquire a more detailed understanding of this problem. We then propose a comprehensive categorization of publicly available techniques and datasets, examine the commonly used evaluation metrics, and perform an in-depth analysis of the effectiveness of the existing models in addressing the challenge of overfitting. Based on our analysis, we provided the difficulties encountered by current methodologies, alongside the possible avenues for future research exploration.
Zhiwei Fei, Jidong Ge, Chuanyi Li, Yuning Li, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.7
2025 An Empirical Study of Code Simplification Methods in Code Intelligence Tasks
abstract
In recent years, pre-trained language models have seen significant success in natural language processing and have been increasingly applied to code-related tasks. Code intelligence tasks have shown promising performance with the support of code pre-trained language models. Pre-processing code simplification methods have been introduced to prune code tokens from the model’s input while maintaining task effectiveness. These methods improve the efficiency of code intelligence tasks while reducing computational costs. Post-prediction code simplification methods provide explanations for code intelligence task outcomes, enhancing the reliability and interpretability of model predictions. However, comprehensive evaluations of these methods across diverse code pre-trained model architectures and code intelligence tasks are lacking. To assess the effectiveness of code simplification methods, we conduct an empirical study integrating these code simplification methods with various pre-trained code models across multiple code intelligence tasks. Our empirical findings suggest that developing task-specific code simplification methods would be beneficial. Then, we recommend leveraging post-prediction methods to summarize prior knowledge, which can pre-process code simplification strategies. Moreover, establishing more evaluation mechanisms for code simplification is crucial. Finally, we propose incorporating code simplification methods into the pre-training phase of code pre-trained models to enhance their program comprehension and code representation capabilities.
Zongwen Shen, Yuning Li, Jidong Ge, Xiang Chen 0005, Chuanyi Li, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.6
2025 Improving Source Code Pre-Training via Type-Specific Masking
abstract
The Masked Language Modeling (MLM) task is widely recognized as one of the most effective pre-training tasks and currently derives many variants in the Software Engineering (SE) field. However, most of these variants mainly focus on code representation without distinguishing between different code token types, while some focus on a specific type, such as code identifiers. Indeed, various code token types exist, and there is no evidence that only identifiers can improve PTMs. Thus, to improve PTMs through different types, we conducted an extensive study to evaluate how different type-specific masking tasks can affect PTMs. First, we extract five code token types, convert them into type-specific masking tasks, and generate their combinations. Second, we pre-train CodeBERT and PLBART using combinations and fine-tuned them on four SE downstream tasks. Experimental results show that type-specific masking tasks can enhance CodeBERT and PLBART on all downstream tasks. Furthermore, we discuss topics related to low-resource datasets, conflicting PTMs that original pre-training tasks conflict with our methods, the cost and performance of our methods, factors that impact the performance of our methods, and applying our methods on state-of-the-art PTMs. These discussions comprehensively analyze the strengths and weaknesses of different type-specific masking tasks.
Wentao Zou, Chuanyi Li, Jidong Ge, Xiang Chen 0005, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.6
2025 Experimental Evaluation of Parameter-Efficient Fine-Tuning for Software Engineering Tasks
abstract
Pre-trained models (PTMs) have succeeded in various software engineering (SE) tasks following the “pre-train then fine-tune” paradigm. As fully fine-tuning all parameters of PTMs can be computationally expensive, a potential solution is parameter-efficient fine-tuning (PEFT), which freezes PTMs while introducing extra parameters. Although PEFT methods have been applied to SE tasks, researchers often focus on specific scenarios and lack a comprehensive comparison of PTMs from different aspects such as field, size, and architecture. To fill this gap, we have conducted an empirical study on six PEFT methods, eight PTMs, and four SE tasks. The experimental results reveal several noteworthy findings. For example, model architecture has little impact on PTM performance when using PEFT methods. Additionally, we provide a comprehensive discussion of PEFT methods from three perspectives. First, we analyze the effectiveness and efficiency of PEFT methods. Second, we explore the impact of the scaling factor hyperparameter. Finally, we investigate the application of PEFT methods on the latest open source large language model, Llama 3.2. These findings provide valuable insights to guide future researchers in effectively applying PEFT methods to SE tasks.
Wentao Zou, Zongwen Shen, Jidong Ge, Chuanyi Li, Xiang Chen 0005, Xiaoyu Shen 0001, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.8
2024 An extensive replication study of the ABLoTS approach for bug localization
Feifei Niu, Enshuo Zhang, Christoph Mayr-Dorn, Wesley K. G. Assunção, LiGuo Huang, Jidong Ge, Bin Luo 0003, Alexander Egyed
Empir. Softw. Eng.5
2024 PassSum: Leveraging paths of abstract syntax trees and self-supervision for code summarization
abstract
Abstract Code summarization is to provide a high‐level comment for a code snippet that typically describes the function and intent of the given code. Recent years have seen the successful application of data‐driven code summarization. To improve the performance of the model, numerous approaches use abstract syntax trees (ASTs) to represent the structural information of the code, which is considered by most researchers to be the main factor that distinguishes code from natural language. Then, such data‐driven methods are trained on large‐scale labeled datasets to obtain a model with strong generalization capabilities that can be applied to new examples. Nevertheless, we argue that state‐of‐the‐art approaches suffer from two key weaknesses: (1) inefficient encoding of ASTs; (2) reliance on a large labeled corpus for model training. As a result, such drawbacks lead to (1) oversized model, slow training, information loss and instability; (2) inability to be applied to programming languages with only a small amount of labeled data. In light of these weaknesses, we propose PassSum, a code summarization approach that addresses the aforementioned weaknesses via (1) a novel input representation which contains an efficient AST encoding method; (2) introducing three pretraining objectives and pretraining our model with a large amount of (easy‐to‐obtain) unlabeled data under the guidance of self‐supervised learning. Experimental results on code summarization for Java, Python, and Ruby methods demonstrate the superiority of PassSum to state‐of‐the‐art methods. Further experiments demonstrate that the input representation we use has both temporal and spatial advantages in addition to performance leadership. In addition, pretraining is also shown to make the model more generalizable with less labeled data, and also to speed up the convergence of the model during training.
Changan Niu, Chuanyi Li, Vincent Ng 0001, Jidong Ge, LiGuo Huang, Bin Luo 0003
J. Softw. Evol. Process.5
2024 Scoping Software Engineering for AI: The TSE Perspective
abstract
Advances in Artificial Intelligence (AI), and in particular in Machine Learning (ML), are introducing profound changes to scholarly submissions across publication venues, affecting in particular the contributions that are being submitted to Software Engineering (SE) conferences and journals. In this context, it is not always clear whether manuscripts submitted to SE venues under the umbrella term SE for AI are indeed relevant to SE, in the sense that they explicitly contain contributions to the SE body of knowledge. This leads to recurring discussions on whether certain AI-related submissions are appropriate to SE venues, or should instead be submitted to other journals and conferences, including AI or ML-specific ones. In this editorial, we discuss the kinds of AI-related contributions that are a better fit-and a less good fit-for publication in the IEEE Transactions on Software Engineering.
Sebastián Uchitel, Marsha Chechik, Massimiliano Di Penta, Bram Adams, Nazareno Aguirre, Gabriele Bavota, Domenico Bianculli, Kelly Blincoe, Ana Cavalcanti 0001, Yvonne Dittrich, Filomena Ferrucci, Rashina Hoda, LiGuo Huang, David Lo 0001, Michael R. Lyu, Lei Ma 0003, Jonathan I. Maletic, Leonardo Mariani, Collin McMillan, Tim Menzies, Martin Monperrus, Ana Moreno, Nachiappan Nagappan, Liliana Pasquale, Patrizio Pelliccione, Michael Pradel, Rahul Purandare, Sukyoung Ryu, Mehrdad Sabetzadeh, Alexander Serebrenik, Jun Sun 0001, Chakkrit Tantithamthavorn, Christoph Treude, Manuel Wimmer, Yingfei Xiong 0001, Tao Yue 0002, Andy Zaidman, Tao Zhang 0001, Hao Zhong 0001
IEEE Trans. Software Eng.13
2023 RAT: A Refactoring-Aware Traceability Model for Bug Localization
abstract
A large number of bug reports are created during the evolution of a software system. Locating the source code files that need to be changed in order to fix these bugs is a challenging task. Information retrieval-based bug localization techniques do so by correlating bug reports with historical information about the source code (e.g., previously resolved bug reports, commit logs). These techniques have shown to be efficient and easy to use. However, one flaw that is nearly omnipresent in all these techniques is that they ignore code refactorings. Code refactorings are common during software system evolution, but from the perspective of typical version control systems, they break the code history. For example, a class when renamed then appears as two separate classes with separate histories. Obviously, this is a problem that affects any technique that leverages code history. This paper proposes a refactoring-aware traceability model to keep track of the code evolution history. With this model, we reconstruct the code history by analyzing the impact of code refactorings to correctly stitch together what would otherwise be a fragmented history. To demonstrate that a refactoring aware history is indeed beneficial, we investigated three widely adopted bug localization techniques that make use of code history, which are important components in existing approaches. Our evaluation on 11 open source projects shows that taking code refactorings into account significantly improves the results of these bug localization techniques without significant changes to the techniques themselves. The more refactorings are used in a project, the stronger the benefit we observed. Based on our findings, we believe that much of the state of the art leveraging code history should benefit from our work.
Feifei Niu, Wesley K. G. Assunção, LiGuo Huang, Christoph Mayr-Dorn, Jidong Ge, Bin Luo 0003, Alexander Egyed
ICSE3
2023 Domain Adaptive Code Completion via Language Models and Decoupled Domain Databases
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in code completion. However, due to the lack of domain-specific knowledge, they may not be optimal in completing code that requires intensive domain knowledge for example completing the library names. Although there are several works that have confirmed the effectiveness of fine-tuning techniques to adapt language models for code completion in specific domains. They are limited by the need for constant fine-tuning of the model when the project is in constant iteration. To address this limitation, in this paper, we propose$k$NM-LM, a retrieval-augmented language model (R-LM), that integrates domain knowledge into language models without fine-tuning. Different from previous techniques, our approach is able to automatically adapt to different language models and domains. Specifically, it utilizes the in-domain code to build the retrieval-based database decoupled from LM, and then combines it with LM through Bayesian inference to complete the code. The extensive experiments on the completion of intra-project and intra-scenario have confirmed that$k$NM-LM brings about appreciable enhancements when compared to CodeGPT and UnixCoder. A deep analysis of our tool including the responding speed, storage usage, specific type code completion, and API invocation completion has confirmed that$k$NM-LM provides satisfactory performance, which renders it highly appropriate for domain adaptive code completion. Furthermore, our approach operates without the requirement for direct access to the language model's parameters. As a result, it can seamlessly integrate with black-box code completion models, making it easy to integrate our approach as a plugin to further enhance the performance of these models.
Ze Tang 0002, Jidong Ge, Shangqing Liu, Tingwei Zhu, Tongtong Xu, LiGuo Huang, Bin Luo 0003
ASE6
2023 The ABLoTS Approach for Bug Localization: is it replicable and generalizable?
abstract
Bug localization is the task of recommending source code locations (typically files) that probably contain the cause of a bug and hence need to be changed to fix the bug. Along these lines, information retrieval-based bug localization (IRBL) approaches have been adopted, which identify the most bug-prone files from the source code space. In current practice, a series of state-of-the-art IRBL techniques leverage the combination of different components, e.g., similar reports, version history, code structure, to achieve better performance. ABLoTS is a recently proposed approach with the core component, TraceScore, that utilizes requirements and traceability information between different issue reports, i.e., feature requests and bug reports, to identify buggy source code snippets with promising results. To evaluate the accuracy of these results and obtain additional insights into the practical applicability of ABLoTS, supporting of future more efficient and rapid replication and comparison, we conducted a replication study of this approach with the original data set and also on an extended data set. The extended data set includes 16 more projects comprising 25,893 bug reports and corresponding source code commits. While we find that the TraceScore component as the core of ABLoTS produces comparable results with the extended data set, we also find that the ABLoTS approach no longer achieves promising results, due to an overlooked side effect of incorrectly choosing a cut-off date that led to training data leaking into test data with significant effects on performance.
Feifei Niu, Christoph Mayr-Dorn, Wesley K. G. Assunção, LiGuo Huang, Jidong Ge, Bin Luo 0003, Alexander Egyed
MSR4
2023 Learning the Relation Between Similarity Loss and Clustering Loss in Self-Supervised Learning
abstract
Self-supervised learning enables networks to learn discriminative features from massive data itself. Most state-of-the-art methods maximize the similarity between two augmentations of one image based on contrastive learning. By utilizing the consistency of two augmentations, the burden of manual annotations can be freed. Contrastive learning exploits instance-level information to learn robust features. However, the learned information is probably confined to different views of the same instance. In this paper, we attempt to leverage the similarity between two distinct images to boost representation in self-supervised learning. In contrast to instance-level information, the similarity between two distinct images may provide more useful information. Besides, we analyze the relation between similarity loss and feature-level cross-entropy loss. These two losses are essential for most deep learning methods. However, the relation between these two losses is not clear. Similarity loss helps obtain instance-level representation, while feature-level cross-entropy loss helps mine the similarity between two distinct images. We provide theoretical analyses and experiments to show that a suitable combination of these two losses can get state-of-the-art results. Code is available at https://github.com/guijiejie/ICCL.
Jidong Ge, Jie Gui, Lanting Fang, Ming Lin 0002, James T. Kwok, LiGuo Huang, Bin Luo 0003
IEEE Trans. Image Process.7
2023 Machine/Deep Learning for Software Engineering: A Systematic Literature Review
abstract
Since 2009, the deep learning revolution, which was triggered by the introduction of ImageNet, has stimulated the synergy between Software Engineering (SE) and Machine Learning (ML)/Deep Learning (DL). Meanwhile, critical reviews have emerged that suggest that ML/DL should be used cautiously. To improve the applicability and generalizability of ML/DL-related SE studies, we conducted a 12-year Systematic Literature Review (SLR) on 1,428 ML/DL-related SE papers published between 2009 and 2020. Our trend analysis demonstrated the impacts that ML/DL brought to SE. We examined the complexity of applying ML/DL solutions to SE problems and how such complexity led to issues concerning the reproducibility and replicability of ML/DL studies in SE. Specifically, we investigated how ML and DL differ in data preprocessing, model training, and evaluation when applied to SE tasks, and what details need to be provided to ensure that a study can be reproduced or replicated. By categorizing the rationales behind the selection of ML/DL techniques into five themes, we analyzed how model performance, robustness, interpretability, complexity, and data simplicity affected the choices of ML/DL models.
LiGuo Huang, Amiao Gao, Jidong Ge, Haitao Feng, Ishna Satyarth, Ming Li 0005, He Zhang 0001, Vincent Ng 0001
IEEE Trans. Software Eng.2
2022 SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations
abstract
Recent years have seen the successful application of large pre-trained models to code representation learning, resulting in substantial improvements on many code-related downstream tasks. But there are issues surrounding their application to SE tasks. First, the majority of the pre-trained models focus on pre-training only the encoder of the Transformer. For generation tasks that are addressed using models with the encoder-decoder architecture, however, there is no reason why the decoder should be left out during pre-training. Second, many existing pre-trained models, including state-of-the-art models such as T5-learning, simply reuse the pretraining tasks designed for natural languages. Moreover, to learn the natural language description of source code needed eventually for code-related tasks such as code summarization, existing pretraining tasks require a bilingual corpus composed of source code and the associated natural language description, which severely limits the amount of data for pre-training. To this end, we propose SPT-Code, a sequence-to-sequence pre-trained model for source code. In order to pre-train SPT-Code in a sequence-to-sequence manner and address the aforementioned weaknesses associated with existing pre-training tasks, we introduce three pre-training tasks that are specifically designed to enable SPT-Code to learn knowledge of source code, the corresponding code structure, as well as a natural language description of the code without relying on any bilingual corpus, and eventually exploit these three sources of information when it is applied to downstream tasks. Experimental results demonstrate that SPT-Code achieves state-of-the-art performance on five code-related downstream tasks after fine-tuning.
Changan Niu, Chuanyi Li, Vincent Ng 0001, Jidong Ge, LiGuo Huang, Bin Luo 0003
ICSE5
2022 AST-Trans: Code Summarization with Efficient Tree-Structured Attention
abstract
Code summarization aims to generate brief natural language descriptions for source codes. The state-of-the-art approaches follow a transformer-based encoder-decoder architecture. As the source code is highly structured and follows strict grammars, its Abstract Syntax Tree (AST) is widely used for encoding structural information. However, ASTs are much longer than the corresponding source code. Existing approaches ignore the size constraint and simply feed the whole linearized AST into the encoders. We argue that such a simple process makes it difficult to extract the truly useful dependency relations from the overlong input sequence. It also incurs significant computational overhead since each node needs to apply self-attention to all other nodes in the AST. To encode the AST more effectively and efficiently, we propose AST-Trans in this paper which exploits two types of node relationships in the AST: ancestor-descendant and sibling relationships. It applies the tree-structured attention to dynamically allocate weights for relevant nodes and exclude irrelevant nodes based on these two relationships. We further propose an efficient implementation to support fast parallel computation for tree-structure attention. On the two code summarization datasets, experimental results show that AST-Trans significantly outperforms the state-of-the-arts while being times more efficient than standard transformers1.
Ze Tang 0002, Xiaoyu Shen 0001, Chuanyi Li, Jidong Ge, LiGuo Huang, Zheling Zhu, Bin Luo 0003
ICSE5
2022 Lighting up supervised learning in user review-based code localization: dataset and benchmark
abstract
As User Reviews (URs) of mobile Apps are proven to provide valuable feedback for maintaining and evolving applications, how to make full use of URs more efficiently in the release cycle of mobile Apps has become a widely concerned and researched topic in the Software Engineering (SE) community. In order to speed up the completion of coding work related to URs to shorten the release cycle as much as possible, the task of User Review-based code localization is proposed and studied in depth. However, due to the lack of large-scale ground truth dataset (i.e., truly related pairs), existing methods are all unsupervised learning-based. In order to light up supervised learning approaches, which are driven by large labeled datasets, for Review2Code, and to compare their performances with unsupervised learning-based methods, we first introduce a large-scale human-labeled ground truth dataset, including the annotation process and statistical analysis. Then, a benchmark consisting of two SOTA unsupervised learning-based and four supervised learning-based Review2Code methods is constructed based on this dataset. We believe that this paper can provide a basis for in-depth exploration of the supervised learning-based Review2Code solutions.
Xinwen Hu, Jianjie Lu, Zheling Zhu, Chuanyi Li, Jidong Ge, LiGuo Huang, Bin Luo 0003
ESEC/SIGSOFT FSE7
2022 Natural Language-Based Automatic Programming for Industrial Robots
Jie Chen 0060, Zhongjin Li, LiGuo Huang
J. Grid Comput.5
2020 A novel completeness definition of event logs and corresponding generation algorithm
abstract
Abstract As the promotion of technologies and applications of Big Data, the research of business process management (BPM) has gradually deepened to consider the impacts and challenges of big business data on existing BPM technologies. Recently, parallel business process mining (e.g. discovering business models from business visual data, integrating runtime business data with interactive business process monitoring visualisation systems and summarising and visualising historical business data for further analysis, etc.) and multi‐perspective business data analytics (e.g. pattern detecting, decision‐making and process behaviour predicting, etc.) have been intensively studied considering the steep increase in business data size and type. However, comprehensive and in‐depth testing is needed to ensure their quality. Testing based solely on existing business processes and their system logs is far from sufficient. Large‐scale randomly generated models and corresponding complete logs should be used in testing. To test parallel algorithms for discovering process models, different log completeness and generation algorithms were proposed. However, they suffer from either state space explosion or non‐full‐covering task dependencies problem. Besides, most existing generation algorithms rely on random executing strategy, which leads to low and unstable efficiency. In this paper, we propose a novel log completeness type, that is, #TAR completeness, as well as its generation algorithm. The experimental results based on a series of randomly generated process models show that the #TAR complete logs outperform the state‐of‐the‐art ones with lower capacity, fuller dependencies covering and higher generating efficiency.
Chuanyi Li, Jidong Ge, Lijie Wen 0001, Victor Chang 0001, LiGuo Huang, Bin Luo 0003
Expert Syst. J. Knowl. Eng.6
2020 Security and performance-aware resource allocation for enterprise multimedia in mobile edge computing
Zhongjin Li, Binbin Huang 0006, Jie Chen 0060, Chuanyi Li, Hua Hu 0001, LiGuo Huang
Multim. Tools Appl.7
2019 Predicting Licenses for Changed Source Code
abstract
Open source software licenses regulate the circumstances under which software can be redistributed, reused and modified. Ensuring license compatibility and preventing license restriction conflicts among source code during software changes are the key to protect their commercial use. However, selecting the appropriate licenses for software changes requires lots of experience and manual effort that involve examining, assimilating and comparing various licenses as well as understanding their relationships with software changes. Worse still, there is no state-of-the-art methodology to provide this capability. Motivated by this observation, we propose in this paper Automatic License Prediction (ALP), a novel learning-based method and tool for predicting licenses as software changes. An extensive evaluation of ALP on predicting licenses in 700 open source projects demonstrate its effectiveness: ALP can achieve not only a high overall prediction accuracy (92.5% in micro F1 score) but also high accuracies across all license types.
LiGuo Huang, Jidong Ge, Vincent Ng 0001
ASE2
2019 Determining relevant training data for effort estimation using Window-based COCOMO calibration
Vu Nguyen 0003, Barry W. Boehm, LiGuo Huang
J. Syst. Softw.3
2019 Investigating the use of duration-based windows and estimation by analogy for COCOMO
abstract
Abstract In model‐based software estimation, using the right training data is a key contributor for making accurate predictions, which is crucial for the success of software projects. This study investigates the use of duration‐based windows and estimation by analogy to calibrate COCOMO and assess their estimation performance. We compare these approaches as well as the use of all available historical data using the COCOMO data set of 341 projects and NASA data set of 93 projects. The results show that timing information exists in the data sets affecting estimation accuracy. Given sufficient data for calibration, using recently completed projects within short durations generates more accurate estimates than retaining all historical data or using k‐nearest neighbors based on estimation by analogy. More training data spanning a long period of time may not lead to improved estimation accuracy. This study offers evidence to support the use of projects completed within recent years for training estimation models.
Vu Nguyen 0003, Thuy Huynh, Barry W. Boehm, LiGuo Huang, Thong Truong
J. Softw. Evol. Process.4
2019 Innovative process paradigms and data driven analytics: A new horizon for software and systems process
abstract
The digital transformation has brought us to a new horizon where devices, products and services are connected, interact with each other and generate massive amounts of data.What are the products and services to be offered in this environment?How will they be developed, delivered, maintained and evolved?Whatever the field, software, military, health care, or business, processes are everywhere.Not only are processes used to design, produce, and deliver enterprise products and services but also they are also used to instantiate enterprise/government practices, policies, and regulations.Fundamentally, processes are the means by which an enterprise satisfies customers and creates value, and they provide the building blocks of all information systems.Given their importance to the success and profitability of companies, it is more important than ever to seriously consider the design and improvement of these processes and to make sure that they are free of flaws and inconsistencies.At the same time, recent advances in the hardware domain have paved the way for more complex and critical systems.The Internet of things (IoT) applications and cloud computing infrastructures have introduced an increasing need for distributed system applications and more intelligent system segmentation.Big data solutions offer the potential to learn from vast amounts of information generated by these systems and processes (when collected).This information can be used to satisfy business needs or simply to learn and to improve systems and processes.The ICSSP 2017 conference aimed to explore how these new trends impact and constrain the way processes should be designed whatever the type of process and whatever the application domain.The conference also aimed to inquire how new visions and tools that incorporate cloud computing, big data, IoT, and DevOps could be incorporated into processes used to design next generation systems.To this end, ICSSP 2017 sought to address the following questions:1. What will the next generation of process paradigms look like? 2. How will the needs of all business and system stakeholders that participate in the development and evolution of systems and processes be addressed?and provided a venue for specific areas of interest including • Big data driven process modeling, assessment, and improvement • Mining software/business process repositories (including code, bug trackers, and etc) to improve processes • Dynamic process evolution, variability, and adaptability for next generation of process paradigms • Predictive monitoring and process deviance mining • Empirical evidence of the effectiveness of Agile/Lean practices and approaches in software, hardware, and hybrid systems development and evolution • Process issues in developing adaptive and evolving software systems Furthermore, we continued to explore topics that are core to the area of software and systems process including • Continuous process improvement in diverse areas and context
David Raffo, Reda Bendraou, LiGuo Huang, Fabrizio Maria Maggi
J. Softw. Evol. Process.3
2019 Monitoring Interactions Across Multi Business Processes with Token Carried Data
abstract
The rapid development of web service provides many opportunities for companies to migrate their business processes to the Internet for wider accessibility and higher collaboration efficiency. However, the open, dynamic and ever-changing Internet also brings challenges in protecting these business processes. There are certain process monitoring methods and the recently proposed ones are based on state changes of process artifacts or places, however, they do not mention defending process interactions from outer tampering, where events could not be detected by process systems, or saving fault-handling time. In this paper, we propose a novel Token-based Interaction Monitoring framework based on token carried data to safeguard process collaboration and reduce problem solving time. Token is a more common data entity in processes than process artifacts and they cover all tasks' executions. Comparing to detecting places' state change, we set security checking points at both when tokens are just produced and to be consumed. This will ensure that even if data is tampered after being created it would be detected before being used. For applying monitoring framework, we develop a collaboration constructing method with token-based process mining techniques to derive global interaction processes as well as organize historical process data in forms of token.
Chuanyi Li, Jidong Ge, Zhongjin Li, LiGuo Huang, Bin Luo 0003
IEEE Trans. Serv. Comput.4
2018 Linking Source Code to Untangled Change Intents
abstract
Previous work [13] suggests that tangled changes (i.e., different change intents aggregated in one single commit message) could complicate tracing to different change tasks when developers manage software changes. Identifying links from changed source code to untangled change intents could help developers solve this problem. Manually identifying such links requires lots of experience and review efforts, however. Unfortunately, there is no automatic method that provides this capability. In this paper, we propose AutoCILink, which automatically identifies code to untangled change intent links with a pattern-based link identification system (AutoCILink-P) and a supervised learning-based link classification system (AutoCILink-ML). Evaluation results demonstrate the effectiveness of both systems: the pattern-based AutoCILink-P and the supervised learning-based AutoCILink-ML achieve average accuracy of 74.6% and 81.2%, respectively.
LiGuo Huang, Chuanyi Liu, Vincent Ng 0001
ICSME2
2018 Effective API recommendation without historical software repositories
abstract
It is time-consuming and labor-intensive to learn and locate the correct API for programming tasks. Thus, it is beneficial to perform API recommendation automatically. The graph-based statistical model has been shown to recommend top-10 API candidates effectively. It falls short, however, in accurately recommending an actual top-1 API. To address this weakness, we propose RecRank, an approach and tool that applies a novel ranking-based discriminative approach leveraging API usage path features to improve top-1 API recommendation. Empirical evaluation on a large corpus of (1385+8) open source projects shows that RecRank significantly improves top-1 API recommendation accuracy and mean reciprocal rank when compared to state-of-the-art API recommendation approaches.
LiGuo Huang, Vincent Ng 0001
ASE2
2018 Automatically classifying user requests in crowdsourcing requirements engineering
Chuanyi Li, LiGuo Huang, Jidong Ge, Bin Luo 0003, Vincent Ng 0001
J. Syst. Softw.2
2018 Do code data sharing dependencies support an early prediction of software actual change impact set?
abstract
Abstract Existing studies have shown that structural dependencies within code are good predictors for code actual change impact set—a set of entities that repeatedly changing together to ensure a consistent and complete change. However, the result is far from ideal, particularly when insufficient historical data are available at an early stage of software development. This paper demonstrates that a better understanding of data dependencies in addition to call dependencies greatly improves actual change impact set prediction. We propose a new approach and tool (namely, CHIP) to predict software actual change impact sets leveraging both call and data sharing dependencies. For this purpose, CHIP employs novel extensions (dependency frequency filtering and shared data type idf filtering) to reduce false positives. CHIP assumes that developers know initial places where to start making changes in the source code even though they may not know all changes. This approach has been empirically evaluated on 4 large‐scale open source systems. Our evaluation demonstrates that data sharing dependencies have a complementary impact on software actual change impact set prediction as compared with predictions based on call dependencies only. CHIP improves the F2‐score compared with the predictors using both Program Dependence Graph and evolutionary couplings.
LiGuo Huang, Alexander Egyed, Jidong Ge
J. Softw. Evol. Process.2
2017 Data driven credit risk management process: a machine learning approach
abstract
Credit scoring process, the most important part in credit risk management, aims at estimating the probability that an applicant will perform bad credit behaviors (e.g., loan default). Managing and developing effective and reliable risk assessment procedures in order to mitigate potential loss caused by new applicants heavily relies on the performance of scoring process. Traditionally this process is manually developed, which is time-consuming. In this paper, we propose an automated credit risk management process based on machine learning to ease the scoring process in order to reduce the human effort. This process is data driven: it leverages machine learning to automatically analyze vast amounts of historical data and build predictive model. We evaluate our process with a real-world proprietary dataset and achieved good performance, which shows the feasibility of using machine learning to facilitate the credit risk management process.
Yann Dautais, LiGuo Huang, Jidong Ge
ICSSP3
2017 Tracing requirements in software design
abstract
Software requirement analysis is an essential step in software development process, which defines what is to be built in a project. Requirements are mostly written in text and will later evolve to fine-grained and actionable artifacts with details about system configurations, technology stacks, etc. Tracing the evolution of requirements enables stakeholders to determine the origin of each requirement and understand how well the software's design reflects to its requirements. Reckoning requirements traceability is not a trivial task, we focus on applying machine learning approach to classify traceability between various associated requirements. In particular, we investigate a 2-learner, ontology-based approach, where we train two classifiers to separately exploit two types of features, lexical features and features derived from a hand-built ontology. In comparison to a supervised baseline system that uses only lexical features, our approach yields a relative error reduction of 25.9%. Most interestingly, results do not deteriorate when the hand-built ontology is replaced with its automatically constructed counterpart.
Zeheng Li, LiGuo Huang, Vincent Ng 0001, Ruili Geng
ICSSP3
2017 Improving Random Test Sets Using a Locally Spreading Approach
abstract
Labeling test cases is expensive. Given a limited number of test cases to be labeled, more evenly spreading test cases over the input domain has a better chance to hit the nonpoint failure patterns. In this paper, we propose a spreading points (test cases) algorithm based on local layout of points (the locally spreading approach is referred to as LS). The LS repositions the initial test set and evolves it to improve the minimum distance among points. During the locally spreading process, for every point, a feasible direction of movement is solved according to its nearest neighbors, and along this direction the point can increase the shortest pairwise distance between it and other points. We investigate the effective of the LS approach, considering it as an add-on to the Adaptive Random Testing (ART) technique. The simulation results show that the LS can improve effectiveness (P-measure) of the ART.
Xiangyang Huang, LiGuo Huang, Shudong Zhang, Minhua Wu
QRS2
2017 Software cybernetics in BPM: Modeling software behavior as feedback for evolution by a novel discovery method based on augmented event logs
Chuanyi Li, Jidong Ge, LiGuo Huang, Budan Wu, Hao Hu 0001, Bin Luo 0003
J. Syst. Softw.3
2016 A security and cost aware scheduling algorithm for heterogeneous tasks of scientific workflow in clouds
Zhongjin Li, Jidong Ge, LiGuo Huang, Hao Hu 0001, Bin Luo 0003
Future Gener. Comput. Syst.4
2016 Process mining with token carried data
Chuanyi Li, Jidong Ge, LiGuo Huang, Budan Wu, Hao Hu 0001, Bin Luo 0003
Inf. Sci.3
2016 Guest editors' introduction
abstract
There was a time when researching software processes meant just that – we were interested in making sure that the process for software development was effective. From a software process perspective, we did not really have to worry about competition, the business nor the domain in which our software was used – because these were not seen as affecting the software processes through which the software was developed. But, things have changed! Software has become more ubiquitous. Software is used in products that are governed by regulation. Software is being developed in organizations that heretofore did not consider themselves software companies – such as automotive and medical device companies. Software is fundamental to many businesses that would not continue to survive if their software is not running effectively. In addition, modern software systems consist of a complex mix of products and services, some are decades old, and some are merely emerging. They may easily suffer from symptoms of aging and need to be continuously adapted to cope with changing requirements and environments. The sources of such changes may be new customers, intense competition, changing organizational structures and regulatory frameworks, changing interacting systems, bug fixing, software degradation and erosion, emerging opportunities and risks in the business environment, as well as emerging software technologies and platforms. In order to overcome or avoid the negative effects of software aging, we have to place change and evolution at the center of the software development and maintenance processes. Evolving process drivers are not necessarily software driven. If we consider drivers in (i) these will be evolved through software engineering, either being developed in industry or by researchers. Drivers in (ii) may possibly evolve through software engineering, and the world of software engineering will have an input. But, often it is the business case that takes priority. And, drivers in (iii) are determined by the business in which we work 5. In these cases, we, as software engineers must consider the business domain to establish what software processes are of interest and will be required to use particular processes 6. However, developing software engineering processes on a one-off basis and then letting them be used ad infinitum without any further changes will not work. Change in the evolving process drivers cause software engineering processes to evolve. In addition, these evolving process drivers can be critical, particularly as the software engineering process increasingly contributes to the company's profitability, be that a direct or indirect contribution. The latest trends of developing and maintaining emerging and evolving software systems also leads to new opportunities and challenges regarding the development processes, including, but not limited to, process evolution, scalability, and process verification and validation. The research presented in the International Conference on Software and Systems Process contributes to this evolution. International Conference on Software and Systems Process is also playing a role in answering these questions. The three papers chosen for this special issue all have an industry focus. It is through publishing such papers that we can bridge the industry-research gap, ensuring that our research will have an impact, not only through our teaching and graduate programs but also by having a direct effect on the improvement of processes within industry. During the evolution history of this conference series, from 1984 to 1996, the International Software Process Workshops (SPW) attracted many academic researchers and industrial practitioners. Then came the International Conference on the Software Process, (ICSP, from 1991 until 1996), the International Workshop on Software Process Simulation and Modeling (from 1998 until 2006), and the SPW (in 2005 and 2006). International Workshop on Software Process Simulation and Modeling and SPW were held together in 2006 and merged in 2007 to form the new International Conference on Software Process. In 2009, the third ICSP conference was held in Vancouver, Canada, on May 16–17, and was co-located with ICSE 2009 (the International Conference on Software Engineering). The ICSSP (International Conference on Software and System Process) conferences continue the successful ICSP conference series, while broadening ICSP's scope of software development processes to system development and explicitly including processes of other domains such as health care, business, and manufacturing. By sharing process development theories and practices from such domains, ICSSP 2014 aimed at investigating novel solutions to today's software and systems process challenges. The theme of ICSSP 2014 was ‘Processes for Emerging and Evolving Software Systems’. There were 55 submissions (including 35 full papers and 20 short papers) to the conference that were received from 18 countries and regions including Australia, Austria, Brazil, Canada, China, Estonia, Denmark, France, Germany, Ireland, Italy, Japan, Netherlands, Portugal, Russian Federation, Sweden, United Kingdom, and United States of America. After a rigorous review process, 23 papers were accepted for the proceedings of the ICSSP 2014 including 12 full papers and 11 short papers. Following the conference, the editors ranked the accepted full papers based on their reviews and sent invitations to authors of five papers to expand and submit their papers for consideration in the special issue. Papers were submitted, peer reviewed, and revised. Based on the reviews, the Guest Editors selected four out of five papers to be included in this special issue of the Journal of Software Evolution and Process. We have also included a paper from keynote speaker, Prof K. Ryan from Lero. In an extension of his talk at ICSSP 2014, he discusses the importance of changing processes because of the changing world of software. He specifically discusses how specialization, industrialization, globalization, and Agile methods have disrupted software development and presents examples of how software process research has evolved new methods to cope with these changes. Software process lines provide a systematic approach to develop and manage software processes. It defines a reference process containing general process assets, whereas a well-defined customization approach allows process engineers to create new process variants. Variability operations are an instrument to realize flexibility by explicitly declaring required modifications, which are applied to create a procedurally generated company-specific process. However, little is known about which variability operations are suitable in practice. M. Kuhrmann, D. Méndez Fernández and T. Ternité, in their paper ‘On the Use of Variability Operations in the V-Modell XT Software Process Line’, study the feasibility of variability operations to support the development of software process lines in the context of the V-Modell XT. They provide an initial catalog of variability operations as an improvement proposal for other process models. The identified variability operations allow for systematically modifying the content of process model elements and the process documentation, and they allow for altering the structure of a process model and its description. High-maturity software development processes, such as the Team Software Process and the accompanying Personal Software Process, can generate significant amounts of data that can be periodically analyzed to identify performance problems, determine their root causes and devise improvement actions. However, there is a lack of tool support for automating that type of analysis and hence diminish the manual effort and expert knowledge required. M. Raza and J. P. Faria, in their paper ‘A Model for Analyzing Performance Problems and Root Causes in the Personal Software Process’, propose a comprehensive performance model, addressing time estimation accuracy, quality, and productivity, to enable the automated (tool based) analysis of performance data produced by Personal Software Process developers, namely, identify, and rank performance problems and their root causes. Software test processes are complex and costly. To reduce testing effort without compromising effectiveness and product quality, automation of test activities has been adopted as a popular approach in software industry. However, because test automation usually requires substantial upfront investments, automation is not always more cost-effective than manual testing. V. Garousi and D. Pfahl, in their paper ‘When to Automate Software Testing? A Decision-Support Approach Based on Process Simulation’, investigate how the simulation model using the System Dynamics modeling technique can help decision-makers decide whether and to what degree the company should automate their test processes. Workflow temporal verification guarantees on-time completion which is one of the most important quality of service dimensions for business processes running in the cloud. However, as today's business systems often need to handle a large number of concurrent customer requests, conventional response time based process monitoring strategies conducted in a one by one fashion cannot be applied efficiently to a large batch of parallel processes because of significant time overhead. To address this problem, X. Liu, D. Wang, D. Yuan, F. Wang, Y/ Yang, in their paper ‘Workflow Temporal Verification for Monitoring ParallelBusiness Processes’, proposes a quality of service-aware throughput based checkpoint selection strategy, which can dynamically select a small number of checkpoints along the system timeline to facilitate the temporal verification of throughput constraints and achieve the target on-time completion rate, based on a novel runtime throughput consistency model. In conclusion, we would like to thank all the authors and reviewers of this special issue. We also would like to thank the ICSSP Steering Committee: Barry Boehm, Ross Jeffrey, Mingshu Li, Leon Osterweil, and David Raffo for their support and help with the conference. We hope you enjoy reading the papers and that, within industry, this research can be useful.
LiGuo Huang, He Zhang 0001, Ita Richardson
J. Softw. Evol. Process.1
2015 Recovering Traceability Links in Requirements Documents
abstract
Software system development is guided by the evolution of requirements.In this paper, we address the task of requirements traceability, which is concerned with providing bi-directional traceability between various requirements, enabling users to find the origin of each requirement and track every change made to it.We propose a knowledge-rich approach to the task, where we extend a supervised baseline system with (1) additional training instances derived from human-provided annotator rationales; and (2) additional features derived from a hand-built ontology.Experiments demonstrate that our approach yields a relative error reduction of 11.1-19.7%.
Zeheng Li, LiGuo Huang, Vincent Ng 0001
CoNLL3
2015 AutoODC: Automated generation of orthogonal defect classifications
LiGuo Huang, Vincent Ng 0001, Isaac Persing, Zeheng Li, Ruili Geng, Jeff Tian
Autom. Softw. Eng.1
2015 SMPLearner: learning to predict software maintainability
LiGuo Huang, Vincent Ng 0001, Jidong Ge
Autom. Softw. Eng.2
2015 Can method data dependencies support the assessment of traceability between requirements and source code?
abstract
Requirements traceability benefits many software engineering activities, such as change impact analysis and risk assessment. However, these activities require complete and correct traceability links which is not trivial, making traceability assessment an important field of study. In recent years, requirements traceability research has focused on using call dependencies within source code to understand how code properties contribute to the implementation of a requirement and to assess whether traceability links are correct and complete. These approaches largely ignore the role of existing data dependencies within the source code. That is, methods may never call each other, but may still depend upon another by sharing data. We identified five research questions and validated them on five software systems, covering 4 to 72 KLOC. We found that data dependencies are as relevant as call dependencies for assessing requirements traceability. Even more interesting, our analyses show that data dependencies complement call dependencies in the assessment. These findings have strong implications on code understanding, including trace capture, maintenance, and validation techniques. Copyright © 2015 John Wiley & Sons, Ltd.
Hongyu Kuang, Patrick Mäder, Hao Hu 0001, Achraf Ghabi, LiGuo Huang, Jian Lu 0001, Alexander Egyed
J. Softw. Evol. Process.5
2014 The incremental commitment spiral model (ICSM): principles and practices for successful systems and software
abstract
This paper summarizes the Incremental Commitment Spiral Model (ICSM), a process model generator that enables organizations to determine which process model, or combination of models, best fits the needs of each system.
Barry W. Boehm, LiGuo Huang
ICSSP2
2014 Reducing view inconsistency by predicting avatars' motion in multi-server distributed virtual environments
Yizhi Ren, LiGuo Huang, Hua Hu 0001
J. Netw. Comput. Appl.4
2014 Guest editors' introduction
LiGuo Huang, Ove Armbrust
J. Softw. Evol. Process.1
2012 Do data dependencies in source code complement call dependencies for understanding requirements traceability?
abstract
It is common practice for requirements traceability research to consider method call dependencies within the source code (e.g., fan-in/fan-out analyses). However, current approaches largely ignore the role of data. The question this paper investigates is whether data dependencies have similar relationships to requirements as do call dependencies. For example, if two methods do not call one another, but do have access to the same data then is this information relevant? We formulated several research questions and validated them on three large software systems, covering about 120 KLOC. Our findings are that data relationships are roughly equally relevant to understanding the relationship to requirements traces than calling dependencies. However, most interestingly, our analyses show that data dependencies complement call dependencies. These findings have strong implications on all forms of code understanding, including trace capture, maintenance, and validation techniques (e.g., information retrieval).
Hongyu Kuang, Patrick Mäder, Hao Hu 0001, Achraf Ghabi, LiGuo Huang, Jian Lu 0001, Alexander Egyed
ICSM5
2012 Discovering process models from event multiset
Dongyi Wang, Jidong Ge, Hao Hu 0001, Bin Luo 0003, LiGuo Huang
Expert Syst. Appl.5
2012 Hybrid modeling and simulation for trustworthy software process management: a stakeholder-oriented approach
abstract
SUMMARY Process Management Model (PMM) and Process Simulation Model (PSM) are the critical infrastructural components of the Trustworthy Process Management Framework (TPMF), which involves a large and heterogeneous group of stakeholders in process modeling and simulation to improve process trustworthiness. Process Modeling Stakeholders (PMS) have different levels of dependency on various process modeling and simulation techniques. They may also possess different perspectives or concerns in modeling. To support trustworthy process management, this paper integrates the stakeholder‐oriented approach and hybrid simulation technique into software process modeling at three levels of abstraction (i.e., activity, sub‐process and system). The hybrid process simulation combinesmicro‐leveldiscrete process models with themacro‐levelcontinuous process models to capture process dynamics. In particular, the stakeholder‐oriented approach addresses the various perspectives of PMS during process modeling and simulation. Finally, a case study with a realistic process model demonstrates that this approach incrementally integrates stakeholders' modeling concerns through hybrid simulation, which is difficult to achieve using discrete or continuous modeling/simulation techniques independently. Copyright © 2010 John Wiley & Sons, Ltd.
LiGuo Huang, He Zhang 0001, Supannika Koolmanojwong
J. Softw. Evol. Process.2
2011 Relevance and alignment of Real-Client Real-Project courses via technology transfer
abstract
It is often claimed Real-Client Real-Project (RCRP) courses are important providers of industry relevant experience and skills to students. How do we know this is so? We cannot prove this or improve RCRP industry relevance without tangible evidence. Here we suggest that the degree an industry partner is willing to accept technology transfer for technologies used within an RCRP course is a strong indicator of relevance. There must be common challenges between RCRP and industry software development for which technology proven in the classroom will also be relevant in industry. We describe our experiences with using an RCRP parallel to the TAME technology transfer model to assess the degree of willingness for technology transfer facilitated by technology use in RCRP courses. We find benefits from realizing envisaged synergies between software engineering research, education, and technology transfer to industry. We further note how the willingness for technology transfer indicator is useful for aligning software engineering courses for industry relevance.
LiGuo Huang, Daniel Port
CSEE&T1
2011 An empirical assessment of a systematic search process for systematic reviews
abstract
Background: Systematic Literature Reviews (SLRs) have been gaining significant attention from Software En-gineering (SE) researchers since 2004. Several researches have also working on improving the scientific and techno-logical infrastructure available to support SLRs in SE. Objective: The study reported in this paper aims to vali-date the QGS-based search process for SLR, i.e. whether a more effective and/or productive search can be achieved by following such a systematic process. Method: We used a dual-case study, in which each case includes two observations of SE literature search for the same SLR but using and not using the QGS-based approach. Results: The overall sensitivity and precision of each observation were calculated for the search cases that im-plemented different search design. Conclusions: A systematic search process (QGS-based search in this paper) appears to gain higher sensitivity and precision in the two cases of SLR. Such a search process may help capture more relevant studies as well as save re-viewers ’ time spent in literature search activities. Our ob-servations also proof that an integrated search strategy is recommended for SLRs in SE to avoid the possible limita-tions of applying single manual or automated search. 1
He Zhang 0001, Muhammad Ali Babar 0001, Juan Li 0001, LiGuo Huang
EASE5
2011 Empirical Research in Software Process Modeling: A Systematic Literature Review
abstract
Recognized as one of the powerful technologies in software process engineering, Software Process Modeling (SPM) has received significant attention over the last three decades. Although empirical research plays a critical role in software engineering, the state-of-the-practice of empirical research in SPM has not been systematically reviewed. This paper serves as a status report of the assessment of empirical research in SPM by analyzing all refereed studies that were published in relevant venues from 1987 to 2008 using systematic review methodology. The primary findings indicate that in current SPM-related empirical studies, (1) software process management and improvement (SPI) was not yet the most popular primary research objectives, (2) exploratory empirical research methods, e.g., case study and action research, were dominantly used, (3) there were common issues in empirical research reports in terms of following rigorous reporting guidelines. Based on the review results, we also suggest the future needs for empirical research in SPM, in terms of research topics, SPM techniques, the strengths of research methodology and the rigors of empirical studies.
He Zhang 0001, LiGuo Huang
ESEM3
2011 Experiences with text mining large collections of unstructured systems development artifacts at jpl
abstract
Often repositories of systems engineering artifacts at NASA's Jet Propulsion Laboratory (JPL) are so large and poorly structured that they have outgrown our capability to effectively manually process their contents to extract useful information. Sophisticated text mining methods and tools seem a quick, low-effort approach to automating our limited manual efforts. Our experiences of exploring such methods mainly in three areas including historical risk analysis, defect identification based on requirements analysis, and over-time analysis of system anomalies at JPL, have shown that obtaining useful results requires substantial unanticipated efforts - from preprocessing the data to transforming the output for practical applications. We have not observed any quick 'wins' or realized benefit from short-term effort avoidance through automation in this area. Surprisingly we have realized a number of unexpected long-term benefits from the process of applying text mining to our repositories. This paper elaborates some of these benefits and our important lessons learned from the process of preparing and applying text mining to large unstructured system artifacts at JPL aiming to benefit future TM applications in similar problem domains and also in hope for being extended to broader areas of applications.
Daniel Port, Allen P. Nikora, Jairus Hihn, LiGuo Huang
ICSE4
2011 Impact of process simulation on software practice: an initial report
abstract
Process simulation has become a powerful technology in support of software project management and process improvement over the past decades. This research, inspired by the Impact Project, intends to investigate the technology transfer of software process simulation to the use in industrial settings, and further identify the best practices to release its full potential in software practice. We collected the reported applications of process simulation in software industry, and identified its wide adoption in the organizations delivering various software intensive systems. This paper, as an initial report of the research, briefs a historical perspective of the impact upon practice based on the documented evidence, and also elaborates the research-practice transition by examining one detailed case study. It is shown that research has a significant impact on practice in this area. The analysis of impact trace also reveals that the success of software process simulation in practice highly relies on the association with other software process techniques or practices and the close collaboration between researchers and practitioners.
He Zhang 0001, D. Ross Jeffery, Dan X. Houston, LiGuo Huang, Liming Zhu 0001
ICSE4
2011 GoPoMoSA: a goal-oriented process modeling and simulation advisor
abstract
This paper presents GoPoMoSA, a Goal-oriented Process Modeling and Simulation Advisor that semi-automatically discovers suitable Software Process Modeling and Simulation (SPMS) techniques for (inexperienced) process modelers to achieve their process modeling goals. GoPoMoSA takes the goal-oriented modeling approach that captures the associations among Process Modeling Stakeholder goals and existing SPMS techniques via Relevant Process Elements modeled in the knowledge graphs. We evaluated the accuracy and feasibility of GoPoMoSA with data collected from 212 published SPMS literatures and a real-world process modeling and simulation case on requirements traceability. Our results show that GoPoMoSA (1) was able to find suitable SPMS techniques based on stakeholder goals with an average of 85.38% accuracy; (2) helped novice process modelers effectively and efficiently achieve their goals.
LiGuo Huang, He Zhang 0001, Alexander Egyed
ICSSP2
2011 AutoODC: Automated generation of Orthogonal Defect Classifications
abstract
Orthogonal Defect Classification (ODC), the most influential framework for software defect classification and analysis, provides valuable in-process feedback to system development and maintenance. Conducting ODC classification on existing organizational defect reports is human intensive and requires experts' knowledge of both ODC and system domains. This paper presents AutoODC, an approach and tool for automating ODC classification by casting it as a supervised text classification problem. Rather than merely apply the standard machine learning framework to this task, we seek to acquire a better ODC classification system by integrating experts' ODC experience and domain knowledge into the learning process via proposing a novel Relevance Annotation Framework. We evaluated AutoODC on an industrial defect report from the social network domain. AutoODC is a promising approach: not only does it leverage minimal human effort beyond the human annotations typically required by standard machine learning approaches, but it achieves an overall accuracy of 80.2% when using manual classifications as a basis of comparison.
LiGuo Huang, Vincent Ng 0001, Isaac Persing, Ruili Geng, Jeff Tian
ASE1
2010 Text mining in supporting software systems risk assurance
abstract
Insufficient risk analysis often leads to software system design defects and system failures. Assurance of software risk documents aims to increase the confidence that identified risks are complete, specific, and correct. Yet assurance methods rely heavily on manual analysis that requires significant knowledge of historical projects and subjective, perhaps biased judgment from domain experts. To address the issue, we have developed RARGen, a text mining-based approach based on well-established methods aiming to automatically create and maintain risk repositories to identify usable risk association rules (RARs) from a corpus of risk analysis documents. RARs are risks that have frequently occurred in historical projects. We evaluate RARGen on 20 publicly available e-service projects. Our evaluation results show that RARGen can effectively reason about RARs, increase confidence and cost-effectiveness of risk assurance, and support difficult-to-perform activities such as assuring complete-risk identification.
LiGuo Huang, Daniel Port, Tao Xie 0001, Tim Menzies
ASE1
2006 Applying the Value/Petri process to ERP software development in China
abstract
Commercial organizations increasingly need software processes sensitive to business value, quick to apply, and capable of early analysis for subprocess consistency and compatibility. This paper presents experience in applying a lightweight synthesis of a Value-Based Software Quality Achievement (VBSQA) process and an Object-Petri-Net-based process model (called VBSQA-OPN) to achieve a manager-satisfactory process for software quality achievement in an on-going ERP software project in China. The results confirmed that 1) the application of value-based approaches was inherently better than value-neutral approaches adopted by most ERP software projects; 2) the VBSQA-OPN model provided project managers with a synchronization and stabilization framework for process activities, success-critical stakeholders and their value propositions; 3) process visualization and simulation tools significantly increased management visibility and controllability for the success of software project.
LiGuo Huang, Barry W. Boehm, Hao Hu 0001, Jidong Ge, Jian Lu 0001
ICSE1
2003 Strategic Architectural Flexibility
abstract
Most projects commit to a set of required features and (at best) a most-likely budget and schedule for developing them. This means that, even before the changes start coming, there is roughly a 50% chance that the most-likely budget and schedule are insufficient, and the project is headed for an overrun. Planning for change in a development project is essential. But how much should be invested in architectural flexibility to accommodate this? Too little will incur a high risk of costly late changes and architecture breakage; too much may not leave enough time to implement a sufficient set of critical capabilities. We have been using and refining a model based approach to assist in determining an appropriate degree of architectural flexibility by introducing a modularity factor for the software architecture based on the core capabilities and a set of anticipated changes. This experience has helped us identify the critical success factors for strategically applying architectural flexibility within tight constraints such as cost, quality, or a fixed schedule. We elaborate the critical success factors, present a case study of their application, and their relation to recent research results in such areas as strategic design.
Daniel Port, LiGuo Huang
ICSM2
2003 A cost driven disk scheduling algorithm for multimedia object retrieval
abstract
This paper describes a novel cost-driven disk scheduling algorithm for environments consisting of multipriority requests. An example application is a video-on-demand (VOD) system that provides high and low quality services, termed priority 2 and 1, respectively. Customers ordering a high quality (priority 2) service pay a higher fee and are assigned a higher priority by the underlying system. Our proposed algorithm minimizes costs by maintaining one-queue and managing requests intelligently in order to meet the deadline of as many priority 1 requests as possible while maximizing the number of priority 2 requests that meet their deadline. Our algorithm is general enough to accommodate an arbitrary number of priority levels. Prior schemes, collectively termed "multiqueue" schemes maintain a separate queue for each priority level in order to optimize the performance of the high priority requests only. When compared with our proposed scheme, in certain cases, our technique provides more than one order of magnitude improvement in total cost.
Shahram Ghandeharizadeh, LiGuo Huang, Ibrahim Kamel
IEEE Trans. Multim.2