Guisheng Fan

dblp:17/5233 · DBLP profile ↗
← Back
118ranked-venue papers
24as first author
68since 2021 · last 2026
0000-0002-2702-0242ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 60 · 16 first-author · 28 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 11 since 2021Systems, architecture and hardware · 13 · 1 first-author · 10 since 2021Computer networks · 13 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
abstract
Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks.However, existing methods generally lack the capability for continuous learning and self-evolution from interactions, limiting the diversity and adaptability of attack strategies.To address this, we propose ASTRA, an automated framework capable of autonomously discovering, retrieving, and evolving attack strategies.ASTRA operates on a closed-loop "attack-evaluate-distillreuse" mechanism, which not only generates attack prompts but also automatically distills reusable strategies from every interaction.To systematically manage these strategies, we introduce a dynamic three-tier strategy library (Effective, Promising, and Ineffective) that categorizes strategies based on performance.This hierarchical memory mechanism enables the framework to enhance efficiency by leveraging successful patterns while optimizing the exploration space by avoiding known failures.Extensive experiments in a black-box setting demonstrate that ASTRA significantly outperforms existing baselines.
Kan Ling, Yichi Zhu, Hengrun Zhang 0004, Guisheng Fan, Huiqun Yu
ACL (1)6
2026 Evo-AA: Evolution-Aware Adaptation of Protein Language Models
Huiqun Yu, Guisheng Fan, Hengrun Zhang 0004
COMPSAC3
2026 Discriminative and semantic-aligned representation learning for just-in-time defect prediction
Yuguo Liang, Chengcheng Wu, Guisheng Fan, Huiqun Yu
Expert Syst. Appl.3
2026 Usage patterns of software product metrics in assessing developers' output: A comprehensive study
Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Yuguo Liang
Inf. Softw. Technol.3
2026 Revisiting pre-trained models and feature fusion strategies for just-in-time defect prediction
Yuguo Liang, Guisheng Fan, Huiqun Yu, Chengcheng Wu, Zijie Huang 0001
Inf. Softw. Technol.2
2026 Automatic identification of extrinsic bug reports for just-in-time bug prediction
Guisheng Fan, Yuguo Liang, Longfei Zu, Huiqun Yu, Zijie Huang 0001
Sci. Comput. Program.1
2026 Guard Against Infringement: An Anti-Distillation Federated Learning Watermarking Framework
abstract
To balance the gap between data privacy and the need for data fusion, federated learning (FL) has been proposed and has become a hot-point method to address data silos and privacy issues. However, AI models exchanged in FL face risks such as illegal copying, redistribution and/or free-riding. To address these risks, FL watermarking frameworks have been proposed to assert and protect the intellectual property (IP) of models, which are resistant to popular watermark removal attacks. Knowledge distillation has recently been of significant contribution to FL convergence performance optimization but brings vulnerability to FL watermark robustness with distillation attack, which enables attackers to maintain high performance on the main task while erasing the watermarks. In response, we introduce a new FL watermarking framework called FedRW, which focuses specifically on anti-distillation. FedRW employs model regularization techniques to bind the main task parameters with the watermark task parameters, thereby enhancing resistance to distillation attacks. Extensive experiments confirm the threat of distillation attacks in FL and demonstrate that FedRW is more resistant to distillation compared to existing FL watermarking frameworks.
Xiao Yi, Hengrun Zhang 0001, Huiqun Yu, Guisheng Fan, Haojin Zhu
IEEE Trans. Dependable Secur. Comput.4
2025 VFProber: A Vulnerability-Fixing Identification Framework Based on Code Changes and Semantic Adjustment
abstract
With the accelerated development of software, developers face the continuous challenge of fixing vulnerabilities but vulnerability-fixing commits often disassociated from the vulnerabilities, and the structural and semantic differences between code changes and natural language present significant challenges in identifying these commits. Existing approaches utilize machine learning and deep learning techniques to address this problem, but they often do not fully leverage the information about code changes. In this paper, we propose VFProber, a method based on a code change pretrained model, aiming to provide a comprehensive and unified framework for identifying vulnerability-fixing commits. VFProber uses semantic adjustment to distinguish between context-sensitive and context-insensitive code units in code changes, thereby enhancing the model’s understanding of code changes during the training process. Secondly, VFProber employs a novel code change pretrained model as a feature extractor. Compared with ordinary code pretrained models, it can better meet the requirements of the vulnerability-fixing identification task. Moreover, we constructed a vulnerability-fixing dataset containing two common programming languages, Java and JavaScript, from industrial projects. In the experimental section, we designed three tasks to evaluate the method. The results show that, compared with the best baseline, VFProber performs better in the vulnerability-fixing identification task and can effectively reduce false positives and false negatives.
Jianan Dong, Guisheng Fan, Yueming Yu, Yuguo Liang, Yujie Ye, Huiqun Yu
COMPSAC2
2025 License Compatibility Detection for OpenJavaWorks Open-Source Projects Based on Code Similarity
abstract
This study investigates the relationship between code similarity and open-source license compatibility in Java projects. While the extensive use of open-source projects promotes code reuse, it also poses challenges concerning license compatibility. Current license detection methods often overlook code similarities and the degree of sharing among projects. To address these issues, we pave a new research path in the field of license compatibility, with a particular focus on the correlation between code similarity and license compatibility. Unlike traditional methods that focus on individual projects, this research offers a comprehensive approach that identifies potential license compatibility issues both within and across multiple projects. Our analysis employs the BigCloneBench (BCB) dataset and conducts a large-scale empirical study on 746,048 Java open-source projects, which we collected and labeled as OJW. Findings indicate that 11.9% of multi-license projects face compatibility issues, while 23.7% of project pairs with over 70% code similarity demonstrate license incompatibility. This innovative approach not only overcomes the limitations of existing research but also provides developers with a practical tool to reduce legal risks during code reuse, ensuring compliance in open-source software.
Huiqun Yu, Guisheng Fan, Yuguo Liang
COMPSAC3
2025 Affinity and Interference-Aware Service Deployment for Energy Efficiency in Cloud Data Centers: A Deep Reinforcement Learning Approach
abstract
Cloud computing has revolutionized data center management by providing scalable and efficient resources for processing and data management. However, deploying containerd-based services in data centers presents significant challenges: (1) Active servers that are underutilized result in high energy consumption, necessitating optimization for energy efficiency; (2) Affinity requirements between services and servers must be considered to ensure appropriate deployments; (3) Quality of Service (QoS) requirements must be met, particularly to avoid performance interference when multiple services are deployed on the same server. To address these challenges, we propose a novel algorithm, Affinity-Interference Energy Deployment (AIED), based on Deep Reinforcement Learning (DRL). This algorithm strategically consolidates services onto fewer servers to optimize energy efficiency while adhering to stringent QoS and affinity constraints. By employing a demand-supply model to quantify QoS requirements and formulating the deployment challenge as a Markov Decision Process (MDP), our algorithm dynamically adapts to fluctuating demands and resource availability. Extensive simulations demonstrate that AIED significantly outperforms existing baseline strategies, reducing energy consumption while ensuring robust compliance with both QoS and affinity constraints.
Huiqun Yu, Guisheng Fan, Shengwei Liu, Hengrun Zhang 0004, Liqiong Chen
COMPSAC3
2025 JIT-Align: A Semantic Alignment-Based Ranking Framework for Just-In-Time Defect Prediction
abstract
To promptly identify software defects and prevent defective code changes from being integrated into the repository, Just-In-Time Software Defect Prediction (JIT-SDP) has demonstrated promising research findings. Recent studies have begun to utilize Pre-trained Models (PTMs) for training and prediction, yet these models inherently impose input length limitations, leading to forced truncation of inputs. However, previous work has largely overlooked the impact of forced truncation, even though it may inadvertently discard critical input information, leading to degraded model performance. Moreover, some existing methods fail to maintain consistency in truncation during each model construction process, leading to unexplainable truncations and unstable model performance. In addition, previous datasets suffer from limitations and incompleteness. To this end, we construct a large-scale and comprehensive dataset, MC4Defect. Moreover, we propose JIT-Align, which prioritizes code changes within a commit using a semantic alignment algorithm to make full use of the limited input space of PTMs. To evaluate the feasibility of JIT-Align, we first assess the classification capability of our method by comparing it against four baselines across five datasets. Then, we conduct ablation studies on the proposed semantic alignment framework to validate its effectiveness. Experimental results show that JIT-Align, along with its semantic alignment framework, outperforms all baselines in JIT-SDP tasks, with average F1 score improvements of 3.1%-9.6% and MCC increases of 3.1%-9.7% across all projects, exhibiting higher stability and better interpretability compared to alternative approaches.
Yujie Ye, Huiqun Yu, Guisheng Fan, Yuguo Liang, Jianan Dong
COMPSAC3
2025 FinPTA: An Effective Model for Financial Sentiment Analysis
Xiao Yi, Guisheng Fan, Huiqun Yu, Hengrun Zhang 0004
ICECCS3
2025 SelectDataset: Enabling Dataset Exploration Through Enriched Descriptive Metadata and Hierarchical Model-Based Application Topic Classification
Ruixin Yuan, Qiantai Peng, Hengrun Zhang 0004, Huiqun Yu, Guisheng Fan
ICIC (8)6
2025 Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
abstract
In-silico prediction of protein mutant stability, measured by the difference in Gibbs free energy change ($\Delta \Delta G$), is fundamental for protein engineering. Current sequence-to-label methods typically employ two-stage pipelines: (i) encoding mutant sequences using neural networks (e.g., transformers), followed by (ii) the $\Delta \Delta G$ regression from the latent representations. Although these methods have demonstrated promising performance, their dependence on specialized neural network encoders significantly increases the complexity. Additionally, the requirement to compute latent representations individually for each mutant sequence negatively impacts computational efficiency and poses the risk of overfitting. This work proposes the Venus-MAXWELL framework, which reformulates mutation $\Delta \Delta G$ prediction as a sequence-to-landscape task. In Venus-MAXWELL, mutations of a protein and their corresponding $\Delta \Delta G$ values are organized into a landscape matrix, allowing our framework to learn the $\Delta \Delta G$ landscape of a protein with a single forward and backward pass during training. To this end, we curated a new $\Delta \Delta G$ benchmark dataset with strict controls on data leakage and redundancy to ensure robust evaluation. Leveraging the zero-shot scoring capability of protein language models (PLMs), Venus-MAXWELL effectively utilizes the evolutionary patterns learned by PLMs during pre-training. More importantly, Venus-MAXWELL is compatible with multiple protein language models. For example, when integrated with the ESM-IF, Venus-MAXWELL achieves higher accuracy than ThermoMPNN with 10$\times$ faster in inference speed (despite having 50$\times$ more parameters than ThermoMPNN). The training codes, model weights, and datasets are publicly available at https://github.com/ai4protein/Venus-MAXWELL.
Yuanxi Yu, Fan Jiang 0013, Xinzhu Ma, Bozitao Zhong, Wanli Ouyang, Guisheng Fan, Huiqun Yu
NeurIPS7
2025 Sequence-only prediction of binding affinity changes: a robust and interpretable model for antibody engineering
abstract
MOTIVATION: A pivotal area of research in antibody engineering is to find effective modifications that enhance antibody-antigen binding affinity. Traditional wet-lab experiments assess mutants in a costly and time-consuming manner. Emerging deep learning solutions offer an alternative by modeling antibody structures to predict binding affinity changes. However, they heavily depend on high-quality complex structures, which are frequently unavailable in practice. Therefore, we propose ProtAttBA, a deep learning model that predicts binding affinity changes based solely on the sequence information of antibody-antigen complexes. RESULTS: ProtAttBA employs a pre-training phase to learn protein sequence patterns, following a supervised training phase using labeled antibody-antigen complex data to train a cross-attention-based regressor for predicting binding affinity changes. We evaluated ProtAttBA on three open benchmarks under different conditions. Compared to both sequence- and structure-based prediction methods, our approach achieves competitive performance, demonstrating notable robustness, especially with uncertain complex structures. Notably, our method possesses interpretability from the attention mechanism. We show that the learned attention scores can identify critical residues with impacts on binding affinity. This work introduces a rapid and cost-effective computational tool for antibody engineering, with the potential to accelerate the development of novel therapeutic antibodies. AVAILABILITY AND IMPLEMENTATION: Source codes and data are available at https://github.com/code4luck/ProtAttBA.
Yang Tan 0001, Wenrui Gou, Guisheng Fan, Bingxin Zhou
Bioinform.5
2025 Tool or Toy: Are SCA tools ready for challenging scenarios?
Congyan Shu, Guisheng Fan, Huiqun Yu, Zijie Huang 0001, Yuguo Liang
Comput. Secur.3
2025 Request Deadline Split and Interference-Aware Request Migration in Edge Cloud
abstract
ABSTRACT Edge computing extends computing resources from the data center to the edge of the network to better handle latency‐sensitive tasks. However, with the rise of the Internet of Things, edge devices with limited processing capabilities face difficulties in executing requests with fluctuating request peaks. In order to meet the deadline constraints of latency‐sensitive tasks, a feasible solution is to offload some latency‐sensitive tasks to other nearby edge devices. This article studies the problem of request migration in edge computing systems and minimizes the request deadline violation rate based on actual online arrival patterns, performance interference phenomena, and deadline constraints. Since a request contains multiple services and request migration will lead to changes in server resource competition pressure, we split the problem into three sub‐problems, dividing the request deadline to determine the maximum response time of the service, determining the performance of the service under different resource pressures and the request migration strategies. To this end, we propose two deadline splitting methods, a performance interference model under multi‐resource pressure, and two heuristic request migration strategies. Since this article considers online edge scenarios, the number and type of requests are black boxes. We conduct simulation experiments and find that our method has only one‐third the number of request violations of other methods.
Huiqun Yu, Guisheng Fan, Jiayin Zhang
Concurr. Comput. Pract. Exp.3
2025 Adaptive Task Scheduling Under Dynamic Edge System Loads: A Deep Reinforcement Learning Approach
abstract
ABSTRACT Edge computing systems are in great need of task scheduling due to resource constraints. However, existing scheduling algorithms typically optimize single objectives and lack adaptability to varying system conditions, failing to balance response time minimization during low workloads with queue balance maintenance under high workloads. This paper proposes an adaptive task scheduling algorithm based on Soft Actor‐Critic (SAC) with a novel workload‐aware reward mechanism, which automatically transitions between response time optimization and queue balance prioritization according to system load conditions. The whole scheduling problem is modeled as a Markov Decision Process (MDP), and a sliding window‐based performance evaluation framework is introduced to provide robust system assessment. Extensive experiments across multiple scenarios demonstrate that our method consistently achieves optimal response time across varying workload conditions, significantly outperforming traditional scheduling algorithms, while maintaining effective queue balance comparable to load‐based approaches under high workload scenarios.
Huiqun Yu, Guisheng Fan, Hengrun Zhang 0004
Concurr. Comput. Pract. Exp.3
2025 Automatic Code Summarization Using Abbreviation Expansion and Subword Segmentation
abstract
ABSTRACT Automatic code summarization refers to generating concise natural language descriptions for code snippets. It is vital for improving the efficiency of program understanding among software developers and maintainers. Despite the impressive strides made by deep learning‐based methods, limitations still exist in their ability to understand and model semantic information due to the unique nature of programming languages. We propose two methods to boost code summarization models: context‐based abbreviation expansion and unigram language model‐based subword segmentation. We use heuristics to expand abbreviations within identifiers, reducing semantic ambiguity and improving the language alignment of code summarization models. Furthermore, we leverage subword segmentation to tokenize code into finer subword sequences, providing more semantic information during training and inference, thereby enhancing program understanding. These methods are model‐agnostic and can be readily integrated into existing automatic code summarization approaches. Experiments conducted on two widely used Java code summarization datasets demonstrated the effectiveness of our approach. Specifically, by fusing original and modified code representations into the Transformer model, our Semantic Enhanced Transformer for Code Summarizsation (SETCS) serves as a robust semantic‐level baseline. By simply modifying the datasets, our methods achieved performance improvements of up to 7.3%, 10.0%, 6.7%, and 3.2% for representative code summarization models in terms of BLEU‐4 , METEOR , ROUGE‐L and SIDE , respectively.
Yuguo Liang, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
Expert Syst. J. Knowl. Eng.2
2025 SRPM-Sol: A Structure Robust Protein Multimodal Model for Solubility Prediction
abstract
The solubility of natural proteins is closely linked to their expression and purification processes. Accurate computational prediction of protein solubility not only aids in functional assessment but also reduces the cost of preliminary wet-lab experiments. The current mainstream deep learning prediction methods have begun to explore the multimodal framework. However, existing multimodal models mainly focus on sequences and structures, overlooking other influential factors. Additionally, inherent errors in predicted structure information pose a significant challenge to model robustness. To solve the above issues, we introduce SRPM-Sol, a novel multimodal protein solubility prediction model. Built upon the state-of-the-art ESM3 model, this framework combines amino acid sequences, structure information, secondary structure sequences, and physicochemical properties for more accurate prediction. This is the most diverse-input model to date in the protein multimodal field for solubility prediction. In order to verify the effectiveness of our method, we have constructed the first hierarchical dataset, PDE-Sol, by organizing data based on Predicted Local Distance Difference Test (pLDDT) scores. The experimental results demonstrate that compared to the baselines, SRPM-Sol achieves stronger robustness and higher accuracy on different levels of PDE-Sol, even in the presence of uncertain structure information.
Wenhui Ge, Yang Tan 0001, Huiqun Yu, Guisheng Fan
IEEE Trans. Comput. Biol. Bioinform.4
2025 Energy, Cost and Reliability-Aware Workflow Scheduling on Multi-Cloud Systems: A Multi-Objective Evolutionary Approach
abstract
Nowadays, cloud computing has become a suitable platform for hosting and executing workflow applications. As the diversity and scale of these applications continue to increase, single-cloud environments are becoming insufficient to meet users’ requirements. Instead, multi-cloud environments have emerged as an ideal solution. However, the complexity of workflow scheduling in multi-cloud environments increases significantly due to the diversified billing mechanisms, heightened reliability demands, and the requirements for reducing energy consumption. To address these challenges, this paper proposes a multi-objective evolutionary algorithm called ECRWSM for workflow scheduling on multi-cloud systems. First, ECRWSM utilizes the population initialization strategy to generate a population with excellent uniformity and sufficient randomness. Then, the diversification strategy is employed to thoroughly explore the solution space. Next, the individual enhancement strategy is used to further improve the solutions. Additionally, an external archive is maintained to store non-dominated solutions throughout the evolutionary process. Comprehensive experiments are conducted to validate the performance of ECRWSM. The experimental results demonstrate that our proposed algorithm ECRWSM outperforms both classical and recent scheduling algorithms.
Zhuoyue Fang, Huiqun Yu, Guisheng Fan, Jiayin Zhang
IEEE Trans. Netw. Serv. Manag.3
2024 Response Time and Energy-Aware Optimization for Co-Locating Microservices and Offline Tasks
abstract
With the widespread application of microservices, data centers use co-location to improve server resource utilization. However, co-location causes performance interference between tasks, which poses challenges to service quality levels. Most of the existing co-location research does not consider the characteristics of microservices and does not take energy consumption as the optimization target of co-location systems. This paper introduces a new indicator to represent the Relative fluctuation of response Time and Energy consumption (RTE). This paper proposes a processor energy optimization mechanism and an offline task scheduling method to solve this problem. This method reduces overall response time by providing a good operating environment for frequently called microservices. The energy consumption optimization mechanism reduces energy consumption by dynamically controlling voltage and frequency. Simulation experiments in practical applications show that our proposed algorithm has the best response time and RTE compared with other algorithms.
Huiqun Yu, Guisheng Fan, Jiayin Zhang
COMPSAC3
2024 Energy-efficient reliability-aware offloading for delay-sensitive tasks in collaborative edge computing
abstract
Summary As a burgeoning paradigm, collaborative mobile edge computing (C‐MEC) can cater to growing computation demand of mobile devices (MDs). However, there are great challenges for joint task offloading and resource allocation. In addition, failures on both MDs and edge servers greatly affect reliable task execution. This paper investigates the joint optimization problem of offloading decision, power allocation, and computation resource allocation in the multi‐user C‐MEC system, where nondivisible tasks may be executed locally, offloaded to nearby collaborative devices, or processed by the edge server. We aim at minimizing the total energy consumption of MDs while satisfying the reliability and delay constraints, and the replication technique is adopted to enhance task execution reliability. To tackle this problem, we first transform it into a bi‐level optimization problem, and then propose an iterative algorithm named JOC. Specifically, the offloading decision is constructed in the upper level on the basis of ant colony system (ACS), and the allocation of working power, transmission power, and computation resources is optimized in the lower level by using the monotonic optimization approach. Simulation experimental results reveal that the proposed algorithm saves more energy and achieves a higher task success rate in comparison with baseline schemes.
Huiqun Yu, Guisheng Fan, Jiayin Zhang
Concurr. Comput. Pract. Exp.3
2024 A code change-oriented approach to just-in-time defect prediction with multiple input semantic fusion
abstract
Abstract Recent research found that fine‐tuning pre‐trained models is superior to training models from scratch in just‐in‐time (JIT) defect prediction. However, existing approaches using pre‐trained models have their limitations. First, the input length is constrained by the pre‐trained models.Secondly, the inputs are change‐agnostic.To address these limitations, we propose JIT‐Block, a JIT defect prediction method that combines multiple input semantics using changed block as the fundamental unit. We restructure the JIT‐Defects4J dataset used in previous research. We then conducted a comprehensive comparison using eleven performance metrics, including both effort‐aware and effort‐agnostic measures, against six state‐of‐the‐art baseline models. The results demonstrate that on the JIT defect prediction task, our approach outperforms the baseline models in all six metrics, showing improvements ranging from 1.5% to 800% in effort‐agnostic metrics and 0.3% to 57% in effort‐aware metrics. For the JIT defect code line localization task, our approach outperforms the baseline models in three out of five metrics, showing improvements of 11% to 140%.
Teng Huang 0004, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Chen-Yu Wu
Expert Syst. J. Knowl. Eng.3
2024 Aligning XAI explanations with software developers' expectations: A case study with code smell prioritization
Zijie Huang 0001, Huiqun Yu, Guisheng Fan, Zhiqing Shao, Yuguo Liang
Expert Syst. Appl.3
2024 Handling hierarchy in cloud data centers: A Hyper-Heuristic approach for resource contention and energy-aware Virtual Machine management
Jiayin Zhang, Huiqun Yu, Guisheng Fan, Jun Li 0153
Expert Syst. Appl.3
2024 Enhancing code summarization with action word prediction
Huiqun Yu, Guisheng Fan, Ziyi Zhou 0002, Zijie Huang 0001
Neurocomputing3
2024 Cold-Start-Aware Cloud-Native Parallel Service Function Chain Caching in Edge-Cloud Network
abstract
Virtualized Network Function (VNF) and Service Function Chain (SFC) are the fundamental components in Network Functions Virtualization (NFV) infrastructure, which supports the evolution of modern 5G networks. For online Internet of Things (IoT) applications, characterized by dynamic and diverse requirements, achieving optimal quality of service hinges on a resource-efficient yet performant SFC caching strategy, which is a critical challenge. Besides, despite the performance boost and flexibility brought by modern cloud-native technology, it brings the cold-start problem due to the requirement for runtime image transmission and booting-up, resulting in a non-negligible launch latency. To tackle these challenges, this paper proposes CPSC (Cloud-Native Parallel SFC Caching framework), a novel approach to address the cloud-native parallel SFC caching problem in edge-cloud networks leveraging Deep Reinforcement Learning (DRL), seeking an efficient resource utilization of the edge-cloud network with consideration of SFC processing performance and cold-start suppressing. Graph Convolutional Network (GCN) -based embeddings are adopted for topology-aware feature extraction of the substrate edge-cloud network as well as the incoming SFC caching requests. Then, a Pointer Network (PN) is utilized for contextual information-aware caching decision-making. Benefiting from the online capability of DRL, CPSC makes caching decisions in an online manner with no prior knowledge requirement on future incoming requests. Extensive simulations show that CPSC manages to outperform the state-of-the-art approaches in edge network acceptance ratio and launch latency, with minimal overhead on the SFC processing performance and decision-making duration.
Jiayin Zhang, Huiqun Yu, Guisheng Fan, Qifeng Tang
IEEE Internet Things J.3
2024 Energy-efficient offloading for DNN-based applications in edge-cloud computing: A hybrid chaotic evolutionary approach
Huiqun Yu, Guisheng Fan, Jiayin Zhang
J. Parallel Distributed Comput.3
2024 On the effectiveness of developer features in code smell prioritization: A replication study
Zijie Huang 0001, Huiqun Yu, Guisheng Fan, Zhiqing Shao, Ziyi Zhou 0002
J. Syst. Softw.3
2024 Bug report priority prediction using social and technical features
abstract
Summary Software stakeholders report bugs in issue tracking system (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time‐insensitive approaches, and they were producing suboptimal results for practical usage. While they ignored the social aspects of ITS, the technical aspects were also limited in surface features of bug reports. To better model the bug report, we extract their topic and most similar code structures. Since ITS bridges users and developers as the main contributors, we also integrate their experience, sentiment, and socio‐technical features to construct a new dataset. Then, we perform two‐classed and multiclassed bug priority prediction based on the dataset. We also introduce adversarial training using generated training data with random word swap and random word deletion. We validate our model in within‐project, cross‐project, and time‐wise scenarios, and it outperforms the two baselines by up to 15% in area under curve‐receiver operating characteristics (AUC‐ROC) and 19% in Matthews correlation coefficient (MCC). We reveal involving contributor (i.e., assignee and reporter) features such as sentiment that could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in social and technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest that researchers should consider both social and technical aspects of ITS in bug report priority prediction and introduce adversarial training to boost model performance.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
J. Softw. Evol. Process.3
2024 Exploring better alternatives to size metrics for explainable software defect prediction
Chenchen Chai, Guisheng Fan, Huiqun Yu, Zijie Huang 0001, Jianshu Ding, Yao Guan
Softw. Qual. J.2
2024 Adaptive edge service deployment in burst load scenarios using deep reinforcement learning
Huiqun Yu, Guisheng Fan, Jiayin Zhang, Qifeng Tang
J. Supercomput.3
2024 Elastic Task Offloading and Resource Allocation Over Hybrid Cloud: A Reinforcement Learning Approach
abstract
Hybrid cloud is an emerging computing cloud solution that leverages the power of the public cloud, without abandoning the computation resources of existing on-premises data-centers. Further, the wide adoption of cloud-native technology, like containers, brings the capability of rapid horizontal and vertical scaling to task workloads. However, the heterogeneity and flexibility can bring more complexity to task processing performance optimization, especially with constrained on-premises energy consumption and public cloud renting cost quota. In this paper, we seek to optimize the task processing performance under long-term on-premises energy consumption and public cloud renting cost constraints via dynamic task offloading and elastic scaling. We formulate the problem as a two-stage mixed integer non-linear programming (MINLP) problem, and propose an online approach named ETHC (elastic task offloading and resource allocation handler over hybrid cloud). For the first stage, we introduce a Lyapunov optimization-assisted Deep Reinforcement Learning (DRL) agent to decompose the long-term optimization problem into per-time-segment sub-problems on making task offloading decisions. In the second stage, based on the M/M/k queuing model, we prove the container instance number configuration and per-instance resource allocation problem as a convex MINLP problem. An efficient bi-section-based algorithm is introduced to obtain the optimal configurations. Extensive simulations show that ETHC manages to stabilize the task processing queue and satisfy the long-term constraints under various environments and parameters setup, with slight overhead on the convergence speed. Besides, optimal resource configuration and instance number can be obtained at each time-segment with low time complexity.
Jiayin Zhang, Huiqun Yu, Guisheng Fan
IEEE Trans. Netw. Serv. Manag.3
2024 Learning to Generate Structured Code Summaries From Hybrid Code Context
abstract
Code summarization aims to automatically generate natural language descriptions for code, and has become a rapidly expanding research area in the past decades. Unfortunately, existing approaches mainly focus on the “one-to-one” mapping from methods to short descriptions, which hinders them from becoming practical tools: 1) The program context is ignored, so they have difficulty in predicting keywords outside the target method; 2) They are typically trained to generate brief function descriptions with only one sentence in length, and therefore have difficulty in providing specific information. These drawbacks are partially due to the limitations of public code summarization datasets. In this paper, we first build a large code summarization dataset including different code contexts and summary content annotations, and then propose a deep learning framework that learns to generate structured code summaries from hybrid program context, named StructCodeSum. It provides both an LLM-based approach and a lightweight approach which are suitable for different scenarios. Given a target method, StructCodeSum predicts its function description, return description, parameter description, and usage description through hybrid code context, and ultimately builds a Javadoc-style code summary. The hybrid code context consists of path context, class context, documentation context and call context of the target method. Extensive experimental results demonstrate: 1) The hybrid context covers more than 70% of the summary tokens in average and significantly boosts the model performance; 2) When generating function descriptions, StructCodeSum outperforms the state-of-the-art approaches by a large margin; 3) According to human evaluation, the quality of the structured summaries generated by our approach is better than the documentation generated by Code Llama.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001
IEEE Trans. Software Eng.4
2023 Cost-efficient security-aware scheduling for dependent tasks with endpoint contention in edge computing
Huiqun Yu, Guisheng Fan, Qifeng Tang, Jiayin Zhang, Liqiong Chen
Comput. Commun.3
2023 Uncertainty-aware scheduling of real-time workflows under deadline constraints on multi-cloud systems
abstract
Summary The elasticity and pay‐as‐you‐go features of cloud computing are popular with customers, and more and more workflow applications are migrating to cloud platforms. Many workflow scheduling algorithms aim to obtain minimal rental costs. However, most of the existing research assumes that task execution time is deterministic. In fact, due to the performance fluctuations of VMs, the task execution time is uncertain before scheduling. Furthermore, many works ignore the cost savings given by multi‐cloud systems. To this end, this paper provides a scheduling framework for real‐time workflows. The framework includes four main components: the workflow analyzer, task pool, task allocation controller, and resource manager. Then based on the framework, we propose the RWSMC heuristic algorithm. The algorithm's goal is to minimize the total rental cost while satisfying the deadline constraints and ensuring the reliability of task execution. The RWSMC algorithm reduces the cost by selecting the appropriate billing mechanism based on the task's execution time and mitigates the impact of uncertain execution time by scheduling the task to the VM with the shortest predicted start time. Simulation experiments demonstrate that our proposed algorithm outperforms three recent state‐of‐the‐art scheduling algorithms in the total rental cost, deadline violation rate, and VM resource utilization.
Huiqun Yu, Guisheng Fan, Jiayin Zhang
Concurr. Comput. Pract. Exp.3
2023 ClassSum: a deep learning model for class-level code summarization
Huiqun Yu, Guisheng Fan, Ziyi Zhou 0002
Neural Comput. Appl.3
2023 Cost-effective approaches for deadline-constrained workflow scheduling in clouds
Huiqun Yu, Guisheng Fan
J. Supercomput.3
2023 An Energy-Efficient Dynamic Scheduling Method of Deadline-Constrained Workflows in a Cloud Environment
abstract
With the rapid development of cloud applications, the computing requests of cloud data centers have increased significantly, consuming a lot of energy, making cloud data centers unsustainable, which is very unfavorable from both the cloud provider’s point of view and the environmental point of view. Therefore, it is crucial to minimize energy consumption and improve resource utilization while ensuring user service quality constraints. In this paper, we propose a hybrid workflow scheduling algorithm (Online Hybrid Dynamic Scheduling, OHDS), which aims to minimize the energy consumption of tasks and maximize service resource utilization while satisfying the sub-deadline and data dependency constraints of workflow tasks. Firstly, the data dependencies between workflow tasks are considered for multi-task merging, and sub-deadline constraints are assigned to workflow tasks based on task priority. Secondly, based on the independent nature of the tasks of different workflows, a hybrid scheduling of multiple workflows is performed to reduce service idle time. Then, the workflow task scheduling priority and its sub-deadlines are dynamically adjusted, and the service status is sensed by the CPU utilization of the service, and the workload on the overloaded/underloaded service is balanced by dynamic migration of virtual machines. Finally, the OHDS method is compared with three existing scheduling methods to verify its better performance in terms of scheduling energy consumption, scheduling success rate and service resource utilization.
Guisheng Fan, Xingpeng Chen, Huiqun Yu, Yingxue Zhang 0007
IEEE Trans. Netw. Serv. Manag.1
2023 Cost-Efficient Fault-Tolerant Workflow Scheduling for Deadline-Constrained Microservice-Based Applications in Clouds
abstract
Microservices are becoming increasingly popular in the construction of cloud applications. On the basis of containers, microservice instances can be implemented with high scalability and maintainability. Due to the need of ensuring various quality of service (QoS) requirements and the two-layer resource structure of containers and virtual machines (VMs), microservice workflow scheduling in clouds is a challenging problem to address. This paper proposes a heuristic algorithm GSMS to minimize execution cost of a microservice-based workflow application while satisfying deadline and reliability constraints. GSMS adopts a greedy fault-tolerant scheduling strategy for replicas of each task to select appropriate resources that meet the sub-deadline and minimize the cost until the sub-reliability is guaranteed. Furthermore, a resource adjustment strategy is incorporated into GSMS to further improve resource utilization. By conducting extensive experiments with several realistic workflow applications, in comparison with existing algorithms, the effectiveness and efficiency of GSMS in achieving lower execution cost and meeting deadline and reliability requirements are validated.
Huiqun Yu, Guisheng Fan, Jiayin Zhang
IEEE Trans. Netw. Serv. Manag.3
2023 Security-Aware and Time-Guaranteed Service Placement in Edge Clouds
abstract
Most of the emerging applications such as the Deep Neural Networks (DNN) based smart Internet of Things (IoT) systems need intensive and high-performance computing, which is contradictory to the limited resources of IoT/terminal devices. It is a big challenge to offload all tasks to the cloud due to the bandwidth limitation, processing overhead, and transmission costs. Edge computing as an extension of cloud computing that can provide abundant computing resources near the edge of the network and thereby can potentially improve the QoS of applications. However, offloading tasks to the edge servers is liable to external security threats. How to balance the response time and the security of application services is a big challenge for realizing good application service placement. This paper proposes a time and security efficient task scheduling framework in the edge-cloud environment. The corresponding computing models are established, such as the security-related model, time model, and the risk probability model. Then, a time-guaranteed and security-aware task scheduling algorithm is proposed including the domain construction, the security-aware task ranking, and the task dispatching. Extensive simulation experiments have been conducted. Results show that the proposed method has better performance than the other four compared methods in general.
Huaiying Sun, Huiqun Yu, Guisheng Fan, Liqiong Chen, Zheng Liu 0023
IEEE Trans. Netw. Serv. Manag.3
2023 Towards Retrieval-Based Neural Code Summarization: A Meta-Learning Approach
abstract
Code summarization aims to generate code summaries automatically, and has attracted a lot of research interest lately. Recent approaches to it commonly adopt neural machine translation techniques, which train a Seq2Seq model on a large corpus and assume it could work on various new code snippets. However, codes are highly varied in practice due to different domains, businesses or programming styles. Therefore, it is challenging to learn such a variety of patterns into a single model. In this paper, we propose a brand-new framework for code summarization based on meta-learning and code retrieval, named MLCS to tackle this issue. In this framework, the summarization of each target code is formalized as a few-shot learning task, where its similar examples are used as training data and the testing example is itself. We retrieve examples similar to the target code in a rank-and-filter manner. Given a neural code summarizer, we optimize it into a meta-learner via Model-Agnostic Meta-Learning (MAML). During inference, the meta-learner first adapts to the retrieved examples and yields an exclusive model for the target code, and then generates its summary. Extensive experiments on real-world datasets show: (1) Utilizing MLCS, a standard Seq2Seq model is able to outperform previous state-of-the-art approaches, including both neural models and retrieval-based neural models; (2) MLCS can flexibly adapt to existing neural code summarizers without modifying their architecture, and could significantly improve their performance with the relative gain of up to 112.7% on BLEU-4, 23.2% on ROUGE-L, and 31.5% on METEOR; (3) Compared to the existing retrieval-based neural approaches, MLCS can better leverage multiple similar examples, and shows better generalization ability on different retrievers, unseen retrieval corpus and low-frequency words.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004
IEEE Trans. Software Eng.3
2022 Dynamic Trust-Based Resource Allocation Mechanism for Secure Edge Computing
Huiqun Yu, Qifeng Tang, Zhiqing Shao, Yiming Yue, Guisheng Fan, Liqiong Chen
CollaborateCom (2)5
2022 Bug Report Priority Prediction Using Developer-Oriented Socio-Technical Features
abstract
Software stakeholders report bugs in Issue Tracking System (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time-insensitive approaches, and they were producing sub-optimal results for practical usage. Moreover, they tend to ignore the developer and social aspects of ITS. Since ITS bridges users and developers, we integrate their sentiment- and community-oriented socio-technical features to perform 2- and multi-classed bug priority prediction and validate our model in within-project, cross-project, and time-wise scenarios. The proposed model outperforms the 2 baselines by up to 10% in AUC-ROC and 13% in MCC, and the significance of improvement is statistically confirmed. We reveal involving assignee and reporter features from socio-technical perspectives such as sentiment could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in socio-technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest researchers should involve contributors’ experience and sentiments in bug report priority prediction.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Internetware3
2022 HQLgen: deep learning based HQL query generation from program context
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Jiayin Zhang
Autom. Softw. Eng.3
2022 Cold-start aware cloud-native service function chain caching in resource-constrained edge: A reinforcement learning approach
Jiayin Zhang, Huiqun Yu, Guisheng Fan
Comput. Commun.3
2022 Time-cost efficient memory configuration for serverless workflow applications
abstract
Summary Recently, workflow applications are increasingly migrated to Function‐as‐a‐Service platforms which are easy to manage, highly‐scalable, and pay‐as‐you‐go. Meanwhile, users face challenges in migration of serverless applications because of the lack of efficient algorithm for workflow memory configuration to optimize the performance. To this end, this article proposes a heuristic urgency‐based algorithm UWC and a meta‐heuristic hybrid algorithm BPSO to tackle the time‐cost tradeoff. UWC sorts functions and allocates each function an appropriate memory size by greedy strategy. BPSO hybridizes particle swarm optimization as well as beetle antennae search algorithm to guide particles to search directionally and utilizes nonlinear inertia weight to avoid local premature convergence. Extensive experiments with classical serverless application demonstrate that UWC and BPSO are very competitive in comparison with existing algorithms as they can find the optimal workflow memory configuration.
Huiqun Yu, Guisheng Fan
Concurr. Comput. Pract. Exp.3
2022 Time-constrained and reliability aware energy minimization scheduling algorithm for heterogeneous multiprocessor environments
abstract
Summary Heterogeneous multiprocessor systems are now widely used in industry by providing high performance and high concurrency. However, with the increasing number of computational nodes, it leads to a dramatic increase in energy consumption of heterogeneous multiprocessor systems. Most of the current research has been conducted to reduce the energy consumption by reducing the processor frequency and extending the task execution time, but these measures often lead to a significant decrease in system reliability. This article addresses the problem of energy‐aware task scheduling in the case of distributed computing systems with deadline and reliability constraints. First, a time and reliability allocation model is built to assign the deadline and reliability constraints to each task, which ensures fairness among tasks. Then, a three‐stage scheduling algorithm is proposed to minimize the system energy consumption. The static energy consumption is minimized by shutting down the inefficient processors. The dynamic energy consumption of the system is reduced by distributing tasks as evenly as possible on each processor through a task redistribution strategy. Finally, experimental results on real scientific workflows and randomly generated graphs show that the proposed algorithm outperforms other algorithms in terms of energy reduction.
Zhaorui Wang 0004, Guisheng Fan, Huiqun Yu
Concurr. Comput. Pract. Exp.2
2022 Modelling and analysing the reliability for microservice-based cloud application based on predicate Petri net
abstract
Abstract Microservice design is a new paradigm of cloud application development. Different from monolithic design, microservice enjoys merits of fine‐grained and loosely coupled services, and it is becoming more and more popular. The application developed with microservice has a good advantage in independent development and flexible deployment, especially for complex distributed systems. However, there is a big gap between the reliability requirements and microservice‐based cloud applications. This article proposes a reliability model of microservice‐based cloud application by using predicate Petri net. First, a microservice reliability requirement is given, some basic concepts of predicate Petri net are defined with syntax and semantics. Second, a microservice reliability strategy is proposed, which uses microservice instances and circuit breaker to improve the reliability of the system. Based on the constructed microservice reliability model, the correctness of predicate Petri net modelling and the effectiveness of the strategies are proven theoretically. Finally, an example is given to illustrate the establishment and analysis process of the model, and several groups of experiments are carried out to verify the effectiveness and feasibility of the method. Experimental results show that the proposed microservice reliability strategy is effective.
Zheng Liu 0023, Guisheng Fan, Huiqun Yu, Liqiong Chen
Expert Syst. J. Knowl. Eng.2
2022 Reliability modelling and optimization for microservice-based cloud application using multi-agent system
abstract
Abstract In the process of the continuous development of the Internet of Things, cloud computing has been applied in many fields, how to guarantee the quality of service, such as low latency, high bandwidth, high reliability etc., has become a challenging problem. This paper proposes a method to model and optimize reliability for microservice‐based cloud applications using multi‐agent system (MAS), thus maximizing the reliability of cloud computing and dynamically scheduling microservices to minimize the delay within the budget. Firstly, a dynamic microservice scheduling scheme is proposed to provide efficient computing services by using MAS. A hierarchical cloud computing model is formed by predicated Petri net (PrT net) and the properties of constructed model are analysed. Secondly, agents have been utilized to describe the essential characteristics of microservice scheduling process in the cloud applications. The partial critical path (PCP) aims to maximize the reliability of cloud applications under the limitation of budget and meet the user‐defined deadline. Finally, the proposed PCPRO algorithm has been applied to cloud environment, which is suitable for different scientific workflows in the cloud computing environment. The effectiveness of this method is verified by simulation, the experiment results show the effectiveness of the proposed method.
Zheng Liu 0023, Huiqun Yu, Guisheng Fan, Liqiong Chen
IET Commun.3
2022 Automatic Identification of High-Impact Bug Report by Product and Test Code Quality
abstract
Bug reports are submitted by the software stakeholders to foster the location and elimination of bugs. However, in large-scale software systems, it may be impossible to track and solve every bug, and thus developers should pay more attention to High-Impact Bugs (HIBs). Previous studies analyzed textual descriptions to automatically identify HIBs, but they ignored the quality of code, which may also indicate the cause of HIBs. To address this issue, we integrate the features reflecting the quality of production (i.e. CK metrics) and test code (i.e. test smells) into our textual similarity based model to identify HIBs. Our model outperforms the compared baseline by up to 39% in terms of AUC-ROC and 64% in terms of F-Measure. Then, we explain the behavior of our model by using SHAP to calculate the importance of each feature, and we apply case studies to empirically demonstrate the relationship between the most important features and HIB. The results show that several test smells (e.g. Assertion Roulette, Conditional Test Logic, Duplicate Assert, Sleepy Test) and product metrics (e.g. NOC, LCC, PF, and ProF) have important contributions to HIB identification.
Jianshu Ding, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
Int. J. Softw. Eng. Knowl. Eng.2
2022 Improving Just-In-Time Comment Updating via AST Edit Sequence
abstract
Code comments are valuable for program comprehension and software maintenance. However, comments can be inconsistent or out-of-date after code changes. To tackle this problem, Just-In-Time (JIT) comment updating aims to automatically update comments with code changes. Existing approaches for this task use edit sequences of source code to model code changes. Meanwhile, recent researches indicate that neural models based on abstract syntax trees (AST) can help represent source code. In this paper, we propose a new method to learn code changes by combining code edit sequences with AST edit sequences, so that the generated new comment can be more accurate. Our approach utilizes three encoders to encode code edit sequences, AST edit sequences and old comment token sequences, respectively. The outputs of the encoders are then decoded into a sequence of edit actions, which is parsed to generate a new comment. The proposed method is evaluated on a public dataset using seven metrics, and the experimental results show that our approach outperforms the baselines. Furthermore, when the new comment has a larger edit distance than the old one, our model shows better performance.
Huiqun Yu, Guisheng Fan, Ziyi Zhou 0002
Int. J. Softw. Eng. Knowl. Eng.3
2022 Code Generation with Hybrid of Structural and Semantic Features Retrieval
abstract
Due to the growing need for faster software delivery, code generation has attracted more and more attention, since it could improve code maintainability by providing suggestions for coding. In the model of generating program source code from natural language (NL), the most effective method is to generate an intermediate architecture (such as Abstract Syntax Tree) combined with a deep learning model. However, these models have the following drawbacks: (1) The data structural information is underutilized and the correlation between samples is not considered. (2) Lack of the ability to memorize large and complex structures, so that complex codes cannot be generated correctly. To address these issues, we propose HRCODE model, a code generation architecture based on Hybrid of structural and semantic features Retrieval CODE model. We transform the NL description into an intermediate structure with structural features. Then, the NL and the intermediate structure are embedded into a vector through weight mixing, and we calculate the similarity score between each vector to retrieve the most relevant samples. Finally, the new input is brought into the PLBART model to generate code. Experiments show that HRCODE is at least 4.7% higher than the state-of-the-art models in the ACC metric and at least 10.3% higher in the BLEU-4 score. We have released our code at https://github.com/jesokang/HRCODE.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Ziyi Zhou 0002
Int. J. Softw. Eng. Knowl. Eng.3
2022 Summarizing source code with hierarchical code representation
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Xingguang Yang
Inf. Softw. Technol.3
2022 Community Smell Occurrence Prediction on Multi-Granularity by Developer-Oriented Features and Process Metrics
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Xingguang Yang, Kang Yang 0004
J. Comput. Sci. Technol.3
2022 HBSniff: A static analysis tool for Java Hibernate object-relational mapping code smell detection
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Sci. Comput. Program.3
2022 A graph sequence neural architecture for code completion with semantic structure features
abstract
Abstract Code completion plays an important role in intelligent software development for accelerating coding efficiency. Recently, the prediction models based on deep learning have achieved good performance in code completion task. However, the existing models cannot avoid three drawbacks: (i) In the existing models, the code representation loses the information (parent–child information between nodes) and lacks many effective features (orientation between nodes). (ii) The known code structure information is not fully utilized, which will cause the model to generate completely irrelevant results. (iii) Simple sequence modeling ignores repeated patterns and structural information. Besides, previous works cannot capture the characteristics of correlation and directionality between nodes. In this paper, we propose a Code Completion approach named CC‐GGNN, which is graph model based on Gated Graph Neural Networks (GGNNs) to address the problems. We introduce a new architecture to obtain the effective code features from code representation. In order to utilize the known information, we propose Classification Mechanism, which classifies the representation of the node using the known parent node and constructs training graph in the model. The experimental results show that our model outperforms the state‐of‐the‐art methods MRR@5 at most 9.2% and ACC at most 11.4% in datasets.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Zijie Huang 0001
J. Softw. Evol. Process.3
2021 Dual-Channel Graph Contextual Self-Attention Network for Session-Based Recommendation
Teng Huang 0004, Huiqun Yu, Guisheng Fan
CollaborateCom (1)3
2021 An Empirical Study of Model-Agnostic Interpretation Technique for Just-in-Time Software Defect Prediction
Xingguang Yang, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Ziyi Zhou 0002
CollaborateCom (1)3
2021 A novel software defect prediction method based on hierarchical neural network
abstract
To ensure software reliability, software defect prediction (SDP) techniques are employed to help developers effectively allocate the testing resources. Recently, researchers utilized deep learning models to extract semantic features from abstract syntax tree (AST) of source code which showed a better prediction performance over metric-based methods. However, the existing file-level SDP models representing the AST as a flattened sequence could jeopardize the preservation of long-term dependency. In this paper, we propose a new Defect Prediction framework based on the Hierarchical Neural Network (DP-HNN). Our method makes use of the hierarchical structure of AST by splitting the large file-level AST into several subtrees according to certain AST nodes crucial to SDP task. These subtrees represented by node-level sequences are encoded separately and then serve as the elements of the subtree-level sequence. Finally, a multi-granularity fusion approach is performed in the subtree-level encoder to obtain the crucial features that represent the code file. Our proposed DP-HNN is aimed at capturing long-term dependency while preserving fine-grained local information. We conducted experiments on 11 open-source projects considering the cross-version and the mixed-version scenario of within-project SDP. Results show that on average, DP-HNN improves the state-of-the-art method by 14% and 3% on MCC and AUC scores respectively.
Huiqun Yu, Xingjie Sun, Ziyi Zhou 0002, Guisheng Fan
COMPSAC4
2021 Predicting Community Smells' Occurrence on Individual Developers by Sentiments
abstract
Community smells appear in sub-optimal software development community structures, causing unforeseen additional project costs, e.g., lower productivity and more technical debt. Previous studies analyzed and predicted community smells in the granularity of community sub-groups using socio-technical factors. However, refactoring such smells requires the effort of developers individually. To eliminate them, supportive measures for every developer should be constructed according to their motifs and working states. Recent work revealed developers' personalities could influence community smells' variation, and their sentiments could impact productivity. Thus, sentiments could be evaluated to predict community smells' occurrence on them. To this aim, this paper builds a developer-oriented and sentiment-aware community smell prediction model considering 3 smells such as Organizational Silo, Lone Wolf, and Bottleneck. Furthermore, it also predicts if a developer quitted the community after being affected by any smell. The proposed model achieves cross- and within-project prediction F-Measure ranging from 76% to 93%. Research also reveals 6 sentimental features having stronger predictive power compared with activeness metrics. Imperative and indicative expressions, politeness, and several emotions are the most powerful predictors. Finally, we test statistically the mean and distribution of sentimental features. Based on our findings, we suggest developers should communicate in a straightforward and polite way.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Ziyi Zhou 0002, Kang Yang 0004, Xingguang Yang
ICPC3
2021 Efficiency-First Fault-Tolerant Replica Scheduling Strategy for Reliability Constrained Cloud Application
Yingxue Zhang 0007, Guisheng Fan, Huiqun Yu, Xingpeng Chen
NPC2
2021 Automatic Identification of High Impact Bug Report by Test Smells of Textual Similar Bug Reports
abstract
Bug reports are written by the software stakeholders to track software defects and vulnerabilities. Since Software Quality Assurance (SQA) resources are limited, developers tend to resolve High-Impact Bugs (HIB) in advance. Prior research identified HIBs by analyzing the textual information in bug reports. However, they only consider textual information instead of the root cause of bugs, such as code quality. Since prior study revealed software test smells (i.e., sub-optimal test code implementation) are related to bug proneness, we intend to measure test smell distribution in textual similar bug reports to identify HIB reports. We first construct an effective model, which outperforms the baseline by 29.3% in terms of AUC-ROC. Secondly, we use SHAP to compute the importance of test smell features. Finally, we conduct an empirical survey to discuss the relationship between test smell and HIB reports. Result shows that Assertion Roulette and Conditional Test Logic test smell are important factors in distinguishing the types of bug reports.
Jianshu Ding, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
QRS2
2021 An Empirical Study on the Impact of Class Overlapin Just-in-Time Software Defect Prediction (S)
abstract
Just-in-time software defect prediction (JIT-SDP) is an active research topic in the field of software engineering, aiming at identifying defect-inducing code changes.Most of the current JIT-SDP work focused on model construction.It is often ignored that the performance of classifiers often depends on high quality data.In this paper, we first investigate the impact of the class overlap problem on the performance of the classifiers in JIT-SDP, and propose a new effective preprocessing method (IKMCCA-TL) combining improved K-Means clustering cleaning approach and Tomek-link method.In order to objectively estimate the impact of class overlap on the classifiers in JIT-SDP, we conduct a large-scale empirical study on the data sets of six open source projects and compare the performance of LR, RF and KNN classifiers by using IKMCCA or KMCCA or NCL and without cleaning data.Experimental results show that after removing overlapping instances, the performance of the classifiers is significantly improved in terms of balance, recall and AUC and our proposed method achieves the best performance.
Minyang Yi, Guisheng Fan, Huiqun Yu, Xingguang Yang
SEKE2
2021 DEJIT: A Differential Evolution Algorithm for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Software defect prediction is an effective approach to save testing resources and improve software quality, which is widely studied in the field of software engineering. The effort-aware just-in-time software defect prediction (JIT-SDP) aims to identify defective software changes in limited software testing resources. Although many methods have been proposed to solve the JIT-SDP, the effort-aware prediction performance of the existing models still needs to be further improved. To this end, we propose a differential evolution (DE) based supervised method DEJIT to build JIT-SDP models. Specifically, first we propose a metric called density-percentile-average (DPA), which is used as optimization objective on the training set. Then, we use logistic regression (LR) to build a prediction model. To make the LR obtain the maximum DPA on the training set, we use the DE algorithm to determine the coefficients of the LR. The experiment uses defect data sets from six open source projects. We compare the proposed method with state-of-the-art four supervised models and four unsupervised models in cross-validation, cross-project-validation and timewise-cross-validation scenarios. The empirical results demonstrate that the DEJIT method can significantly improve the effort-aware prediction performance in the three evaluation scenarios. Therefore, the DEJIT method is promising for the effort-aware JIT-SDP.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
Int. J. Softw. Eng. Knowl. Eng.3
2021 Adversarial training and ensemble learning for automatic code summarization
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan
Neural Comput. Appl.3
2021 An Approach to Modeling and Analyzing Reliability for Microservice-Oriented Cloud Applications
abstract
Microservice architecture is a cloud‐native architectural style, which has attracted extensive attention from the scientific research and industry communities to benefit independent development and deployment. However, due to the complexity of cloud‐based platforms, the design of fault‐tolerant strategies for microservice‐oriented cloud applications becomes challenging. In order to improve the quality of service, it is essential to focus on the microservice with more criticality and maximize the reliability of the entire cloud application. This paper studies the modeling and analysis of service reliability in the cloud environment. Firstly, a formal description language is defined to model microservice, user request, and container accurately. Secondly, the reliability analysis is conducted to measure a critical microservice’s fluctuation and vibration attributes within a period, and the related properties of the constructed model are analyzed. Thirdly, a fault‐tolerant strategy with redundancy operation has been proposed to optimize cloud application reliability. Finally, the effectiveness of the method is verified by experiments. The simulation results show that the algorithm obtains the maximum benefits and has high performance through several experiments.
Zheng Liu 0023, Guisheng Fan, Huiqun Yu, Liqiong Chen
Wirel. Commun. Mob. Comput.2
2020 WSN Coverage Optimization Based on Two-Stage PSO
Huiqun Yu, Guisheng Fan, Xinxiu Wen
CollaborateCom (1)3
2020 Code Prediction Based on Graph Embedding Model
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Liqiong Chen
CollaborateCom (2)3
2020 EFMLP: A Novel Model for Web Service QoS Prediction
Kailing Ye, Huiqun Yu, Guisheng Fan, Liqiong Chen
CollaborateCom (2)3
2020 Location-Based Service Recommendation for Cold-Start in Mobile Edge Computing
Mengshan Yu, Guisheng Fan, Huiqun Yu
NPC2
2020 Energy and time efficient task offloading and resource allocation on the generic IoT-fog-cloud architecture
Huaiying Sun, Huiqun Yu, Guisheng Fan, Liqiong Chen
Peer-to-Peer Netw. Appl.3
2020 Effective approaches to combining lexical and syntactical information for code summarization
abstract
Summary Natural language summaries of source codes are important during software development and maintenance. Recently, deep learning based models have achieved good performance on the task of automatic code summarization, which encode token sequence or abstract syntax tree (AST) of code with neural networks. However, there has been little work on the efficient combination of lexical and syntactical information of code for better summarization quality. In this paper, we propose two general and effective approaches to leveraging both types of information: a convolutional neural network that aims to better extract vector representation of AST node for downstream models; and a Switch Network that learns an adaptive weight vector to combine different code representations for summary generation. We integrate these approaches into a comprehensive code summarization model, which includes a sequential encoder for token sequence of code and a tree based encoder for its AST. We evaluate our model on a large Java dataset. The experimental results show that our model outperforms several state‐of‐the‐art models on various metrics, and the proposed approaches contribute a lot to the improvements.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan
Softw. Pract. Exp.3
2020 Contract-Based Resource Sharing for Time Effective Task Scheduling in Fog-Cloud Environment
abstract
Fog computing as an extension of the cloud based infrastructure, provides a better computing platform than cloud computing for mobile computing, Internet of Things, etc. One of the problems is how to make full use of the resources of the fog so that more requests of applications can be executed on the edge, reducing the pressure on the network and ensuring the time requirement of tasks. The high mobility of fog nodes also has a great impact on the task completion time and user satisfaction. Thus, a general IoT-Fog-Cloud computing architecture with a contract-based resource sharing mechanism is proposed in this paper. The contract establishment problem of resource sharing mechanism among fog clusters is modeled as a sealed-bid bilateral auction in order to take full advantage of the fog resources and ensure that more tasks could be executed on the fog. Then, we propose a scheduling method based on functional domain construction to mitigate the influence of mobility of fog nodes. It includes the selection of critical fog nodes and the construction of fog function domains based on spectral clustering. The selection of critical fog nodes is used to find the best fog nodes in each fog cluster with respect to the betweenness centrality, computing performance and communication delay to the IoT nodes. The critical nodes are responsible for building the functional domains of the remaining fog nodes in each fog cluster. Functional domain construction is used to determine the set of fog nodes contained in the corresponding functional domain. Finally, through extensive simulation experiments, the performance difference between the proposed method and the other four methods in terms of average service time, average utilization of fog nodes, success rate of tasks, average WLAN delay and the average cost of successful tasks are evaluated. Results show that our method generally outperforms the other four methods in these metrics.
Huaiying Sun, Huiqun Yu, Guisheng Fan
IEEE Trans. Netw. Serv. Manag.3
2020 Modeling and Analyzing Dynamic Fault-Tolerant Strategy for Deadline Constrained Task Scheduling in Cloud Computing
abstract
Cloud computing has been increasingly concerned in scientific computing area. More and more enterprises and research institutes have migrated their applications to the clouds. Due to the complexity of cloud computing system in structural and behavioral aspects, how to design the fault tolerant cloud computing system becomes a challenging problem. This paper investigates the modeling and analysis of fault tolerant strategy for deadline constrained task scheduling in cloud computing. First, a formal description language is defined to accurately model the different components of cloud application, and use it to characterize the operational mechanisms and fault behaviors. Second, we propose a fault tolerant strategy, which includes the scheduling mechanism, synchronization mechanism, and exception mechanism, to dynamically compute the execution mode and required virtual machine for tasks, thus ensuring the reliability and real-time requirement of cloud application. An enforcement algorithm is also designed to realize the proposed strategy. Third, the techniques of Petri nets are provided to analyze and validate the correctness of proposed method. Finally, several experiments are done to illustrate that the reliability of cloud application is improved and its deadline is met.
Guisheng Fan, Liqiong Chen, Huiqun Yu
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Deep Semantic Feature Learning with Embedded Static Metrics for Software Defect Prediction
abstract
Software defect prediction, which locates defective code snippets, can assist developers in finding potential bugs and assigning their testing efforts. Traditional defect prediction features are static code metrics, which only contain statistic information of programs and fail to capture semantics in programs, leading to the degradation of defect prediction performance. To take full advantage of the semantics and static metrics of programs, we propose a framework called Defect Prediction via Attention Mechanism (DP-AM) in this paper. Specifically, DPAM first extracts vectors which are then encoded as digital vectors by mapping and word embedding from abstract syntax trees (ASTs) of programs. Then it feeds these numerical vectors into Recurrent Neural Network to automatically learn semantic features of programs. After that, it applies self-attention mechanism to further build relationship among these features. Furthermore, it employs global attention mechanism to generate significant features among them. Finally, we combine these semantic features with traditional static metrics for accurate software defect prediction. We evaluate our method in terms of F1-measure on seven open-source Java projects in Apache. Our experimental results show that DP-AM improves F1-measure by 11% in average, compared with the state-of-the-art methods.
Guisheng Fan, Xuyang Diao, Huiqun Yu, Kang Yang 0004, Liqiong Chen
APSEC1
2019 Energy-Aware Resource Scheduling with Fault-Tolerance in Edge Computing
Yanfen Xue, Guisheng Fan, Huiqun Yu, Huaiying Sun
NPC2
2019 An Empirical Studies on Optimal Solutions Selection Strategies for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Just-in-time software defect prediction (JIT-SDP) is an active topic in the filed of software engineering, and many methods have been proposed to solve this problem.Stateof-the-art method MULTI applies multi-objective optimization algorithm to the effort-aware JIT-SDP problem, and obtains good average performance.Although the average performance of the MULTI method is high, there are many optimal solutions with poor performance.If an optimal solution is randomly selected, a poor prediction model may be obtained.In order to further improve the performance of the MULTI method, we propose three optimal solutions selection strategies: benefit priority (BP), cost priority (CP), and a compromise between cost and benefit (CCB).In order to compare and validate the effectiveness of the strategies, we conduct a large-scale empirical study on data sets of six open source projects.The experimental results show that, compared with the average performance of MULTI, the optimal solutions selection strategy based on BP has a significant improvement in ACC and Popt indicators.Therefore, we recommend using the BP-based optimal solutions selection strategy to improve the performance of MULTI when using the MULTI method to solve the effort-aware JIT-SDP problem.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
SEKE3
2019 Mutation with Local Searching and Elite Inheritance Mechanism in Multi-Objective Optimization Algorithm: A Case Study in Software Product Line
abstract
An effective method for addressing the configuration optimization problem (COP) in Software Product Lines (SPLs) is to deploy a multi-objective evolutionary algorithm, for example, the state-of-the-art SATIBEA. In this paper, an improved hybrid algorithm, called SATIBEA-LSSF, is proposed to further improve the algorithm performance of SATIBEA, which is composed of a multi-children generating strategy, an enhanced mutation strategy with local searching and an elite inheritance mechanism. Empirical results on the same case studies demonstrate that our algorithm significantly outperforms the state-of-the-art for four out of five SPLs on a quality Hypervolume indicator and the convergence speed. To verify the effectiveness and robustness of our algorithm, the parameter sensitivity analysis is discussed and three observations are reported in detail.
Kai Shi 0006, Huiqun Yu, Guisheng Fan, Jianmei Guo, Liqiong Chen, Xingguang Yang, Huaiying Sun
Int. J. Softw. Eng. Knowl. Eng.3
2019 A Parallel Framework of Combining Satisfiability Modulo Theory with Indicator-Based Evolutionary Algorithm for Configuring Large and Real Software Product Lines
abstract
Multi-objective evolutionary algorithm (MOEA) has been widely applied to software product lines (SPLs) for addressing the configuration optimization problems. For example, the state-of-the-art SMTIBEA algorithm extends the constraint expressiveness and supports richer constraints to better address these problems. However, it just works better than the competitor for four out of five SPLs in five objectives and the convergence speed is not significantly increased for largest Linux SPL from 5 to 30[Formula: see text]min. To further improve the optimization efficiency, we propose a parallel framework SMTPORT, which combines four corresponding SMTIBEA variants and performs these variants by utilizing parallelization techniques within the limited time budget. For case studies in LVAT repository, we conduct a series of experiments on seven real-world and highly-constrained SPLs. Empirical results demonstrate that our approach significantly outperforms the state-of-the-art for all the seven SPLs in terms of a quality Hypervolume metric and a diversity Pareto Front Size indicator.
Kai Shi 0006, Huiqun Yu, Jianmei Guo, Guisheng Fan, Liqiong Chen, Xingguang Yang
Int. J. Softw. Eng. Knowl. Eng.4
2018 A Load-Balanced Approach to Time Efficient Resource Scheduling in SDN-Enabled Data Center
abstract
Nowadays it is common for applications to run on data centers and deliver services to users. With the increase of tasks of multiple applications, it is a challenge for data center providers to make full use of the available resources, and improve task response time without too much computational cost. This paper focuses on load-balance based time efficient resource scheduling. A resource allocation architecture for SDN-enabled data center and a load-balance based resource allocation approach(LBA) are proposed. LBA is mainly used to maintain the the whole resource in a balancing state and assign appropriate resources to tasks, majorly consisting of three parts: Load-balance, VM-selection and Path-selection. Comprehensive simulation experiments are conducted to evaluate the effectiveness of LBA. Experiment results show that LBA can take full advantage of the available resources and improve task response time on the basis of load-balance, making both SLA violation rate and average cost as small as possible.
Huaiying Sun, Huiqun Yu, Guisheng Fan, Liqiong Chen
COMPSAC (2)3
2018 An Efficient Approach to Forecasting Monthly Calls for Repair from Gas Consumers
abstract
Forecasting monthly calls for repair from gas consumers is an important part of the gas company to improve the level of service, optimize the allocation of resources and improve the living level of people. In this paper, through the study of historical data of monthly calls for repair from gas consumers, we find that it has the characteristics of seasonal periodic variation. A hybridization methodology based on Seasonal Autoregressive Integrated Moving Average (SARIMA) and back propagation(BP) neural network is proposed, which is used to forecast monthly calls for repair from gas consumers. The time series of monthly calls for repair from gas consumers is decomposed into linear autocorrelation and non-linear structure of two parts. The SARIMA model is used to predict the linear part of the sequence, and the BP neural network model is used to predict the non-linear residual part. Finally, the forecast results of two parts are synthesized into the final result. The case study shows that the hybrid model outperforms either of the models used separately. Moreover, the hybrid model can balance the deviation of a single model with better applicability and higher accuracy.
Huiqun Yu, Cunbin Deng, Guisheng Fan, Liqiong Chen, Huaiying Sun
COMPSAC (2)3
2018 Combining Constraint Solving with Different MOEAs for Configuring Large Software Product Lines: A Case Study
abstract
Multi-objective evolutionary algorithm (MOEA) with the constraint solving has been successfully applied to address the configuration optimization problem in software product line (SPL), for example, the state-of-the-art SATIBEA algorithm. However, each different MOEA with special search operator demonstrates the different strength and weakness in terms of optimality and convergence speed. The SATIBEA just combines the SAT (Boolean satisfiability problem) constraint solving with the Indicator-Based Evolutionary Algorithm (IBEA) for evaluating the algorithm performance. In this paper, we propose six hybrid algorithms which combine the SAT solving with different MOEAs. Case study is based on five large-scale, rich-constrained and real-world SPLs. Empirical results demonstrate that SATMOCell algorithm obtains a competitive optimization performance to the state-of-the-art that outperforms the SATIBEA in terms of quality Hypervolume metric for 2 out of 5 SPLs within the same time budget. Moreover, the convergence speed of SATMOCell and SATssNSGA2 is comparable after 10min terminal times. Particularly, the Hypervolume value of SATssNSGA2 reports the average improvement of 1.33% after 20min terminal times.
Huiqun Yu, Kai Shi 0006, Jianmei Guo, Guisheng Fan, Xingguang Yang, Liqiong Chen
COMPSAC (1)4
2018 imBBO: An Improved Biogeography-Based Optimization Algorithm
Kai Shi 0006, Huiqun Yu, Guisheng Fan, Xingguang Yang
GPC3
2018 Formally modeling and analyzing cost-aware job scheduling for cloud data center
abstract
Summary With the rapid development of cloud computing, many distributed data centers have been deployed. This means larger energy consumption requirements from the data center. How to reduce the cost of data center has received significant attention recently. Although there are several efforts in studying energy consumption of the data center, very few have considered modeling and analyzing cost‐aware job scheduling for the cloud data center. To address this emerging problem, we propose a systematic approach that considers both basic elements and their relationships in cloud data center. First, we present a formal language to describe the cloud data center, and a job scheduling net is proposed to formally model the basic elements such as user request, Web portal, data center, and server. Second, we minimize the total cost of the cloud data center by considering the multidimensional resource and local electricity price on the basis of the state space of constructed model. The dynamic job scheduling algorithm and its specific execution steps are proposed based on the alternating direction method of multipliers algorithm. Third, the operational semantics and related theories of Petri nets for establishing the correctness of our proposed method are presented. Finally, a series of simulations are performed to illustrate that the proposed method can guarantee the correct behavior of job scheduling in the cloud data center while meeting the required cost.
Guisheng Fan, Liqiong Chen, Huiqun Yu
Softw. Pract. Exp.1
2018 A parallel portfolio approach to configuration optimization for large software product lines
abstract
Summary Software product line (SPL) engineering demands for optimal or near‐optimal products that balance multiple often competing and conflicting objectives. A major challenge for large SPLs is to efficiently explore a huge space of various products and satisfy a large number of predefined constraints simultaneously. To improve the optimality and convergence speed, we propose a parallel portfolio approach, called IBEAPORT, which designs three algorithm variants by incorporating constraint solving into the indicator‐based evolutionary algorithm in different ways and performs these variants by utilizing parallelization techniques. Our approach utilizes the exploration capabilities of different algorithms and improves optimality as far as possible within a limited time budget. We evaluate our approach on five large‐scale real‐world SPLs. Empirical results demonstrate that our approach significantly outperforms the state of the art for all five SPLs on a quality indicator and a diversity indicator. Moreover, IBEAPORT quickly converges to a relatively stable hypervolume value even for the largest SPL with 6888 features.
Kai Shi 0006, Huiqun Yu, Jianmei Guo, Guisheng Fan, Xingguang Yang
Softw. Pract. Exp.4
2017 Hierarchy attribute-based encryption scheme to support direct revocation in cloud storage
abstract
Attribute-Based Encryption Scheme solves the security problems faced by cloud storage. In many practical applications, the files encrypted by Attribute-Based Encryption have the characteristics of multiple levels. However, CPABE cannot encrypt the data according to the level. This paper proposes a hierarchy attribute-based encryption algorithm to support the direct revocation. First, the encryption process of algorithm is introduced, second, the model of hierarchy attribute-based encryption algorithm is proposed to support the direct revocation, and the security proof of the algorithm is also given. Finally, we design two experiments to compare the proposed algorithm with existing algorithms from the perspective of time.
Shuci Jiang, Weibin Guo, Guisheng Fan
ICIS3
2017 Mini-XML: An efficient mapping approach between XML and relational database
abstract
In recent years, XML technology has won wide attention from both industry and academic. It can be used to mark data, define the data type and their own markup language. It is a cross-platform, context-dependent technology in the Internet environment and an effective tool for todays distributed structure information. The S-XML is a new approach for storing semi-structured data, and it supports query of the node in XML with SQL statements, which has shown impressive performance on many classic data sets. However, it is difficult to store XML data into a relational database, and the S-XML spends much more time and space to store the data. In this paper, we propose an efficient mapping approach, the mini-XML, to mapping XML into the relational database. In addition, path technique and position information are used to indicate the complex node relationship. Finally, two experiments are conducted to prove that the proposed method can achieve better performance in the decreasing of the storage time and storage space, especially dealing with the large amount of data.
Huchao Zhu, Huiqun Yu, Guisheng Fan, Huaiying Sun
ICIS3
2017 A Regression Test Case Prioritization Algorithm Based on Program Changes and Method Invocation Relationship
abstract
Regression testing is essential for assuring the quality of a software product. Because rerunning all test cases in regression testing may be impractical under limited resources, test case prioritization is a feasible solution to optimize regression testing by reordering test cases for the current testing version. In this paper, we propose a new test case prioritization algorithm based on program changes and method (function) invocation relationship. Combining the estimated risk value of each program method (function) and the method (function) coverage information, the fault detection capability of each test case can be calculated. The algorithm reduces the prioritization problem to an integer linear programming (ILP) problem, and finally prioritizes test cases according to their fault detection capabilities. Experiments are conducted on 11 programs to validate the effectiveness of our proposed algorithm. Experimental results show that our approach is more effective than some well studied test case prioritization techniques in terms of average percentage of fault detected (APFD) values.
Huiqun Yu, Guisheng Fan, Xiang Ji 0002, Xin Pei
APSEC3
2017 A game theoretic method to model and analyze attack-defense strategy of resource service in cloud application
abstract
Summary Cloud computing has attracted much attention recently in both industry field and academic research area. More and more Internet applications are moving to the cloud environment. However, it is difficult to construct perfectly secure mechanisms facing up with complex and various attacks in cloud computing, the efficient attack‐defense strategy is highly demanded. In this paper, a stochastic game model is proposed based on combining stochastic Petri nets with game theory, which is used to describe the attack‐defense behaviors in cloud computing. The physical machine, attack‐defense behavior, and their attributes are also modeled by stochastic game model thus forming the attack‐defense game model of cloud computing. On this basis, the Nash equilibrium of attack‐defense process in physical machine is computed to get the optimal defense strategy. The related theories of Petri nets and the reachable states of attack‐defense game model are used to formally verify the correctness and effectiveness of the proposed method. The enforcement algorithm is proposed to make cloud computing dynamically evaluate and select the defense strategy to against attack behavior as quickly as possible. Both case study and simulation results show that the proposed method can adapt quickly to the changes in cloud application thus improving the security of cloud computing.
Guisheng Fan, Liqiong Chen, Huiqun Yu
Concurr. Comput. Pract. Exp.1
2016 Modeling and Analyzing Cost and Utilization Based Task Scheduling for Cloud Application
abstract
Cloud computing has attracted much interest recently from both industry and academic. However, it is difficult to model and analyze cost and utilization based task scheduling due to complex and heterogeneous environment in cloud computing. In this paper, we propose the systematic approach to modeling and analyzing cost and utilization based task scheduling for cloud application. A Game theory based task scheduling scheme is proposed based on analyzing the factors affecting the cost and utilization of virtual machine. Then a formal description language is presented to model the different components of cloud application. The backward induction game algorithm is proposed to ensure that cloud application can dynamically meet the customers' request while improving the utilization of virtual machine. The operational semantics and related theories of Petri nets help establish the correctness of our proposed method. Finally, a series of simulations are performed to evaluate the efficiency of our proposed approach.
Guisheng Fan, Huiqun Yu, Liqiong Chen
COMPSAC1
2016 Modeling and analyzing cost-aware fault tolerant strategy for cloud application
abstract
In this paper, we propose a method to model and analyze cost-aware fault tolerant strategy for cloud computing.First, Petri nets are used to describe the structure of cloud computing, including component, cloud service and cloud application, thus forming the fault tolerant model of cloud computing.Second, a dynamic fault tolerant strategy is proposed, which can dynamically make fault tolerant strategy with the lowest cost based on the current state and failed component.Third, we present operational semantics and related theories of Petri nets for establishing the correctness of our proposed method.We have also performed a series of simulations to evaluate our proposed approach.Results show that it can help reveal the structural and behavioral characteristics of cloud computing, and reduce the fault tolerant cost.
Liqiong Chen, Guisheng Fan
SEKE2
2016 Multi-Objective Biogeography-Based Method to Optimize Virtual Machine Consolidation
abstract
Virtual machine consolidation (VMC) is an important issue in cloud computing, which can be used to reduce power consumption and achieve reasonable resource allocation.In this paper, an IMBBO algorithm is proposed to solve the multi-objective optimization problem of VMC through improving the classical Biogeography-Based Optimization (BBO).An improved Cosine migration model and an improved mutation model are presented to increase the efficiency of achieving the optimal solution.Meanwhile, three optimization objectives for server power consumption, load balancing and migration resource overhead are mainly addressed.Finally, several experiments are done to evaluate the performance of IMBBO by comparing with Gravitational Search Algorithm (GSA) based on the synthetic and real VM running data.The results show that the IMBBO optimizes VM consolidation with higher efficiency.
Kai Shi 0006, Huiqun Yu, Fei Luo 0002, Guisheng Fan
SEKE4
2016 Attack-defense trees based cyber security analysis for CPSs
abstract
Cyber-physical system (CPS) is the fuse of cyber world and the dynamic physical world and it is being widely used in areas closely related to people's livelihood. Therefore, the security issues of CPS have drawn a global attention and an appropriate risk assessment for CPS is in urgent need. The existing proposals using attack trees for risk assessment mainly focus on depicting the possible intrusions, not for interactions between threats and defenses. In this paper, a risk assessment idea for cyber-physical system with the use of attack-defense tree (ADTree) is proposed, considering the effect of both the attack cost and defense cost. The effectiveness of the proposed approach is evaluated by a set of metrics like probability of success, attack and defense cost and the impact of an attack. In addition, we introduce two economic factors (ROA and ROI) to evaluate the performance of ADTree. Finally, an illustration case of threat risk analysis in SCADA system is given to demonstrate our approach. Overall, our approach provides an effective means of risk assessment and countermeasures evaluation in the evolutional process of security management for cyber-physical system security.
Xiang Ji 0002, Huiqun Yu, Guisheng Fan
SNPD3
2016 A novel approach for efficient accessing of small files in HDFS: TLB-MapFile
abstract
Hadoop distributed file system (HDFS) was originally designed for streaming access large files, but the access and storage efficiency is low for the mass small files. This paper presents an access optimization approach for HDFS small file based on MapFile: TLB-MapFile. TLB-MapFile merges massive small files into large files by MapFile mechanism to reduce NameNode memory consumption and add fast table structure (TLB) in DataNode, and to improve retrieval efficiency of small files. First, according to MapFile mechanism, small files are merged into large files and stored in HDFS. Second, the access frequency and the ordered queue of small files (per unit time) can be obtained through accessing system audit logs in HDFS, and the mapping information between block and small files are stored in the TLB table with regularly being updated. TLB-MapFile improves access efficiency of small files through the prefetching of priori strategies based on TLB table. Experiment results show that this method can effectively reduce NameNode memory consumption and improve the reading speed of small files.
Bing Meng, Weibin Guo, Guisheng Fan, Neng-wu Qian
SNPD3
2016 Formally Modeling and Analyzing the Reliability of Cloud Applications
abstract
Cloud computing has become an important, useful paradigm for building applications with cloud services. However, cloud services exist in heterogeneous environments on the Internet. It is challenging to guarantee the reliability of cloud applications. Although there are efforts studying cloud and grid service reliability, very few have considered the modeling and analysis of the reliability of cloud applications. To address this emerging, important problem, we propose the first systematic approach that considers both cloud application elements and their running environment so as to faithfully model the dynamics of cloud computing. First, we present a formal description language to model the different components of a cloud application, and use it to analyze the static and dynamic factors affecting the reliability of cloud applications. Second, we propose reliability assurance strategies to ensure that cloud applications dynamically meet their required reliability. Third, Computation Tree Logic (CTL) is used to convert the reliability assurance strategy into the CTL formulas. We present operational semantics and related theories of Petri nets for establishing the correctness of our proposed method. Finally, a series of simulations are performed to evaluate the efficiency of our proposed approach.
Guisheng Fan, Huiqun Yu, Liqiong Chen
Int. J. Softw. Eng. Knowl. Eng.1
2016 A Formal Aspect-Oriented Method for Modeling and Analyzing Adaptive Resource Scheduling in Cloud Computing
abstract
Cloud computing has attracted much interest recently from both industry and academia. However, the scale and highly dynamic nature of cloud application imposes significant new challenges to resource management, and efficient resource scheduling schemes are highly demanded. In this paper, we propose a systematic method to address the reliability, running time, and failure processing of resource scheduling in cloud computing. A reflection mechanism is used to abstract the resource scheduling process as a metaobject. Petri nets are used to construct the base layer model, meta layer model, metaobject protocol, and other components, thus forming the resource scheduling model. The adaptive resource scheduling strategy is converted into CTL formulas, and the properties are analyzed. Meanwhile, an enforcement algorithm is proposed, which can guarantee the correct behavior of cloud computing while meeting the required reliability within deadline constraints. The operational semantics and related theories of Petri nets help prove its effectiveness and correctness. We have also performed a series of simulations to evaluate our approach. Results show that it can help reveal the structural and behavioral characteristics of cloud computing and improve the efficiency of resource management.
Guisheng Fan, Huiqun Yu, Liqiong Chen
IEEE Trans. Netw. Serv. Manag.1
2015 Geographical Job Scheduling in Data Centers with Heterogeneous Demands and Servers
abstract
The fast proliferation of cloud computing promotes the rapid development of large-scale commercial data centers. Tens or even hundreds of geographically distributed data centers have been deployed for better reliability and quality of services. This brings huge energy consumption for data centers. Previous research has proved that the geographical load balancing technique can achieve significant energy cost savings for geographically distributed data centers. However, existing methods for geographical load balancing often assume data centers with homogeneous servers, and workloads with single-dimension or uniform resource demands. This is an over-simplification in reality, especially when modern data centers are typically constructed from a variety of server classes. In this paper, we systematically study the problem of job scheduling for geographically distributed data centers to embrace the heterogeneity of underlying platforms and workloads. We develop a novel distributed algorithm to solve the problem efficiently based on the alternating direction method of multipliers. Extensive evaluations based on real-life data center topology, traffic traces, and electricity price data show high efficiency and efficacy of our method.
Xingjian Lu, Fanxin Kong, Jianwei Yin, Xue (Steve) Liu, Huiqun Yu, Guisheng Fan
CLOUD6
2015 Modeling and Analyzing Adaptive Energy Consumption for Service Composition
abstract
In this paper, Petri nets are used to model the different components of service composition, and form the energy consumption model of service composition based on the relationship between components, Agent is also introduced in the energy consumption management process.Then, an adaptive energy consumption strategy are proposed to dynamically ensure that service composition can get the lowest energy consumption.The operational semantics and related theories of Petri nets help establish the correctness of our proposed method.We have also performed two simulations to evaluate our proposed approach.Results show that it can help reveal the structural and behavioral characteristics of energy consumption in service composition.
Guisheng Fan, Huiqun Yu, Liqiong Chen
SEKE1
2015 Achieving Efficient Access Control via XACML Policy in Cloud Computing
abstract
One primary challenge of applying access control methods in cloud computing is to ensure data security while supporting access efficiency, particularly when adopting multiple access control policies.Many existing works attempt to propose suitable frameworks and schemes to solve the problems, however, these proposals only satisfy specified use cases.In this paper, we take XACML as the policy language and build up a logical model.Based on this, we introduce the fine-grained data fragment algorithm to optimize the policies, whose resource property represents physical meaningful data blocks.Data are organized in a tree structure, where each leaf node represents a minimal physical meaningful data block, and internal nodes are combined data types.This method can eliminate conflicts and redundancies among rules and policies, thus to refine the policy set and achieve fine-grained access control.Our approach can also be applied to processing multi-types of data, and experiments are carried out to show the improvements of efficiencies.
Xin Pei, Huiqun Yu, Guisheng Fan
SEKE3
2015 Formally Modeling and Analyzing the Reliability of Composite Service Evolution
abstract
Service composition is an important means for integrating the individual Web services for creating new value added systems. However, Web service exists in the heterogeneous environments on the Internet, thus it is challenging to guarantee the reliability of composite service evolution. To address this problem, we propose the approach to modeling and analyzing the reliability of composite service evolution. First, we present a formal description language to model the different components of service composition, and use it to analyze the reliability of composite service evolution. Second, we propose an evolution mechanism to ensure that service composition can dynamically meet the required reliability. Third, we present the operational semantics and related theories of Petri nets for establishing the consistency in the evolution process. We have also performed a series of simulations to evaluate our proposed method. Results show that it can help reveal the structural and behavioral characteristics of service composition, and improve the reliability of composite service evolution.
Guisheng Fan, Liqiong Chen, Huiqun Yu
TASE1
2015 Fine-Grained Access Control via XACML Policy Optimization in Cloud Computing
abstract
One primary challenge of enforcing access control in cloud computing is how to ensure access with high efficiency while preserving data security. This paper proposes a fine-grained access control method for cloud resources. The basic idea is to use XACML as access control language and to optimize policies by data fragmentation and policy refinement algorithms. Through data fragmentation, the accessible resources are divided into disjoint data blocks, and each of them will be combined with a set of policy rules. This helps to refine the policy and to avoid data leakage caused by rule conflicting on the resource intersections. Finally, the disjoint data blocks and the optimized policy are distributed in the three-layered cloud, and the decision to a request is made by rule matching on a specific resource rather than traversing the whole policy rules. Experiments show that our proposal enjoys higher efficiency in cloud-based access control.
Xin Pei, Huiqun Yu, Guisheng Fan
Int. J. Softw. Eng. Knowl. Eng.3
2014 Formal Modeling and Analyzing the Reliability for Service Composition
abstract
Service composition is an important means for integrating the individual Web services for creating new value added systems. However, Web service runs in the heterogeneous environments on the Internet, it is difficult to guarantee the reliability of service composition. To address this problem, we propose a systematic method to model and analyze reliability for service composition. First, we present a formal description language to model the different components of service composition, and use it to analyze the reliability of service composition. Second, we propose reliability assurance strategy to ensure that service composition dynamically meet the required reliability. Third, we present operational semantics and related theories of Petri nets for establishing the correctness of our proposed method. We have also performed a series of simulations to evaluate our proposed approach. Results show that the method can help reveal the structural and behavioral characteristics of service composition, and improve the reliability of service composition.
Guisheng Fan, Huiqun Yu, Liqiong Chen
APSEC (1)1
2013 Modeling and Optimizing Resource Scheduling for Service Composition Based on Queuing Petri Nets
abstract
Service composition is an important means for integrating the individual Web services to create new value added systems. However, because highly dynamic nature of service composition poses new challenges to resource management, efficient resource scheduling schemes are highly demanded. In this paper, a hierarchal service scheduling net is proposed to model different components of service composition, queuing theory is used to describe the competition process of available service, thus forming the scheduling model of service composition. On this basis, the evaluation function and resource scheduling strategy of service composition are proposed by considering the preference, the price and response time of available service. The related theories of Petri net are used to formally verify the correctness of proposed method. Both case study and simulation results show that the method can optimize the resource scheduling process of service composition, which has the merits of rich expressivity, while improving the performance.
Guisheng Fan, Huiqun Yu, Liqiong Chen
COMPSAC1
2013 Modeling and Analyzing Attack-Defense Strategy of Resource Service in Cloud Computing
Huiqun Yu, Guisheng Fan, Liqiong Chen
SEKE2
2013 Aspect Orientation Based Test Case Selection Strategy for Service Composition
abstract
Software testing is an important part of software maintenance, but it can also be very expensive. To reduce this expense, software testers may select part of their test cases so that those that are more important are run earlier in the testing process. However, the methods that can be used to select test cases for service composition and its analysis are still lacking at present. This paper proposes an aspect orientation based test case selection strategy for service composition. Aspect-orientation is used to weave testing crosscutting concerns of service composition, which includes component testing concern and testing concern of service composition, the weaving mechanism dynamically integrates these schemas into a testing enforcement model. Based on this, the test cases selection strategy for service composition is given, and abstract it as a crosscutting concern to weave into testing model, the corresponding enforcement algorithm is also given, the operation semantics and related theories of Petri nets help prove its effectiveness and feasibility. A case study explains the testing process of service composition, and a series of experiments are done to explain that the use of aspects for testing Web service is more efficient than conventional techniques, which can improve the testing quality and efficiency.
Guisheng Fan, Huiqun Yu, Liqiong Chen
TASE1
2013 Petri net based techniques for constructing reliable service composition
Guisheng Fan, Huiqun Yu, Liqiong Chen
J. Syst. Softw.1
2012 Model Based Byzantine Fault Detection Technique for Cloud Computing
abstract
Cloud computing has attracted much interest recently from both industry and academic. More and more Internet applications are moving to the cloud environment. However, fault detection technique in cloud application is a crucial issue. This issue is especially difficult since cloud computing relies by nature on a highly dynamic environment. In this paper, we propose a model based Byzantine fault detection technique for cloud computing. A cloud computing fault net (CFN) is used to precisely model the different components of cloud computing, such as service resources, cloud module, the detection and failure process, etc, the basic properties of the constructed model are analyzed. Based on this, the fault detection strategy is proposed, which can dynamically detect the fault of cloud application in the execution process. The operational semantics and related theories of Petri nets help prove its effectiveness and correctness. An example is used to simulate the modeling and analyzing process, and a series of experiments are done to explain the effectiveness of proposed method.
Guisheng Fan, Huiqun Yu, Liqiong Chen
APSCC1
2012 A Petri Net-Based Byzantine Fault Diagnosis Method for Service Composition
abstract
Service composition is an important means for integrating the individual Web services to create new value added systems that can satisfy complex requirements. However, it is a challenge to enforce fault diagnosis mechanism for those applications due to the uncertainty of service quality in distributive and heterogeneous environment. In this paper, a Byzantine fault diagnosis method for service composition based on Petri nets is proposed. The reliability of service are taken into account for the appropriate selection of required services. And a service composition fault net (SCFN) is proposed, which can be used to model different components of service composition. Finally, the fault detection strategy is provided for processing fault of service composition in dynamic environment. Theories of Petri nets help prove its correctness and effectiveness, thus guarantee the reliability of service composition. A case study illustrates the applicability of proposed method, and its feasibility has been demonstrated by simulation.
Guisheng Fan, Huiqun Yu, Liqiong Chen
COMPSAC1
2011 An Approach to Modeling and Analyzing Security Requirements of Service Composition
abstract
Service composition is an important means for integrating the individual Web services to create new value added systems that can satisfy complex requirements. However, it is a challenge to analyze security requirements for those applications due to the uncertainty factors in distributive environment. This paper proposes an approach to modeling and analyzing security requirements of service composition. Petri nets are used to model the different components of service composition, the dynamic matching strategy of service composition is proposed. Aspect-orientation is used to weave the security requirements into service composition, which includes evaluation concern, authorization, security level outputting and access outputting. The operation semantics and related theories of Petri nets help prove its effectiveness and correctness. An example explains the modeling and analyzing process of service composition, and a series of experiments are done to explain that the use of aspects for analyzing security requirements of service composition is more efficient than conventional techniques.
Guisheng Fan, Huiqun Yu, Liqiong Chen
APSCC1
2011 A Regression Test Technique for Analyzing the Functionalities of Service Composition
Huiqun Yu, Guisheng Fan, Liqiong Chen
SEKE3
2011 An Approach to Handling Failure Recovery in Service Composition and Its Analysis
abstract
Service composition is an effective way to build complex Web service applications. However, it is a challenge to handle failure recovery due to the uncertainty of service in distributed and heterogeneous environment. This paper proposes an approach to handling failure recovery in service composition. Petri nets are used to model the different components of service composition, failure recovery rules and service selection strategies are given. Based on these, aspect-orientation is used to weave failure recovery concern into service composition, which includes failure warning concern, service selection concern and recovery concern, the weaving mechanism dynamically integrates these schemas into a failure recovery model. The operation semantics and related theories of Petri nets help prove its effectiveness and correctness. A case study and experimental results demonstrate the approach can simplify the failure recovery process, and improve the design quality of service composition.
Guisheng Fan, Huiqun Yu, Liqiong Chen, Chunhua Gu
TASE1
2011 A Certificate Driven Access Control Strategy for Service Composition and Its Analysis
abstract
Service composition is an effective way to achieve value-added service, which has found wide application in various key areas. However, most access control techniques for service composition were in ad hoc fashion and fell short in precise notations. In this paper, we propose a certificate driven access control strategy for service composition. Petri nets are used to precisely define and model the different components of service composition. The access control strategy for service composition are proposed, which can dynamically adjust available service to meet the actual requirements. Based on this, theories of Petri nets help prove correctness of the access control strategy and the enforcement algorithm is given, thereby getting the service composition which can meet the functional requirements while meets the required security. The proposed method is applied to a real-world domain to show the feasibility and effectiveness.
Guisheng Fan, Huiqun Yu, Liqiong Chen
TrustCom1
2010 An Aspect Oriented Approach to Analyzing Fault of Service Composition
abstract
Service composition is an effective way to achieve value-added service, which has found wide application in software system. Fault handling is critical to achieve high reliability for these applications. However, the existing service composition methods seldom consider services' fault handling, which results in high risk of runtime failure. This paper proposes a formal aspect-oriented approach to designing and analyzing fault of service composition. The underlying formalism is Petri net and its corresponding modeling method. The fault handling process is encapsulated into aspect net and base net, and Petri net is used to model the core concerns and crosscutting concerns, the weaving mechanism systematically integrates these schemas into a complete service composition model. Based on the model, the related theories of Petri net help prove the correctness of fault handling. Finally, an Export Service and simulation results show that our method can ensure the high reliability and design quantity of service composition.
Guisheng Fan, Huiqun Yu, Chunhua Gu, Liqiong Chen
APSCC1
2010 Aspect Oriented Approach to Building Secure Service Composition
abstract
Service composition is an effective way to achieve value-added service, which has found wide application in various areas. security design at architecture level is critical to achieve high assurance for these applications. However, most security design techniques for service composition were in ad hoc fashion and fell short in precise notations. This paper proposes a formal aspect-oriented approach to designing and analyzing secure service composition. The underlying formalism is Petri net and its modeling method, and focuses on the service authorization, implementation trace ability, data protection and fault handling. Aspect specification provides means to observe behaviors of basic aspect schema, and to describe their interrelationship, while the weaving mechanism systematically integrates these schemas into a complete service composition model. Based on this, the security and fault recovery mechanism of service composition are analyzed, and its correctness and effectiveness are proved. A case study of Export Service demonstrates the approach can simplify the modeling process and improve the design quality.
Guisheng Fan, Huiqun Yu, Liqiong Chen
APSEC1
2009 A Method for Modeling and Analyzing Fault-Tolerant Service Composition
abstract
Reliability is a key issue of the service-oriented architecture (SOA) that is widely employed in distributed systems such as e-commerce and e-government. Redundancy based technologies are usually employed for building reliable service composition on top of unreliable Web services. This paper proposes a strategy for modeling and analyzing fault tolerant service composition. The strategy consists of service selection mechanism, service synchronization mechanism and task exception mechanism. Petri nets are used to construct different components of service composition. Once the model is constructed, theories of Petri nets help prove the consistency of processing states and reliability of the strategy. The corresponding enforcement method for constructing fault-tolerant service composition is proposed. Experiments are conducted to demonstrate the applicability and effectiveness of the fault tolerant strategy.
Guisheng Fan, Huiqun Yu, Liqiong Chen
APSEC1
2008 Analyzing BPEL Compositionality Based on Petri Nets
abstract
Process of service composition is complex and error-prone, which makes a formal modeling and analysis method highly desirable. This paper presents a Petri net-based approach to analyzing the soundness and compositionality of services in BPEL. A set of translation rules is proposed to transform BPEL processes into Petri nets, by which behaviors of the BPEL processes are articulated. The instantiation net of target services are used to capture all of the possible implementation flows of composition processes. Based on theories of Petri nets, the principles for analyzing soundness and compositionality of Web services are provided. A detailed example is given to demonstrate the applicability of our method.
Guisheng Fan, Huiqun Yu, Liqiong Chen
COMPSAC1