VLDB 2026 Research / reviewers in the wild / expert
Xiaobing Sun 0001
dblp:30/4077-1
· DBLP profile ↗
145ranked-venue papers
26as first author
91since 2021 · last 2026
0000-0001-5165-5080ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 83 · 19 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 17 since 2021Security and privacy · 11 · 2 first-author · 11 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiKD: Backdoor Defense in Federated Graph Learning via Attention-Guided Multi-Teacher DistillationabstractBackdoor attacks pose a severe threat to federated graph learning (FGL), where malicious clients can inject hidden triggers into the global model without being detected. Defending against such attacks is particularly challenging due to the complex graph structures and the stealthy nature of trigger patterns. In this work, we propose MultiKD, a novel backdoor mitigation method based on attention-guided multi-teacher distillation. Unlike existing defenses that focus on detecting suspicious clients or restricting backdoor activation, MultiKD directly purifies the global model on the server side by exploiting intermediate representations. It integrates knowledge from multiple client models and guides the global model to suppress backdoor behaviors by aligning attention maps and preserving inter-layer relational consistency. Our defensive intuition enables MultiKD to retain task-relevant information while mitigating malicious patterns, even when some teacher models are compromised. Extensive experiments on four real-world datasets demonstrate the effectiveness of our approach in significantly reducing attack success rate (≤ 8%) with minimal impact on utility (≤ 5%). Jiale Zhang 0001, Bosen Rao, Xiaobing Sun 0001, Yu Li 0022 |
AAAI | 5 |
| 2026 | PIONEER: improving the robustness of student models when compressing pre-trained models of code
Xiangyue Liu 0002, Lili Bo, Xiaoxue Wu 0001, Yun Yang 0003, Xiaobing Sun 0001 |
Autom. Softw. Eng. | 6 |
| 2026 | AdvGen-X: Transferability driven adversarial example generation for pre-trained models of code
Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Bin Li 0006, Xiaoxue Wu 0001, Sicong Cao, Yufei Hu |
Empir. Softw. Eng. | 3 |
| 2026 | Aspect-centric vulnerability understanding via semantics-aware commit representation learning
Xiaobing Sun 0001, Sicong Cao, Zhenlei Ye |
Empir. Softw. Eng. | 1 |
| 2026 | Multimodal heterogeneous graph neural networks for recommendation via large language models' guidance
Jianyi Chen, Kai Yang 0031, Zijuan Zhao, Xiaobing Sun 0001, Zhao-Long Hu |
Expert Syst. Appl. | 4 |
| 2026 | Heterogeneous graph neural networks with weak information
Ce Na, Kai Yang 0031, Xiaobing Sun 0001 |
Expert Syst. Appl. | 3 |
| 2026 | GFedHPR: Graph Federated Learning for Hybrid Privacy Recommendation
Ce Na, Qingqing Ye 0001, Dengzhao Fang, Jingtong Gao, Xiaobing Sun 0001, Yi Chang 0001 |
Pattern Recognit. | 9 |
| 2026 | Missing visual modality graph transformer for multi-modal entity alignment
Kai Yang 0031, Junyan Guo, Xiaobing Sun 0001 |
Pattern Recognit. | 4 |
| 2026 | Code Language Models for Security Patch Management: How Far are We?abstractThe rapid expansion of open-source software has also brought significant security challenges to cloud infrastructure, particularly introducing and propagating vulnerabilities. In response, effective security patch management establishes a continuous, structured pipeline by systematically identifying, testing, and deploying security patches to fix vulnerabilities. However, manually managing a large number of security patches (i.e., any update is approved and installed by hand) is time-consuming, leading to a great motivation for automating this process. Although Code Language Models (CodeLMs) have shown potential in various code-centric tasks, there remains an open question as to how well CodeLMs perform within the context of security patch management. To bridge this gap, we performed the first comprehensive empirical study on fine-tuning or prompting nine state-of-the-art CodeLMs for three security-patch-related downstream tasks, including silent patch identification (distinguishing security patches from normal commits), record-patch linking (connecting authoritative vulnerability records, e.g., CVE, to the corresponding fixing commits), and vulnerability description generation (providing a piece of text summarizing the vulnerability fixed by the patch), covering classification, ranking, and generation problems. Our findings reveal that there is no “one-size-fits-all” model that can always perform the best. Furthermore, due to the lack of task-specific knowledge, naively prompting LLMs with the basic strategies is not consistently reliable and may even underperform smaller PTMs. Additionally, existing automated evaluation metrics cannot fully reflect the capability of LLMs in considered tasks. These findings underscore the considerable gap between current capabilities and the practical requirements for deploying CodeLMs in automating security patch management. Xingwei Lin, Sicong Cao, Le Yu 0002, Xiaobing Sun 0001, Fu Xiao 0001, Lei Xue 0001, Chunming Wu 0001, Kui Ren 0001, David Lo 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI Ecosystem
Xingan Gao, Xiaobing Sun 0001, Sicong Cao, Kaifeng Huang 0001, Xingwei Lin |
USENIX Security Symposium | 2 |
| 2025 | Beyond Dataset Watermarking: Model-Level Copyright Protection for Code Summarization Models
Jiale Zhang 0001, Di Wu 0050, Xiaobing Sun 0001, Qinghua Lu 0001, Guodong Long |
WWW | 4 |
| 2025 | Fixer-level supervised contrastive learning for bug assignment
Rongcun Wang, Xingyu Ji, Yuan Tian 0008, Senlei Xu, Xiaobing Sun 0001, Shujuan Jiang |
Empir. Softw. Eng. | 5 |
| 2025 | EPAD: Ethereum phishing scam detection via graph contrastive learning
Hao Sui 0003, Jiale Zhang 0001, Bing Chen 0002, Di Wu 0050, Xiaobing Sun 0001, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 5 |
| 2025 | AI-Enhanced Resource Allocation for LPWAN-Based LoRaWAN:A Hybrid TinyML and Deep Learning ApproachabstractThe integration of Artificial Intelligence (AI) with Low Power Wide Area Networks (LPWAN) offers a promising approach to address resource constraints and dynamic network conditions inherent in these networks. However, deploying complex AI algorithms on resource-limited edge devices presents significant challenges due to their limited computational capabilities. In this study, we propose a hybrid Tiny Machine Learning (TinyML) and Deep Neural Network (DNN)-based solution for optimizing resource allocation in LPWAN-based LoRaWAN networks, targeting both static and mobile applications. Our approach leverages the strengths of a 1-D Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) model implemented on the network server, combined with TinyML models deployed on edge devices. The CNN-LSTM model predicts optimal spreading factor and transmission power by analyzing spatial and temporal patterns from real-time data, while the TinyML models enable edge devices to autonomously adjust communication parameters in resource-constrained and disconnected scenarios. This hybrid framework enhances network performance by improving the packet success ratio (PSR), maximizing energy efficiency, and addressing the challenges posed by dynamic IoT environments. Muhammad Ali Lodhi, Xiaobing Sun 0001, Khalid Mahmood 0002, Anum Lodhi, Youngho Park 0005, Majid Hussain |
IEEE Internet Things J. | 2 |
| 2025 | Meta Computing-Driven Optimization of AoI in Industrial IoT: A Hybrid Scheme With Self-Organizing Maps and Reinforcement LearningabstractIn the domain of Industrial Internet of Things (IIoT) applications, methodologies such as meta computing are essential for ensuring timely and efficient data acquisition within intricate environments characterized by multiple edge devices and isolated sensor networks. The Age of Information (AoI), a key metric for evaluating data timeliness and relevance, has emerged as a focal point for improving decision-making and system responsiveness. However, managing these systems in resourcelimited edge environments is challenging, as traditional methods struggle to link sink node selection with data collection path planning, leading to inefficiency and poor AoI performance. This paper develops a collaborative optimization framework based on meta computing, which dynamically coordinates distributed computational resources as a unified virtual system. It combines self-organizing mapping (SOM) with reinforcement learning (RL) to optimize sink node selection and data collection paths using AoI metrics. Experimental results show the method significantly reduces AoI, improves data freshness, and optimizes energy use, highlighting meta computings potential for scalable, efficient IIoT solutions. Min Wang 0026, Nuanlai Wang, Xiaobing Sun 0001, Ming Li 0042 |
IEEE Internet Things J. | 5 |
| 2025 | Evaluating the Test Adequacy of Benchmarks for LLMs on Code GenerationabstractABSTRACT Code generation for users' intent has become increasingly prevalent with the large language models (LLMs). To automatically evaluate the effectiveness of these models, multiple execution‐based benchmarks are proposed, including specially crafted tasks, accompanied by some test cases and a ground truth solution. LLMs are regarded as well‐performed in code generation tasks if they can pass the test cases corresponding to most tasks in these benchmarks. However, it is unknown whether the test cases have sufficient test adequacy and whether the test adequacy can affect the evaluation. In this paper, we conducted an empirical study to evaluate the test adequacy of the execution‐based benchmarks and to explore their effects during evaluation for LLMs. Based on the evaluation of the widely used benchmarks, HumanEval, MBPP, and two enhanced benchmarks HumanEval+ and MBPP+, we obtained the following results: (1) All the evaluated benchmarks have high statement coverage (above 99.16%), low branch coverage (74.39%) and low mutation score (87.69%). Especially for the tasks with higher cyclomatic complexities in the HumanEval and MBPP, the mutation score of test cases is lower. (2) No significant correlation exists between test adequacy (statement coverage, branch coverage and mutation score) of benchmarks and evaluating results on LLMs at the individual task level. (3) There is a significant positive correlation between mutation score‐based evaluation and another execution‐based evaluation metric () on LLMs at the individual task level. (4) The existing test case augmentation techniques have limited improvement in the coverage of test cases in the benchmark, while significantly improving the mutation score by approximately 34.60% and also can bring a more rigorous evaluation to LLMs on code generation. (5) The LLM‐based test case generation technique (EvalPlus) performs better than the traditional search‐based technique (Pynguin) in improving the benchmarks' test quality and evaluation ability of code generation. Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Yufei Hu, Zhenlei Ye |
J. Softw. Evol. Process. | 2 |
| 2025 | From Data to Knowledge: Mining Linux Vulnerability Characteristics and Evolution With Knowledge GraphsabstractABSTRACT An operating system is the essence of software, serving as the foundation for the operation of various application software. The security of the operating system is crucial for national informatization construction. Data indicate that many cybersecurity incidents result from exploiting security vulnerabilities in the operating system. Linux is currently the most widely used open‐source operating system, with thousands of Common Vulnerabilities and Exposures (CVEs) related to Linux systems reported each year. Therefore, research and prevention of vulnerabilities in the Linux system are particularly important. To gain a better understanding of the characteristics of Linux system vulnerabilities, this paper leverages knowledge in the field of software security to analyze nearly 10,000 historical vulnerability data in two core systems of Linux: Linux Kernel and Debian Linux. The study explores the evolutionary patterns of vulnerability characteristics. Specific research contents include the following: (1) data collection and cleaning of vulnerability data in Linux Kernel and Debian Linux systems; (2) cross‐statistical analysis of structured data features in vulnerability reports; (3) unstructured data characteristics mining in vulnerability reports based on domain knowledge; (4) analysis of the evolution of vulnerability characteristics. This paper provides empirical lessons and guidance for Linux system vulnerabilities to assist practitioners and researchers in better preventing and detecting vulnerabilities in Linux and Linux‐based systems. Shiyu Weng, Xiaoxue Wu 0001, Wenjing Shan, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 6 |
| 2025 | NG_MDERANK: A software vulnerability feature knowledge extraction method based on N-gram similarityabstractAbstract As software grows in size and complexity, software vulnerabilities are increasing, leading to a range of serious insecurity issues. Open‐source software vulnerability reports and documentation can provide researchers with great convenience for analysis and detection. However, the quality of different data sources varies, the data are duplicated and lack of correlation, which often requires a lot of manual management and analysis. In order to solve the problems of scattered and heterogeneous data and lack of correlation in traditional vulnerability repositories, this paper proposes a software vulnerability feature knowledge extraction method that combines the N‐gram model and mask similarity. The method generates mask text data based on the extraction of N‐gram candidate keywords and extracts vulnerability feature knowledge by calculating the similarity of mask text. This method analyzes the samples efficiently and stably in the environment of large sample size and complex samples and can obtain high‐value semi‐structured data. Then, the final node, relationship, and attribute information are obtained by secondary knowledge cleaning and extraction of the extracted semi‐structured data results. And based on the extraction results, the corresponding software vulnerability domain knowledge graph is constructed to deeply explore the semantic information features and entity relationships of vulnerabilities, which can help to efficiently study software security problems and solve vulnerability problems. The effectiveness and superiority of the proposed method is verified by comparing it with several traditional keyword extraction algorithms on Common Weakness Enumeration (CWE) and Common Vulnerabilities and Exposures (CVE) vulnerability data. Xiaoxue Wu 0001, Shiyu Weng, Wei Zheng 0006, Xiang Chen 0005, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 6 |
| 2025 | HgtJIT: Just-in-Time Vulnerability Detection Based on Heterogeneous Graph TransformerabstractVulnerability detection plays a crucial role in the software development lifecycle. Commit-level vulnerability detection aims to detect whether the changed code contributed to potential vulnerabilities by the developer when submitting the code, which is also referred to as Just-In-Time (JIT) vulnerability detection. Previous JIT vulnerability detection approaches relied on code metrics and textual features, which were unable to effectively characterize vulnerability-contributing commits (VCCs). Recently, CodeJIT (a code-centric learning-based approach) has been proposed to detect vulnerability at the commit-level. However, CodeJIT still has its limitations: imprecise feature representation, static code embedding, and underutilized heterogeneous information. In this paper, we propose HgtJIT, a JIT vulnerability detection approach based on a Heterogeneous Graph Transformer (HGT) in order to address several limitations of the state-of-the-art CodeJIT approach. We propose diffPDG to represent code changes and use the CCT5 model (the latest feature encoder pre-trained on a large-scale code change corpus) to embed graph nodes to generate the most meaningful vector representations. In addition, we employ HGT to adequately utilize heterogeneous information of the graph to learn vulnerability features. Extensive experiments have shown that HgtJIT is the best-performing model, with F1 and AUC improvement of 14.6%-37.5% and 12.2%-53.7% compared to the baseline model Xiaobing Sun 0001, Mingxuan Zhou, Sicong Cao, Xiaoxue Wu 0001, Lili Bo, Di Wu 0050, Bin Li 0006, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Detecting Reentrancy Vulnerabilities for Solidity Smart Contracts With Contract Standards-Based RulesabstractThe reentrancy vulnerability is one of the most notorious vulnerabilities of smart contracts. It enables attackers to hijack the control flow of a smart contract by invoking a function as the entry point and then re-invoking a function as the reentry point before the execution of the entry point ends. Although several approaches have been proposed to detect this vulnerability, they still face two main limitations. Firstly, existing approaches oversimplify the rules for identifying entry and reentry points, and many even neglect reentry point identification during vulnerability detection. Secondly, most existing approaches overlook the flow of state variables that are not promptly updated, a critical aspect of the reentrancy vulnerability. To address the limitations mentioned above, this article proposes a novel static analysis framework for reentry vulnerability detection. We formulate the reentrancy vulnerability detection as entry and reentry point identification with the state variable flow tracking. Based on the insight that most smart contracts are implemented following various technical standards, we utilize static analysis with standard-based rules to identify potential entry and reentry points. This is achieved by detecting the presence of hijackable and exploitable operations inside the smart contract. Meanwhile, we also conduct state variable flow tracking by the static taint analysis. To verify the effectiveness of our proposed approach, we construct three different datasets. Then We compare our approach with eight state-of-the-art smart contract vulnerability detectors, and our tool outperforms these baselines in detecting more vulnerable samples with fewer false positive samples. Meanwhile, our approach achieves a relatively shorter detection time with better detection results, striking a trade-off between effectiveness and efficiency. Jie Cai 0006, Jiachi Chen, Tao Zhang 0001, Xiapu Luo, Xiaobing Sun 0001, Bin Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | GraphCleanse: Defending Backdoor Attacks in Graph Learning via Contrastive TrainingabstractGraph Neural Networks (GNNs) are highly susceptible to numerous adversarial attacks, among which the backdoor attack is one of the toughest to deal with due to the fact that it can lead to misclassification of the model. Similar to Deep Neural Networks (DNNs), backdoor attacks in GNNs work by an attacker changing a portion of the graph data with a hidden trigger and modifying their labels to target labels, which induces the model to learn the trigger feature during its training phase. Although recent defense techniques have emerged, approaches based on explainability and data isolation often fail to detect malicious samples with covert triggers, while discrepancy learning methods tend to degrade performance by removing useful features. To overcome these limitations, we propose a novel backdoor defense method, namedGraphCleanse, on GNNs that can effectively eliminate the possible backdoor features during the training process. Specifically,GraphCleansecan easily break the strong correlation between backdoor features and target labels based on graph contrastive training. To further improve the model accuracy, we present a mutual information maximization method to learn the important feature information in the labeled credible samples and unlabeled suspicious samples by clustering the features obtained from the graph contrastive encoder. Compared with the potential solutions, such as randomized smoothing,GraphCleanseeffectively avoids the negative influence of backdoored samples while maintaining a high model performance. Extensive experimental evaluations on four benchmark datasets demonstrate thatGraphCleansecan reduce the attack success rate to 10% with less performance degradation (within 7%). Jiale Zhang 0001, Hao Sui 0003, Wanquan Zhu, Xiaobing Sun 0001, Chunpeng Ge 0001, Bing Chen 0002, Mingsheng Cao 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | SSLDefender: Backdoor Defense in Self-Supervised Learning via Distillation-Guided UnlearningabstractSelf-supervised learning utilizes unlabelled data to train encoders, acquiring high-quality representations of input data, significantly advancing the field of computer vision. However, recent studies have demonstrated that self-supervised learning suffers from numerous adversarial attacks. Among them, backdoor attack is one of the focal issues, where downstream classifiers inherit the backdoor behavior of the pre-trained encoder. Existing defense methods against backdoor attacks primarily focus on supervised learning, which heavily relies on labeled data and cannot be directly migrated to self-supervised scenarios. Furthermore, defense methods for self-supervised backdoor aims to separate poisoned samples on assumed small-scale datasets and retraining to obtain a clean encoder. However, these approaches are useless against encoders that have been implanted with a backdoor. To address these issues, we propose SSLDefender, a novel image-based backdoor mitigation method specially designed for self-supervised learning, which can remove backdoor attributes directly from the backdoor encoder. Specifically, we employ a trigger recovery method based on mutual information maximization to efficiently obtain trigger that resembles the target backdoor’s influence. Additionally, we design a distillation-guided unlearning strategy to purify backdoor features steadily and ensure the retention of clean knowledge to prevent overforgetting. Extensive experimental evaluations on six benchmark datasets demonstrate that SSLDefender can successfully reduce the attack success rate of Badencoder to around 2% while maintaining high model accuracy on the main task. Its performance surpasses state-of-the-art methods. Jiale Zhang 0001, Wanquan Zhu, Kai Wang 0062, Xiaobing Sun 0001, Weizhi Meng 0001, Xiapu Luo |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | New RNN Algorithms for Different Time-Variant Matrix Inequalities Solving Under Discrete-Time FrameworkabstractA series of discrete time-variant matrix inequalities is generally regarded as one of the challenging problems in science and engineering fields. As a discrete time-variant problem, the existing solving schemes generally need the theoretical support under the continuous-time framework, and there is no independent solving scheme under the discrete-time framework. The theoretical deficiency of solving scheme greatly limits the theoretical research and practical application of discrete time-variant matrix inequalities. In this article, new discrete-time recurrent neural network (RNN) algorithms are proposed, analyzed, and investigated for solving different time-variant matrix inequalities under the discrete-time framework, including discrete time-variant matrix vector inequality (discrete time-variant MVI), discrete time-variant generalized matrix inequality (discrete time-variant GMI), discrete time-variant generalized-Sylvester matrix inequality (discrete time-variant GSMI), and discrete time-variant complicated-Sylvester matrix inequality (discrete time-variant CSMI), and all solving processes are based on the direct discretization thought. Specifically, first of all, four discrete time-variant matrix inequalities are presented as the target problems of these researches. Second, for solving such problems, we propose corresponding discrete-time recurrent neural network (RNN) (DT-RNN) algorithms (termed DT-RNN-MVI algorithm, DT-RNN-GMI algorithm, DT-RNN-GSMI algorithm, and DT-RNN-CSMI algorithm), which are different from the traditional DT-RNN design thought because second-order Taylor expansion is applied to derive the DT-RNN algorithms. This creative process avoids the intervention of continuous-time framework. Then, theoretical analyses are presented, which show the convergence and precision of the DT-RNN algorithms. Abundant numerical experiments are further carried out, which further confirm the excellent properties of the DT-RNN algorithms. Yang Shi 0003, Chenling Ding, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Large Language Model for Vulnerability Detection and Repair: Literature Review and the Road AheadabstractThe significant advancements in Large Language Models (LLMs) have resulted in their widespread adoption across various tasks within Software Engineering (SE), including vulnerability detection and repair. Numerous studies have investigated the application of LLMs to enhance vulnerability detection and repair tasks. Despite the increasing research interest, there is currently no existing survey that focuses on the utilization of LLMs for vulnerability detection and repair. In this paper, we aim to bridge this gap by offering a systematic literature review of approaches aimed at improving vulnerability detection and repair through the utilization of LLMs. The review encompasses research work from leading SE, AI, and Security conferences and journals, encompassing 43 papers published across 25 distinct venues, along with 15 high-quality preprint papers, bringing the total to 58 papers. By answering three key research questions, we aim to (1) summarize the LLMs employed in the relevant literature, (2) categorize various LLM adaptation techniques in vulnerability detection, and (3) classify various LLM adaptation techniques in vulnerability repair. Based on our findings, we have identified a series of limitations of existing studies. Additionally, we have outlined a roadmap highlighting potential opportunities that we believe are pertinent and crucial for future research endeavors. Xin Zhou 0014, Sicong Cao, Xiaobing Sun 0001, David Lo 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | KG4VA: Constructing Vulnerability Knowledge Graph for Software Vulnerability AssessmentabstractSoftware vulnerabilities pose serious threats to software security. When faced with multiple software vulnerabilities at the same time, it is urgent to determine whether the vulnerabilities are high-risk. Existing vulnerability assessment approaches only learn the mapping relationships between vulnerability descriptions and severity levels, while ignoring the sharing of the same or similar elements between vulnerabilities. Furthermore, solely focusing on vulnerability descriptions fails to accurately characterize the vulnerability behavior. In this paper, we propose a novel vulnerability knowledge graph (KG) to capture the relationships between vulnerabilities. To construct the vulnerability KG automatically, we propose to leverage vulnerability elements extracted from vulnerability descriptions to link different vulnerabilities. Based on the constructed KG, we further propose a novel KG-based vulnerability assessment (VA) approach KG4VA, which precisely finds the similar vulnerability for an encountered vulnerability description by analyzing and matching the elements entities based on the vulnerability KG. The experiment results show that KG4VA outperforms the baselines in almost all metrics (e.g., 3.27%-10.83% accuracy improvements). Moreover, our ablation experiments demonstrate that the vulnerability knowledge graph can indeed offer valuable information for vulnerability assessment. Zhenlei Ye, Xiaobing Sun 0001, Lili Bo, Sicong Cao, Xiaoxue Ren, Lianyong Qi, Jiale Zhang 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | Misactivation-Aware Stealthy Backdoor Attacks on Neural Code Understanding ModelsabstractNeural code models (NCMs) play a crucial role in helping developers solve code understanding tasks. Recent studies have exposed that NCMs are vulnerable to several security threats, among which backdoor attack is one of the toughest. It is usually achieved through data poisoning. Specifically, backdoored NCMs work normally on the clean example but produce attacker-expected output on the example injected with backdoor triggers. However, existing backdoor attacks against NCMs face two significant drawbacks: 1) lack of stealthiness, that is trigger tokens are easily detected by defense techniques/humans when they appear in excessive numbers; 2) damage to the model’s normal performance, that is partial trigger tokens may frequently appear as benign features in the clean samples, resulting in clean samples containing them may falsely activate the backdoor. To address these drawbacks, we propose a misactivation-aware stealthy backdoor attack against NCMs through data poisoning called MISNCM. MISNCM features target-biased trigger generation, thus achieving stealthy backdoor attacks. Moreover, we utilize misactivation-aware data poisoning to create calibration samples with partial trigger tokens to reduce false activations and ensure the regular performance of the model. We conduct comprehensive experiments to evaluate the effectiveness of MISNCM in attacking NCMs used for three code understanding tasks: defect detection, clone detection, and authorship attribution. The experimental results demonstrate that the triggers generated by MISNCM achieve an average attack success rate increase of 12.67% over IR and 8.38% over AFRAIDOOR. Furthermore, MISNCM achieves a 3.64% improvement in F1 score on the code clone detection task, and an average of 5.91% improvement in accuracy on the defect detection and authorship attribution tasks, compared with the two baselines. Xiaobing Sun 0001, Yiran Xiao, Lili Bo, Weisong Sun, Xiangyue Liu 0002, Bin Li 0006, Jiale Zhang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2024 | Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsabstractRecently, Graph Neural Network (GNN)-based vulnerability detection systems have achieved remarkable success. However, the lack of explainability poses a critical challenge to deploy black-box models in security-related domains. For this reason, several approaches have been proposed to explain the decision logic of the detection model by providing a set of crucial statements positively contributing to its predictions. Unfortunately, due to the weakly-robust detection models and suboptimal explanation strategy, they have the danger of revealing spurious correlations and redundancy issue. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Wei Liu 0010 |
ICSE | 2 |
| 2024 | BADFSS: Backdoor Attacks on Federated Self-Supervised Learning
Jiale Zhang 0001, Di Wu 0050, Xiaobing Sun 0001, Jianming Yong, Guodong Long |
IJCAI | 4 |
| 2024 | 1+1>2: Integrating Deep Code Behaviors with Metadata Features for Malicious PyPI Package DetectionabstractPyPI, the official package registry for Python, has seen a surge in the number of malicious package uploads in recent years. Prior studies have demonstrated the effectiveness of learning-based solutions in malicious package detection. However, manually-crafted expert rules are expensive and struggle to keep pace with the rapidly evolving malicious behaviors, while deep features automatically extracted from code are still inaccurate in certain cases. To mitigate these issues, in this paper, we propose Ea4mp, a novel approach which integrates deep code behaviors with metadata features to detect malicious PyPI packages. Specifically, Ea4mp extracts code behavior sequences from all script files and fine-tunes a BERT model to learn deep semantic features of malicious code. In addition, we realize the value of metadata information and construct an ensemble classifier to combine the strengths of deep code behavior features and metadata features for more effective detection. We evaluated Ea4mp against three state-of-the-art baselines on a newly constructed dataset. The experimental results show that Ea4mp improves precision by 6.9%-24.6% and recall by 10.5%-18.4%. With Ea4mp, we successfully identified 119 previously unknown malicious packages from a pool of 46,573 newly-uploaded packages over a three-week period, and 82 out of them have been removed by the PyPI official. Xiaobing Sun 0001, Xingan Gao, Sicong Cao, Lili Bo, Xiaoxue Wu 0001, Kaifeng Huang 0001 |
ASE | 1 |
| 2024 | ChatBR: Automated assessment and improvement of bug report quality using ChatGPTabstractBug reports, containing crucial information such as the Observed Behavior (OB), the Expected Behavior (EB), and the Steps to Reproduce (S2R), can help developers localize and fix bugs efficiently. However, due to the increasing complexity of some bugs and the limited experience of some reporters, large numbers of bug reports miss this crucial information. Although machine learning (ML)-based and information retrieval (IR)-based approaches are proposed to detect and supplement the missing information in bug reports, the performance of these approaches depends heavily on the size and quality of bug report datasets. Lili Bo, Wangjie Ji, Xiaobing Sun 0001, Ting Zhang 0011, Xiaoxue Wu 0001, Ying Wei 0012 |
ASE | 3 |
| 2024 | Snopy: Bridging Sample Denoising with Causal Graph Learning for Effective Vulnerability DetectionabstractDeep Learning (DL) has emerged as a promising means for vulnerability detection due to its ability to automatically derive features from vulnerable code. Unfortunately, current solutions struggle to focus on vulnerability-related parts of vulnerable functions, and tend to exploit spurious correlations for prediction, thus undermining their effectiveness in practice. In this paper, we propose Snopy, a novel DL-based approach, which bridges sample denoising with causal graph learning to capture real vulnerability patterns from vulnerable samples with numerous noise for effective detection. Specifically, Snopy adopts a change-based sample denoising approach to automatically weed out vulnerability-irrelevant code elements in the vulnerable functions without sacrificing the label accuracy. Then, Snopy constructs a novel Causality-Aware Graph Attention Network (CA-GAT) with Feature Caching Scheme (FCS) to learn causal vulnerability features while maintaining efficiency. Experiments on the three public benchmark datasets show that Snopy outperforms the state-of-the-art baselines by an average of 27.22%, 85.89%, and 75.50% in terms of F1-score, respectively. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Xiaolei Liu 0001, Xingwei Lin, Wei Liu 0010 |
ASE | 2 |
| 2024 | A Software Bug Fixing Approach Based on Knowledge-Enhanced Large Language ModelsabstractSoftware Bug Fixing is a time-consuming task in software development and maintenance. Despite the success of Large Language Models (LLMs) using in Automatic Program Repair (APR), they still have the limitations of generating patches with low accuracy and explainability. In this paper, we propose a software bug-fixing approach based on knowledge-enhanced large language models. First, we collect bugs as well as their fix information from bug tracking systems, such as Github and Stack Overflow. Then, we extract bug entities and inter-entity relationships using Named Entity Recognition (NER) to construct a Bug Knowledge Graph (BKG). Finally, we utilize LLMs (e.g., GPT-4) which is enhanced by the knowledge of the similar historical bugs as well as fix information from BKG to generate patches for new bugs. The experimental results show that the our approach can fix 28.52% (85\298) bugs correctly, which is significantly better than the state-of-the-art approaches. Furthermore, the generated patches are explainable and more credible. Lili Bo, Xiaobing Sun 0001, Wangjie Ji |
QRS | 3 |
| 2024 | SCL-CVD: Supervised contrastive learning for code vulnerability detection via GraphCodeBERT
Rongcun Wang, Senlei Xu, Yuan Tian 0008, Xingyu Ji, Xiaobing Sun 0001, Shujuan Jiang |
Comput. Secur. | 5 |
| 2024 | Deterministic fuzzy two-dimensional on-line tessellation automata and their languages
Xiaobing Sun 0001, Qingyu He, Yongming Li 0001 |
Fuzzy Sets Syst. | 2 |
| 2024 | Relative approximate bisimulations for fuzzy picture automata
Ruiling Wu, Xiaobing Sun 0001, Yongming Li 0001 |
Inf. Comput. | 3 |
| 2024 | EXVul: Toward Effective and Explainable Vulnerability Detection for IoT DevicesabstractAs with anything connected to the internet, Internet of Things (IoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of Deep Learning (DL)-based approaches, the lack of well-labeled IoT vulnerability samples available for training and explainability pose a critical challenge to deploy them in practice. In this paper, we propose, a novel DL-based approach for Effective and eXplainable IoT VULnerability detection. Specifically, inspired by recent advances of self-supervised learning in label-expensive tasks, we propose a new combinatorial contrastive loss to combine the strengths of large-scale unlabeled code corpus and limited IoT vulnerability samples. Then, given a binary detection result, provides a set of faithful and stable code statements positively contributing to the model’s predictions as understandable explanations. Experimental results indicate that outperforms state-of-the-art baselines by 33.44%-72.91% and 19.52%-98.78% with respect to the accuracy and F1 score metrics, respectively. For vulnerability explanation, improves over the best-performing baseline explainer PGExplainer by 22.97% in MSP, 49.55% in MSR, and 48.40% in MIoU, demonstrating that the explanations provided by can correctly point out the vulnerable statements relevant to the detected vulnerabilities. Sicong Cao, Xiaobing Sun 0001, Wei Liu 0010, Di Wu 0050, Jiale Zhang 0001, Yan Li 0002, Tom H. Luan, Longxiang Gao |
IEEE Internet Things J. | 2 |
| 2024 | A new recurrent neural network based on direct discretization method for solving discrete time-variant matrix inversion with application
Yang Shi 0003, Wei Chong, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001 |
Inf. Sci. | 6 |
| 2024 | Fine-grained smart contract vulnerability detection by heterogeneous code feature learning and automated dataset construction
Jie Cai 0006, Bin Li 0006, Tao Zhang 0001, Jiale Zhang 0001, Xiaobing Sun 0001 |
J. Syst. Softw. | 5 |
| 2024 | Real-time and screen-cam robust screen watermarking
Weitong Chen 0002, Zhenhao Niu, Yanyan Xu 0003, Anja Keskinarkaus, Tapio Seppänen, Xiaobing Sun 0001 |
Knowl. Based Syst. | 7 |
| 2024 | TDFix: A lightweight tool for fixing deadlocks based on templates
Wangjie Ji, Lili Bo, Yanchi Yuan, Xiaobing Sun 0001 |
Sci. Comput. Program. | 4 |
| 2024 | Software bug localization based on optimized and ensembled deep learning modelsabstractAbstract An automated task for finding the essential buggy files among software projects with the help of a given bug report is termed bug localization. The conventional approaches suffer from the challenges of performing lexical matching. Particularly, the terms utilized for describing the bugs in the bug reports are observed to be irrelevant to the terms used in the source code files. To resolve these problems, we propose an optimized and ensemble deep learning model for software bug localization. These features are reduced by the principle component analysis (PCA). Then, they are selected by the weighted convolutional neural network (CNN) model with the support of the Modified Scatter Probability‐based Coyote Optimization Algorithm (MSP‐COA). Finally, the optimal features are subjected to the ensemble deep neural network and long short‐term memory (DNN‐LSTM), with parameter tuning by the MSP‐COA. Experimental results show that the proposed approach can achieve higher bug localization accuracy than individual models. Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Aakash Ali, Ying Wei 0012 |
J. Softw. Evol. Process. | 3 |
| 2024 | Application programming interface recommendation for smart contract using deep learning from augmented code representationabstractAbstract Application programming interface (API) recommendation plays a crucial role in facilitating smart contract development by providing developers with a ranked list of candidate APIs for specific recommendation points. Deep learning‐based approaches have shown promising results in this field. However, existing approaches mainly rely on token sequences or abstract syntax trees (ASTs) for learning recommendation point‐related features, which may overlook the essential knowledge implied in the relations between or within statements and may include task‐irrelevant components during feature learning. To address these limitations, we propose a novel code graph called pruned and augmented AST (pa‐AST). Our approach enhances the AST by incorporating additional knowledge derived from the control and data flow relations between and within statements in the smart contract code. Through this augmentation, the pa‐AST can better represent the semantic features of the code. Furthermore, we conduct AST pruning to eliminate task‐irrelevant components based on the identified flow relations. This step helps mitigate the interference caused by these irrelevant parts during the model feature learning process. Additionally, we extract the API sequence surrounding the recommendation point to provide supplementary knowledge for the model learning. The experimental results demonstrate our proposed approach achieving an average mean reciprocal rank (MRR) of 68.02%, outperforming the baselines' performance. Furthermore, through ablation experiments, we explore the effectiveness of our proposed code representation approach. The results indicate that combining pa‐AST with the API sequence yields improved performance compared with using them individually. Moreover, our AST augmentation and pruning techniques significantly contribute to the overall results. Jie Cai 0006, Qian Cai, Bin Li 0006, Jiale Zhang 0001, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 5 |
| 2024 | Automatic software vulnerability classification by extracting vulnerability triggersabstractAbstract Vulnerability classification is a significant activity in software development and software maintenance. Natural Language Processing (NLP) techniques, which utilize the descriptions in public repositories, are widely used in automatic software vulnerability classification. However, vulnerability descriptions are ordinarily short and contain many technical terms, making them difficult for machines to automatically comprehend. In this paper, we present an approach based on vulnerability triggers to automatically classify vulnerabilities. First, we extract vulnerability triggers with Bert Question and Answer (Bert Q&A). Then, we use Recurrent Convolutional Neural Networks for Text classification (TextRCNN) to classify vulnerabilities based on Common Weakness Enumeration (CWE). We statistically perform an analysis of vulnerability triggers and comprehensively evaluate the classification performance of our approach on a set of 4769 prelabeled vulnerability entries, as well as compare it with state‐of‐the‐art vulnerability classification approaches. Experiment results show that our approach can achieve a F1‐measure of 95% on extraction and 80.8% on classification. Xiaobing Sun 0001, Lili Bo, Xiaojun Wu 0001, Ying Wei 0012, Bin Li 0006 |
J. Softw. Evol. Process. | 1 |
| 2024 | FATS: Feature Distribution Analysis-Based Test Selection for Deep Learning EnhancementabstractDeep Learning has been applied to many applications across different domains. However, the distribution shift between the test data and training data is a major factor impacting the quality of deep neural networks (DNNs). To address this issue, existing research mainly focuses on enhancing DNN models by retraining them using labeled test data. However, labeling test data is costly, which seriously reduces the efficiency of DNN testing. To solve this problem, test selection strategically selected a small set of tests to label. Unfortunately, existing test selection methods seldom focus on the data distribution shift. To address the issue, this paper proposes an approach for test selection named Feature Distribution Analysis-Based Test Selection (FATS). FATS analyzes the distributions of test data and training data and then adopts learning to rank (a kind of supervised machine learning to solve ranking tasks) to intelligently combine the results of analysis for test selection. We conduct an empirical study on popular datasets and DNN models, and then compare FATS with seven test selection methods. Experiment results show that FATS effectively alleviates the impact of distribution shifts and outperforms the compared methods with the average accuracy improvement of 19.6%$\sim$69.7% for DNN model enhancement. Li Li 0124, Chuanqi Tao, Hongjing Guo, Xiaobing Sun 0001 |
IEEE Trans. Big Data | 5 |
| 2024 | BadCleaner: Defending Backdoor Attacks in Federated Learning via Attention-Based Multi-Teacher DistillationabstractAs a privacy-preserving distributed learning paradigm, federated learning (FL) has been proven to be vulnerable to various attacks, among which backdoor attack is one of the toughest. In this attack, malicious users attempt to embed backdoor triggers into local models, resulting in the crafted inputs being misclassified as the targeted labels. To address such attack, several defense mechanisms are proposed, but may lose the effectiveness due to the following drawbacks. First, current methods heavily rely on massive labeled clean data, which is an impractical setting in FL. Moreover, an in-avoidable performance degradation usually occurs in the defensive procedure. To alleviate such concerns, we proposeBadCleaner, a lossless and efficient backdoor defense scheme via attention-based federated multi-teacher distillation. Firstly,BadCleanercan effectively tune the backdoored joint model without performance degradation, by distilling the in-depth knowledge from multiple teachers with only a small part of unlabeled clean data. Secondly, to fully eliminate the hidden backdoor patterns, we present an attention transfer method to alleviate the attention of models to the trigger regions. The extensive evaluation demonstrates thatBadCleanercan reduce the success rates of state-of-the-art backdoor attacks without compromising the model performance. Jiale Zhang 0001, Chunpeng Ge 0001, Chuan Ma 0001, Yanchao Zhao, Xiaobing Sun 0001, Bing Chen 0002 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | FLPurifier: Backdoor Defense in Federated Learning via Decoupled Contrastive TrainingabstractRecent studies have demonstrated that backdoor attacks can cause a significant security threat to federated learning. Existing defense methods mainly focus on detecting or eliminating the backdoor patterns after the model is backdoored. However, these methods either cause model performance degradation or heavily rely on impractical assumptions, such as labeled clean data, which exhibit limited effectiveness in federated learning. To this end, we proposeFLPurifier, a novel backdoor defense method in federated learning that can effectively purify the possible backdoor attributes before federated aggregation. Specifically,FLPurifiersplits a complete model into a feature extractor and classifier, in which the extractor is trained in a decoupled contrastive manner to break the strong correlation between trigger features and the target label. Compared with existing backdoor mitigation methods,FLPurifierdoesn’t rely on impractical assumptions since it can effectively purify the backdoor effects in the training process rather than an already trained model. Moreover, to decrease the negative impact of backdoored classifiers and improve global model accuracy, we further design an adaptive classifier aggregation strategy to dynamically adjust the weight coefficients. Extensive experimental evaluations on six benchmark datasets demonstrate thatFLPurifieris effective against known backdoor attacks in federated learning with negligible performance degradation and outperforms the state-of-the-art defense methods. Jiale Zhang 0001, Xiaobing Sun 0001, Chunpeng Ge 0001, Bing Chen 0002, Willy Susilo, Shui Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Hierarchy-Aware Representation Learning for Industrial IoT Vulnerability ClassificationabstractAs with anything connected to the internet, industrial Internet of Things (IIoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of deep learning-based approaches, most solutions can only detect the presence of a vulnerability but fail to pinpoint its corresponding type. Recently, TreeVul formalizes the task as a hierarchical multilabel classification problem to predict complete coarse-to-fine vulnerability type hierarchy. Yet, the TreeVul approach is still inaccurate and neglects samples labeled at coarse categories. In this article, we proposeHierVul, a novel hierarchy-aware representation learning approach for IIoT vulnerability classification. Specifically, to make full use of vulnerable samples labeled at any granularity,HierVulconstructs hierarchy-specific extractors as well as classifiers to disentangle level-wise vulnerability features from the code representation learning network backbone, and maximizes their marginal probability in the probability space constrained by the Common Weakness Enumeration tree hierarchy. Furthermore, considering that the distinction between two vulnerability types at the same level of abstraction becomes smaller and smaller as the refinement of classification granularity,HierVulleverages residual connections to add parent-level coarser-grained features to child-level finer-grained features to transfer hierarchical knowledge across levels. The experimental results show thatHierVulachieves 15.25%, 45.16%, and 14.52% relative improvement over TreeVul on Weight F1, Macro F1, and PF, respectively, indicating the effectiveness ofHierVulin the practical scenario. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Wei Liu 0010, Bin Li 0006 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Neurodynamics for Equality-Constrained Time-Variant Nonlinear Optimization Using DiscretizationabstractTime-variant problems are widespread in science and engineering, and discrete-time recurrent neurodynamics (DTRN) method has been proved to be an effective way to deal with a variety of discrete time-variant problems. However, this DTRN method is usually based on the study of continuous time-variant problems and lacks a direct study of discrete time-variant problems. To solve the abovementioned problem, based on a pioneering direct discretization technique, we study and develop a new DTRN method to solve equality-constrained discrete time-variant nonlinear optimization (EC-DTVNO) problem. Specifically, first, to solve the EC-DTVNO problem, the recent method widely used by researchers is Lagrange multiplier method. By introducing Lagrange multiplier to construct Lagrange function, the objective function and equality constraint are integrated into a discrete time-variant nonlinear system. Then, the corresponding error function is defined, and the corresponding DTRN method for solving the EC-DTVNO problem can be obtained by direct discretization technique. Thereafter, this DTRN method is analyzed theoretically and its convergence is proved. In addition, numerical experiments and application experiments further confirm the effectiveness and superiority of DTRN method. Yang Shi 0003, Wangrong Sheng, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Learning to Detect Memory-related VulnerabilitiesabstractMemory-related vulnerabilities can result in performance degradation or even program crashes, constituting severe threats to the security of modern software. Despite the promising results of deep learning (DL)-based vulnerability detectors, there exist three main limitations: (1) rich contextual program semantics related to vulnerabilities have not yet been fully modeled; (2) multi-granularity vulnerability features in hierarchical code structure are still hard to be captured; and (3) heterogeneous flow information is not well utilized. To address these limitations, in this article, we propose a novel DL-based approach, called MVD+ , to detect memory-related vulnerabilities at the statement-level. Specifically, it conducts both intraprocedural and interprocedural analysis to model vulnerability features, and adopts a hierarchical representation learning strategy, which performs syntax-aware neural embedding within statements and captures structured context information across statements based on a novel Flow-Sensitive Graph Neural Networks, to learn both syntactic and semantic features of vulnerable code. To demonstrate the performance, we conducted extensive experiments against eight state-of-the-art DL-based approaches as well as five well-known static analyzers on our constructed dataset with 6,879 vulnerabilities in 12 popular C/C++ applications. The experimental results confirmed that MVD+ can significantly outperform current state-of-the-art baselines and make a great trade-off between effectiveness and efficiency. Sicong Cao, Xiaobing Sun 0001, Lili Bo, Rongxin Wu, Bin Li 0006, Xiaoxue Wu 0001, Chuanqi Tao, Tao Zhang 0001, Wei Liu 0010 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | Ponzi Scheme Detection in Smart Contract via Transaction Semantic Representation LearningabstractThe Ponzi scheme implemented through smart contracts is one of the most common scams on the blockchain platform. Although various learning-based Ponzi smart contract detection approaches have been proposed, they still suffer from several limitations, i.e., 1) extracting insufficient semantics and gathering Ponzi irrelevant components from the smart contract during feature engineering, and 2) underutilizing structured semantic features during model training. As the Ponzi scheme is an economic crime with the typical Rob-Peter-to-Pay-Paul transaction pattern, we propose a transaction semantic learning based approach to mitigate the above limitations. The fundamental idea of our approach is to represent the transaction-related semantics of a smart contract as a graph and utilize a graph convolutional network (GCN) to learn the potential Ponzi-like transaction pattern from it. We define a novel code representation named slice transaction property graph (sTPG) to represent the transaction-related semantics, which can encode multiple transaction-related semantics inside a smart contract function into a graph and eliminate other irrelevant fragments. Then, we propose a relation-sensitive GCN as the learning model to identify potential Ponzi-scheme-like transaction patterns from sTPG by considering both nodes and edges features in sTPG. We evaluate our approach on two datasets: 1) smart contracts collected from Forum and Public datasets, and 2) really deployed smart contracts on the Ethereum blockchain. The experiment results show that our approach outperforms the state-of-the-art learning-based approaches. Jie Cai 0006, Bin Li 0006, Jiale Zhang 0001, Xiaobing Sun 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | GrabPhisher: Phishing Scams Detection in Ethereum via Temporally Evolving GNNsabstractPhishing scams are one of Ethereum's most representative security risks that can defraud many transactions in a short period and severely threaten network security. Existing deep learning-based phishing scam detection methods mainly rely on constructing static transaction graphs which are assumed to be accessible before model training. However, static methods that have a high false positive rate to detect newly generated phishing scams by adding this newly generated data to existing algorithms for execution, due to new accounts and transactions constantly appearing in the real-world Ethereum network. Therefore, this article, for the first time, proposes a novel evolve-based phishing scams detection method (named GrabPhisher) that extracts temporal features of accounts and captures information about the dynamic topology of the graph as it evolves. Specifically, GrabPhisher can build the evolutionary pattern of accounts trading on Ethereum as a diffusion network graph in continuous time. It can continue to capture new transaction features based on existing transactions, which facilitates the identification of phishing accounts. Additionally, we implement GrabPhisher on the real-world Ethereum phishing scams datasets. Extensive experimental results demonstrate that GrabPhisher can effectively extract dynamic temporal features and outperform state-of-the-art methods (95% Recall, and 88% F1-score). Jiale Zhang 0001, Hao Sui 0003, Xiaobing Sun 0001, Chunpeng Ge 0001, Lu Zhou 0002, Willy Susilo |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Real-Time Tracking Control and Efficiency Analyses for Stewart Platform Based on Discrete-Time Recurrent Neural Networkabstractrgb0.00,0.00,0.00 In recent years, the discrete-time recurrent neural network (DTRNN) model has received growing attention. This fully benefits from the recurrent neural networks (RNNs) that not only have plenty of advantages for solving computing problems in the real-time tracking control but also have the remarkable potential of parallel processing and nonlinear processing. However, there is a general lack of research on the applicability of DTRNN model to handle parallel robot. In addition, the precision is always an important point in real-time tracking control, and most of existing studies generally lack the elaborate researches on the precision analyses. In this article, the corresponding DTRNN model (i.e., general five-instant discretization (FID) formula DTRNN model) with parameter selection method is established. As one of the important theoretical contributions, the dominant term of truncation error of discretization formula and the conditions of maintaining precision of corresponding DTRNN model are proved from the mathematical view strictly. Besides, the influence of the selected parameter for the precision of such a DTRNN model is also analyzed. Finally, the above theoretical analyses are verified in the tracking control experiments of the Stewart platform, which is a widely used and representative parallel robot. Yang Shi 0003, Wangrong Sheng, Jie Wang 0091, Long Jin 0001, Bin Li 0006, Xiaobing Sun 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2023 | An Intelligent Duplicate Bug Report Detection Method Based on Technical Term ExtractionabstractAs the bug description data generated during the software maintenance cycle, bug reports are usually hastily written by different users, resulting in many redundant and duplicate bug reports (DBRs). Once the DBRs are repeatedly assigned to developers, it will inevitably lead to a serious waste of human resources, especially for large-scale open-source projects. Recently, many experts and scholars have devoted themselves to researching the detection of DBRs and put forward a series of detection methods for DBRs. However, there is still much room for improvement in the performance of DBR prediction. Therefore, this paper proposes a new method for detecting DBR based on technical term extraction, CTEDB (Combination of Term Extraction and DeBERTaV3) for short. This method first extracts technical terms from the text information of bug reports based on Word2Vec and TextRank algorithms. Then it calculates the semantic similarity of technical terms between different bug reports by combining Word2Vec and SBERT models. Finally, it completes the DBR detection task by combining the DeBERTaV3 model. The experimental results show that CTEDB has achieved good results in detecting DBR, and has obviously improved the accuracy, F1-score, recall and precision compared with the baseline approaches. Xiaoxue Wu 0001, Wenjing Shan, Wei Zheng 0006, Xiaobing Sun 0001 |
AST | 6 |
| 2023 | Improving Java Deserialization Gadget Chain Mining via Overriding-Guided Object GenerationabstractJava (de)serialization is prone to causing security-critical vulnerabilities that attackers can invoke existing methods (gadgets) on the application's classpath to construct a gadget chain to perform malicious behaviors. Several techniques have been proposed to statically identify suspicious gadget chains and dynamically generate injection objects for fuzzing. However, due to their incomplete support for dynamic program features (e.g., Java runtime polymorphism) and ineffective injection object generation for fuzzing, the existing techniques are still far from satisfactory. In this paper, we first performed an empirical study to investigate the characteristics of Java deserialization vulnerabilities based on our manually collected 86 publicly known gadget chains. The empirical results show that 1) Java deserialization gadgets are usually exploited by abusing runtime polymorphism, which enables attackers to reuse serializable overridden methods; and 2) attackers usually invoke exploitable overridden methods (gadgets) via dynamic binding to generate injection objects for gadget chain construction. Based on our empirical findings, we propose a novel gadget chain mining approach, GCMiner, which captures both explicit and implicit method calls to identify more gadget chains, and adopts an overriding-guided object generation approach to generate valid injection objects for fuzzing. The evaluation results show that GCMiner significantly outperforms the state-of-the-art techniques, and discovers 56 unique gadget chains that cannot be identified by the baseline approaches. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Lili Bo, Bin Li 0006, Rongxin Wu, Wei Liu 0010, Biao He 0002, Yu Ouyang |
ICSE | 2 |
| 2023 | RepresentThemAll: A Universal Learning Representation of Bug ReportsabstractDeep learning techniques have shown promising performance in automated software maintenance tasks associated with bug reports. Currently, all existing studies learn the customized representation of bug reports for a specific downstream task. Despite early success, training multiple models for multiple downstream tasks faces three issues: complexity, cost, and compatibility, due to the customization, disparity, and uniqueness of these automated approaches. To resolve the above challenges, we propose RepresentThemAll, a pre-trained approach that can learn the universal representation of bug reports and handle multiple downstream tasks. Specifically, RepresentThemAll is a universal bug report framework that is pre-trained with two carefully designed learning objectives: one is the dynamic masked language model and another one is a contrastive learning objective, “find yourself”. We evaluate the performance of RepresentThemAll on four downstream tasks, including duplicate bug report detection, bug report summarization, bug priority prediction, and bug severity prediction. Our experimental results show that RepresentThemAll outperforms all baseline approaches on all considered downstream tasks after well-designed fine-tuning. Sen Fang, Tao Zhang 0001, Youshuai Tan, He Jiang 0001, Xin Xia 0001, Xiaobing Sun 0001 |
ICSE | 6 |
| 2023 | ODDFuzz: Discovering Java Deserialization Vulnerabilities via Structure-Aware Directed Greybox FuzzingabstractJava deserialization vulnerability is a severe threat in practice. Researchers have proposed static analysis solutions to locate candidate vulnerabilities and fuzzing solutions to generate proof-of-concept (PoC) serialized objects to trigger them. However, existing solutions have limited effectiveness and efficiency.In this paper, we propose a novel hybrid solution ODDFuzz to efficiently discover Java deserialization vulnerabilities. First, ODDFuzz performs lightweight static taint analysis to identify candidate gadget chains that may cause deserialization vulnerabilities. In this step, ODDFuzz tries to locate all candidates and avoid false negatives. Then, ODDFuzz performs directed greybox fuzzing (DGF) to explore those candidates and generate PoC testcases to mitigate false positives. Specifically, ODDFuzz applies a structure-aware seed generation method to guarantee the validity of the testcases, and adopts a novel hybrid feedback and a step-forward strategy to guide the directed fuzzing.We implemented a prototype of ODDFuzz and evaluated it on the popular Java deserialization repository ysoserial. Results show that, ODDFuzz could discover 16 out of 34 known gadget chains, while two state-of-the-art baselines only identify three of them. In addition, we evaluated ODDFuzz on real-world applications including Oracle WebLogic Server, Apache Dubbo, Sonatype Nexus, and protostuff, and found six previously unreported exploitable gadget chains with five CVEs assigned. Sicong Cao, Biao He 0002, Xiaobing Sun 0001, Yu Ouyang, Chao Zhang 0008, Xiaoxue Wu 0001, Ting Su 0001, Lili Bo, Bin Li 0006, Chuanlei Ma, Tao Wei 0002 |
SP | 3 |
| 2023 | TemLock: A Lightweight Template-based Approach for Fixing Deadlocks Caused by ReentrantLockabstractReentrantLock, an alternative to Synchronized, is provided in Java5 to handle the conflicts of memory accesses in concurrent programs. However, falsely using ReentrantLock may introduce deadlocks. To fix deadlocks caused by ReentrantLock, in this paper, we propose TemLock, an approach that can detect and fix deadlocks in Java programs based on the fix templates. We detect and fix deadlocks based on the predefined templates by searching and modifying the node information in AST of the program. Experimental results show that TemLock can fix 156 out of 177 deadlocks caused by ReentrantLock in Java projects, indicating its effectiveness. The URL of this tool is https://github.com/yyc36/TemLock_ReentrantLock/tree/master. The video of our demo is available at https://www.youtube.com/watch?v=LIRcRF99ApY. Lili Bo, Yanchi Yuan, Xiaobing Sun 0001, Bin Li 0006 |
SANER | 3 |
| 2023 | Extended Abstract of Combine Sliced Joint Graph with Graph Neural Networks for Smart Contract Vulnerability DetectionabstractExisting smart contract vulnerability detection efforts heavily rely on fixed rules defined by experts, which are inefficient and inflexible. To overcome the limitations of existing vulnerability detection approaches, we propose a GNN based approach. First, we construct a graph representation for a smart contract function with syntactic and semantic features by combining abstract syntax tree (AST), control flow graph (CFG), and program dependency graph (PDG). To further strengthen the presentation ability of our approach, we perform program slicing to normalize the graph and eliminate the redundant information unrelated to vulnerabilities. Then, we use a Bidirectional Gated Graph Neural-Network model with hybrid attention pooling to identify potential vulnerabilities in smart contract functions. Experiment results show that our approach can achieve 89.2% precision and 92.9% recall in smart contract vulnerability detection on our dataset and reveal the effectiveness and efficiency of our approach. Jie Cai 0006, Bin Li 0006, Jiale Zhang 0001, Xiaobing Sun 0001, Bing Chen 0002 |
SANER | 4 |
| 2023 | ADFL: Defending backdoor attacks in federated learning via adversarial distillation
Jiale Zhang 0001, Xiaobing Sun 0001, Bing Chen 0002, Weizhi Meng 0001 |
Comput. Secur. | 3 |
| 2023 | Automated software bug localization enabled by meta-heuristic-based convolutional neural network and improved deep neural network
Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Saifullah Memon, Saima Siraj, Ann Suwaree Ashton |
Expert Syst. Appl. | 3 |
| 2023 | VulLoc: vulnerability localization based on inducing commits and fixing commits
Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Bin Li 0006 |
Frontiers Comput. Sci. | 3 |
| 2023 | DLRegion: Coverage-guided fuzz testing of deep neural networks with region-based neuron selection strategies
Chuanqi Tao, Yali Tao, Hongjing Guo, Xiaobing Sun 0001 |
Inf. Softw. Technol. | 5 |
| 2023 | Automated event extraction of CVE descriptions
Ying Wei 0012, Lili Bo, Xiaobing Sun 0001, Bin Li 0006, Tao Zhang 0001, Chuanqi Tao |
Inf. Softw. Technol. | 3 |
| 2023 | ASSBert: Active and semi-supervised bert for smart contract vulnerability detection
Xiaobing Sun 0001, Liangqiong Tu, Jiale Zhang 0001, Jie Cai 0006, Bin Li 0006, Yu Wang 0017 |
J. Inf. Secur. Appl. | 1 |
| 2023 | Combine sliced joint graph with graph neural networks for smart contract vulnerability detection
Jie Cai 0006, Bin Li 0006, Jiale Zhang 0001, Xiaobing Sun 0001, Bing Chen 0002 |
J. Syst. Softw. | 4 |
| 2023 | Automatic software vulnerability assessment by extracting vulnerability elements
Xiaobing Sun 0001, Zhenlei Ye, Lili Bo, Xiaoxue Wu 0001, Ying Wei 0012, Tao Zhang 0001, Bin Li 0006 |
J. Syst. Softw. | 1 |
| 2023 | Leveraging multi-level embeddings for knowledge-aware bug report reformulation
Bin Li 0006, Xiaobing Sun 0001 |
J. Syst. Softw. | 3 |
| 2023 | Multi-level membership inference attacks in federated Learning based on active GAN
Hao Sui 0003, Xiaobing Sun 0001, Jiale Zhang 0001, Bing Chen 0002, Wenjuan Li 0001 |
Neural Comput. Appl. | 2 |
| 2023 | A direct discretization recurrent neurodynamics method for time-variant nonlinear optimization with redundant robot manipulators
Yang Shi 0003, Wangrong Sheng, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001, Dimitrios Gerontitis |
Neural Networks | 5 |
| 2023 | Tracking Control of Cable-Driven Planar Robot Based on Discrete-Time Recurrent Neural Network With Immediate Discretization MethodabstractIn recent years, the cable-driven planar robot has made fruitful achievements in many fields, but the related researches are scarce yet in the industrial engineering field. In this article, as a powerful tool for solving discrete time-varying problems, the discrete-time recurrent neural network (DTRNN) is extended to drive the cable-driven planar robot for discrete real-time tracking control, which is derived by a new immediate discretization method, and thus, is termed as ID-DTRNN model. Specifically, first, we present the physical structure and mathematical model of the cable-driven planar robot. Then, the new ID-DTRNN model is proposed and applied for driving such cable-driven planar robot, which bases on the a different way of construction of the traditional DTRNN model. Through numerical experiments, the feasibility, validity, and physical reliability of the ID-DTRNN model for discrete real-time tracking control of the cable-driven planar robot are fully verified. In addition, in the real world, physical experiments of the cable-driven planar robot are presented, which successfully promote the development of physical application of the ID-DTRNN model, and fill the gap of such model in the industrial engineering field. Yang Shi 0003, Jie Wang 0091, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Novel Discrete-Time Recurrent Neural Network for Robot Manipulator: A Direct Discretization Technical RouteabstractControlling and processing of time-variant problem is universal in the fields of engineering and science, and the discrete-time recurrent neural network (RNN) model has been proven as an effective method for handling a variety of discrete time-variant problems. However, such model usually originates from the discretization research of continuous time-variant problem, and there is little research on the direct discretization method. To address the aforementioned problem, this article introduces a novel discrete-time RNN model for solving the discrete time-variant problem in a pioneering manner. Specifically, a discrete time-variant nonlinear system, which originates from the mathematical modeling of serial robot manipulator, is presented as a target problem. For solving the problem, first, the technique of second-order Taylor expansion is used to deal with the discrete time-variant nonlinear system, and the novel discrete-time RNN model is proposed subsequently. Second, the theoretical analyses are investigated and developed, which shows the convergence and precision of the proposed discrete-time RNN model. Furthermore, three distinct numerical experiments verify the excellent performance of the proposed discrete-time RNN model. In addition, a robot manipulator example further verifies the effectiveness and practicability of the proposed novel discrete-time RNN model. Yang Shi 0003, Shuai Li 0002, Bin Li 0006, Xiaobing Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | CRAM: Code Recommendation With Programming Context Based on Self-Attention MechanismabstractCode recommendation with programming context is to use the contextual code surrounding the missing code to automatically find that which of the code snippets would be useful to assist in the completion of the program. In this way, the developers need not take time to formulate explicit queries or write descriptions. Existing work only treats code as textual documents and use information retrieval techniques to retrieve relevant code snippets, which is difficult to capture the semantics of code adequately. The self-attention mechanism have achieved promising progress in various natural language processing tasks, especially for extracting deep semantic information from long sequences. Inspired from this, we propose a novel code recommendation with programming context based on self-attention mechanism (CRAM). The proposed approach first builds a small-scale candidate set from codebase. Then, it utilizes self-attention networks in the abstract syntax tree to capture the deep semantics of code, and finally recommend the relevant code to developers. We conduct several experiments to evaluate our approach in a large-scale codebase containing 741 148 code snippets. The experimental results show that CRAM can effectively recommend code and outperforms related work in recall, precision, and NDCG. Chuanqi Tao, Xiaobing Sun 0001 |
IEEE Trans. Reliab. | 4 |
| 2022 | MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural NetworksabstractMemory-related vulnerabilities constitute severe threats to the security of modern software. Despite the success of deep learning-based approaches to generic vulnerability detection, they are still limited by the underutilization of flow information when applied for detecting memory-related vulnerabilities, leading to high false positives. Sicong Cao, Xiaobing Sun 0001, Lili Bo, Rongxin Wu, Bin Li 0006, Chuanqi Tao |
ICSE | 2 |
| 2022 | A Comprehensive Analysis of NVD Concurrency VulnerabilitiesabstractConcurrency vulnerabilities caused by synchronization problems will occur in the execution of multi-threaded programs, and the emergence of concurrency vulnerabilities often cause great threats to the system. Once the concurrency vulnerabilities are exploited, the system will suffer various attacks, seriously affecting its availability, confidentiality and security. In this paper, we extract 839 concurrency vulnerabilities from Common Vulnerabilities and Exposures (CVE), and conduct a comprehensive analysis of the trend, classifications, causes, severity, and impact. Finally, we obtained some findings: 1) From 1999 to 2021, the number of concurrency vulnerabilities disclosures show an overall upward trend. 2) In the distribution of concurrency vulnerability, race condition accounts for the largest proportion. 3) The overall severity of concurrency vulnerabilities is medium risk. 4) The number of concurrency vulnerabilities that can be exploited for local access and network access is almost equal, and nearly half of the concurrency vulnerabilities (377/839) can be accessed remotely. 5) The access complexity of 571 concurrency vulnerabilities is medium, and the number of concurrency vulnerabilities with high or low access complexity is almost equal. The results obtained through the empirical study can provide more support and guidance for research in the field of concurrency vulnerabilities. Lili Bo, Xing Meng, Xiaobing Sun 0001, Jingli Xia, Xiaoxue Wu 0001 |
QRS | 3 |
| 2022 | KVS: a tool for knowledge-driven vulnerability searchingabstractIt is difficult to quickly locate and search for specific vulnerabilities and their solutions because vulnerability information is scattered in the existing vulnerability management library. To alleviate this problem, we extract knowledge from vulnerability reports and organize the vulnerability information into the form of a knowledge graph. Then, we implement a tool for knowledge-driven vulnerability searching, KVS. This tool mainly uses the BERT model to realize the vulnerability named entity recognition and construct the vulnerability knowledge graph (VulKG). Finally, we can search vulnerabilities of interest-based on VulKG. The URL of this tool is https://cinnqi.github.io/Neo4j-D3-VKG/. Video of our demo is available at https://youtu.be/FT1BaLUGPk0. Xingqi Cheng, Xiaobing Sun 0001, Lili Bo, Ying Wei 0012 |
ESEC/SIGSOFT FSE | 2 |
| 2022 | Towards the identification of bug entities and relations in bug reports
Bin Li 0006, Ying Wei 0012, Xiaobing Sun 0001, Lili Bo, Dingshan Chen, Chuanqi Tao |
Autom. Softw. Eng. | 3 |
| 2022 | SPVF: security property assisted vulnerability fixing via attention-based models
Lili Bo, Xiaoxue Wu 0001, Xiaobing Sun 0001, Tao Zhang 0001, Bin Li 0006, Jiale Zhang 0001, Sicong Cao |
Empir. Softw. Eng. | 4 |
| 2022 | An approach of method-level bug localizationabstractAbstract Bug localization is an important field in software engineering research. The traditional bug localization approaches based on information retrieval separate words through lexical analysis. In this way, the comments of the source code are ignored or treated as plain text, which will lose some semantic information. In this paper, MBL_SHL, an automatic Method‐level Bug Localization approach, which utilises code Summarization, Historical fixed bugs and code Length, is presented. Based on the code summarization technology, this approach first supplements the comment for uncommented code, and then calculates the Word2vec vector and Term Frequency–Inverse Document Frequency vector for the bug report, methods and comments, respectively. After that the authors calculate separately the similarity between the bug report and each method, the bug report and each comment. The code length information and historical fix information are also considered as a weight and a part of the score, respectively, to calculate the final score of each method. Finally, the scores are sorted to determine the list of methods that may need to be modified when fixing the software bugs. We built a method‐granular bug localization dataset, which contains five open‐source projects. The experimental results show that the proposed approach significantly outperforms the existing approaches on the method level. Zhen Ni, Lili Bo, Bin Li 0006, Tianhao Chen, Xiaobing Sun 0001, Xiaoxue Wu 0001 |
IET Softw. | 5 |
| 2022 | A deep learning-based approach for software vulnerability detection using code metricsabstractAbstract Vulnerabilities can have devastating effects on information security, affecting the economy, social stability, and national security. The idea of automatic vulnerability detection has always attracted researchers. From traditional manual vulnerability mining techniques to static and dynamic detection, all rely on human experts for feature definition. The rapid development of machine learning and deep learning has alleviated the tedious task of manually defining features by human experts while reducing the lack of objectivity caused by human subjective awareness. However, it is still necessary to find an objective characterisation method to define the features of vulnerabilities. Therefore, the authors use code metrics for code characterisation, sequences of metrics representing code. To use code metrics for vulnerability detection, a deep learning‐based vulnerability detection approach that uses a composite neural network of convolutional neural network (CNN) with long short‐term memory (LSTM) is proposed. The authors conduct experiments independently using the proposed approach for CNN‐LSTM CNN, LSTM, gated recurrent units (GRU), and deep neural network (DNN). The authors’ experimental results show that CNN‐LSTM has a high precision of 92%, a recall of 99%, and an accuracy of 91%. In terms of the F1‐score, it is 95%, compared to previous research results, which indicated an improvement of 18%. Compared to other deep learning‐based vulnerability detection models, the authors’ proposed model produced a lower false‐positive rate, a lower miss rate, and improved accuracy. Fazli Subhan, Xiaojun Wu 0001, Lili Bo, Xiaobing Sun 0001, Muhammad Rahman 0005 |
IET Softw. | 4 |
| 2022 | Intelligent analysis for software data: research and applications
Tao Zhang 0001, Xiaobing Sun 0001, Zibin Zheng |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | Domain knowledge-based security bug reports prediction
Wei Zheng 0006, Jingyuan Cheng, Xiaoxue Wu 0001, Ruiyang Sun, Xiaobing Sun 0001 |
Knowl. Based Syst. | 6 |
| 2021 | GrasP: Graph-to-Sequence Learning for Automated Program RepairabstractMany deep learning models, for example, neural machine translation (NMT) models, have been developed for Automated Program Repair (APR). Due to the advantages of NMT model's strong generalization ability and less manual in-tervention, NMT-based methods perform well in APR. However, previous NMT-based APR approaches regard a code snippet as a sequence of tokens, which ignores the inherent structure of code. In this paper, we propose a novel end-to-end approach with Graph-to-Sequence learning, GrasP, to generate patches for buggy methods. To better represent the buggy method, we use a graph based on abstract syntax tree (AST) to represent the source code. In order to learn complex graph representation, we introduce the attention-based encoder-decoder model for graph-to-sequence learning. The empirical evaluation on the popular benchmark Defects4J shows that GrasP can generate compilable patches for 75 bugs, of which 34 patches are correct. Ben Tang, Bin Li 0006, Lili Bo, Xiaoxue Wu 0001, Sicong Cao, Xiaobing Sun 0001 |
QRS | 6 |
| 2021 | Why and what happened? Aiding bug comprehension with automated category and causal link identification
Bin Li 0006, Xiaobing Sun 0001, Lili Bo |
Empir. Softw. Eng. | 3 |
| 2021 | Experience report: investigating bug fixes in machine learning frameworks/libraries
Xiaobing Sun 0001, Tianchi Zhou, Rongcun Wang, Yucong Duan, Lili Bo, Jianming Chang |
Frontiers Comput. Sci. | 1 |
| 2021 | BGNN4VD: Constructing Bidirectional Graph Neural-Network for Vulnerability Detection
Sicong Cao, Xiaobing Sun 0001, Lili Bo, Ying Wei 0012, Bin Li 0006 |
Inf. Softw. Technol. | 2 |
| 2021 | Special Issue on New Generation of Bug Fixing
Xiapu Luo, Weiyi Shang, Xiaobing Sun 0001, Tao Zhang 0001 |
J. Syst. Softw. | 3 |
| 2021 | Node deletion-based algorithm for blocking maximizing on negative influence from uncertain sources
Weijia Ju, Ling Chen 0005, Bin Li 0006, Yixin Chen 0001, Xiaobing Sun 0001 |
Knowl. Based Syst. | 5 |
| 2021 | BEAT: Considering question types for bug question answering via templates
Jinting Lu, Xiaobing Sun 0001, Bin Li 0006, Lili Bo, Tao Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2021 | A comprehensive study on security bug characteristicsabstractAbstract Security bugs can catastrophically impact our increasingly digital lives. Designing effective tools for detecting and fixing software security bugs requires a deep understanding of security bug characteristics. In this paper, we conducted a comprehensive study on security bugs and proposed the classification criteria for security bug category, that is, root cause, consequence, and location. In addition, we selected 1076 bug reports from five projects (i.e., Apache Tomcat, Apache HTTP Server, Mozilla Firefox, Linux Kernel, and Eclipse) in the NVD for investigation. Finally, we investigated the correlation between the classification results and obtained some findings: (1) memory operation is the most common security bug; (2) the primary root causes of security bugs are CON (Configuration Error), INP (Input Validation Error), and MEM (Memory Error); (3) the severity of more than 40% of security bugs is high; (4) security bugs caused by INP mainly occur on web; and (5) security bugs caused by LOG (Logic Resource Error) usually lead to DoS (Denial of Service). We discussed these findings through data analysis, which can also help developers better understand the characteristics of security bugs. Ying Wei 0012, Xiaobing Sun 0001, Lili Bo, Sicong Cao, Xin Xia 0001, Bin Li 0006 |
J. Softw. Evol. Process. | 2 |
| 2021 | image2emmet: Automatic code generation from web user interface imageabstractAbstract Web development usually follows with analyzing the functionality, designing the user interface (UI) prototype, implementing the UI by front‐end (FE) developers and implementing the REpresentational State Transfer (RESTful) application programming interface (API) by back‐end (BE) programmers. Unfortunately, web development is a tedious, cumbersome, and time‐consuming task, which makes it a challenge for the FE programmers to work in an efficient way. In this paper, we propose an approach, image2emmet, to assist FE programmers in implementing the UI. First, we collect HyperText Markup Language, Cascading Style Sheets (HTML‐CSS) dataset in an automatic and efficient way. The HTML‐CSS dataset used for model training consists of HTML‐CSS code and its display images. Second, the faster region‐based convolutional neural network (CNN) (R‐CNN) is utilized to detect the UI component. Finally, we build a model combining CNN and long short‐term memory (LSTM) to transform the UI component into the HTML‐CSS code. The empirical study demonstrates that image2emmet can achieve a precision of 80% on the UI component detection and 60% on the transformation of UI component into HTML‐CSS code. Lili Bo, Xiaobing Sun 0001, Bin Li 0006, Jing Jiang 0005 |
J. Softw. Evol. Process. | 3 |
| 2021 | Transformation-based processing of typed resources for multimedia sources in the IoT environment
Honghao Gao, Yucong Duan, Lixu Shao, Xiaobing Sun 0001 |
Wirel. Networks | 4 |
| 2020 | Analyzing bug fix for automatic bug cause classificationabstractDuring the bug fixing process, developers usually need to analyze the source code to induce the bug cause, which is useful for bug understanding and localization. The bug fixes of historical bugs usually reflects the bug causes when fixing them. This paper aims at exploiting the corresponding relationship between bug causes and bug fixes to automatically classify bugs into their cause categories. First, we define the code-related bug classification criterion from the perspective of the cause of bugs. Then, we propose a new model to exploit the knowledge in the bug fix by constructing fix trees from the diff source code at Abstract Syntax Tree (AST) level, and representing each fix tree based on the encoding method of Tree-based Convolutional Neural Network (TBCNN). Finally, the corresponding relationship between bug causes and bug fixes is analyzed by automatically classifying bugs into their cause categories. We collected 2000 real-world bugs from two open source projects Mozilla and Radare2 to evaluate our approach. The experimental results show the existence of observational correlation between the bug fix and the cause of the historical bugs, and the proposed fix tree can effectively express the characteristics of the historical bugs for bug cause classification. Zhen Ni, Bin Li 0006, Xiaobing Sun 0001, Tianhao Chen, Ben Tang, Xinchen Shi |
J. Syst. Softw. | 3 |
| 2020 | Improving software bug-specific named entity recognition with deep neural network
Bin Li 0006, Xiaobing Sun 0001 |
J. Syst. Softw. | 3 |
| 2019 | How security bugs are fixed and what can be improved: an empirical study with Mozilla
Xiaobing Sun 0001, Xin Peng 0001, Yang Liu 0003, Yuanfang Cai |
Sci. China Inf. Sci. | 1 |
| 2019 | Data Privacy Protection for Edge Computing of Smart City in a DIKW Architecture
Yucong Duan, Zhihui Lu 0002, Zhangbing Zhou, Xiaobing Sun 0001, Jie Wu 0003 |
Eng. Appl. Artif. Intell. | 4 |
| 2019 | Improving defect prediction with deep forest
Tianchi Zhou, Xiaobing Sun 0001, Xin Xia 0001, Bin Li 0006, Xiang Chen 0005 |
Inf. Softw. Technol. | 2 |
| 2018 | Recognizing software bug-specific named entity in software bug repositoryabstractSoftware bug issues are unavoidable in software development and maintenance. In order to manage bugs effectively, bug tracking systems are developed to help to record, manage and track the bugs of each project. The rich information in the bug repository provides the possibility of establishment of entity-centric knowledge bases to help understand and fix the bugs. However, existing named entity recognition (NER) systems deal with text that is structured, formal, well written, with a good grammatical structure and few spelling errors, which cannot be directly used for bug-specific named entity recognition. For bug data, they are free-form texts, which include a mixed language studded with code, abbreviations and software-specific vocabularies. In this paper, we summarize the characteristics of bug entities, propose a classification method for bug entities, and build a baseline corpus on two open source projects (Mozilla and Eclipse). On this basis, we propose an approach for bug-specific entity recognition called BNER with the Conditional Random Fields (CRF) model and word embedding technique. An empirical study is conducted to evaluate the accuracy of our BNER technique, and the results show that the two designed baseline corpus are suitable for bug-specific named entity recognition, and our BNER approach is effective on cross-projects NER. Bin Li 0006, Xiaobing Sun 0001, Hongjing Guo |
ICPC | 3 |
| 2018 | Towards Cost Effective Privacy Provision for Typed Resources in IoT Environment (S)abstractWe present privacy resources in IoT as data, information, and knowledge.We construct a privacy protection architecture on our previously proposed DIKW graphs: Data Graph, Information Graph, and Knowledge Graph.On this architecture, we search privacy protection target resources both as they appear explicitly in their original types and as they appear implicitly which means that they are expressed not in their original types.For a single privacy protection target, it may have various concrete compositions in various layers of DIKW Graph.It becomes more complex since the implementation of a privacy target might also be intertwined with the implementation of other privacy targets.We propose to protect target resources according to their types by either isolating the elements comprising an implementation, or weakening relationships among elements comprising an implement.To optimize among several choices of implementing a protection in a business environment, we introduced the tradeoff between customers' expectations/investment and privacy providers' expectation.Thereafter we proposed to prioritize implementation according to their ratio of cost/benefit. Yucong Duan, Zhengyang Song, Xiaoxian Yang, Quan Zou 0001, Xiaobing Sun 0001 |
SEKE | 5 |
| 2018 | Personalized project recommendation on GitHub
Xiaobing Sun 0001, Wenyuan Xu 0006, Xin Xia 0001, Xiang Chen 0005, Bin Li 0006 |
Sci. China Inf. Sci. | 1 |
| 2018 | Effectiveness of exploring historical commits for developer recommendation: an empirical study
Xiaobing Sun 0001, Hareton K. N. Leung, Bin Li 0006, Hanchao Jerry Li, Lingzhi Liao |
Frontiers Comput. Sci. | 1 |
| 2018 | MULAPI: Improving API method recommendation with API usage location
Congying Xu, Xiaobing Sun 0001, Bin Li 0006, Hongjing Guo |
J. Syst. Softw. | 2 |
| 2018 | Processing Optimization of Typed Resources with Synchronized Storage and Computation Adaptation in Fog ComputingabstractWide application of the Internet of Things (IoT) system has been increasingly demanding more hardware facilities for processing various resources including data, information, and knowledge. With the rapid growth of generated resource quantity, it is difficult to adapt to this situation by using traditional cloud computing models. Fog computing enables storage and computing services to perform at the edge of the network to extend cloud computing. However, there are some problems such as restricted computation, limited storage, and expensive network bandwidth in Fog computing applications. It is a challenge to balance the distribution of network resources. We propose a processing optimization mechanism of typed resources with synchronized storage and computation adaptation in Fog computing. In this mechanism, we process typed resources in a wireless‐network‐based three‐tier architecture consisting of Data Graph, Information Graph, and Knowledge Graph. The proposed mechanism aims to minimize processing cost over network, computation, and storage while maximizing the performance of processing in a business value driven manner. Simulation results show that the proposed approach improves the ratio of performance over user investment. Meanwhile, conversions between resource types deliver support for dynamically allocating network resources. Zhengyang Song, Yucong Duan, Shixiang Wan, Xiaobing Sun 0001, Quan Zou 0001, Honghao Gao, Donghai Zhu |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | An Empirical Study on Real Bugs for Machine Learning ProgramsabstractDue to the availability of various open source Machine Learning (ML) tools and libraries, developers nowadays can easily implement their purposes by just invoking machine learning APIs without knowing the details of the algorithm. However, the owners of ML tools and libraries usually pay more attention to the correctness and functionality of their algorithm, while spending much less effort on maintaining their code and keeping their code at a high quality level. Considering the popularity of machine learning in today's world, low quality ML tools and libraries can have a huge impact on the software products that use ML algorithms. So in this paper, we conduct an empirical study on real machine learning bugs to examine their patterns and how they evolve over time. We collect three popular machine learning projects on Github, and manually analyzed 329 closed bugs from the perspectives of their bug category, fix pattern, fix scale, fix duration, and type of software maintenance. The results show that (1) there are seven categories of bugs in machine learning programs; (2) twelve different fix patterns are commonly used to fix the bugs; (3) 63.83% of the patches belong to micro-scale-fix and small-scale-fix, and 68.39% of the bugs are fixed within one month; (4) 47.77% of the bug fixes belong to corrective activity from the view of software maintenance. Xiaobing Sun 0001, Tianchi Zhou, Gengjie Li, Jiajun Hu, Bin Li 0006 |
APSEC | 1 |
| 2017 | Constructing Search as a Service Towards Non-deterministic and Not Validated Resource Environment with a Positive-Negative Strategy
Yucong Duan, Lixu Shao, Xiaobing Sun 0001, Li-Zhen Cui 0001, Donghai Zhu, Zhengyang Song |
CollaborateCom | 3 |
| 2017 | REPERSP: Recommending Personalized Software Projects on GitHubabstractIn the open source community such as GitHub, developers usually need to find projects similar to their work, with the aim to reuse their functions and explore ideas of features that could be possibly added into their project at hand. Traditional text search engine can help detect similar resources. However, it is difficult for developers to use in open source community because a few query words cannot describe the whole features of a project. In this paper, we present a practical software recommendation system, REPERSP, which is used to recommend personalized software projects in GitHub. According to the features of projects created by developers and their behavior to other known projects, REPERSP recommends the top N relevant and personalized software projects. Moreover, REPERSP is implemented with the MapReduce parallel processing frame - Apache Spark for large-scale data, which can be scaled to a large number of users and projects for practical usage. Empirical results show that REPERSP can recommend more accurate results compared with other two recommendation algorithms, i.e., UserCF (user collaborative filtering) and ItemCF (item collaborative filtering). Video of our demo is available at https://youtu.be/WKigSUV4UA0. Wenyuan Xu 0006, Xiaobing Sun 0001, Jiajun Hu, Bin Li 0006 |
ICSME | 2 |
| 2017 | An Investment Defined Transaction Processing Towards Temporal and Spatial Optimization with Collaborative Storage and Computation Adaptation
Yucong Duan, Lixu Shao, Xiaobing Sun 0001, Donghai Zhu, Xiaoxian Yang, Abdelrahman Osman Elfaki |
IDEAL | 3 |
| 2017 | A Pay as You Use Resource Security Provision Approach Based on Data Graph, Information Graph and Knowledge Graph
Lixu Shao, Yucong Duan, Li-Zhen Cui 0001, Quan Zou 0001, Xiaobing Sun 0001 |
IDEAL | 5 |
| 2017 | Scalable Relevant Project Recommendation on GitHubabstractGitHub, one of the largest social coding platforms, fosters a flexible and collaborative development process. In practice, developers in the open source software platform need to find projects relevant to their development work to reuse their function, explore ideas of possible features, or analyze the requirements for their projects. Recommending relevant projects to a developer is a difficult problem considering that there are millions of projects hosted on GitHub, and different developers may have different requirements on relevant projects. In this paper, we propose a scalable and personalized approach to recommend projects by leveraging both developers' behaviors and project features. Based on the features of projects created by developers and their behaviors to other projects, our approach automatically recommends top N most relevant software projects to developers. Moreover, to improve the scalability of our approach, we implement our approach in a parallel processing frame (i.e., Apache Spark) to analyze large-scale data on GitHub for efficient recommendation. We perform an empirical study on the data crawled from GitHub, and the results show that our approach can efficiently recommend relevant software projects with a relatively high precision fit for developers' interests. Wenyuan Xu 0006, Xiaobing Sun 0001, Xin Xia 0001, Xiang Chen 0005 |
Internetware | 2 |
| 2017 | Answering Who/When, What, How, Why through Constructing Data Graph, Information Graph, Knowledge Graph and Wisdom GraphabstractKnowledge graphs have been widely adopted, in large part owing to their schema-less nature.It enables knowledge graphs to grow seamlessly and allows for new relationships and entities as needed.Natural language questions are the most intuitive way of formulating an information need.People can formulate questions to express their information needs.Natural language questions as a query language present an ideal compromise between keyword and structured querying.Questions can be used to express complex information needs that cannot be expressed as keywords without a significant loss in structure and semantics.Knowledge graph has abundant natural semantics and can contain various and more complete information.Its expression mechanism is closer to natural language.We propose to clarify the expression of knowledge graph as a whole.We use knowledge graph to solve the Five Ws problems respectively which are guided by interrogative words such as who/when, what, how and why.We also propose to specify knowledge graph in a progressive manner as four basic forms including data graph, information graph, knowledge graph and wisdom graph. Lixu Shao, Yucong Duan, Xiaobing Sun 0001, Honghao Gao, Donghai Zhu, Weikai Miao |
SEKE | 3 |
| 2017 | Bidirectional value driven design between economical planning and technical implementation based on data graph, information graph and knowledge graphabstractValue-Driven Design enables rational decisions to be made in terms of the optimum business and technical solution at every level of engineering design by employing economics in decision making. In order to maximize the business profitability, we propose to bridge bidirectional value driven design between economic planning and technology implementation on the basis of the data graph, information graph and knowledge graph. We use data graph, information graph and knowledge graph to analyze problems that have negative impact on activities of software development including requirement analysis, summary design and detail design. We propose to improve system reliability and robustness by managing data and information reuse, redundancy as well as structure. Lixu Shao, Yucong Duan, Xiaobing Sun 0001, Quan Zou 0001, Rongqi Jing, Jiami Lin |
SERA | 3 |
| 2017 | Enhancing developer recommendation with supplementary information via mining historical commits
Xiaobing Sun 0001, Xin Xia 0001, Bin Li 0006 |
J. Syst. Softw. | 1 |
| 2016 | Refinement from service economics planning to ubiquitous services implementationabstractService Economics has been successfully proposed and implemented for promoting the macro service market especially in Global value chains which however does not provide guidance and integrate directly with the IT side implementation. Therefore there leaves a gap of the refinement from the strategic business planning to ubiquitous services implementation in the API (Application Program Interface) economy. Recently Gio Wiederhold attributed this gap to the knowledge difference between economists and the IT professionals and proposed to solve it from the IT side. As an instance of this guidance in the practice of Value Driven Design, we proposed a systemic formalization from the value calculation to the design quality measurement which binds the modification and change on the design artifacts with the business value strategy through a framework of managed quality properties in a service design process. Yucong Duan, Gongzhu Hu, Xiaobing Sun 0001 |
ICIS | 4 |
| 2016 | On Automatic Summarization of What and Why Information in Source Code ChangesabstractAccurate and complete commit messages summarizing software changes are important to support various software maintenance activities. In practice, these commit messages are often manually submitted by individual software developer to provide information about the changes involved in the incremental changes. Hence, the content and quality of these commit messages may be different. For example, some commit messages are too short and lack of essential information while others with too much detailed information can be time-consuming to read. What's more, most of the commit messages focus on what has been changed by developers in a commit, but why they changed and the motivation behind the code changes (which can assist developers in understanding code changes), are usually ignored. In this paper, we present an approach that can automatically generate the commit messages related to the code changes, including not only what have been changed but also why they were changed. Our approach uses method stereotypes and the type of changes to generate commit messages. We evaluate our approach by comparing the quality of generated messages with the original commit messages written by the original developers and those generated by a state-of-art technique, i.e., ChangeScribe. The results demonstrate that the messages generated by our approach are preferred in about 69% of the cases. Jinfeng Shen, Xiaobing Sun 0001, Bin Li 0006, Jiajun Hu |
COMPSAC | 2 |
| 2016 | DR_PSF: Enhancing Developer Recommendation by Leveraging Personalized Source-Code FilesabstractGiven a new issue request, suitable developers should be arranged to implement it. Technologies such as developer recommendation have been proposed to tackle this issue. These techniques tend to recommend experienced developers, i.e., the more experienced the developer is, the more possible he/she is recommended. However, if the experienced developers are hectic, the junior developers may be employed to finish the incoming issue. But they may have difficulty in finishing these tasks for lack of developing experience. In this paper, we propose a novel approach, DR_PSF (Developer Recommendation with Personalized Source-code Files), to enhance developer recommendation by leveraging personalized source-code files. DR_PSF uses the collaborative topic modeling (CTM) technique to analyze developer expertise and triage some personalized files for the recommended developers. An empirical study is conducted, and the results show that DR_PSF can effectively recommend useful personalized source-code files for them to refer when they implement the incoming issue. Xiaobing Sun 0001, Bin Li 0006, Yucong Duan |
COMPSAC | 2 |
| 2016 | Enhancing UML Class Diagram Abstraction with Knowledge Graph
Yucong Duan, Xiaobing Sun 0001, Zhaoxin Lin, Chuanpu Zhu |
IDEAL | 3 |
| 2016 | On Expanding Abbreviated Identifiers in the Source Code
Xiaobing Sun 0001, Yucong Duan, Bin Li 0006 |
IDEAL | 2 |
| 2016 | WB4SP: A tool to build the word base for specific programsabstractSoftware becomes increasingly complex with its continuous maintenance activities. Given a system under maintenance, developers used to employing code search techniques to locate the code of their interests. However, they may have difficulties in understanding the source code elements and the relationship among them in the searching results. If there is a word base for a specific system, the developers can refer it to help locate and recover the source code elements and their relationships, which can improve the maintenance efficiency. In this paper, we present a tool, WB4SP(Word Base for Specific Programs), which focuses on building the word base for a specific system. WB4SP can retrieve the words, recover the relationship between them, and display the evolution of these words during the software evolution. Weisong Sun, Xiaobing Sun 0001, Bin Li 0006 |
ICPC | 2 |
| 2016 | Exploring topic models in software engineering data analysis: A surveyabstractTopic models are shown to be effective to mine unstructured software engineering (SE) data. In this paper, we give a simple survey of exploring topic models to support various SE tasks between 2003 and 2015. The survey results show that there is an increasing concern in this area. Among the SE tasks, source code comprehension and software history comprehension are the mostly studied, followed by software defects prediction. However, there is still only a few studies on other SE tasks, such as feature location and regression testing. Xiaobing Sun 0001, Xiangyue Liu 0002, Bin Li 0006, Yucong Duan, Jiajun Hu |
SNPD | 1 |
| 2016 | IPSETFUL: an iterative process of selecting test cases for effective fault localization by exploring concept lattice of program spectra
Xiaobing Sun 0001, Xin Peng 0001, Bin Li 0006, Bixin Li, Wanzhi Wen |
Frontiers Comput. Sci. | 1 |
| 2016 | Code Comment Quality Analysis and Improvement Recommendation: An Automated ApproachabstractProgram comprehension is one of the first and most frequently performed activities during software maintenance and evolution. In a program, there are not only source code, but also comments. Comments in a program is one of the main sources of information for program comprehension. If a program has good comments, it will be easier for developers to understand it. Unfortunately, for many software systems, due to developers’ poor coding style or hectic work schedule, it is often the case that a number of methods and classes are not written with good comments. This can make it difficult for developers to understand the methods and classes, when they are performing future software maintenance tasks. To deal with this problem, in this paper we propose an approach which assesses the quality of a code comment and generates suggestions to improve comment quality. A user study is conducted to assess the effectiveness of our approach and the results show that our comment quality assessments are similar to the assessments made by our user study participants, the suggestions provided by our approach are useful to improve comment quality, and our approach can improve the accuracy of the previous comment quality analysis approaches. Xiaobing Sun 0001, David Lo 0001, Yucong Duan, Xiangyue Liu 0002, Bin Li 0006 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2016 | ComboRT: A New Approach for Generating Regression Test Cases for Evolving ProgramsabstractRegression testing is essential to ensure software quality during software evolution. Two widely-used regression testing techniques, test case selection and prioritization, are used to maximize the value of the continuously enlarging test suite. However, few works consider both these two techniques together, which decreases the usefulness of the independently studied techniques in practice. In the presence of changes during program evolution, regression testing is usually conducted by selecting the test cases that cover the impact results of the changes. It seldom considers the false-positives in the information covered. Hence, the effectiveness of such regression testing techniques is decreased. In this paper, we propose an approach, ComboRT, which combines test case selection and prioritization together to directly generate a ranked list of test cases. It is based on the impact results predicted by the change impact analysis (CIA) technique, FCA–CIA, which generates a ranked list of impacted methods. Test cases which cover these impacted methods are included in the new test suite. As each method predicted by FCA–CIA is assigned with an impact factor value corresponding to the probability of this method to be impacted, test cases are then ordered according to the impact factor values of the impacted methods. Empirical studies on four Java based software systems demonstrate that ComboRT can be effectively used for regression testing in object-oriented Java-based software systems during their evolution. Xiaobing Sun 0001, Xin Peng 0001, Hareton K. N. Leung, Bin Li 0006 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2015 | Everything as a Service (XaaS) on the Cloud: Origins, Current and Future TrendsabstractFor several years now, scientists have been proposing numerous models for defining anything "as a service (aaS)", including discussions of products, processes, data & information management, and security as a service. In this paper, based on a thorough literature survey, we investigate the vast stream of the state of the art in Everything as a Service (XaaS). We then use this investigation to explore an integrated view of XaaS that will help propose approaches for migrating applications to the cloud and exposing them as services. Yucong Duan, Guohua Fu, Nianjun Zhou, Xiaobing Sun 0001, Nanjangud C. Narendra |
CLOUD | 4 |
| 2015 | Analyzing program readability based on WordNetabstractComments to describe the intent of the code is crucial to measure the program readability, especially for the methods and their comments in a program. Existing program readability techniques mainly focus on matching method and its comments on whether there is the same content between them. But these techniques cannot accurately analyze polysemy and synonyms in the program. In this paper, we propose an approach to analyze program readability based on WordNet, which is able to expand the range of keyword search and solve the problem of semantic ambiguity. Based on the same semantic query function of WordNet, we match keywords between comments and methods, and analyze the readability of the classes and packages in a program. Yangchao Liu, Xiaobing Sun 0001, Yucong Duan |
EASE | 2 |
| 2015 | A Problem-Value-Constraint Framework for Minimizing Under Design and Over Design in Web Service Based System DevelopmentabstractWith the increasing of the popularity of Service oriented Software Development, we have identified there is a need to systemically reduce the complexity and increase the robustness of developed Service systems through a guided development process. We choose under design (UD) and over design (OD) as core value assets to uniformly represent the target of managing the human errors and deficiencies during a development process. In the study of a Web Service architectural selection case, based on the ideology of Value Driven Design, we proposed the Problem-Value-Constraint (PVC) approach as the framework of minimizing under design and over design covering both IT concerns including functionalities and qualities and Business concerns including value, cost, price and usage. Our PVC solution connects the characteristics including knowledge management, information transformation and business value. Partially through service design patterns, our study showed that PVC is able to outline both business and technical characteristics at the same time keep the simplicity of model structures. Yucong Duan, Chengxiang Ren, Nianjun Zhou, Xiaobing Sun 0001, Mingdong Tang, Honghao Gao |
ICSS | 4 |
| 2015 | Various "aaS" of everything as a serviceabstractNumerous service models have been proposed in the form of ”as a Service” or ”aaS” in the past. This is especially eminent in the era of the Cloud under the term of anything as a service or everything as a service(XaaS). An unified view of XaaS can aid efficient classification of services for service registration, discovery, and composition. In view of that there lacks such an unified view which is demanded by systematic application of XaaS in the Web Service Ecosystem, towards proposing a taxonomy of ”aaS”, we present in this work a collection of ”aaS” based on a throughout literature survey. Yucong Duan, Xiaobing Sun 0001 |
SNPD | 3 |
| 2015 | Explore the evolution of development topics via on-line LDAabstractSoftware repositories such as revision control systems and bug tracking systems are usually used to manage the changes of software projects. During software maintenance and evolution, software developers and stakeholders need to investigate these repositories to identify what tasks were worked on in a particular time interval and how much effort was devoted to them. A typical way of mining software repositories is to use topic analysis models, e.g., Latent Dirichlet Allocation (LDA), to identify and organize the underlying structure in software documents to understand the evolution of development topics. These previously LDA-based topic analysis models can capture either changes on the strength (popularity) of various development topics over time (i.e., strength evolution) or changes in the content (the words that form the topic) of existing topics over time (i.e., content evolution). Unfortunately, few techniques can capture both strength and content evolution simultaneously. However, both pieces of information are necessary for developers to fully understand how software evolves. In this paper, we propose a novel approach to analyze commit messages within a project's lifetime to capture both strength and content evolution simultaneously via Online Latent Dirichlet Allocation (On-Line LDA). Moreover, the proposed approach also provides an efficient way to detect emerging topics in real development iteration when a new feature request arrives at a particular time, thus helping project stakeholds progress their projects smoothly. Jiajun Hu, Xiaobing Sun 0001, Bin Li 0006 |
SANER | 2 |
| 2015 | Modeling the evolution of development topics using Dynamic Topic ModelsabstractAs the development of a software project progresses, its complexity grows accordingly, making it difficult to understand and maintain. During software maintenance and evolution, software developers and stakeholders constantly shift their focus between different tasks and topics. They need to investigate into software repositories (e.g., revision control systems) to know what tasks have recently been worked on and how much effort has been devoted to them. For example, if an important new feature request is received, an amount of work that developers perform on ought to be relevant to the addition of the incoming feature. If this does not happen, project managers might wonder what kind of work developers are currently working on. Several topic analysis tools based on Latent Dirichlet Allocation (LDA) have been proposed to analyze information stored in software repositories to model software evolution, thus helping software stakeholders to be aware of the focus of development efforts at various time during software evolution. Previous LDA-based topic analysis tools can capture either changes on the strengths of various development topics over time (i.e., strength evolution) or changes in the content of existing topics over time (i.e., content evolution). Unfortunately, none of the existing techniques can capture both strength and content evolution. In this paper, we use Dynamic Topic Models (DTM) to analyze commit messages within a project's lifetime to capture both strength and content evolution simultaneously. We evaluate our approach by conducting a case study on commit messages of two well-known open source software systems, jEdit and PostgreSQL. The results show that our approach could capture not only how the strengths of various development topics change over time, but also how the content of each topic (i.e., words that form the topic) changes over time. Compared with existing topic analysis approaches, our approach can provide a more complete and valuable view of software evolution to help developers better understand the evolution of their projects. Jiajun Hu, Xiaobing Sun 0001, David Lo 0001, Bin Li 0006 |
SANER | 2 |
| 2015 | Query expansion via WordNet for effective code searchabstractSource code search plays an important role in software maintenance. The effectiveness of source code search not only relies on the search technique, but also on the quality of the query. In practice, software systems are large, thus it is difficult for a developer to format an accurate query to express what really in her/his mind, especially when the maintainer and the original developer are not the same person. When a query performs poorly, it has to be reformulated. But the words used in a query may be different from those that have similar semantics in the source code, i.e., the synonyms, which will affect the accuracy of code search results. To address this issue, we propose an approach that extends a query with synonyms generated from WordNet. Our approach extracts natural language phrases from source code identifiers, matches expanded queries with these phrases, and sorts the search results. It allows developers to explore word usage in a piece of software, helps them quickly identify relevant program elements for investigation or quickly recognize alternative words for query reformulation. Our initial empirical study on search tasks performed on the JavaScript/ECMAScript interpreter and compiler, Rhino, shows that the synonyms used to expand the queries help recommend good alternative queries. Our approach also improves the precision and recall of Conquer, a state-of-the-art query expansion/reformulation technique, by 5% and 8% respectively. Meili Lu, Xiaobing Sun 0001, Shaowei Wang 0002, David Lo 0001, Yucong Duan |
SANER | 2 |
| 2015 | MSR4SM: Using topic models to effectively mining software repositories for software maintenance tasks
Xiaobing Sun 0001, Bixin Li, Hareton K. N. Leung, Bin Li 0006, Yun Li 0010 |
Inf. Softw. Technol. | 1 |
| 2015 | Static change impact analysis techniques: A comparative study
Xiaobing Sun 0001, Bixin Li, Hareton K. N. Leung, Bin Li 0006, Junwu Zhu |
J. Syst. Softw. | 1 |
| 2014 | Automatic generation of package diagram to understand Java packagesabstractProgram comprehension is a prerequisite in most software maintenance and evolution tasks. Given an unfamiliar system, it is difficult for practitioners to determine which software artifacts are relevant to the current task. Generally, there are a variety of packages in a Java software system. These packages often have different intents and different relationships between each other. Different information of packages and the relationships between different stereotypes packages form a signature of the system. This paper proposes a novel approach to automatically generate the description of the packages and its diagram to show relationships between the packages. The generated description and diagram can allow developers to more easily understand the main intent and structure of the system. Xiaobing Sun 0001, Yun Li 0010, Xiangyue Liu 0002 |
ICIS | 2 |
| 2014 | Supporting program comprehension with program summarizationabstractA large amount of software maintenance effort is spent on program comprehension. How to accurately and quickly get the functional features in a program becomes a hot issue in program comprehension. Some studies in this area are focused on extracting the topics by analyzing linguistic information in the source code based on the textual mining techniques. However, the extracted topics are usually composed of some standalone words and difficult to understand. In this paper, we attempt to solve this problem based on a novel program summarization technique. First, we propose to use latent semantic indexing and clustering to group source artifacts with similar vocabulary to analyze the composition of each package in the program. Then, some topics composed of a vector of independent words can be extracted based on latent semantic indexing. Finally, we employ Minipar, a nature language parser, to help generate the summaries. The summaries can effectively organize the words from the topics in the form of the predefined sentence based on some rules. With such form of summaries, developers can understand what the features the program has and their corresponding source artifacts. Xiaobing Sun 0001, Xiangyue Liu 0002, Yun Li 0010 |
ICIS | 2 |
| 2014 | PFN: A novel program feature network for program comprehensionabstractProgram comprehension is one of the most frequently performed activities during software maintenance and evolution. In order to facilitate program comprehension, a variety of graphical models have been proposed in software engineering community to construct relationships between program elements. These graphical models are mostly used for understanding the system based on structural syntax dependencies between program elements. However, these graphical models fail to extract the functional or semantic features of the system. Thus, developers still cannot effectively identify the functional part in source code fit for their needs. This paper tries to fill this gap, and proposes a novel representation, program feature network (PFN), to identify the semantic features of the program at class level. PFN is generated based on the relational topic model, a hierarchical probabilistic model of networks. Based on PFN, the semantic features and the links between pairs of two classes in the program can be clearly shown. In addition, PFN can predict the possible links between the newly change request in existing program feature network rather than reconstructing the representation from the start. Xiangyue Liu 0002, Xiaobing Sun 0001, Bin Li 0006, Junwu Zhu |
ICIS | 2 |
| 2014 | Change impact analysis and changeability assessment for a change proposal: An empirical study ☆☆
Xiaobing Sun 0001, Hareton K. N. Leung, Bin Li 0006, Bixin Li |
J. Syst. Softw. | 1 |
| 2013 | Analyzing Impact Rules of Different Change Types to Support Change Impact AnalysisabstractSoftware change impact analysis (CIA) is a key technique for identifying unpredicted and potential effects caused by changes made to software. Different changes have different ripple effects to other parts in the program, even some changes do not affect other entities in spite of some dependencies existing between these entities and the modified one. This induces imprecision if such a factor is neglected. This article proposes a static CIA technique which considers the impact rules of different change types to predict the change effects. Input of our CIA includes changed classes, class methods and class fields, and the output is composed of potentially affected classes, class methods, and class fields. Precision improvement of the CIA technique relies on three aspects: change types of a modified entity, dependencies between the modified entity and other entities, and a precise initial impact set (IIS), on which the final impact set (FIS) is computed. Experimental case studies demonstrate the effectiveness of our technique, and present its potential applications in software maintenance. Xiaobing Sun 0001, Bixin Li, Wanzhi Wen, Sai Zhang 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2013 | FCA-CIA: An approach of using FCA to support cross-level change impact analysis for object oriented Java programs
Bixin Li, Xiaobing Sun 0001, Jacky W. Keung |
Inf. Softw. Technol. | 2 |
| 2013 | A survey of code-based change impact analysis techniquesabstractSUMMARY Software change impact analysis (CIA) is a technique for identifying the effects of a change, or estimating what needs to be modified to accomplish a change. Since the 1980s, there have been many investigations on CIA, especially for code‐based CIA techniques. However, there have been very few surveys on this topic. This article tries to fill this gap. And 30 papers that provide empirical evaluation on 23 code‐based CIA techniques are identified. Then, data was synthesized against four research questions. The study presents a comparative framework including seven properties, which characterize the CIA techniques, and identifies key applications of CIA techniques in software maintenance. In addition, the need for further research is also presented in the following areas: evaluating existing CIA techniques and proposing new CIA techniques under the proposed framework, developing more mature tools to support CIA, comparing current CIA techniques empirically with unified metrics and common benchmarks, and applying the CIA more extensively and effectively in the software maintenance phase. Copyright © 2012 John Wiley & Sons, Ltd. Bixin Li, Xiaobing Sun 0001, Hareton K. N. Leung, Sai Zhang 0001 |
Softw. Test. Verification Reliab. | 2 |
| 2012 | A Change Proposal Driven Approach for Changeability Assessment Using FCA-Based Impact AnalysisabstractGiven a change proposal, how can we evaluate the changeability of the original system to absorb this change proposal before change implementation? Changes to software often have unexpected ripple effects. To avoid this and alleviate the risk of performing undesirable changes, a predictive measurement of these ripple effects should be conducted and a decision of acceptance or rejection should be made on this change proposal. In this paper, we propose an approach to evaluate a software system's changeability with two steps. First, our approach uses formal concept analysis to perform change impact analysis ($CIA$), which estimates the ripple effects of the change proposal. Then, we propose a novel impactness metric to indicate the system's changeability to absorb this change proposal. Case studies on three real-world programs show the effectiveness of our changeability assessment approach. Xiaobing Sun 0001, Bixin Li, Qiandong Zhang |
COMPSAC | 1 |
| 2012 | A comparative study of static CIA techniquesabstractSoftware Change Impact Analysis (CIA) is an essential technique to identify the unpredicted and potential effects caused by software changes. A rich body of different CIA techniques, especially static CIA techniques, have continuously emerged in recent years. However, it is difficult for researchers or practitioners to decide which technique is most appropriate for their needs, or which CIA technique is more effective. Unfortunately, there was only a few work on the comparison of the CIA techniques. This paper presents a comparison study of different types of popular static CIA approaches, i.e., structural static analysis, textual analysis, and historical analysis. For each kind of static CIA approach, we introduce a representative technique, that is FCA -- CIA, ROSE, and IRC2M, respectively. Finally, some empirical studies are conducted on three real-world programs to compare the accuracy of these CIA techniques based on the precision and recall metrics. The results show that the accuracy of these three CIA techniques is different, and FCA - CIA has the best precision while the IRC2M has the best recall. Xiaobing Sun 0001, Bin Li 0006, Bixin Li, Wanzhi Wen |
Internetware | 1 |
| 2012 | Using FCA-based Change Impact Analysis for Regression Testing
Xiaobing Sun 0001, Bixin Li, Chuanqi Tao, Qiandong Zhang |
SEKE | 1 |
| 2012 | Mining Call Graph for Change Impact Analysis
Qiandong Zhang, Bixin Li, Xiaobing Sun 0001 |
SEKE | 3 |
| 2011 | Using Formal Concept Analysis to support change analysisabstractSoftware needs to be maintained and changed to cope with new requirement, existing faults and change requests as software evolves. One particular issue in software maintenance is how to deal with a change proposal before change implementation? Changes to software often cause unexpected ripple effects. To avoid this and alleviate the risk of performing undesirable changes, some predictive measurement should be conducted and a change scheme of the change proposal should be presented. This research intends to provide a unified framework for change analysis, which includes dependencies extraction, change impact analysis, changeability assessment, etc. We expect that our change analysis framework will contribute directly to the improvement of the accuracy of these predictive measures before change implementation, and thus provide more accurate change analysis results for software maintainers, improve quality of software evolution and reduce the software maintenance effort and cost. Xiaobing Sun 0001, Bixin Li |
ASE | 1 |
| 2011 | Program slicing spectrum-based software fault localization
Wanzhi Wen, Bixin Li, Xiaobing Sun 0001, Jiakai Li |
SEKE | 3 |
| 2010 | A Hierarchical Model for Regression Test Selection and Cost Analysis of Java ProgramsabstractRegression testing is an important but expensive stage of software maintenance. Regression test selection addresses the problem of reducing testing cost through selecting a subset of the existing test cases or rerun. Cost-effectiveness is an indispensable factor to consider when developing a regression testing technique. Cost models are created for the purpose of assessing cost-effectiveness of these techniques. The current regression test selection strategies or cost analysis seldom consider hierarchy, which is an inherent characteristic of object-oriented program. In addition, selecting test cases at different levels influences the precision and efficiency of selection. This paper presents a hierarchical regression test selection technique for Java programs to stepwise select test cases from high level of program to low level of program. To effectively evaluate the correlative overall cost, this paper also proposes a hierarchical cost model for analyzing the cost-effectiveness of selection level according to hierarchy step by step. The empirical studies show that our approach can reduce the total cost and achieve more cost-effective results. Chuanqi Tao, Bixin Li, Xiaobing Sun 0001 |
APSEC | 3 |
| 2010 | Change Impact Analysis Based on a Taxonomy of Change TypesabstractSoftware change impact analysis (CIA) is a key technique for identifying unpredicted and potential effects caused by changes made to software. Different change types often have different impact mechanisms, even some changes do not impact other entities in programs in spite of some dependences existed between these entities and the modified entity. In this paper, we propose a static CIA technique, which considers different impact mechanisms and rules of different change types, to calculate the impact sets. Precision improvement of the impact sets relies on 3 aspects: change types of a modified entity, dependences between the modified entity and other entities, and the intuition that to win at the start -- if the initial impact set is estimated more accurately, then the final impact set depending on this initial impact set will be more precise. Experimental case study demonstrates the effectiveness of our technique, and its potential applications in software maintenance. Xiaobing Sun 0001, Bixin Li, Chuanqi Tao, Wanzhi Wen, Sai Zhang 0001 |
COMPSAC | 1 |