VLDB 2026 Research / reviewers in the wild / expert
Yan Xiao 0002
dblp:00/6220-2
· DBLP profile ↗
53ranked-venue papers
8as first author
42since 2021 · last 2026
0000-0002-2563-083XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 36 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modulation-Based Backdoors: Leveraging Amplitude and Frequency Patterns to Attack Speaker RecognitionabstractDeep neural networks (DNNs) are widely and successfully applied in the field of speaker recognition. However, recent studies reveal that these models are vulnerable to backdoor attacks, where adversaries inject malicious behaviors into victim models by poisoning the training process. Existing attack methods often rely on environmental noise or complex voice transformations, which are typically difficult to implement and exhibit poor stealthiness. To address these issues, this paper proposes two modulation-based backdoor attacks that leverage frequency modulation (FM) and amplitude modulation (AM) to construct audio triggers. In real-world scenarios, regular variations in frequency and amplitude are often imperceptible to human listeners, making the proposed attacks more covert. Experimental results show that our methods achieve high attack success rates in both digital and physical settings, while also demonstrating strong resistance to various state-of-the-art backdoor defenses. Hanbo Cai, Pengcheng Zhang 0001, Yan Xiao 0002, Hanting Chu |
AAAI | 3 |
| 2026 | Inverting the Shield: Systematically Generating Safety Tests from Policy SpecificationsabstractXiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu, Kuntai Cai, Yan Xiao, Jin Song Dong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoyue Lu, Xianglin Yang, Kuntai Cai, Yan Xiao 0002, Jin Song Dong 0001 |
ACL (1) | 6 |
| 2026 | Incomplete In-context LearningabstractWenqiang Wang, Wen Yujia, Yan Xiao, Zhifeng Chen, Yangshijie Zhang, Peng Chen, Mingbo Yang, Xiaochun Cao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yujia Wen, Yan Xiao 0002, Yangshijie Zhang, Mingbo Yang, Xiaochun Cao |
ACL (1) | 3 |
| 2026 | Beyond Lexical: Functional Semantics and Fusion for Precise Architecture Recovery
Bixin Li, Yan Xiao 0002 |
SANER | 3 |
| 2026 | TIAFuzz: Transferable fuzzing via distillation for image-based deep learning systems
Shunhui Ji, Hai Dong 0001, Yan Xiao 0002, Mingxuan Xiao, Pengcheng Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Context-aware smart contract comment generation using information retrieval and scenario-driven chain-of-thought
Yanxiang Tong, Hai Dong 0001, Yan Xiao 0002, Pengcheng Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Automated robustness testing for LLM-based natural language processing software
Mingxuan Xiao, Yan Xiao 0002, Shunhui Ji, Hanbo Cai, Lei Xue 0001, Pengcheng Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Multi-task Adversarial Attacks against Black-box Model with Few-shot QueriesabstractCurrent multi-task adversarial text attacks rely on abundant access to shared internal features and numerous queries, often limited to a single task type.As a result, these attacks are less effective against practical scenarios involving black-box feedback APIs, limited queries, or multiple task types.To bridge this gap, we propose Cluster and Ensemble Multi-task Text Adversarial Attack (CEMA), an effective blackbox attack that exploits the transferability of adversarial texts across different tasks.CEMA simplifies complex multi-task scenarios by using a deep-level substitute model trained in a plug-and-play manner for text classification, enabling attacks without mimicking the victim model.This approach requires only a few queries for training, converting multi-task attacks into classification attacks and allowing attacks across various tasks.CEMA generates multiple adversarial candidates using different text classification methods and selects the one that most effectively attacks substitute models.In experiments involving multi-task models with two, three, or six tasks-spanning classification, translation, summarization, and text-toimage generation-CEMA demonstrates significant attack success with as few as 100 queries.Furthermore, CEMA can target commercial APIs (e.g., Baidu and Google Translate), large language models (e.g., ChatGPT 4o), and image-generation models (e.g., Stable Diffusion V2), showcasing its versatility and effectiveness in real-world applications. Yan Xiao 0002, Yangshijie Zhang, Xiaochun Cao |
ACL (1) | 2 |
| 2025 | LLM-Based Semantic Modeling and Cooperative Evolutionary Fuzzing for Traffic Violation Scenario GenerationabstractEnsuring the safety of autonomous driving systems (ADS) in a cost-effective and efficient manner remains a critical challenge. Existing law-guided scenario generation approaches are typically limited to a narrow subset of legal rules, resulting in insufficient scenario diversity, and search-based methods often struggle with large and sparse search spaces. To address these limitations, we propose SLaFE (Semantic Law Modeling and Fuzzing based on Cooperative Evolution), a novel traffic violation scenario generation framework designed to systematically evaluate the safety of ADS. SLaFE harnesses the reasoning capabilities of large language models (LLMs) to convert traffic laws into structured scenario constraints. These scenarios are then optimized via a cooperative evolutionary fuzzing algorithm that explores the parameter space to identify boundary cases likely to trigger abnormal ADS behaviors. We evaluate SLaFE on the Apollo platform within the LGSVL simulator using ten realworld traffic regulations. Experimental results show that SLaFE successfully triggered all 10 types of traffic law violations (10/10), outperforming the best existing method, VioHawk (9/10), while others detected no more than 3. Moreover, SLaFE achieved an average triggering time of 5.1 minutes per law type, significantly faster than VioHawk (9.0 minutes) and other baselines. These results highlight SLaFE’s effectiveness in discovering diverse and critical law-violating scenarios for ADS testing. Yan Xiao 0002, Miao Zhang 0025, Pengcheng Zhang 0001 |
APSEC | 3 |
| 2025 | An Empirical Study of Reinforcement Learning-based Class Integration Test Order GenerationabstractComplicated class dependencies in object-oriented systems challenge traditional integration testing methods. Given varying efforts required to construct test stubs, determining optimal class test orders to minimize stubbing complexity is critical in integration testing. Reinforcement learning (RL), with its strengths in solving complex optimizations, has attracted substantial research attention. To fill the gap in the existing research that lacks of cross-strategy comparisons of RL, we systematically compare five RL algorithms’ performance using varied stubbing complexity weighting methods and critical class identification. Experimental results indicate that when fixed weights are used to calculate stubbing complexity, D3QN achieved the minimum overall stubbing complexity in seven out of nine tested programs. However, PPO attained the lowest mean overall stubbing complexity in five programs, demonstrating the best overall performance. DQN exhibited improved performance when stub complexity was calculated using the entropy-weighted method. After incorporating importance scores as a factor, D3QN achieved the best performance. The RL reward function with introduced importance values is more suitable for systems that are highly centralized and rely on core classes. This work provides robust empirical support and selection guidelines for the practical application of RL algorithms in CITO generation, as well as targeted insights for optimizing RL applications in this field. Shuxiang Zheng, Miao Zhang 0025, Yan Xiao 0002, Peihong Chen, Xiaoxing Yang |
APSEC | 4 |
| 2025 | On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning ImplementationsabstractDeep Reinforcement Learning (DRL) is a paradigm of artificial intelligence where an agent uses a neural network to learn which actions to take in a given environment. DRL has recently gained traction from being able to solve complex environments like driving simulators, 3D robotic control, and multiplayer-online-battle-arena video games. Numerous implementations of the state-of-the-art algorithms responsible for training these agents, like the Deep Q-Network (DQN) and Proximal Policy Optimization (PPO) algorithms, currently exist. However, studies make the mistake of assuming implementations of the same algorithm to be consistent and thus, interchangeable. In this paper, through a differential testing lens, we present the results of studying the extent of implementation inconsistencies, their effect on the implementations' performance, as well as their impact on the conclusions of prior studies under the assumption of interchangeable implementations. The outcomes of our differential tests showed significant discrepancies between the tested algorithm implementations, indicating that they are not interchangeable. In particular, out of the five PPO implementations tested on 56 games, three implementations achieved superhuman performance for 50% of their total trials while the other two implementations only achieved superhuman performance for less than 15% of their total trials. Furthermore, the performance among the high-performing PPO implementations was found to differ significantly in nine games. As part of a meticulous manual analysis of the implementations' source code, we analyzed implementation discrepancies and determined that code-level inconsistencies primarily caused these discrepancies. Lastly, we replicated a study and showed that this assumption of implementation interchangeability was sufficient to flip experiment outcomes. Therefore, this calls for a shift in how implementations are being used. In addition, we recommend for (1) replicability studies for studies mistakenly assuming implementation interchangeability, (2) DRL researchers and practitioners to adopt the differential testing methodology proposed in this paper to combat implementation inconsistencies, and (3) the use of large environment suites. Rajdeep Singh Hundal, Yan Xiao 0002, Xiaochun Cao, Jin Song Dong 0001, Manuel Rigger |
ICSE | 2 |
| 2025 | Clean-label backdoor attack based on robust feature attenuation for speech recognition
Hanbo Cai, Pengcheng Zhang 0001, Yan Xiao 0002, Shunhui Ji, Mingxuan Xiao, Letian Cheng |
Expert Syst. Appl. | 3 |
| 2025 | TPGRec: Text-enhanced and popularity-smoothing graph collaborative filtering for long-tail item recommendation
Chenyun Yu, Yingle Luo, Yan Xiao 0002 |
Neurocomputing | 5 |
| 2025 | Advancing autonomous driving system testing: Demands, challenges, and future directions
Yihan Liao, Jacky W. Keung, Yan Xiao 0002, Yurou Dai |
Inf. Softw. Technol. | 4 |
| 2025 | Promises and perils of using Transformer-based models for SE researchabstractMany Transformer-based pre-trained models for code have been developed and applied to code-related tasks. In this paper, we analyze 519 papers published on this topic during 2017-2023, examine the suitability of model architectures for different tasks, summarize their resource consumption, and look at the generalization ability of models on different datasets. We examine three representative pre-trained models for code: CodeBERT, CodeGPT, and CodeT5, and conduct experiments on the four topmost targeted software engineering tasks from the literature: Bug Fixing, Bug Detection, Code Summarization, and Code Search. We make four important empirical contributions to the field. First, we demonstrate that encoder-only models (CodeBERT) can outperform encoder-decoder models for general-purpose coding tasks, and showcase the capability of decoder-only models (CodeGPT) for certain generation tasks. Second, we study the most frequently used model-task combinations in the literature and find that less popular models can provide higher performance. Third, we find that CodeBERT is efficient in understanding tasks while CodeT5's efficiency is unreliable on generation tasks due to its high resource consumption. Fourth, we report on poor model generalization for the most popular benchmarks and datasets on Bug Fixing and Code Summarization tasks. We frame our contributions in terms of promises and perils, and document the numerous practical issues in advancing future research on transformer-based models for code-related tasks. Yan Xiao 0002, Xinyue Zuo, Xiaoyue Lu, Jin Song Dong 0001, Xiaochun Cao, Ivan Beschastnikh |
Neural Networks | 1 |
| 2025 | TAEFuzz: Automatic Fuzzing for Image-based Deep Learning Systems via Transferable Adversarial ExamplesabstractDeep learning (DL) components have been broadly applied in diverse applications. Similar to traditional software engineering, effective test case generation methods are needed by industry to enhance the quality and robustness of these deep learning components. To this end, we propose a novel automatic software testing technique, TAEFuzz (Automatic Fuzz -Testing via T ransferable A dversarial E xamples), which aims to automatically assess and enhance the robustness of image-based deep learning (DL) systems based on test cases generated by transferable adversarial examples. TAEFuzz alleviates the over-fitting problem during optimized test case generation and prevents test cases from prematurely falling into local optima. In addition, TAEFuzz enhances the visual quality of test cases through constraining perturbations inserted into sensitive areas of the images. For a system with low robustness, TAEFuzz trains a low-cost denoising module to reduce the impact of perturbations in transferable adversarial examples on the system. Experimental results demonstrate that the test cases generated by TAEFuzz can discover up to 46.1% more errors in the targeted systems, and ensure the visual quality of test cases. Compared to existing techniques, TAEFuzz also enhances the robustness of the target systems against transferable adversarial examples with the perturbation denoising module. Shunhui Ji, Changrong Huang, Hai Dong 0001, Lars Grunske, Yan Xiao 0002, Pengcheng Zhang 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | PonziFinder: Attention-Based Edge-Enhanced Ponzi Contract DetectionabstractPonzi contractsare fraudulent investment scams that promise high returns with little risk to investors. However, existing methods for detecting Ponzi contracts have several limitations. For example, they struggle to deal with the class imbalance problem, and their analysis of function call transactions is inadequate, resulting in redundant features. To tackle the challenges of detecting Ponzi contracts, we present PonziFinder, a novel approach that leverages convolutional-based edge-enhanced graph neural network and attention mechanism for the classification of contract transaction graphs. In contrast to previous methods, we not only consider transaction value and timestamp but also analyze transaction input to standardize and sort transactions. We extract node and edge features that capture the unique characteristics of Ponzi contracts. The edge feature, reflecting interaccount correlation, enhances the propagation and updating of node features for effective Ponzi contract detection. To prevent oversmoothing of node embedding caused by the shallow transaction graph and extract important account node information, we introduce an attention-based global layerwise aggregation mechanism (ALGA) for generating the final contract graph representation for classification. Moreover, we optimize the node feature set and use an effective strategy based on undersampling and ensemble learning to address the issue of class imbalance. Experimental results show that PonziFinder can detect all types of Ponzi contracts (100%) with 97% accuracy when there is sufficient transaction data, outperforming other models. The analysis of input values and the ALGA mechanism are experimentally shown to improve accuracy by 4% and 2%, respectively. In summary, PonziFinder is a novel and effective method for detecting Ponzi contracts. Our approach addresses the limitations of existing methods and demonstrates significant improvements in accuracy and efficiency. Bixin Li, Yan Xiao 0002, Xiaoning Du 0001 |
IEEE Trans. Reliab. | 3 |
| 2025 | DeepFusion: Smart Contract Vulnerability Detection Via Deep Learning and Data FusionabstractGiven that smart contracts execute transactions worth hundreds of millions of dollars daily, the issue of smart contract security has attracted considerable attention over the past few years. Traditional methods for detecting vulnerabilities heavily rely on manually developed rules and features, leading to the problems of low accuracy, high false positives, and poor scalability. Although deep learning-inspired approaches were designed to alleviate the problem, most of them rely on monothetic features, which may result in information incompetence during the learning process. Furthermore, the lack of available labeled vulnerability datasets is also a major limitation. To address these issues, we collect and construct a dataset of five labeled smart contract vulnerabilities, and proposeDeepFusion, a vulnerability detection method that fuses code representation information, including program slice information and abstraction syntax tree (AST) structured information. First, we develop automated tools to extract contract vulnerability slicing information from source code, and extract structured information from source code-converted AST. Second, code features and global structured features are fused into the data. Finally, the fused data are input into the Bidirectional Long Short-Term Memory+ Attention (BiLSTM+ATT) model for smart contract vulnerability detection. The BiLSTM model can capture long-term dependencies in both directions and is more suitable for processing serialized information generated byDeepFusion, while the attention mechanism can highlight the characteristic information of vulnerabilities. We conducted experiments via collecting a real smart contract dataset. The experimental results show that our method significantly outperforms the existing methods in detecting the vulnerabilities ofreentrancy,timestamp dependence,integer overflow and underflow,Use tx.origin for authentication, andUnprotected Self-destruct Instructionby 6.36%, 6.42%, 16.5%, 21.29%, and 25.05%, respectively. To the best of our knowledge, the latter two vulnerabilities are the first to be detected using deep learning methods. Hanting Chu, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Shunhui Ji |
IEEE Trans. Reliab. | 4 |
| 2025 | DT4LM: Differential Testing for Reliable Language Model Updates in Classification Tasks
Xinyue Zuo, Yan Xiao 0002, Xiaochun Cao, Wenya Wang 0001, Jin Song Dong 0001 |
IEEE Trans. Software Eng. | 2 |
| 2024 | Enhancing Valid Test Input Generation with Distribution Awareness for Deep Neural NetworksabstractComprehensive testing is important in improving the reliability of Deep Learning (DL)-based systems. Various Test Input Generators (TIGs) have been proposed to generate misbehavior-inducing test inputs. However, the lack of validity checking in TIGs often results in the generation of invalid inputs (i.e., out of the learned distribution), leading to unreliable testing. To save the effort of manually checking the validity and improve test efficiency, it is important to assess the effectiveness and reliability of automated validators. In this study, we comprehensively assess four automated Input Validators (IV s), Our findings show that the accuracy of IVs ranges from 49% to 77%. Distance-based IVs generally outperform reconstruction-based and density-based IVs for both classification and regression tasks. Based on the findings, we enhance existing testing frameworks by incorporating distribution awareness through joint optimization. The results demonstrate our framework leads to a 2 % to 10% increase in the number of valid inputs, which establishes our method as an effective technique for valid test input generation. Jacky W. Keung, Yan Xiao 0002, Yishu Li, Wing Kwong Chan |
COMPSAC | 5 |
| 2024 | Improving the undersampling technique by optimizing the termination condition for software defect predictionabstractThe class imbalance problem significantly hinders the ability of the software defect prediction (SDP) models to distinguish between defective (minority class) and non-defective (majority class) software instances. Recent studies on the data resampling technique have shown that Random UnderSampling (RUS) is more effective than several complex oversampling techniques at alleviating this problem. However, RUS blindly removes majority class instances, leading to significant information loss. These studies have also pointed out that the conventional termination condition (i.e., terminating the data resampling technique when the number of instances for both the minority and majority classes are the same) of the data resampling technique can result in suboptimal performance. In fact, the undersampling technique can be likened to a recommender system or a web search engine that recommends majority class instances to SDP models. Therefore, we propose the Learning-To-Rank Undersampling technique (LTRUS). Our work is novel in two aspects: (1) We consider the undersampling process as a learning-to-rank task, optimizing a linear model to rank majority class instances and remove them from the bottom of the rank to alleviate the class imbalance problem . (2) We propose two termination conditions for the undersampling technique, which differ from the conventional termination condition. LTRUS significantly outperforms RUS, the clustering-based undersampling technique, the complexity-based oversampling technique, SMOTUNED, and Borderline-SMOTE in terms of F-measure, AUC, and MCC by 8.9%, 7.6%, and 18.0% on average under the conventional termination condition. Furthermore, LTRUS under the two termination conditions we propose yield similar performance, and both outperform LTRUS and all the other baselines under the conventional termination condition. The experimental results demonstrate the effectiveness of LTRUS and indicate that the conventional termination condition for the data resampling technique is improper. Shuo Feng 0003, Jacky W. Keung, Yan Xiao 0002, Peichang Zhang, Xiao Yu 0008, Xiaochun Cao |
Expert Syst. Appl. | 3 |
| 2024 | SGDL: Smart contract vulnerability generation via deep learningabstractAbstract The growing popularity of smart contracts in various areas, such as digital payments and the Internet of Things, has led to an increase in smart contract security challenges. Researchers have responded by developing vulnerability detection tools. However, the effectiveness of these tools is limited due to the lack of authentic smart contract vulnerability datasets to comprehensively assess their capacity for diverse vulnerabilities. This paper proposes a Deep Learning‐based Smart contract vulnerability Generation approach (SGDL) to overcome this challenge. SGDL utilizes static analysis techniques to extract both syntactic and semantic information from the contracts. It then uses a classification technique to match injected vulnerabilities with contracts. A generative adversarial network is employed to generate smart contract vulnerability fragments, creating a diverse and authentic pool of fragments. The vulnerability fragments are then injected into the smart contracts using an abstract syntax tree to ensure their syntactic correctness. Our experimental results demonstrate that our method is more effective than existing vulnerability injection methods in evaluating the contract vulnerability detection capacity of existing detection tools. Overall, SGDL provides a comprehensive and innovative solution to address the critical issue of authentic and diverse smart contract vulnerability datasets. Hanting Chu, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Shunhui Ji |
J. Softw. Evol. Process. | 4 |
| 2024 | Toward Stealthy Backdoor Attacks Against Speech Recognition via Elements of SoundabstractDeep neural networks (DNNs) have been widely and successfully adopted and deployed in various applications of speech recognition. Recently, a few works revealed that these models are vulnerable to backdoor attacks, where the adversaries can implant malicious prediction behaviors into victim models by poisoning their training process. In this paper, we revisit poison-only backdoor attacks against speech recognition. We reveal that existing methods are not stealthy since their trigger patterns are perceptible to humans or machine detection. This limitation is mostly because their trigger patterns are simple noises or separable and distinctive clips. Motivated by these findings, we propose to exploit elements of sound (e.g., pitch and timbre) to design more stealthy yet effective poison-only backdoor attacks. Specifically, we insert a short-duration high-pitched signal as the trigger and increase the pitch of remaining audio clips to ‘mask’ it for designing stealthy pitch-based triggers. We manipulate timbre features of victim audio to design the stealthy timbre-based attack and design a voiceprint selection module to facilitate the multi-backdoor attack. Our attacks can generate more ‘natural’ poisoned samples and therefore are more stealthy. Extensive experiments are conducted on benchmark datasets, which verify the effectiveness of our attacks under different settings (e.g., all-to-one, all-to-all, clean-label, physical, and multi-backdoor settings) and their stealthiness. Our methods achieve attack success rates of over 95% in most cases and are nearly undetectable. The code for reproducing main experiments are available at https://github.com/HanboCai/BadSpeech_SoE. Hanbo Cai, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Stefanos Koffas, Yiming Li 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | A Semisupervised Approach for Industrial Anomaly Detection via Self-Adaptive ClusteringabstractWith the rapid development of the Industrial Internet of Things, log-based anomaly detection has become vital for smart industrial construction that has prompted many researchers to contribute. To detect anomalies based on log data, semisupervised approaches stand out from supervised and unsupervised approaches because they only require a portion of labeled data and are relatively stable. However, the state-of-the-art semisupervised approaches still suffer from two main problems: manual parameter setting and unsatisfactory performance with high false positives. We propose AdaLog, an integrated semisupervised approach based on self-adaptive clustering, for industrial anomaly detection. In particular, the clustering step performs automatic label probability estimation by distinguishing 12 situations so that the label probability of each unlabeled data can be carefully calculated, leading to high accuracy. In addition, AdaLog employs a pretrained model to learn contextual information comprehensively and a transformer-based model to detect anomalies efficiently. To alleviate class imbalance, an undersampling method is incorporated. The results on three popular datasets demonstrate that AdaLog significantly outperforms three state-of-the-art semisupervised approaches by 17.8%–2489.8% on average in terms of F1-score, and is even superior to two supervised approaches in most cases with average improvements of 10.9%–23.8%. Jacky W. Keung, Pinjia He, Yan Xiao 0002, Xiao Yu 0008, Yishu Li |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | UniAda: Universal Adaptive Multiobjective Adversarial Attack for End-to-End Autonomous Driving SystemsabstractAdversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated that subtle perturbations to the input can elicit erroneous outcomes, thereby substantially compromising the security of DL systems. This has emerged as a critical concern in the development of DL-based safety–critical systems like autonomous driving systems (ADSs). The focus of existing adversarial attack methods on end-to-end (E2E) ADSs has predominantly centered on misbehaviors of steering angle, which overlooks speed-related controls or imperceptible perturbations. To address these challenges, we introduce UniAda–a multiobjective white-box attack technique with a core function that revolves around crafting an image-agnostic adversarial perturbation capable of simultaneously influencing both steering and speed controls. UniAda capitalizes on an intricately designed multiobjective optimization function with the adaptive weighting scheme (AWS), enabling the concurrent optimization of diverse objectives. Validated with both simulated and real-world driving data, UniAda outperforms five benchmarks across two metrics, inducing steering and speed deviations from 3.54$^{\circ }$to 29$^{\circ }$and 11 to 22 km/h on average. This systematic approach establishes UniAda as a proven technique for adversarial attacks on modern DL-based E2E ADSs. Jacky W. Keung, Yan Xiao 0002, Yihan Liao, Yishu Li |
IEEE Trans. Reliab. | 3 |
| 2023 | Ponzi Scheme Detection Based on Control Flow Graph Feature ExtractionabstractThe blockchain ecosystem is expanding as a result of advancements in blockchain technology and the emergence of BaaS (Blockchain as a Service) platforms. Smart contracts are designed to carry out diverse business operations, but there is a risk of Ponzi schemes being concealed within them. These schemes masquerade as investment agreements and deceive users, resulting in substantial losses for the blockchain community. Detecting Ponzi schemes in smart contracts is crucial. This study introduces a machine learning approach to identify Ponzi schemes by extracting features from smart contracts using the control flow graph. During the construction of the control flow graph for the smart contract’s bytecode, elements unrelated to its functionality are identified and eliminated. We utilize the control flow graph to extract n-gram Term Frequency and n-gram Term Frequency-Inverse Document Frequency features. These features are respectively employed to construct a Random Forest model for Ponzi scheme detection. To address the issue of imbalanced samples, the SVM_SMOTE oversampling algorithm is applied to balance the number of positive and negative samples. The results from experiments conducted on a real-world dataset demonstrate the effectiveness of our approach. The feature extraction method based on the control flow graph outperforms the method based on continuous text. Additionally, the Random Forest model utilizing SVM_SMOTE outperforms four existing models. Shunhui Ji, Congxiong Huang, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002 |
ICWS | 5 |
| 2023 | LEAP: Efficient and Automated Test Method for NLP SoftwareabstractThe widespread adoption of DNNs in NLP software has highlighted the need for robustness. Researchers proposed various automatic testing techniques for adversarial test cases. However, existing methods suffer from two limitations: weak error-discovering capabilities, with success rates ranging from 0% to 24.6% for BERT-based NLP software, and time inefficiency, taking 177.8s to 205.28s per test case, making them challenging for time-constrained scenarios. To address these issues, this paper proposes LEAP, an automated test method that uses LEvy flight-based Adaptive Particle swarm optimization integrated with textual features to generate adversarial test cases. Specifically, we adopt Levy flight for population initialization to increase the diversity of generated test cases. We also design an inertial weight adaptive update operator to improve the efficiency of LEAP's global optimization of high-dimensional text examples and a mutation operator based on the greedy strategy to reduce the search time. We conducted a series of experiments to validate LEAP's ability to test NLP software and found that the average success rate of LEAP in generating adversarial test cases is 79.1%, which is 6.1% higher than the next best approach (PSOattack). While ensuring high success rates, LEAP significantly reduces time overhead by up to 147.6s compared to other heuristic-based methods. Additionally, the experimental results demonstrate that LEAP can generate more transferable test cases and significantly enhance the robustness of DNN-based systems. Mingxuan Xiao, Yan Xiao 0002, Hai Dong 0001, Shunhui Ji, Pengcheng Zhang 0001 |
ASE | 2 |
| 2023 | A survey on smart contract vulnerabilities: Data sources, detection and repair
Hanting Chu, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Shunhui Ji, Wenrui Li 0002 |
Inf. Softw. Technol. | 4 |
| 2023 | On the relative value of imbalanced learning for code smell detectionabstractSummary Machine learning‐based code smell detection (CSD) has been demonstrated to be a valuable approach for improving software quality and enabling developers to identify problematic patterns in code. However, previous researches have shown that the code smell datasets commonly used to train these models are heavily imbalanced. While some recent studies have explored the use of imbalanced learning techniques for CSD, they have only evaluated a limited number of techniques and thus their conclusions about the most effective methods may be biased and inconclusive. To thoroughly evaluate the effect of imbalanced learning techniques for machine learning‐based CSD, we examine 31 imbalanced learning techniques with seven classifiers to build CSD models on four code smell data sets. We employ four evaluation metrics to assess the detection performance with the Wilcoxon signed‐rank test and Cliff's . The results show that (1) Not all imbalanced learning techniques significantly improve detection performance, but deep forest significantly outperforms the other techniques on all code smell data sets. (2) SMOTE (Synthetic Minority Over‐sampling TEchnique) is not the most effective technique for resampling code smell data sets. (3) The best‐performing imbalanced learning techniques and the top‐3 data resampling techniques have little time cost for code smell detection. Therefore, we provide some practical guidelines. First, researchers and practitioners should select the appropriate imbalanced learning techniques (e.g., deep forest) to ameliorate the class imbalance problem. In contrast, the blind application of imbalanced learning techniques could be harmful. Then, better data resampling techniques than SMOTE should be selected to preprocess the code smell data sets. Kuan Zou, Jacky W. Keung, Xiao Yu 0008, Shuo Feng 0003, Yan Xiao 0002 |
Softw. Pract. Exp. | 6 |
| 2023 | On the Significance of Category Prediction for Code-Comment SynchronizationabstractSoftware comments sometimes are not promptly updated in sync when the associated code is changed. The inconsistency between code and comments may mislead the developers and result in future bugs. Thus, studies concerning code-comment synchronization have become highly important, which aims to automatically synchronize comments with code changes. Existing code-comment synchronization approaches mainly contain two types, i.e., (1) deep learning-based (e.g., CUP), and (2) heuristic-based (e.g., HebCUP). The former constructs a neural machine translation-structured semantic model, which has a more generalized capability on synchronizing comments with software evolution and growth. However, the latter designs a series of rules for performing token-level replacements on old comments, which can generate the completely correct comments for the samples fully covered by their fine-designed heuristic rules. In this article, we propose a composite approach named CBS (i.e., Classifying Before Synchronizing ) to further improve the code-comment synchronization performance, which combines the advantages of CUP and HebCUP with the assistance of inferred categories of Code-Comment Inconsistent (CCI) samples. Specifically, we firstly define two categories (i.e., heuristic-prone and non-heuristic-prone) for CCI samples and propose five features to assist category prediction. The samples whose comments can be correctly synchronized by HebCUP are heuristic-prone, while others are non-heuristic-prone. Then, CBS employs our proposed Multi-Subsets Ensemble Learning (MSEL) classification algorithm to alleviate the class imbalance problem and construct the category prediction model. Next, CBS uses the trained MSEL to predict the category of the new sample. If the predicted category is heuristic-prone, CBS employs HebCUP to conduct the code-comment synchronization for the sample, otherwise, CBS allocates CUP to handle it. Our extensive experiments demonstrate that CBS statistically significantly outperforms CUP and HebCUP, and obtains an average improvement of 23.47%, 22.84%, 3.04%, 3.04%, 1.64%, and 19.39% in terms of Accuracy, Recall@5, Average Edit Distance (AED) , Relative Edit Distance (RED) , BLEU-4, and Effective Synchronized Sample (ESS) ratio, respectively, which highlights that category prediction for CCI samples can boost the code-comment synchronization performance. Zhen Yang 0022, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Zhi Jin 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | BiAn: Smart Contract Source Code ObfuscationabstractWith the rising prominence of smart contracts, security attacks targeting them have increased, posing severe threats to their security and intellectual property rights. Existing simplistic datasets hinder effective vulnerability detection, raising security concerns. To address these challenges, we proposeBiAn, a source code level smart contract obfuscation method that generates complex vulnerability test datasets.BiAnprotects contracts by obfuscating data flows, control flows, and code layouts, increasing complexity and making it harder for attackers to discover vulnerabilities. Our experiments with buggy contracts showed an average complexity enhancement of approximately 174% after obfuscation. Decompilers Vandal and Gigahorse had total failure rate increments of 38.8% and 40.5% respectively. Obfuscated contracts also decreased vulnerability detection rates in more than 50% of cases for ten widely-used static analysis detection tools. Pengcheng Zhang 0001, Yan Xiao 0002, Hai Dong 0001, Xiapu Luo |
IEEE Trans. Software Eng. | 3 |
| 2022 | Bytecode Obfuscation for Smart ContractsabstractEthereum smart contracts face serious security problems, which not only cause huge economic losses, but also destroy the Ethereum credit system. To solve this problem, code obfuscation techniques are applied to smart contracts to improve their complexity and security. However, the current source code obfuscation methods have insufficient anti-decompilation ability. Therefore, we propose a novel bytecode obfuscation approach called BOSC based on four kinds of bytecode obfuscation techniques, which is directed at solidity. The experimental results show that, after the bytecode obfuscation, the failure rate of decompilation tools is over 99% and only a small amount of gas is consumed. Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Shunhui Ji |
APSEC | 4 |
| 2022 | Repairing Failure-inducing Inputs with Input ReflectionabstractTrained with a sufficiently large training and testing dataset, Deep Neural Networks (DNNs) are expected to generalize. However, inputs may deviate from the training dataset distribution in real deployments. This is a fundamental issue with using a finite dataset, which may lead deployed DNNs to mis-predict in production. Yan Xiao 0002, Yun Lin 0001, Ivan Beschastnikh, Changsheng Sun, David S. Rosenblum, Jin Song Dong 0001 |
ASE | 1 |
| 2022 | The impact of the distance metric and measure on SMOTE-based techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Peichang Zhang, Yan Xiao 0002, Miao Zhang 0025 |
Inf. Softw. Technol. | 4 |
| 2022 | Predicting the precise number of software defects: Are we there yet?abstractContext: Defect Number Prediction (DNP) models can offer more benefits than classification-based defect prediction . Recently, many researchers proposed to employ regression algorithms for DNP, and found that the algorithms achieve low Average Absolute Error (AAE) and high Pred(0.3) values. However, since the defect datasets generally contain many non-defective modules, even if a DNP model predicts the number of defects in all modules as zero, the AAE value of the model will be low and Pred(0.3) value will be high. Therefore, the good performance of the regression algorithms in terms of AAE and Pred(0.3) may be questioned due to the imbalanced distribution of the number of defects. Objective: To revisit the impact of regression algorithms for predicting the precise number of defects. Method: We examine the practical effects of 12 widely-used regression algorithms, two data resampling algorithm (SmoteR and ROS), and three ensemble learning algorithms (gradient boosting regression, AdaBoost .R2, and Bagging), one feature selection method (information gain) and one parameter optimization method (grid search) for predicting the precise number of defects on the 18 PROMISE datasets. We propose to evaluate the AAE and Pred(0.3) values for the modules with different numbers of defects separately. Results: The AAE values for defective modules are very high and the Pred(0.3) values are very low, i.e., the regression algorithms are very inaccurate for predicting the precise number of defects in defective modules. Conclusion: The problem of predicting the precise number of defects via regression algorithms is far from being solved. We recommend that software testers use regression algorithms to rank modules for testing resource allocation , rather than predict the precise number of defects to evaluate the software reliability and maintenance effort. In addition, most existing DNP studies employing the whole AAE and Pred(0.3) values of all modules as the evaluation metrics for the proposed DNP algorithms should be revisited. Xiao Yu 0008, Jacky W. Keung, Yan Xiao 0002, Shuo Feng 0003, Heng Dai |
Inf. Softw. Technol. | 3 |
| 2021 | ROCT: Radius-based Class Overlap Cleaning Technique to Alleviate the Class Overlap Problem in Software Defect PredictionabstractThe training data commonly used in software defect prediction (SDP) usually contains some instances that have similar values on features but are in different classes, which significantly degrades the performance of prediction models trained using these instances. This is referred to as the class overlap problem (COP). Previous studies have concluded that COP has a more negative impact on the performance of prediction models than the class imbalance problem (CIP). However, less research has been conducted on COP than CIP. Moreover, the performance of the existing class overlap cleaning techniques heavily relies on the settings of hyperparameters such as the value of K in the K-nearest neighbor algorithm or the K-means algorithm, but how to find those optimal hyperparameters is still a challenge. In this study, we propose a novel technique named the radius-based class overlap cleaning technique (ROCT) to better alleviate COP without tuning hyperparameters in SDP. The basic idea of ROCT is to take each instance as the center of a hypersphere and directly optimize the radius of the hypersphere. Then ROCT identifies those instances with the opposite label of the center instance as the overlapping instance and removes them. To investigate the performance of ROCT, we conduct the empirical experiment across 29 datasets collected from various software repositories on the K-nearest neighbor, random forest, logistic regression, and naive Bayes classifiers measured by AUC, balance, pd, and pf. The experimental results show that ROCT performs the best and significantly improves the performance of prediction models by as much as 15.2% and 29.9% in terms of AUC and balance compared with the existing class overlap cleaning techniques. The superior performance of ROCT indicates that ROCT should be recommended as an efficient alternative to alleviate COP in SDP. Shuo Feng 0003, Jacky W. Keung, Jie Liu 0016, Yan Xiao 0002, Xiao Yu 0008, Miao Zhang 0025 |
COMPSAC | 4 |
| 2021 | Self-Checking Deep Neural Networks in DeploymentabstractThe widespread adoption of Deep Neural Networks (DNNs) in important domains raises questions about the trustworthiness of DNN outputs. Even a highly accurate DNN will make mistakes some of the time, and in settings like self-driving vehicles these mistakes must be quickly detected and properly dealt with in deployment. Just as our community has developed effective techniques and mechanisms to monitor and check programmed components, we believe it is now necessary to do the same for DNNs. In this paper we present DNN self-checking as a process by which internal DNN layer features are used to check DNN predictions. We detail SelfChecker, a self-checking system that monitors DNN outputs and triggers an alarm if the internal layer features of the model are inconsistent with the final prediction. SelfChecker also provides advice in the form of an alternative prediction. We evaluated SelfChecker on four popular image datasets and three DNN models and found that SelfChecker triggers correct alarms on 60.56% of wrong DNN predictions, and false alarms on 2.04% of correct DNN predictions. This is a substantial improvement over prior work (SelfOracle, Dissector, and ConfidNet). In experiments with self-driving car scenarios, SelfChecker triggers more correct alarms than SelfOracle for two DNN models (DAVE-2 and Chauffeur) with comparable false alarms. Our implementation is available as open source. Yan Xiao 0002, Ivan Beschastnikh, David S. Rosenblum, Changsheng Sun, Sebastian G. Elbaum, Yun Lin 0001, Jin Song Dong 0001 |
ICSE | 1 |
| 2021 | COSTE: Complexity-based OverSampling TEchnique to alleviate the class imbalance problem in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Kwabena Ebo Bennin, Md. Alamgir Kabir, Miao Zhang 0025 |
Inf. Softw. Technol. | 4 |
| 2021 | Investigation on the stability of SMOTE-based oversampling techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Miao Zhang 0025 |
Inf. Softw. Technol. | 4 |
| 2021 | The effectiveness of data augmentation in code readability classification
Qing Mi, Yan Xiao 0002, Zhi Cai, Xibin Jia |
Inf. Softw. Technol. | 2 |
| 2021 | Validating class integration test order generation systems with Metamorphic Testing
Miao Zhang 0025, Jacky W. Keung, Tsong Yueh Chen, Yan Xiao 0002 |
Inf. Softw. Technol. | 4 |
| 2021 | Evaluating the effects of similar-class combination on class integration test order generation
Miao Zhang 0025, Jacky W. Keung, Yan Xiao 0002, Md. Alamgir Kabir |
Inf. Softw. Technol. | 3 |
| 2020 | Smart Contracts Vulnerability Auditing with Multi-semanticsabstractSmart contracts vulnerability auditing is vitally critical to ensure transaction execution in normal on blockchain. The current data-driven approaches normally tokenize smart contracts into a series of sequences according to only one tokenization standard for vulnerability detection purpose, resulting some of the semantic contexts could not be reflected within restricted sequence length. To address this limitation, we generate sequences from smart contracts in three tokenization standards for which we utilize n-gram language model to capture semantic contexts respectively, and finally exploiting our effective combination strategy of Intersection or Union to integrate the audited results from multiple semantic contexts. In order to evaluate the proposed approach, we applied it on over 7200 Ethereum smart contract samples. Experimental result shows our proposed method is capable of detecting vulnerabilities and competitive with the baseline in test sets, with improved precision of over 44% when Intersection is applied in their results, as well as improved Recall measure up by over 300% and F-measure up by 220% when Union is applied. Our proposed method for smart contract vulnerability detection, an important tool for developing quality decentralized software applications, is able to analyze multiple semantic contexts and successfully detects more true vulnerabilities with high precision, outperforming that of the baseline approaches. Zhen Yang 0022, Jacky W. Keung, Miao Zhang 0025, Yan Xiao 0002, Yangyang Huang, Tik Hui |
COMPSAC | 4 |
| 2019 | A Heuristic Approach to Break Cycles for the Class Integration Test Order GenerationabstractIt is a general objective to minimize overall stubbing cost when performing class integration test order generation. Existing approaches are unable to obtain a cost-optimal class test order, this is largely due to the lack of a comprehensive analysis on the factors that affect overall stubbing cost, i.e., the number of required test stubs and the corresponding stubbing complexity. To address this issue, we propose an approach called HBCITO (Heuristic approach to Break Cycles for the class Integration Test Order generation). Given a set of removed dependencies, a heuristic algorithm is employed to search for a near ideal set of class dependencies. Such dependencies break the same or greater number of cycles as the initialized dependencies but attract less stubbing cost. The experimental results show that HBCITO is capable of generating class test orders with significantly lower stubbing cost compared with other approaches. Miao Zhang 0025, Jacky W. Keung, Yan Xiao 0002, Md. Alamgir Kabir, Shuo Feng 0003 |
COMPSAC (1) | 3 |
| 2019 | Improving bug localization with word embedding and enhanced convolutional neural networks
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 1 |
| 2019 | Node-Constrained Traffic Engineering: Theory and ApplicationsabstractTraffic engineering (TE) is a fundamental task in networking. Conventionally, traffic can take any path connecting the source and destination. Emerging technologies such as segment routing, however, use logical paths that are composed of shortest paths going through a predetermined set of middlepoints in order to reduce the flow table overhead of TE implementation. Inspired by this, in this paper, we introduce the problem of node-constrained TE, where the traffic must go through a set of middlepoints, and study its theoretical fundamentals. We show that the general node-constrained TE that allows the traffic to take any path going through one or more middlepoints is NP-hard for directed graphs but strongly polynomial for undirected graphs, unveiling a profound dichotomy between the two cases. We also investigate a variant of node-constrained TE that uses only shortest paths between middlepoints, and prove that the problem can now be solved in weakly polynomial time for a fixed number of middlepoints, which explains why existing work focuses on this variant. Yet, if we constrain the end-to-end paths to be acyclic, the problem can become NP-hard. An important application of our work concerns flow centrality, for which we are able to derive complexity results. Furthermore, we investigate the middlepoint selection problem in general node-constrained TE. We introduce and study group flow centrality as a solution concept, and show that it is monotone but not submodular. Our work provides a thorough theoretical treatment of node-constrained TE and sheds light on the development of the emerging node-constrained TE in practice. George Trimponias, Yan Xiao 0002, Xiaorui Wu, Hong Xu 0001, Yanhui Geng |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Improving Bug Localization with Character-Level Convolutional Neural Network and Recurrent Neural NetworkabstractBackground: Automated bug localization in large amounts of source files for bug reports is a crucial task in software engineering. However, the different representations of bug reports and source files limited the accuracy of the existing bug localization techniques. Aims: We propose a novel deep learning-based model to improve the accuracy of bug localization for bug reports by expressing them in character and analyzing them with a language model. Method: The proposed model is composed of two main parts: character-level convolutional neural network (CNN) and recurrent neural network (RNN) language model. Both bug reports and source files are expressed in a character level and then input into a CNN, whose output is given to an RNN encoder-decoder architecture. Results: The results of preliminary experiments show that the proposed model achieves comparable or even higher accuracy than the existing machine translation-based bug localization technique. Conclusion: The proposed model is capable of automatically localizing buggy files for bug reports and achieves better accuracy by analyzing them in character level where both bug reports and source code can be expressed. Yan Xiao 0002, Jacky W. Keung |
APSEC | 1 |
| 2018 | An Inception Architecture-Based Model for Improving Code Readability ClassificationabstractThe process of classifying a piece of source code into a Readable or Unreadable class is referred to as Code Readability Classification. To build accurate classification models, existing studies focus on handcrafting features from different aspects that intuitively seem to correlate with code readability, and then exploring various machine learning algorithms based on the newly proposed features. On the contrary, our work opens up a new way to tackle the problem by using the technique of deep learning. Specifically, we propose IncepCRM, a novel model based on the Inception architecture that can learn multi-scale features automatically from source code with little manual intervention. We apply the information of human annotators as the auxiliary input for training IncepCRM and empirically verify the performance of IncepCRM on three publicly available datasets. The results show that: 1) Annotator information is beneficial for model performance as confirmed by robust statistical tests (i.e., the Brunner-Munzel test and Cliff's delta); 2) IncepCRM can achieve an improved accuracy against previously reported models across all datasets. The findings of our study confirm the feasibility and effectiveness of deep learning for code readability classification. Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Xiupei Mei |
EASE | 3 |
| 2018 | Bug Localization with Semantic and Structural Features using Convolutional Neural Network and Cascade ForestabstractBackground: Correctly localizing buggy files for bug reports together with their semantic and structural information is a crucial task, which would essentially improve the accuracy of bug localization techniques. Aims: To empirically evaluate and demonstrate the effects of both semantic and structural information in bug reports and source files on improving the performance of bug localization, we propose CNN_Forest involving convolutional neural network and ensemble of random forests that have excellent performance in the tasks of semantic parsing and structural information extraction. Method: We first employ convolutional neural network with multiple filters and an ensemble of random forests with multi-grained scanning to extract semantic and structural features from the word vectors derived from bug reports and source files. And a subsequent cascade forest (a cascade of ensembles of random forests) is used to further extract deeper features and observe the correlated relationships between bug reports and source files. CNNLForest is then empirically evaluated over 10,754 bug reports extracted from AspectJ, Eclipse UI, JDT, SWT, and Tomcat projects. Results: The experiments empirically demonstrate the significance of including semantic and structural information in bug localization, and further show that the proposed CNN_Forest achieves higher Mean Average Precision and Mean Reciprocal Rank measures than the best results of the four current state-of-the-art approaches (NPCNN, LR+WE, DNNLOC, and BugLocator). Conclusion: CNNLForest is capable of defining the correlated relationships between bug reports and source files, and we empirically show that semantic and structural information in bug reports and source files are crucial in improving bug localization. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
EASE | 1 |
| 2018 | Improving code readability classification using convolutional neural networks
Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Yujin Gao |
Inf. Softw. Technol. | 3 |
| 2018 | Machine translation-based bug localization technique for bridging lexical gap
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 1 |
| 2017 | Identifying Textual Features of High-Quality Questions: An Empirical Study on Stack OverflowabstractBackground: Stack Overflow (SO) is a programming-specific Q&A website that serves as a valuable repository of software engineering knowledge. For SO members, formulating a good question is the first step towards eliciting satisfactory responses. Aims: To guide SO members on how to make a good question, we conduct an empirical study using the publicly available Stack Overflow Data Dump for the period of 2008-2016. Method: We first choose 25 features along 5 dimensions to represent the textual characteristics that we are interested in. Making use of the Boruta algorithm, we then capture all features that are either strongly or weakly relevant to the question quality. Results: The results show that the number of tags and code snippets are the most discriminative features, whereas there is only a weak correlation between the question quality and the sentiment-related factors. Based on the empirical evidence, we provide useful and usable suggestions to SO members on how to optimize their questions. Conclusions: We consider that our findings will provide SO members with a better understanding of the patterns behind high-quality questions, this is to support effective and efficient utilization of Q&A websites as the ultimate goal. Qing Mi, Yujin Gao, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah |
APSEC | 4 |
| 2017 | Improving Bug Localization with an Enhanced Convolutional Neural NetworkabstractBackground: Localizing buggy files automatically speeds up the process of bug fixing so as to improve the efficiency and productivity of software quality teams. There are other useful semantic information available in bug reports and source code, but are mostly underutilized by existing bug localization approaches. Aims: We propose DeepLocator, a novel deep learning based model to improve the performance of bug localization by making full use of semantic information. Method: DeepLocator is composed of an enhanced CNN (Convolutional Neural Network) proposed in this study considering bug-fixing experience, together with a new rTF-IDuF method and pretrained word2vec technique. DeepLocator is then evaluated on over 18,500 bug reports extracted from AspectJ, Eclipse, JDT, SWT and Tomcat projects. Results: The experimental results show that DeepLocator achieves 9.77% to 26.65% higher Fmeasure than the conventional CNN and 3.8% higher MAP than a state-of-the-art method HyLoc using less computation time. Conclusion: DeepLocator is capable of automatically connecting bug reports to the corresponding buggy files and successfully achieves better performance based on a deep understanding of semantics in bug reports and source code. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
APSEC | 1 |