VLDB 2026 Research / reviewers in the wild / expert
Yinxing Xue
dblp:73/7055
· DBLP profile ↗
73ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0002-2979-7151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 52 · 8 first-author · 27 since 2021Security and privacy · 9 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HalluTrigger: Triggering input-conflicting hallucinations in large language models with semantic-guided metamorphic relations
Xiaoning Ren, Yinxing Xue |
Expert Syst. Appl. | 3 |
| 2026 | NeuroPatch: Lightweight diffusion model repair for backdoor attack mitigation based on neuron-level patching
Chengze Wu, Xiaoning Ren, Yinxing Xue |
Neurocomputing | 4 |
| 2026 | Dynamic Assessment of Mission Reliability for Autonomous Vehicles Considering Mission Criticality and Environmental DependenceabstractWith the growing concern for autonomous vehicle (AV) reliability, various statistical metrics have been developed to measure their long-term and average behaviors. However, these metrics overlook the characteristics during the phased-mission operations of AVs. This article proposes a novel method to dynamically assess the mission reliability of AV systems. We first establish a dedicated mission reliability metric specifically tailored for AV applications. A probabilistic assessment model is then developed to consider time-varying mission demands, mission criticality, environmental dependence, and measurement noises. Integrating the local linear regression model and sample average approximation approach, the distribution of mission performance is analyzed to address nonlinear environmental dependencies. The theoretical proof for the finite-sample performance guarantee of our model is rigorously established. Numerical simulations demonstrate the method’s superior performance in finite-sample scenarios over Monte Carlo approaches. A real-world case study on lateral vehicle control further confirms the effectiveness of the method. Yan-Fu Li, Yinxing Xue |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | RP-PGD: Boosting Segmentation Robustness with a Region-and-Prototype Based Adversarial AttackabstractAdversarial attack and defense have been extensively explored in classification tasks, but their study in semantic segmentation remains limited. Moreover, current attacks fail to act as strong underlying attacks for adversarial training (AT), making it difficult to achieve segmentation robustness against strong attacks. In this paper, we present RP-PGD, a novel Region-and-Prototype based Projected Gradient Descent attack tailored to fool segmentation models. In particular, we propose a region-based attack, which leverages a spatial-temporal way to separate the pixels into three disjoint regions, and highlights the attack on the crucial True Region and Boundary Region. Moreover, we introduce a prototype-based attack to disrupt the feature space, further enhancing the attack capability. To boost the robustness of segmentation models, we inject adversaries generated by RP-PGD into the clean data and perform AT. Extensive experiments on multiple datasets showcase that RP-PGD generates adversaries with faster convergence and stronger attack effectiveness, surpassing state-of-the-art attacks by a large margin. Consequently, RP-PGD serves as a strong underlying attack for segmentation models to perform AT, assisting them in defending against a variety of strong attacks without incurring additional computational costs during inference. Yuxuan Zhang 0007, Zhenbo Shi, Shuchang Wang, Wei Yang 0011, Shaowei Wang 0003, Yinxing Xue |
AAAI | 6 |
| 2025 | Network Protocol Security Evaluation via LLM-Enhanced Fuzzing in Extended ProFuzzBench
Hanwen Gong, Yinxing Xue |
ICIC (22) | 3 |
| 2025 | LOFT: An LLM-Enhanced Multi-Objective Search Framework for Fault Injection Testing of Autonomous Driving SystemsabstractAutonomous Driving Systems (ADS) are considered safety-critical, as even a minor fault may lead to catastrophic consequences. To evaluate their reliability and robustness under failure conditions, Fault Injection (FI) techniques have been widely adopted. Most existing FI methods employ data-driven approaches, such as surrogate modeling and reinforcement learning, to generate test cases. While these techniques have shown promise, they often incur substantial costs in terms of data collection and training time. Moreover, their performance is highly sensitive to the quality and quantity of training data, which can limit their applicability in diverse or unseen scenarios. In this paper, we propose LOFT, an efficient multi-objective search-based FI testing framework that leverages Large Language Models (LLMs) to identify diverse and realistic critical faults. To accommodate the structured and non-linguistic nature of raw simulation data, LOFT adopts a two-stage LLM-based fault injection pipeline. In the first stage, an LLM converts singleframe simulation data into natural language descriptions and suggests appropriate fault types. In the second stage, a separate LLM examines the broader scenario context to determine the optimal time window for fault injection. The outputs from the two LLMs are then used to initialize and guide a multi-objective search procedure aiming at discovering a diverse set of critical faults. We implement LOFT and evaluate on an ADS provided by our industrial partner. Experimental results show that, compared with two baseline approaches, LOFT detects over $90 \%$ more critical faults and identifies an average of 2.2 additional fault types within an equivalent number of simulations. Guangdong You, Shuncheng Tang, Jixiang Zhou, Hezhen Liu, Junfang Jiang, Yan-Fu Li, Yinxing Xue |
ISSRE | 7 |
| 2025 | Function Clustering-Based Fuzzing Termination: Toward Smarter Early Stopping
Wenzhang Yang, Yinxing Xue |
ASE | 3 |
| 2025 | Demystifying the Evolution of Neural Networks with BOM Analysis: Insights from a Large-Scale Study of 55,997 GitHub RepositoriesabstractNeural networks have become integral to many fields due to their exceptional performance. The open-source community has witnessed a rapid influx of neural network (NN) repositories with fast-paced iterations, making it crucial for practitioners to analyze their evolution to guide development and stay ahead of trends. While extensive research has explored traditional software evolution using Software Bill of Materials (SBOMs), these are ill-suited for NN software, which relies on pre-defined modules and pre-trained models (PTMs) with distinct component structures and reuse patterns. Conceptual AI Bills of Materials (AIBOMs) also lack practical implementations for large-scale evolutionary analysis. To fill this gap, we introduce the Neural Network Bill of Material (NNBOM), a comprehensive dataset construct tailored for NN software. We create a large-scale NNBOM database from 55,997 curated PyTorch GitHub repositories, cataloging their TPLs, PTMs, and modules. Leveraging this database, we conduct a comprehensive empirical study of neural network software evolution across software scale, component reuse, and inter-domain dependency, providing maintainers and developers with a holistic view of its long-term trends. Building on these findings, we develop two prototype applications, Multi repository Evolution Analyzer and Single repository Component Assessor and Recommender, to demonstrate the practical value of our analysis. Xiaoning Ren, Yuhang Ye 0004, Xiongfei Wu, Yueming Wu 0001, Yinxing Xue |
ASE | 5 |
| 2025 | Towards Automated and Accurate Understanding of ARINC Standard in Heterogeneous Data FormatsabstractAccuracy and rigor are vital indicators of the specification document, especially for the ARINC653 aviation industry standard. A high-quality standard or specification should clearly depict the system behaviors yet leave no fatal vulnerability. Formal verification could definitely help achieve this goal, but it requires intensive professional domain knowledge and overwhelming manpower. Recently, fast-growing natural language processing (NLP) techniques do well in harvesting knowledge extraction for the downstream tasks. However, since knowledge about an entity is scattered over heterogeneous contents (plain text, pseudocode, XML, etc.) for almost all such standard documents, a single content or not all contents cannot account for the entire knowledge. To this end, we propose a novel and practical approach to construct the Ontology of ARINC653 and extract the logical guards. Technically, we combine the NLP techniques with domainspecific naming and lexical rules for entity recognition in Ontology and then apply information extraction and relation formalization for relation extraction (in terms of guards). We evaluate the quality of our Ontology against that induced by the domain professor. We further apply this approach to the historical ARINC653 standards and evaluate the performance. Results show that our approach indeed helps construct knowledge integration and aid for specification understanding. Cuifeng Gao, Wenzhang Yang, Xianchang Luo, Yinxing Xue |
QRS | 4 |
| 2025 | Flash Loan Attack is More Than Just Price Oracle Manipulation: A Comprehensive Empirical StudyabstractThe rapid growth of the decentralized finance (DeFi) ecosystem has given rise to flash loan, a type of uncollateralized loan service that enables users to easily borrow substantial amounts of funds. However, this has prompted attackers to conduct malicious arbitrage within DeFi protocols, known as notorious flash loan attacks, resulting in significant asset losses. Existing works primarily focus on investigating price oracle manipulation, a common tactic in flash loan attacks, but lack a comprehensive understanding regarding the entire process of flash loan attacks and the diverse range of attack methods. In this paper, we empirically study 155 real-world flash loan attack incidents, representing the largest-scale study to date. We first categorize these incidents into five types based on their root causes and compile statistics on their distribution, then elucidate the vulnerable code and finance mechanisms exploited in each category. Subsequently, we identify the symptoms of codebased vulnerabilities and summarize the abstract attack models for the entire process. Finally, we evaluate the effectiveness of state-of-the-art off-chain tools in detecting code-based vulnerabilities within their scope of capabilities. We find that Slither performs the best in detecting 22 % of temporal reentrancy vulnerabilities, and DeFiTainter has a 52% false negative rate in detecting price oracle manipulation, mainly attributed to three limitations. Cuifeng Gao, Jiajun Ye, Wenzhang Yang, Yinxing Xue |
QRS | 4 |
| 2025 | White-box structure analysis of pre-trained language models of code for effective attacking
Xiaoning Ren, Yinxing Xue |
Inf. Softw. Technol. | 3 |
| 2024 | Rust-lancet: Automated Ownership-Rule-Violation Fixing with Behavior PreservationabstractAs a relatively new programming language, Rust is designed to provide both memory safety and runtime performance. To achieve this goal, Rust conducts rigorous static checks against its safety rules during compilation, effectively eliminating memory safety issues that plague C/C++ programs. Although useful, the safety rules pose programming challenges to Rust programmers, since programmers can easily violate safety rules when coding in Rust, leading their code to be rejected by the Rust compiler, a fact underscored by a recent user study. There exists a desire to automate the process of fixing safety-rule violations to enhance Rust's programmability. Wenzhang Yang, Linhai Song, Yinxing Xue |
ICSE | 3 |
| 2024 | GenSeg: On Generating Unified Adversary for Segmentation
Yuxuan Zhang 0007, Zhenbo Shi, Wei Yang 0011, Shuchang Wang, Shaowei Wang 0003, Yinxing Xue |
IJCAI | 6 |
| 2024 | LeGEND: A Top-Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language ModelsabstractAutonomous driving systems (ADS) are safety-critical and require comprehensive testing before their deployment on public roads. While existing testing approaches primarily aim at the criticality of scenarios, they often overlook the diversity of the generated scenarios that is also important to reflect system defects in different aspects. To bridge the gap, we propose LeGEND, that features a top-down fashion of scenario generation: it starts with abstract functional scenarios, and then steps downwards to logical and concrete scenarios, such that scenario diversity can be controlled at the functional level. However, unlike logical scenarios that can be formally described, functional scenarios are often documented in natural languages (e.g., accident reports) and thus cannot be precisely parsed and processed by computers. To tackle that issue, LeGEND leverages the recent advances of large language models (LLMs) to transform textual functional scenarios to formal logical scenarios. To mitigate the distraction of useless information in functional scenario description, we devise a two-phase transformation that features the use of an intermediate language; consequently, we adopt two LLMs in LeGEND, one for extracting information from functional scenarios, the other for converting the extracted information to formal logical scenarios. We experimentally evaluate LeGEND on Apollo, an industry-grade ADS from Baidu. Evaluation results show that LeGEND can effectively identify critical scenarios, and compared to baseline approaches, LeGEND exhibits evident superiority in diversity of generated scenarios. Moreover, we also demonstrate the advantages of our two-phase transformation framework, and the accuracy of the adopted LLMs. Shuncheng Tang, Zhenya Zhang 0001, Jixiang Zhou, Yuan Zhou 0005, Yinxing Xue |
ASE | 6 |
| 2024 | Rust-twins: Automatic Rust Compiler Testing through Program Mutation and Dual Macros GenerationabstractRust is a relatively new programming language known for its memory safety and numerous advanced features. It has been widely used in system software in recent years. Thus, ensuring the reliability and robustness of the only implementation of the Rust compiler, rustc, is critical. However, compiler testing, as one of the most effective techniques to detect bugs, faces difficulties in generating valid Rust programs with sufficient diversity due to its stringent memory safety mechanisms. Furthermore, existing research primarily focuses on testing rustc to trigger crash errors, neglecting incorrect compilation results - miscompilation. Detecting miscompilation remains a challenge in the absence of multiple implementations of the Rust compiler to serve as a test oracle. Wenzhang Yang, Cuifeng Gao, Yuekang Li, Yinxing Xue |
ASE | 5 |
| 2024 | Making vulnerability prediction more practical: Prediction, categorization, and localization
Xiang Chen 0005, Xiangwei Li, Yinxing Xue |
Inf. Softw. Technol. | 4 |
| 2024 | Semantic Conformance Testing of Relational DBMSabstractRelational DBMS implementations are expected to adhere to SQL standards. However, there are currently no tools available that can automatically verify this conformance. The main reasons are twofold. First, the SQL standard specification, documented in natural language, tends to be ambiguous and is not directly executable. Second, it is difficult to generate test queries that thoroughly cover all aspects, e.g., keywords and parameters, defined in the SQL specification. In this work, we introduce the first method for semantic conformance testing of RDBMSs. Our contributions are threefold. Firstly, we formally define the denotational semantics of SQL and implement them in Prolog, creating an executable reference RDBMS for differential testing against existing RDBMSs. Secondly, we propose three coverage criteria based on these formal semantics, along with a coverage-guided query generation algorithm that effectively generates queries achieving high semantic coverage. Lastly, we apply our approach to six widely-used and thoroughly tested RDBMSs, e.g., MySQL, PostgreSQL and OceanBase, uncovering 19 bugs and 13 inconsistencies, all of which are confirmed by RDBMS developers. Shuang Liu 0007, Chenglin Tian, Jun Sun 0001, Wei Lu 0015, Yinxing Xue, Junjie Wang 0007, Xiaoyong Du 0001 |
Proc. VLDB Endow. | 7 |
| 2024 | xFuzz: Machine Learning Guided Cross-Contract FuzzingabstractSmart contract transactions are increasingly interleaved by cross-contract calls. While many tools have been developed to identify a common set of vulnerabilities, the cross-contract vulnerability is overlooked by existing tools. Cross-contract vulnerabilities are exploitable bugs that manifest in the presence of more than two interacting contracts. Existing methods are however limited to analyze a maximum of two contracts at the same time. Detecting cross-contract vulnerabilities is highly non-trivial. With multiple interacting contracts, the search space is much larger than that of a single contract. To address this problem, we presentxFuzz, a machine learning guided smart contract fuzzing framework. The machine learning models are trained with novel features (e.g., word vectors and instructions) and are used to filter likely benign program paths. Comparing with existing static tools, machine learning model is proven to be more robust, avoiding directly adopting manually-defined rules in specific tools. We comparexFuzzwith three state-of-the-art tools on 7,391 contracts.xFuzzdetects 18 exploitable cross-contract vulnerabilities, of which 15 vulnerabilities are exposed for the first time. Furthermore, our approach is shown to be efficient in detecting non-cross-contract vulnerabilities as well—using less than 20% time as that of other fuzzing tools,xFuzzdetects twice as many vulnerabilities. Yinxing Xue, Jiaming Ye, Jun Sun 0001, Lei Ma 0003, Haijun Wang 0002, Jianjun Zhao 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Like Attracts Like: Personalized Federated Learning in Decentralized Edge ComputingabstractThe emerging Personalized Federated Learning (PFL) methods aim to produce personalized models for different users, so as to keep track of their individualized requirements in Edge Computing (EC). The centralized PFL methods may suffer from the communication bottleneck and single point of failure. As an alternative solution, the decentralized PFL (DPFL) methods are performed in a Peer-to-Peer (P2P) manner, and collaboratively train personalized models by model aggregation among congenial devices. However, these DPFL methods may incur high communication cost and low resource utilization induced by large-scale models. Herein, we take the communication constraint and heterogeneity into consideration and propose to realize communication-efficient DPFL with adaptive model pruning and neighbor selection. We theoretically analyze the convergence of the proposed DPFL method, and study the impacts of both model pruning and neighbor selection on training performance. Furthermore, we propose an efficient algorithm that combines model pruning and neighbor selection to achieve a trade-off between model quality and communication cost. Extensive simulation and testbed experiments on real-world datasets are conducted. The experimental results demonstrate that the proposed algorithm can improve the test accuracy by at most 13% and save the traffic consumption by 45.4% on average compared with the existing PFL methods. Zhen-guo Ma, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Yinxing Xue |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | FedLC: Accelerating Asynchronous Federated Learning in Edge ComputingabstractFederated Learning (FL) has been widely adopted to process the enormous data in the application scenarios like Edge Computing (EC). However, the commonly-used synchronous mechanism in FL may incur unacceptable waiting time for heterogeneous devices, leading to a great strain on the devices' constrained resources. In addition, the alternative asynchronous FL is known to suffer from the model staleness, which will lead to performance degradation of the trained model, especially onnon-i.i.d.data. In this paper, we design a novel asynchronous FL mechanism, named FedLC, to handle thenon-i.i.d.issue in EC by enabling the local collaboration among edge devices. Specifically, apart from uploading the local model directly to the server, each device will transmit its gradient to the other devices with different data distributions for local collaboration, which can improve the model generality. We theoretically analyze the convergence rate of FedLC and obtain the quantitative relationship between convergence bound and local collaboration. We design an efficient algorithm utilizing demand-list to determine the set of devices receiving gradients from each device. To handle the model staleness, we further assign different learning rates for various devices according to their participation frequency. The extensive experimental results demonstrate the effectiveness of our proposed mechanism. Yang Xu 0020, Zhen-guo Ma, Hongli Xu 0001, Suo Chen, Jianchun Liu, Yinxing Xue |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | sGuard+: Machine Learning Guided Rule-Based Automated Vulnerability Repair on Smart ContractsabstractSmart contracts are becoming appealing targets for hackers because of the vast amount of cryptocurrencies under their control. Asset loss due to the exploitation of smart contract codes has increased significantly in recent years. To guarantee that smart contracts are vulnerability-free, there are many works to detect the vulnerabilities of smart contracts, but only a few vulnerability repair works have been proposed. Repairing smart contract vulnerabilities at the source code level is attractive as it is transparent to users, whereas existing repair tools, such as SCRepair and sGuard , suffer from many limitations: (1) ignoring the code of vulnerability prevention; (2) possibly applying the repair to the wrong statements and changing the original business logic of smart contracts; and (3) showing poor performance in terms of time and gas overhead. In this work, we propose machine learning guided rule-based automated vulnerability repair on smart contracts to improve the effectiveness and efficiency of sGuard . To address the limitations mentioned above, we design the features that characterize both the symptoms of vulnerabilities and the methods of vulnerability prevention to learn various vulnerability patterns and reduce false positives. Additionally, a fine-grained localization algorithm is designed by traversing the nodes of the abstract syntax tree, and we refine and extend the repair rules of sGuard to preserve the original business logic of smart contracts and support new vulnerability types. Our tool, named sGuard+ , reduces time overhead based on machine learning models, and reduces gas overhead by fewer code changes and precise patching. In our experiment, we collect a publicly available vulnerability dataset from CVE, SWC, and SmartBugs Curated as a ground truth for evaluations. Overall, sGuard+ repairs more vulnerabilities with less time and gas overhead than state-of-the-art tools. Furthermore, we reproduce about 9,000 historical transactions for regression testing. It is shown that sGuard+ has no impact on the original business logic of smart contracts. Cuifeng Gao, Wenzhang Yang, Jiaming Ye, Yinxing Xue, Jun Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Read It, Don't Watch It: Captioning Bug Recordings AutomaticallyabstractScreen recordings of mobile applications are easy to capture and include a wealth of information, making them a popular mechanism for users to inform developers of the problems encountered in the bug reports. However, watching the bug recordings and efficiently understanding the semantics of user actions can be time-consuming and tedious for developers. Inspired by the conception of the video subtitle in movie industry, we present a lightweight approach CAPdroid to caption bug recordings automatically. CAPdroid is a purely image-based and non-intrusive approach by using image processing and convolutional deep learning models to segment bug recordings, infer user action attributes, and generate subtitle descriptions. The automated experiments demonstrate the good performance of CAPdroid in inferring user actions from the recordings, and a user study confirms the usefulness of our generated step descriptions in assisting developers with bug replay. Sidong Feng, Mulong Xie, Yinxing Xue, Chunyang Chen 0001 |
ICSE | 3 |
| 2023 | DeepArc: Modularizing Neural Networks for the Model MaintenanceabstractNeural networks are an emerging data-driven programming paradigm widely used in many areas. Unlike traditional software systems consisting of decomposable modules, a neural network is usually delivered as a monolithic package, raising challenges for some maintenance tasks such as model restructure and re-adaption. In this work, we propose DeepArc, a novel modularization method for neural networks, to reduce the cost of model maintenance tasks. Specifically, DeepArc decomposes a neural network into several consecutive modules, each of which encapsulates consecutive layers with similar semantics. The network modularization facilitates practical tasks such as refactoring the model to preserve existing features (e.g., model compression) and enhancing the model with new features (e.g., fitting new samples). The modularization and encapsulation allow us to restructure or retrain the model by only pruning and tuning a few localized neurons and layers. Our experiments show that (1) DeepArc can boost the runtime efficiency of the state-of-the-art model compression techniques by 14.8%; (2) compared to the traditional model retraining, DeepArc only needs to train less than 20% of the neurons on average to fit adversarial samples and repair under-performing models, leading to 32.85% faster training performance while achieving similar model prediction performance. Xiaoning Ren, Yun Lin 0001, Yinxing Xue, Jun Sun 0001, Zhiyong Feng 0002, Jin Song Dong 0001 |
ICSE | 3 |
| 2023 | Comparison and Evaluation of Clone Detection Techniques with Different Code RepresentationsabstractAs one of bad smells in code, code clones may increase the cost of software maintenance and the risk of vulnerability propagation. In the past two decades, numerous clone detection technologies have been proposed. They can be divided into text-based, token-based, tree-based, and graph-based approaches according to their code representations. Different code representations abstract the code details from different perspectives. However, it is unclear which code representation is more effective in detecting code clones and how to combine different code representations to achieve ideal performance. In this paper, we present an empirical study to compare the clone detection ability of different code representations. Specifically, we reproduce 12 clone detection algorithms and divide them into different groups according to their code representations. After analyzing the empirical results, we find that token and tree representations can perform better than graph representation when detecting simple code clones. However, when the code complexity of a code pair increases, graph representation becomes more effective. To make our findings more practical, we perform manual analysis on open-source projects to seek a possible distribution of different clone types in the open-source community. Through the results, we observe that most clone pairs belong to simple code clones. Based on this observation, we discard heavyweight graph-based clone detection algorithms and conduct combination experiments to find out a suitable combination of token-based and tree-based approaches for achieving scalable and effective code clone detection. We develop the suitable combination into a tool called TACC and evaluate it with other state-of-the-art code clone detectors. Experimental results indicate that TACC performs better and has the ability to detect large-scale code clones. Yuekun Wang, Yuhang Ye 0004, Yueming Wu 0001, Yinxing Xue, Yang Liu 0003 |
ICSE | 5 |
| 2023 | EvoScenario: Integrating Road Structures into Critical Scenario Generation for Autonomous Driving System TestingabstractAutonomous Driving Systems (ADS) are safety-critical and require comprehensive testing before their deployment on public roads. Most existing testing approaches consist in generating scenarios that vary the behaviors of dynamic objects, while leaving a predefined road environment unchanged. Consequently, these approaches overlook the influence of different road structures on ADS safety, e.g., collisions can happen more frequently than usual on a merging road, because of the specific road structure. In this paper, we propose EvoScenario, a novel approach that integrates road structures into the generation of critical scenarios for exposing safety risks of ADS. Specifically, EvoScenario models a driving road as a sequence of road segments characterized in different aspects, such as their shapes and widths. Then, a test case is defined by concatenating the sequence of road segments and the sequence of dynamic object maneuvers. Inspired by EvoSuite that generates sequential method calls for Java unit testing, EvoScenario leverages the sequential models of test cases and constructs a multi-objective optimization framework to search for critical scenarios. We implement and demonstrate EvoScenario on an ADS provided by our industrial partner. Evaluation results show that EvoScenario can identify 6 types of safety violations, and outperform existing baseline testing approaches. Shuncheng Tang, Zhenya Zhang 0001, Jixiang Zhou, Yuan Zhou 0005, Yan-Fu Li, Yinxing Xue |
ISSRE | 6 |
| 2023 | From Collision to Verdict: Responsibility Attribution for Autonomous Driving Systems TestingabstractAutonomous driving systems (ADS) are safety-critical systems that require thorough testing to ensure their safety. Current testing methods for ADS primarily focus on finding crash scenarios involving ADS. However, most of these scenarios are unavoidable by ADS, such as collisions caused by the reckless behavior of other vehicles. To address this limitation, we propose CollVer, a framework designed to generate and identify scenarios in which ADS violate driving rules. Specifically, CollVer utilizes multi-modal technology by taking the violation scenario and the corresponding accident description as inputs to judge whether the accident can be attributed to the ADS. Moreover, CollVer introduces a metric called collision position coverage (CPC), to quantify and guide the selection of test cases. Finally, CollVer integrates the multi-modal model and the CPC metric into a multi-objective genetic algorithm to explore more diverse and challenging scenarios. We evaluate CollVer on an industrial-grade ADS, Baidu Apollo, and experimental results show that CollVer can identify 10 distinct types of safety violations, with 4 of them resulting from ADS violating driving rules. Jixiang Zhou, Shuncheng Tang, Yan-Fu Li, Yinxing Xue |
ISSRE | 5 |
| 2023 | Prediction of Vulnerability Characteristics Based on Vulnerability Description and Prompt LearningabstractIdentifying which vulnerabilities need to be prioritized is a long-term challenge in IT security, especially as the number of vulnerabilities grows. Faced with a large number of vulnerability reports, there is an urgent need for automated tools or models to assess the potential severity and exploitability of vulnerabilities. This will help security experts screen vulnerabilities that should be focused on. In this study, we aim to predict vulnerability severity and exploitability characteristics using only vulnerability descriptions. Some previous studies are based on traditional deep learning models, and their performance is relatively backward in the current era of pre-trained language models (PLMs). Therefore, we introduce a prompt learning method based on PLMs to predict vulnerability characteristics. The conventional fine-tuning PLMs method is difficult to make full use of the domain knowledge in PLMs and performs poorly with less training data. Unlike the fine-tuning paradigm, prompt learning imitates the pre-training process of PLM by reconstructing the task input and adding prompts, and uses the output of PLM itself as the prediction output. Combined with prompt ensembling and transfer learning, the performance of prompt learning in the above tasks is further improved. Our experiments show that prompt learning can make more effective use of the knowledge in PLMs. Compared with fine-tuning PLMs and other deep learning models, prompt learning based on BERT or RoBERTa achieves better performance in the above tasks. This advantage is more significant in predicting exploitability with few samples, which proves the ability of prompt learning in few-sample scenarios. In addition, prompt learning also shows the transferability between different tasks in the domain. Xiangwei Li, Xiaoning Ren, Yinxing Xue, Zhenchang Xing, Jiamou Sun |
SANER | 3 |
| 2023 | CCStokener: Fast yet accurate code clone detection with semantic token
Zihan Deng, Yinxing Xue |
J. Syst. Softw. | 3 |
| 2023 | Adaptive Batch Size for Federated Learning in Resource-Constrained Edge ComputingabstractThe emerging Federated Learning (FL) enables IoT devices to collaboratively learn a shared model based on their local datasets. However, due to end devices’ heterogeneity, it will magnify the inherent synchronization barrier issue of FL and result in non-negligible waiting time when local models are trained with the identical batch size. Moreover, the useless waiting time will further lead to a great strain on devices’ limited battery life. Herein, we aim to alleviate the negative impact of synchronization barrier through adaptive batch size during model training. When using different batch sizes, stability and convergence of the global model should be enforced by assigning appropriate learning rates on different devices. Therefore, we first study the relationship between batch size and learning rate, and formulate a scaling rule to guide the setting of learning rate in terms of batch size. Then we theoretically analyze the convergence rate of global model and obtain a convergence upper bound. On these bases, we propose an efficient algorithm that adaptively adjusts batch size with scaled learning rate for heterogeneous devices to reduce the waiting time and save battery life. We conduct extensive simulations and testbed experiments, and the experimental results demonstrate the effectiveness of our method. Zhen-guo Ma, Yang Xu 0020, Hongli Xu 0001, Zeyu Meng, Liusheng Huang, Yinxing Xue |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | A Survey on Automated Driving System Testing: Landscapes and TrendsabstractAutomated Driving Systems ( ADS ) have made great achievements in recent years thanks to the efforts from both academia and industry. A typical ADS is composed of multiple modules, including sensing, perception, planning, and control, which brings together the latest advances in different domains. Despite these achievements, safety assurance of ADS is of great significance, since unsafe behavior of ADS can bring catastrophic consequences. Testing has been recognized as an important system validation approach that aims to expose unsafe system behavior; however, in the context of ADS, it is extremely challenging to devise effective testing techniques, due to the high complexity and multidisciplinarity of the systems. There has been great much literature that focuses on the testing of ADS, and a number of surveys have also emerged to summarize the technical advances. Most of the surveys focus on the system-level testing performed within software simulators, and they thereby ignore the distinct features of different modules. In this article, we provide a comprehensive survey on the existing ADS testing literature, which takes into account both module-level and system-level testing. Specifically, we make the following contributions: (1) We survey the module-level testing techniques for ADS and highlight the technical differences affected by the features of different modules; (2) we also survey the system-level testing techniques, with focuses on the empirical studies that summarize the issues occurring in system development or deployment, the problems due to the collaborations between different modules, and the gap between ADS testing in simulators and the real world; and (3) we identify the challenges and opportunities in ADS testing, which pave the path to the future research in this field. Shuncheng Tang, Zhenya Zhang 0001, Jixiang Zhou, Shuang Liu 0007, Shengjian Guo, Yan-Fu Li, Lei Ma 0003, Yinxing Xue, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 10 |
| 2023 | Challenging Machine Learning-Based Clone Detectors via Semantic-Preserving Code TransformationsabstractSoftware clone detection identifies similar or identical code snippets. It has been an active research topic that attracts extensive attention over the last two decades. In recent years, machine learning (ML) based detectors, especially deep learning-based ones, have demonstrated impressive capability on clone detection. It seems that this longstanding problem has already been tamed owing to the advances in ML techniques. In this work, we would like to challenge the robustness of the recent ML-based clone detectors through code semantic-preserving transformations. We first utilize fifteen simple code transformation operators combined with commonly-used heuristics (i.e., Random Search, Genetic Algorithm, and Markov Chain Monte Carlo) to perform equivalent program transformation. Furthermore, we propose a deep reinforcement learning-based sequence generation (DRLSG) strategy to effectively guide the search process of generating clones that could escape from the detection. We then evaluate the ML-based detectors with the pairs of original and generated clones. We realize our method in a framework named CloneGen (stands for Clone Generator). CloneGen In evaluation, we challenge the three state-of-the-art ML-based detectors and four traditional detectors with the code clones after semantic-preserving transformations via the aid of CloneGen. Surprisingly, our experiments show that, despite the notable successes achieved by existing clone detectors, the ML models inside these detectors still cannot distinguish numerous clones produced by the code transformations in CloneGen. In addition, adversarial training of ML-based clone detectors using clones generated by CloneGen can improve their robustness and accuracy. Meanwhile, compared with the commonly-used heuristics, the DRLSG strategy has shown the best effectiveness in generating code clones to decrease the detection accuracy of the ML-based detectors. Our investigation reveals an explicable but always ignored robustness issue of the latest ML-based detectors. Therefore, we call for more attention to the robustness of these new ML-based detectors. Shengjian Guo, Hongyu Zhang 0002, Yulei Sui, Yinxing Xue |
IEEE Trans. Software Eng. | 5 |
| 2022 | PRCBERT: Prompt Learning for Requirement Classification using BERT-based Pretrained Language ModelsabstractSoftware requirement classification is a longstanding and important problem in requirement engineering. Previous studies have applied various machine learning techniques for this problem, including Support Vector Machine (SVM) and decision trees. With the recent popularity of NLP technique, the state-of-the-art approach NoRBERT utilizes the pre-trained language model BERT and achieves a satisfactory performance. However, the dataset PROMISE used by the existing approaches for this problem consists of only hundreds of requirements that are outdated according to today’s technology and market trends. Besides, the NLP technique applied in these approaches might be obsolete. In this paper, we propose an approach of prompt learning for requirement classification using BERT-based pretrained language models (PRCBERT), which applies flexible prompt templates to achieve accurate requirements classification. Experiments conducted on two existing small-size requirement datasets (PROMISE and NFR-Review) and our collected large-scale requirement dataset NFR-SO prove that PRCBERT exhibits moderately better classification performance than NoRBERT and MLM-BERT (BERT with the standard prompt template). On the de-labeled NFR-Review and NFR-SO datasets, Trans_PRCBERT (the version of PRCBERT which is fine-tuned on PROMISE) is able to have a satisfactory zero-shot performance with 53.27% and 72.96% F1-score when enabling a self-learning strategy. Xianchang Luo, Yinxing Xue, Zhenchang Xing, Jiamou Sun |
ASE | 2 |
| 2022 | Unleashing the power of pseudo-code for binary code similarity analysisabstractAbstract Code similarity analysis has become more popular due to its significant applicantions, including vulnerability detection, malware detection, and patch analysis. Since the source code of the software is difficult to obtain under most circumstances, binary-level code similarity analysis (BCSA) has been paid much attention to. In recent years, many BCSA studies incorporating AI techniques focus on deriving semantic information from binary functions with code representations such as assembly code, intermediate representations, and control flow graphs to measure the similarity. However, due to the impacts of different compilers, architectures, and obfuscations, binaries compiled from the same source code may vary considerably, which becomes the major obstacle for these works to obtain robust features. In this paper, we propose a solution, named UPPC (Unleashing the Power of Pseudo-code), which leverages the pseudo-code of binary function as input, to address the binary code similarity analysis challenge, since pseudo-code has higher abstraction and is platform-independent compared to binary instructions. UPPC selectively inlines the functions to capture the full function semantics across different compiler optimization levels and uses a deep pyramidal convolutional neural network to obtain the semantic embedding of the function. We evaluated UPPC on a data set containing vulnerabilities and a data set including different architectures (X86, ARM), different optimization options (O0-O3), different compilers (GCC, Clang), and four obfuscation strategies. The experimental results show that the accuracy of UPPC in function search is 33.2% higher than that of existing methods. Zhengzi Xu, Yang Xiao 0011, Yinxing Xue |
Cybersecur. | 4 |
| 2022 | On the usage and development of deep learning compilers: an empirical study on TVM
Xiongfei Wu, Jinqiu Yang 0001, Lei Ma 0003, Yinxing Xue, Jianjun Zhao 0001 |
Empir. Softw. Eng. | 4 |
| 2022 | Multi-objective integer programming approaches to Next Release Problem - Enhancing exact methods for finding whole pareto front
Shi Dong 0006, Yinxing Xue, Sjaak Brinkkemper, Yan-Fu Li |
Inf. Softw. Technol. | 2 |
| 2022 | Vulpedia: Detecting vulnerable ethereum smart contracts via abstracted vulnerability signatures
Jiaming Ye, Mingliang Ma, Yun Lin 0001, Lei Ma 0003, Yinxing Xue, Jianjun Zhao 0001 |
J. Syst. Softw. | 5 |
| 2021 | A lightweight framework for function name reassignment based on large-scale stripped binariesabstractSoftware in the wild is usually released as stripped binaries that contain no debug information (e.g., function names). This paper studies the issue of reassigning descriptive names for functions to help facilitate reverse engineering. Since the essence of this issue is a data-driven prediction task, persuasive research should be based on sufficiently large-scale and diverse data. However, prior studies can only be based on small-scale datasets because their techniques suffer from heavyweight binary analysis, making them powerless in the face of big-size and large-scale binaries. Han Gao 0014, Shaoyin Cheng, Yinxing Xue, Weiming Zhang 0001 |
ISSTA | 3 |
| 2021 | An empirical study of GUI widget detection for industrial mobile gamesabstractWith the widespread adoption of smartphones in our daily life, mobile games experienced increasing demand over the past years. Meanwhile, the quality of mobile games has been continuously drawing more and more attention, which can greatly affect the player experience. For better quality assurance, general-purpose testing has been extensively studied for mobile apps. However, due to the unique characteristic of mobile games, existing mobile testing techniques may not be directly suitable and applicable. To better understand the challenges in mobile game testing, in this paper, we first initiate an early step to conduct an empirical study towards understanding the challenges and pain points of mobile game testing process at our industrial partner NetEase Games. Specifically, we first conduct a survey from the mobile test development team at NetEase Games via both scrum interviews and questionnaires. We found that accurate and effective GUI widget detection for mobile games could be the pillar to boost the automation of mobile game testing and other downstream analysis tasks in practice. Jiaming Ye, Xiaofei Xie, Lei Ma 0003, Ruochen Huang, Yinxing Xue, Jianjun Zhao 0001 |
ESEC/SIGSOFT FSE | 7 |
| 2021 | APICraft: Fuzz Driver Generation for Closed-source SDK Libraries
Cen Zhang, Xingwei Lin, Yuekang Li, Yinxing Xue, Jundong Xie, Hongxu Chen 0001, Xinlei Ying, Jiashui Wang, Yang Liu 0003 |
USENIX Security Symposium | 4 |
| 2021 | Erratum to "Accurate and Scalable Cross-Architecture Cross-OS Binary Code Search With Emulation"
Yinxing Xue, Zhengzi Xu, Mahinthan Chandramohan, Yang Liu 0003 |
IEEE Trans. Software Eng. | 1 |
| 2020 | An empirical assessment of security risks of global Android banking appsabstractMobile banking apps, belonging to the most security-critical app category, render massive and dynamic transactions susceptible to security risks. Given huge potential financial loss caused by vulnerabilities, existing research lacks a comprehensive empirical study on the security risks of global banking apps to provide useful insights and improve the security of banking apps. Sen Chen 0001, Lingling Fan 0003, Guozhu Meng, Ting Su 0001, Minhui Xue 0001, Yinxing Xue, Yang Liu 0003, Lihua Xu |
ICSE | 6 |
| 2020 | How are Deep Learning Models Similar?: An Empirical Study on Clone Analysis of Deep Learning SoftwareabstractDeep learning (DL) has been successfully applied to many cutting-edge applications, e.g., image processing, speech recognition, and natural language processing. As more and more DL software is made open-sourced, publicly available, and organized in model repositories and stores (Model Zoo, ModelDepot), there comes a need to understand the relationships of these DL models regarding their maintenance and evolution tasks. Although clone analysis has been extensively studied for traditional software, up to the present, clone analysis has not been investigated for DL software. Since DL software adopts the data-driven development paradigm, it is still not clear whether and to what extent the clone analysis techniques of traditional software could be adapted to DL software. Xiongfei Wu, Liangyu Qin, Xiaofei Xie, Lei Ma 0003, Yinxing Xue, Yang Liu 0003, Jianjun Zhao 0001 |
ICPC | 6 |
| 2020 | Cross-Contract Static Analysis for Detecting Practical Reentrancy Vulnerabilities in Smart ContractsabstractReentrancy bugs, one of the most severe vulnerabilities in smart contracts, have caused huge financial loss in recent years. Researchers have proposed many approaches to detecting them. However, empirical studies have shown that these approaches suffer from undesirable false positives and false negatives, when the code under detection involves the interaction between multiple smart contracts. Yinxing Xue, Mingliang Ma, Yun Lin 0001, Yulei Sui, Jiaming Ye, Tianyong Peng |
ASE | 1 |
| 2020 | CCGraph: a PDG-based code clone detector with approximate graph matchingabstractSoftware clone detection is an active research area, which is very important for software maintenance, bug detection, etc. The two pieces of cloned code reflect some similarities or equivalents in the syntax or structure of the code representations. There are many representations of code like AST, token, PDG, etc. The PDG (Program Dependency Graph) of source code can contain both syntactic and structural information. However, most existing PDG-based tools are quite time-consuming and miss many clones because they detect code clones with exact graph matching by using subgraph isomorphism. In this paper, we propose a novel PDG-based code clone detector, CCGraph, that uses graph kernels. Firstly, we normalize the structure of PDGs and design a two-stage filtering strategy by measuring the characteristic vectors of codes. Then we detect the code clones by using an approximate graph matching algorithm based on the reforming WL (Weisfeiler-Lehman) graph kernel. Experiment results show that CCGraph retains a high accuracy, has both better recall and F1-score values, and detects more semantic clones than other two related state-of-the-art tools. Besides, CCGraph is much more efficient than the existing PDG-based tools. Yue Zou, Bihuan Ban, Yinxing Xue |
ASE | 3 |
| 2020 | MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded Programs
Hongxu Chen 0001, Shengjian Guo, Yinxing Xue, Yulei Sui, Cen Zhang, Yuekang Li, Haijun Wang 0002, Yang Liu 0003 |
USENIX Security Symposium | 3 |
| 2020 | Multi-objective Integer Programming Approaches for Solving the Multi-criteria Test-suite Minimization Problem: Towards Sound and Complete Solutions of a Particular Search-based Software-engineering ProblemabstractTest-suite minimization is one key technique for optimizing the software testing process. Due to the need to balance multiple factors, multi-criteria test-suite minimization (MCTSM) becomes a popular research topic in the recent decade. The MCTSM problem is typically modeled as integer linear programming (ILP) problem and solved with weighted-sum single objective approach. However, there is no existing approach that can generate sound (i.e., being Pareto-optimal) and complete (i.e., covering the entire Pareto front) Pareto-optimal solution set, to the knowledge of the authors. In this work, we first prove that the ILP formulation can accurately model the MCTSM problem and then propose the multi-objective integer programming (MOIP) approaches to solve it. We apply our MOIP approaches on three specific MCTSM problems and compare the results with those of the cutting-edge methods, namely, NonlinearFormulation_LinearSolver (NF_LS) and two Multi-Objective Evolutionary Algorithms (MOEAs). The results show that our MOIP approaches can always find sound and complete solutions on five subject programs, using similar or significantly less time than NF_LS and two MOEAs do. The current experimental results are quite promising, and our approaches have the potential to be applied for other similar search-based software engineering problems. Yinxing Xue, Yan-Fu Li |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2020 | Ant Colony System With Sorting-Based Local Search for Coverage-Based Test Case PrioritizationabstractTest case prioritization (TCP) is a popular regression testing technique in software engineering field. The task of TCP is to schedule the execution order of test cases so that certain objective (e.g., code coverage) can be achieved quickly. In this article, we propose an efficient ant colony system framework for the TCP problem, with the aim of maximizing the code coverage as soon as possible. In the proposed framework, an effective heuristic function is proposed to guide the ants to construct solutions based on additional statement coverage among remaining test cases. Besides, a sorting-based local search mechanism is proposed to further accelerate the convergence speed of the algorithm. Experimental results on different benchmark problems, and a real-world application, have shown that the proposed framework can outperform several state-of-the-art methods, in terms of solution quality and search efficiency. Chengyu Lu, Jinghui Zhong, Yinxing Xue, Liang Feng 0001, Jun Zhang 0003 |
IEEE Trans. Reliab. | 3 |
| 2019 | Cerebro: context-aware adaptive fuzzing for effective vulnerability detectionabstractExisting greybox fuzzers mainly utilize program coverage as the goal to guide the fuzzing process. To maximize their outputs, coverage-based greybox fuzzers need to evaluate the quality of seeds properly, which involves making two decisions: 1) which is the most promising seed to fuzz next (seed prioritization), and 2) how many efforts should be made to the current seed (power scheduling). In this paper, we present our fuzzer, Cerebro, to address the above challenges. For the seed prioritization problem, we propose an online multi-objective based algorithm to balance various metrics such as code complexity, coverage, execution time, etc. To address the power scheduling problem, we introduce the concept of input potential to measure the complexity of uncovered code and propose a cost-effective algorithm to update it dynamically. Unlike previous approaches where the fuzzer evaluates an input solely based on the execution traces that it has covered, Cerebro is able to foresee the benefits of fuzzing the input by adaptively evaluating its input potential. We perform a thorough evaluation for Cerebro on 8 different real-world programs. The experiments show that Cerebro can find more vulnerabilities and achieve better coverage than state-of-the-art fuzzers such as AFL and AFLFast. Yuekang Li, Yinxing Xue, Hongxu Chen 0001, Xiuheng Wu, Cen Zhang, Xiaofei Xie, Haijun Wang 0002, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 2 |
| 2019 | Securing Android App Markets via Modeling and Predicting Malware Spread Between MarketsabstractThe Android ecosystem has recently dominated mobile devices. Android app markets, including official Google Play and other third party markets, are becoming hotbeds, where malware originates and spreads. Android malware has been observed to both propagate within markets and spread between markets. If the spread of Android malware between markets can be predicted, market administrators can take appropriate measures to prevent the outbreak of malware and minimize the damages caused by malware. In this paper, we make the first attempt to protect the Android ecosystem by modeling and predicting the spread of Android malware between markets. To this end, we study the social behaviors that affect the spread of malware, model these spread behaviors with multiple epidemic models, and predict the infection time and order among markets for well-known malware families. To achieve an accurate prediction of malware spread, we model spread behaviors in the following fashion: 1) for a single market, we model the within-market malware growth by considering both the creation and removal of malware; 2) for multiple markets, we determine market relevance by calculating the mutual information among them; and 3) based on the previous two steps, we simulate a susceptible infected model stochastically for spread among markets. The model inference is performed using a publicly available well-labeled dataset AndRadar. To conduct extensive experiments to evaluate our approach, we collected a large number (334,782) of malware samples from 25 Android markets around the world. The experimental results show our approach can depict and simulate the growth of Android malware on a large scale, and predict the infection time and order among markets with 0.89 and 0.66 precision, respectively. Guozhu Meng, Matthew Patrick, Yinxing Xue, Yang Liu 0003, Jie Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Accurate and Scalable Cross-Architecture Cross-OS Binary Code Search with EmulationabstractDifferent from source code clone detection, clone detection (similar code search) in binary executables faces big challenges due to the gigantic differences in the syntax and the structure of binary code that result from different configurations of compilers, architectures and OSs. Existing studies have proposed different categories of features for detecting binary code clones, including CFG structures, n-gram in CFG, input/output values, etc. In our previous study and the tool BinGo, to mitigate the huge gaps in CFG structures due to different compilation scenarios, we propose a selective inlining technique to capture the complete function semantics by inlining relevant library and user-defined functions. However, only features of input/output values are considered in BinGo. In this study, we propose to incorporate features from different categories (e.g., structural features and high-level semantic features) for accuracy improvement and emulation for efficiency improvement. We empirically compare our tool, BinGo-E, with the pervious tool BinGo and the available state-of-the-art tools of binary code search in terms of search accuracy and performance. Results show that BinGo-E achieves significantly better accuracies than BinGo for cross-architecture matching, cross-OS matching, cross-compiler matching and intra-compiler matching. Additionally, in the new task of matching binaries of forked projects, BinGo-E also exhibits a better accuracy than the existing benchmark tool. Meanwhile, BinGo-E takes less time than BinGo during the process of matching. Yinxing Xue, Zhengzi Xu, Mahinthan Chandramohan, Yang Liu 0003 |
IEEE Trans. Software Eng. | 1 |
| 2018 | Hawkeye: Towards a Desired Directed Grey-box FuzzerabstractGrey-box fuzzing is a practically effective approach to test real-world programs. However, most existing grey-box fuzzers lack directedness, i.e. the capability of executing towards user-specified target sites in the program. To emphasize existing challenges in directed fuzzing, we propose Hawkeye to feature four desired properties of directed grey-box fuzzers. Owing to a novel static analysis on the program under test and the target sites, Hawkeye precisely collects the information such as the call graph, function and basic block level distances to the targets. During fuzzing, Hawkeye evaluates exercised seeds based on both static information and the execution traces to generate the dynamic metrics, which are then used for seed prioritization, power scheduling and adaptive mutating. These strategies help Hawkeye to achieve better directedness and gravitate towards the target sites. We implemented Hawkeye as a fuzzing framework and evaluated it on various real-world programs under different scenarios. The experimental results showed that Hawkeye can reach the target sites and reproduce the crashes much faster than state-of-the-art grey-box fuzzers such as AFL and AFLGo. Specially, Hawkeye can reduce the time to exposure for certain vulnerabilities from about 3.5 hours to 0.5 hour. By now, Hawkeye has detected more than 41 previously unknown crashes in projects such as Oniguruma, MJS with the target sites provided by vulnerability prediction tools; all these crashes are confirmed and 15 of them have been assigned CVE IDs. Hongxu Chen 0001, Yinxing Xue, Yuekang Li, Bihuan Chen 0001, Xiaofei Xie, Xiuheng Wu, Yang Liu 0003 |
CCS | 2 |
| 2018 | Multi-objective integer programming approaches for solving optimal feature selection problem: a new perspective on multi-objective optimization problems in SBSEabstractThe optimal feature selection problem in software product line is typically addressed by the approaches based on Indicator-based Evolutionary Algorithm (IBEA). In this study we first expose the mathematical nature of this problem --- multi-objective binary integer linear programming. Then, we implement/propose three mathematical programming approaches to solve this problem at different scales. For small-scale problems (roughly less than 100 features), we implement two established approaches to find all exact solutions. For medium-to-large problems (roughly, more than 100 features), we propose one efficient approach that can generate a representation of the entire Pareto front in linear time complexity. The empirical results show that our proposed method can find significantly more non-dominated solutions in similar or less execution time, in comparison with IBEA and its recent enhancement (i.e., IBED that combines IBEA and Differential Evolution). Yinxing Xue, Yan-Fu Li |
ICSE | 1 |
| 2018 | FOT: a versatile, configurable, extensible fuzzing frameworkabstractGreybox fuzzing is one of the most effective approaches for detecting software vulnerabilities. Various new techniques have been continuously emerging to enhance the effectiveness and/or efficiency by incorporating novel ideas into different components of a greybox fuzzer. However, there lacks a modularized fuzzing framework that can easily plugin new techniques and hence facilitate the reuse, integration and comparison of different techniques. To address this problem, we propose a fuzzing framework, namely Fuzzing Orchestration Toolkit (FOT). FOT is designed to be versatile, configurable and extensible. With FOT and its extensions, we have found 111 new bugs from 11 projects. Among these bugs, 18 CVEs have been assigned. Video link: https://youtu.be/O6Qu7BJ8RP0. Hongxu Chen 0001, Yuekang Li, Bihuan Chen 0001, Yinxing Xue, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 4 |
| 2017 | Feedback-based debuggingabstractSoftware debugging has long been regarded as a time and effort consuming task. In the process of debugging, developers usually need to manually inspect many program steps to see whether they deviate from their intended behaviors. Given that intended behaviors usually exist nowhere but in human mind, the automation of debugging turns out to be extremely hard, if not impossible. In this work, we propose a feedback-based debugging approach, which (1) builds on light-weight human feedbacks on a buggy program and (2) regards the feedbacks as partial program specification to infer suspicious steps of the buggy execution. Given a buggy program, we record its execution trace and allow developers to provide light-weight feedback on trace steps. Based on the feedbacks, we recommend suspicious steps on the trace. Moreover, our approach can further learn and approximate bug-free paths, which helps reduce required feedbacks to expedite the debugging process. We conduct an experiment to evaluate our approach with simulated feedbacks on 3409 mutated bugs across 3 open source projects. The results show that our feedback-based approach can detect 92.8% of the bugs and 65% of the detected bugs require less than 20 feedbacks. In addition, we implement our proof-of-concept tool, Microbat, and conduct a user study involving 16 participants on 3 debugging tasks. The results show that, compared to the participants using the baseline tool, Whyline, the ones using Microbat can spend on average 55.8% less time to locate the bugs. Yun Lin 0001, Jun Sun 0001, Yinxing Xue, Yang Liu 0003, Jin Song Dong 0001 |
ICSE | 3 |
| 2017 | Mining implicit design templates for actionable code reuseabstractIn this paper, we propose an approach to detecting project-specific recurring designs in code base and abstracting them into design templates as reuse opportunities. The mined templates allow programmers to make further customization for generating new code. The generated code involves the code skeleton of recurring design as well as the semi-implemented code bodies annotated with comments to remind programmers of necessary modification. We implemented our approach as an Eclipse plugin called MICoDe. We evaluated our approach with a reuse simulation experiment and a user study involving 16 participants. The results of our simulation experiment on 10 open source Java projects show that, to create a new similar feature with a design template, (1) on average 69% of the elements in the template can be reused and (2) on average 60% code of the new feature can be adopted from the template. Our user study further shows that, compared to the participants adopting the copy-paste-modify strategy, the ones using MICoDe are more effective to understand a big design picture and more efficient to accomplish the code reuse task. Yun Lin 0001, Guozhu Meng, Yinxing Xue, Zhenchang Xing, Jun Sun 0001, Xin Peng 0001, Yang Liu 0003, Wenyun Zhao, Jin Song Dong 0001 |
ASE | 3 |
| 2017 | Auditing Anti-Malware Tools by Evolving Android Malware and Dynamic Loading TechniqueabstractAlthough a previous paper shows that existing anti-malware tools (AMTs) may have high detection rate, the report is based on existing malware and thus it does not imply that AMTs can effectively deal with future malware. It is desirable to have an alternative way of auditing AMTs. In our previous paper, we use malware samples from android malware collection Genome to summarize a malware meta-model for modularizing the common attack behaviors and evasion techniques in reusable features. We then combine different features with an evolutionary algorithm, in which way we evolve malware for variants. Previous results have shown that the existing AMTs only exhibit detection rate of 20%-30% for 10 000 evolved malware variants. In this paper, based on the modularized attack features, we apply the dynamic code generation and loading techniques to produce malware, so that we can audit the AMTs at runtime. We implement our approach, named Mystique-S, as a service-oriented malware generation system. Mystique-S automatically selects attack features under various user scenarios and delivers the corresponding malicious payloads at runtime. Relying on dynamic code binding (via service) and loading (via reflection) techniques, Mystique-S enables dynamic execution of payloads on user devices at runtime. Experimental results on real-world devices show that existing AMTs are incapable of detecting most of our generated malware. Last, we propose the enhancements for existing AMTs. Yinxing Xue, Guozhu Meng, Yang Liu 0003, Tian Huat Tan, Hongxu Chen 0001, Jun Sun 0001, Jie Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Mystique: Evolving Android Malware for Auditing Anti-Malware ToolsabstractIn the arms race of attackers and defenders, the defense is usually more challenging than the attack due to the unpredicted vulnerabilities and newly emerging attacks every day. Currently, most of existing malware detection solutions are individually proposed to address certain types of attacks or certain evasion techniques. Thus, it is desired to conduct a systematic investigation and evaluation of anti-malware solutions and tools based on different attacks and evasion techniques. In this paper, we first propose a meta model for Android malware to capture the common attack features and evasion features in the malware. Based on this model, we develop a framework, MYSTIQUE, to automatically generate malware covering four attack features and two evasion features, by adopting the software product line engineering approach. With the help of MYSTIQUE, we conduct experiments to 1) understand Android malware and the associated attack features as well as evasion techniques; 2) evaluate and compare the 57 off-the-shelf anti-malware tools, 9 academic solutions and 4 App market vetting processes in terms of accuracy in detecting attack features and capability in addressing evasion. Last but not least, we provide a benchmark of Android malware with proper labeling of contained attack and evasion features. Guozhu Meng, Yinxing Xue, Mahinthan Chandramohan, Annamalai Narayanan, Yang Liu 0003, Jie Zhang 0002, Tieming Chen |
AsiaCCS | 2 |
| 2016 | Optimizing selection of competing services with probabilistic hierarchical refinementabstractRecently, many large enterprises (e.g., Netflix, Amazon) have decomposed their monolithic application into services, and composed them to fulfill their business functionalities. Many hosting services on the cloud, with different Quality of Service (QoS) (e.g., availability, cost), can be used to host the services. This is an example of competing services. QoS is crucial for the satisfaction of users. It is important to choose a set of services that maximize the overall QoS, and satisfy all QoS requirements for the service composition. This problem, known as optimal service selection, is NP-hard. Therefore, an effective method for reducing the search space and guiding the search process is highly desirable. To this end, we introduce a novel technique, called Probabilistic Hierarchical Refinement (ProHR). ProHR effectively reduces the search space by removing competing services that cannot be part of the selection. ProHR provides two methods, probabilistic ranking and hierarchical refinement, that enable smart exploration of the reduced search space. Unlike existing approaches that perform poorly when QoS requirements become stricter, ProHR maintains high performance and accuracy, independent of the strictness of the QoS requirements. ProHR has been evaluated on a publicly available dataset, and has shown significant improvement over existing approaches. Tian Huat Tan, Manman Chen, Jun Sun 0001, Yang Liu 0003, Étienne André 0001, Yinxing Xue, Jin Song Dong 0001 |
ICSE | 6 |
| 2016 | Semantic modelling of Android malware for effective malware comprehension, detection, and classificationabstractMalware has posed a major threat to the Android ecosystem. Existing malware detection tools mainly rely on signature- or feature- based approaches, failing to provide detailed information beyond the mere detection. In this work, we propose a precise semantic model of Android malware based on Deterministic Symbolic Automaton (DSA) for the purpose of malware comprehension, detection and classification. It shows that DSA can capture the common malicious behaviors of a malware family, as well as the malware variants. Based on DSA, we develop an automatic analysis framework, named SMART, which learns DSA by detecting and summarizing semantic clones from malware families, and then extracts semantic features from the learned DSA to classify malware according to the attack patterns. We conduct the experiments in both malware benchmark and 223,170 real-world apps. The results show that SMART builds meaningful semantic models and outperforms both state-of-the-art approaches and anti-virus tools in malware detection. SMART identifies 4583 new malware in real-world apps that are missed by most anti-virus tools. The classification step further identifies new malware variants and unknown families. Guozhu Meng, Yinxing Xue, Zhengzi Xu, Yang Liu 0003, Jie Zhang 0002, Annamalai Narayanan |
ISSTA | 2 |
| 2016 | BinGo: cross-architecture cross-OS binary searchabstractBinary code search has received much attention recently due to its impactful applications, e.g., plagiarism detection, malware detection and software vulnerability auditing. However, developing an effective binary code search tool is challenging due to the gigantic syntax and structural differences in binaries resulted from different compilers, architectures and OSs. In this paper, we propose BINGO — a scalable and robust binary search engine supporting various architectures and OSs. The key contribution is a selective inlining technique to capture the complete function semantics by inlining relevant library and user-defined functions. In addition, architecture and OS neutral function filtering is proposed to dramatically reduce the irrelevant target functions. Besides, we introduce length variant partial traces to model binary functions in a program structure agnostic fashion. The experimental results show that BINGO can find semantic similar functions across architecture and OS boundaries, even with the presence of program structure distortion, in a scalable manner. Using BINGO, we also discovered a zero-day vulnerability in Adobe PDF Reader, a COTS binary. Mahinthan Chandramohan, Yinxing Xue, Zhengzi Xu, Yang Liu 0003, Chia Yuan Cho, Hee Beng Kuan Tan |
SIGSOFT FSE | 2 |
| 2015 | JSDC: A Hybrid Approach for JavaScript Malware Detection and ClassificationabstractMalicious JavaScript is one of the biggest threats in cyber security. Existing research and anti-virus products mainly focus on detection of JavaScript malware rather than classification. Usually, the detection will simply report the malware family name without elaborating details about attacks conducted by the malware. Worse yet, the reported family name may differ from one tool to another due to the different naming conventions. In this paper, we propose a hybrid approach to perform JavaScript malware detection and classification in an accurate and efficient way, which could not only explain the attack model but also potentially discover new malware variants and new vulnerabilities. Our approach starts with machine learning techniques to detect JavaScript malware using predicative features of textual information, program structures and risky function calls. For the detected malware, we classify them into eight known attack types according to their attack feature vector or dynamic execution traces by using machine learning and dynamic program analysis respectively. We implement our approach in a tool named JSDC, and conduct large-scale evaluations to show its effectiveness. The controlled experiments (with 942 malware) show that JSDC gives low false positive rate (0.2123%) and low false negative rate (0.8492%), compared with other tools. We further apply JSDC on 1,400,000 real-world JavaScript with over 1,500 malware reported, for which many anti-virus tools failed. Lastly, JSDC can effectively and accurately classify these detected malwares into either attack types. Junjie Wang 0007, Yinxing Xue, Yang Liu 0003, Tian Huat Tan |
AsiaCCS | 2 |
| 2015 | An Adaptive Markov Strategy for Effective Network Intrusion DetectionabstractNetwork monitoring is an important way to ensure the security of hosts from being attacked by malicious attackers. One challenging problem for network operators is how to distribute the limited monitoring resources (e.g., intrusion detectors) among the network to detect attacks in a cost-effective manner, especially when the attacking strategies can be changing dynamically and unpredictable. To this end, we adopt Markov game to model the interactions between the network operator and the attacker and propose an adaptive Markov strategy (AMS) to determine how the detectors should be placed on the network against possible attacks to minimize the network's accumulated cost over time. The AMS is guaranteed to converge to the best response strategy when the attacker's strategy is fixed (rationality), converge to a fixed strategy under self-play (convergence) and obtain a payoff no less than that under the precomputed Nash equilibrium strategy of the Markov game (safety). The experimental results show that the AMS can achieve better protection for the network compared with both previous approaches based on the prediction of attack paths (equivalent to a graph coloring problem) and Nash equilibrium strategy. Jianye Hao, Yinxing Xue, Mahinthan Chandramohan, Yang Liu 0003, Jun Sun 0001 |
ICTAI | 2 |
| 2015 | Optimizing selection of competing features via feedback-directed evolutionary algorithmsabstractSoftware that support various groups of customers usually require complicated configurations to attain different functionalities. To model the configuration options, feature model is proposed to capture the commonalities and competing variabilities of the product variants in software family or Software Product Line (SPL). A key challenge for deriving a new product is to find a set of features that do not have inconsistencies or conflicts, yet optimize multiple objectives (e.g., minimizing cost and maximizing number of features), which are often competing with each other. Existing works have attempted to make use of evolutionary algorithms (EAs) to address this problem. In this work, we incorporated a novel feedback-directed mechanism into existing EAs. Our empirical results have shown that our method has improved noticeably over all unguided version of EAs on the optimal feature selection. In particular, for case studies in SPLOT and LVAT repositories, the feedback-directed Indicator-Based EA (IBEA) has increased the number of correct solutions found by 72.33% and 75%, compared to unguided IBEA. In addition, by leveraging a pre-computed solution, we have found 34 sound solutions for Linux X86, which contains 6888 features, in less than 40 seconds. Tian Huat Tan, Yinxing Xue, Manman Chen, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001 |
ISSTA | 2 |
| 2015 | Detection and classification of malicious JavaScript via attack behavior modellingabstractExisting malicious JavaScript (JS) detection tools and commercial anti-virus tools mostly use feature-based or signature-based approaches to detect JS malware. These tools are weak in resistance to obfuscation and JS malware variants, not mentioning about providing detailed information of attack behaviors. Such limitations root in the incapability of capturing attack behaviors in these approches. In this paper, we propose to use Deterministic Finite Automaton (DFA) to abstract and summarize common behaviors of malicious JS of the same attack type. We propose an automatic behavior learning framework, named JS*, to learn DFAs from dynamic execution traces of JS malware, where we implement an effective online teacher by combining data dependency analysis, defense rules and trace replay mechanism. We evaluate JS* using real world data of 10000 benign and 276 malicious JS samples to cover 8 most-infectious attack types. The results demonstrate the scalability and effectiveness of our approach in the malware detection and classification, compared with commercial anti-virus tools. We also show how to use our DFAs to detect variants and new attacks. Yinxing Xue, Junjie Wang 0007, Yang Liu 0003, Jun Sun 0001, Mahinthan Chandramohan |
ISSTA | 1 |
| 2014 | Detecting differences across multiple instances of code clonesabstractClone detectors find similar code fragments (i.e., instances of code clones) and report large numbers of them for industrial systems. To maintain or manage code clones, developers often have to investigate differences of multiple cloned code fragments. However,existing program differencing techniques compare only two code fragments at a time. Developers then have to manually combine several pairwise differencing results. In this paper, we present an approach to automatically detecting differences across multiple clone instances. We have implemented our approach as an Eclipse plugin and evaluated its accuracy with three Java software systems. Our evaluation shows that our algorithm has precision over 97.66% and recall over 95.63% in three open source Java projects. We also conducted a user study of 18 developers to evaluate the usefulness of our approach for eight clone-related refactoring tasks. Our study shows that our approach can significantly improve developers’performance in refactoring decisions, refactoring details, and task completion time on clone-related refactoring tasks. Automatically detecting differences across multiple clone instances also opens opportunities for building practical applications of code clones in software maintenance, such as auto-generation of application skeleton, intelligent simultaneous code editing. Yun Lin 0001, Zhenchang Xing, Yinxing Xue, Yang Liu 0003, Xin Peng 0001, Jun Sun 0001, Wenyun Zhao |
ICSE | 3 |
| 2013 | A large scale Linux-kernel based benchmark for feature location researchabstractMany software maintenance tasks require locating code units that implement a certain feature (termed as feature location). Feature location has been an active research area for more than two decades. However, there is lack of publicly available, large scale benchmarks for e valuating and comparing feature location approaches. In this paper, we present a LinuxKernel based benchmark for feature location research. This benchmark is large scale and extensible. By providing rich feature and program information and accurate ground-truth links between features and code units, it supports the e valuation of a wide range of feature location approaches. It allows researchers to gain deeper insights into existing approaches and how they can be improved. It also enables communication and collaboration among different researchers. (video: http://www.youtube.com/watch?v=3D_HihwRNeK3I). Zhenchang Xing, Yinxing Xue, Stan Jarzabek |
ICSE | 2 |
| 2012 | Client-Side Rendering Mechanism: A Double-Edged Sword for Browser-Based Web Applications
Yinxing Xue, Keizo Oyama |
SEKE | 2 |
| 2011 | Reengineering legacy software products into software product line based on automatic variability analysisabstractIn order to deliver the various and short time-to-market software products to customers, the paradigm of Software Product Line (SPL) represents a new endeavor to the software development. To migrate a family of legacy software products into SPL for effective reuse, one has to understand commonality and variability among existing products variants. The existing techniques rely on manual identification and modeling of variability, and the analysis based on those techniques is performed at several mutually independent levels of abstraction. We propose a sandwich approach that consolidates feature knowledge from top-down domain analysis with bottom-up analysis of code similarities in subject software products. Our proposed method integrates model differencing, clone detection, and information retrieval techniques, which can provide a systematic means to reengineer the legacy software products into SPL based on automatic variability analysis. Yinxing Xue |
ICSE | 1 |
| 2011 | Improving Product Line Architecture Design and Customization by Raising the Level of Variability Modeling
Xin Peng 0001, Stan Jarzabek, Zhenchang Xing, Yinxing Xue, Wenyun Zhao |
ICSR | 5 |
| 2011 | CloneDifferentiator: Analyzing clones by differentiationabstractClone detection provides a scalable and efficient way to detect similar code fragments. But it offers limited explanation of differences of functions performed by clones and variations of control and data flows of clones. We refer to such differences as semantic differences of clones. Understanding these semantic differences is essential to correctly interpret cloning information and perform maintenance tasks on clones. Manual analysis of semantic differences of clones is complicated and error-prone. In the paper, we present our clone analysis tool, called Clone-Differentiator. Our tool automatically characterizes clones returned by a clone detector by differentiating Program Dependence Graphs (PDGs) of clones. CloneDifferentiator is able to provide a precise characterization of semantic differences of clones. It can provide an effective means of analyzing clones in a task oriented manner. Zhenchang Xing, Yinxing Xue, Stan Jarzabek |
ASE | 2 |
| 2011 | Scalability of Variability Management: An Example of Industrial Practice and Some Improvements
Yinxing Xue, Stan Jarzabek, Pengfei Ye, Xin Peng 0001, Wenyun Zhao |
SEKE | 1 |
| 2009 | Avoiding Some Common Preprocessing Pitfalls with Feature QueriesabstractPreprocessors (e.g., cpp) provide simple means to manage software product variants by including/excluding required feature code to/from base program. Feature-related customizations occur at variation points in base program marked with preprocessing directives. Problems emerge when the number of inter-dependent features grows, and each feature maps to many variation points in many base program components. Component-based and architecture-centric techniques promoted by a Software Product Line approach to reuse help us contain the impact of some features in small number of base components. Still, accommodating other features into product variants requires fine granular code changes in many components, at many variation points. Fine granular code level changes are often handled by preprocessors, which becomes a source of well-known complications during component customization for reuse. In this paper, we show how some of the common preprocessing problems can be alleviated with a query-based environment that assists programmers in analysis of features handled with preprocessor's directives. We describe problems of preprocessing that can be aided by tool like ours, and problems that we believe are inherent in approaches that attempt to manage features in the base code. Stan Jarzabek, Yinxing Xue, Hongyu Zhang 0002, Youpeng Lee |
APSEC | 2 |
| 2009 | A Case Study of Variation Mechanism in an Industrial Product Line
Pengfei Ye, Xin Peng 0001, Yinxing Xue, Stan Jarzabek |
ICSR | 3 |