EDBT 2026 Demo / reviewers in the wild / expert
Xiaoning Du 0001
dblp:97/6476-1
· DBLP profile ↗
44ranked-venue papers
7as first author
36since 2021 · last 2026
0000-0003-3728-9541ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 27 · 6 first-author · 20 since 2021Security and privacy · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSG: Logit-Balanced Vocabulary Partitioning for LLM WatermarkingabstractWatermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs).Among existing approaches, the KGW scheme is particularly attractive due to its versatility, efficiency, and effectiveness in natural language generation.However, KGW's effectiveness degrades significantly under low-entropy settings such as code generation and mathematical reasoning.A crucial step in the KGW method is random vocabulary partitioning, which enables adjustments to token selection based on specific preferences.Our study revealed that the nexttoken probability distribution plays an critical role in determining how much, or even whether, we can modify token selection and, consequently, the effectiveness of watermarking.We refer to this characteristic, associated with the probability distribution of each token prediction, as watermark strength.In cases of random vocabulary partitioning, the lower bound of watermark strength is dictated by the nexttoken probability distribution.However, we found that, by redesigning the vocabulary partitioning algorithm, we can potentially raise this lower bound.In this paper, we propose SSG (Sort-then-Split by Groups), a method that partitions the vocabulary into two logit-balanced subsets.This design lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability.Experiments on code generation and mathematical reasoning datasets demonstrate the effectiveness of SSG.The source code is available at https://github.com/AllenG-L/SSG. Chenxi Gu, Xiaoning Du 0001, John C. Grundy |
ACL (1) | 2 |
| 2026 | An Empirical Study of Vulnerabilities in Python Packages and Their DetectionabstractContextIn the rapidly evolving software development landscape, Python stands out for its simplicity, versatility, and extensive ecosystem. Python packages, as units of organization, reusability, and distribution, have become a pressing concern, highlighted by the considerable number of vulnerability reports. As a scripting language, Python often cooperates with other programming languages for performance or interoperability. This also adds complexity to the vulnerabilities inherent to Python packages, and the effectiveness of current vulnerability detection tools remains underexplored in the research community.ObjectivesTo bridge this gap, we present PyVul, the first comprehensive benchmark suite of Pythonpackage vulnerabilities. We use this benchmark to conduct an empirical study that characterizes these vulnerabilities and evaluates the limitations of state-of-the-art detection tools.MethodsWe collect real-world vulnerability reports from GitHub Advisories, Snyk, and Huntr, and curate our benchmark at both the commit level and function level. To improve accuracy, we propose LLM-VDC, a large language model–assisted cleansing method. Based on PyVul, we systematically analyze vulnerabilities and assess the capabilities of both rule-based and machine learning–based detectors.ResultsAfter cleansing, PyVul achieves an accuracy of 100% at the commit level with 1,157 repository snapshots, and 94.0% at the function level with 2,082 vulnerable functions, establishing it as the most precise automatically collected Python vulnerability benchmark. Our empirical analysis reveals that current rule-based vulnerability detectors suffer from mismatches between their assumptions and real-world security scenarios, and limited support for high-order vulnerabilities, cross-language interactions, and Python’s unique language features. On the other hand, ML-based detectors suffer from their inability to reach the necessary context.ConclusionA significant discrepancy exists between the capabilities of existing tools and the demands of effectively identifying real-world security issues in Python packages. PyVul provides a solid foundation for advancing vulnerability research and tool development in this domain. Haowei Quan, Junjie Wang 0007, Terry Yue Zhuo, Xiao Chen 0002, Xiaoning Du 0001 |
MSR | 6 |
| 2026 | The arts and crafts of android adware across a decade
Chao Wang 0097, Tianming Liu 0001, Yanjie Zhao 0001, Lin Zhang 0062, Xiaoning Du 0001, Li Li 0029, Haoyu Wang 0001 |
Autom. Softw. Eng. | 5 |
| 2026 | PatchFuzz: Patch fuzzing for JavaScript enginesabstractPatch fuzzing is a technique aimed at identifying vulnerabilities that arise from newly patched code. While researchers have made efforts to apply patch fuzzing to testing JavaScript (JS) engines with considerable success, these efforts have been limited to using ordinary test cases or publicly available vulnerability PoCs (Proof of Concepts) as seeds, and the sustainability of these approaches is hindered by the challenges associated with automating the PoC collection. To address these limitations, we propose an end-to-end sustainable approach for JS engine patch fuzzing, named PatchFuzz. It automates the collection of PoCs of a broader range of historical vulnerabilities and leverages both the PoCs and their corresponding patches to uncover new vulnerabilities more effectively. PatchFuzz starts by recognizing git commits which intend to fix security bugs. Subsequently, it extracts and processes PoCs from these commits to form the seeds for fuzzing, while utilizing code revisions to focus limited fuzzing resources on the more vulnerable code areas through selective instrumentation. The mutation strategy of PatchFuzz is also optimized to maximize the potential of the PoCs. Experimental results demonstrate the effectiveness of PatchFuzz. Notably, 54 bugs across six popular JS engines have been exposed and a total of $62,500 bounties has been received. PatchFuzz effectively enables sustainable and automated patch fuzzing for JavaScript engines by leveraging historical PoCs and selective instrumentation to focus on vulnerable code regions. Junjie Wang 0007, Xiaofei Xie, Xiaoning Du 0001, Xiangwei Zhang |
Inf. Softw. Technol. | 4 |
| 2026 | TRAGIC: Test Oracle Generation for ISA Compliance Testing via Large Language ModelabstractRISC-V is one of the latest Instruction Set Architectures (ISAs). It features an extremely modular and extensible design that makes it suitable for a variety of applications, from embedded systems to large computing systems. However, its inherent customizability and open nature also present challenges in terms of compliance and standardization. Traditionally, the validation relies on Spike as the reference simulator to generate compliance test oracles. Nevertheless, research has demonstrated that Spike may contain bugs, which may compromise the reliability of the generated test oracles. Therefore, developing new cross-validation methods has become increasingly important. To address this challenge, we propose TRAGIC, a method for generating compliance test oracles for RISC-V compliance testing based on Large Language Models (LLMs). Different from traditional software testing, ISA compliance testing demands a finer-grained multidimensional test oracle, including not only the program outputs but also the detailed program states, such as key registers and memory addresses. Furthermore, long instruction sequences frequently exceed the context limitations of LLMs, and direct reasoning under such conditions leads to a notable degradation in accuracy. Therefore, TRAGIC employs a hierarchical strategy with two key components. First, the Block Segmentation Component (BSC) decomposes complex test cases into manageable sub-tasks by analyzing the control flow and applying block slicing techniques. The segmentation both preserves the original program semantics and reduces the reasoning context, thereby enhancing inference accuracy. Second, the Graph of Inference Component (GIC) performs structured reasoning on these sub-tasks using explicitly designed Chain-of-Thought prompts. We utilize LLM to dynamically infer the output of each subtask and determine the next block to execute, continuing this iterative process until the entire task is completed. Meanwhile, key register and memory address tables are maintained and integrated into the final test oracle. By combining the BSC and the GIC, TRAGIC effectively mitigates the accuracy loss associated with long context information and enhances the accuracy of test oracle generation. TRAGIC passes evaluation on the official RISC-V compliance test suite with manually-written test cases. Furthermore, we also show that when integrated with automated test generation tools, our method found 6 bugs, 2 of which were previously unknown. Bixin Li, Xiaoning Du 0001, Lulu Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPTabstractEthical and social risks persist as a crucial yet challenging topic in human-AI interactions, especially in ensuring the safe usage of natural language processing (NLP). The emergence of large language models (LLMs) like ChatGPT introduces the potential for exacerbating this concern. However, prior works on the ethics and risks of emergent LLMs either overlook the practical implications in real-world scenarios, lag behind rapid NLP advancements, lack user consensus on ethical risks, or fail to holistically address the entire spectrum of ethical considerations. In this article, we comprehensively evaluate, qualitatively explore, and catalog ethical dilemmas and risks in ChatGPT through benchmarking with eight representative datasets and red teaming involving diverse case studies. Our findings show that while ChatGPT demonstrates superior safety performance on benchmark datasets, its guardrails can be bypassed via our manually curated examples, revealing not only the limitations of current benchmarks for risk assessment but also unexplored risks in five distinct scenarios, including social bias in code generation, bias in cross-lingual question answering, toxic language in personalized dialogue, misleading information from hallucination, and prompt injections for unethical behaviors. We conclude with implications from red teaming ChatGPT and recommendations for designing future responsible large language models. Terry Yue Zhuo, Yujin Huang, Chunyang Chen 0001, Xiaoning Du 0001, Zhenchang Xing |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | A Reinforcement Learning-Driven Adversarial Attack Methods With Dynamic Perturbation Optimization
Lulu Wang 0001, Xiaoning Du 0001, Jianming Chang, Bixin Li |
IEEE Trans. Reliab. | 3 |
| 2026 | A Reinforcement Learning-Driven Adversarial Attack Methods With Dynamic Perturbation Optimization
Lulu Wang 0001, Xiaoning Du 0001, Jianming Chang, Bixin Li |
IEEE Trans. Reliab. | 3 |
| 2026 | Optimizing Knowledge Utilization for Multi-Intent Comment Generation With Large Language Models
Shuochuan Li, Xiaoning Du 0001, Jiuqiao Yu, Junjie Chen 0003 |
IEEE Trans. Software Eng. | 3 |
| 2026 | UntrustVul: Automated Untrustworthy Alert Identification in Vulnerability Detection ModelsabstractMachine learning (ML) has shown promising results in detecting software vulnerabilities. However, ML detectors are not guaranteed to make predictions based on the right indicators. Studies have revealed that they can rely onirrelevantcode features, such as identifiers or function signatures, particularly those that commonly appear in vulnerable code, yet are not related to the actual vulnerabilities. As a result, the lines of code that the detectors depend on and flag as suspicious are not always genuinely vulnerable. Consequently, developers must manually review these suspicious lines, which is time-consuming and error-prone. If the suspicious lines are wrong, developers may be misled, spend unnecessary effort, or even reach incorrect patching strategies. This highlights the need for automated approaches to identify untrustworthy vulnerability predictions.In this paper, we introduce UNTRUSTVUL, a new approach for identifying untrustworthy vulnerability predictions. Specifically, we focus on cases where a model highlights suspicious lines that would not appear in reliable predictions, i.e., lines that are inherently non-vulnerable and unrelated to any vulnerabilities. To achieve this, we leverage patterns of vulnerable lines observed in historical data. UNTRUSTVUL automatically rules out as untrustworthy any predictions that highlight suspicious lines neither observed in history nor influential to those that have been observed. We refer to such lines as vulnerability-irrelevant. A line is deemed vulnerability-irrelevant if ① it does not match any known patterns of historical vulnerabilities, and ② all its successors in the data and control dependency graph are also vulnerability-irrelevant. Intuitively, a vulnerability-irrelevant line shows low similarity to known vulnerabilities and has no dependency paths to any lines outside the vulnerability-irrelevant category. Notably, these rules are designed to be conservative, as mislabeling a trustworthy prediction as untrustworthy is also undesired. We evaluate UNTRUSTVULon 115K vulnerability predictions made by four models across BigVul, MegaVul, SARD, and PrimeVul datasets, with ground-truth trustworthiness labeled based on the overlap between actual denoised vulnerable lines and model-annotated suspicious lines. UNTRUSTVULeffectively detects untrustworthy predictions with AUC of 70%–88% and F1–score of 82%–94%, outperforming existing approaches by 6%–59% in AUC and 13%–92% in F1–score. Lam Nguyen Tung, Xiaoning Du 0001, Neelofar, Aldeida Aleti |
IEEE Trans. Software Eng. | 2 |
| 2026 | Identifying and Mitigating API Misuse in Large Language ModelsabstractAPI misuse in code generated by large language models (LLMs) presents a serious and growing challenge in software development. While LLMs demonstrate impressive code generation capabilities, their interactions with complex library APIs are often error-prone, potentially leading to software failures and vulnerabilities. In this paper, we conduct a large-scale study of API misuse patterns in LLM-generated code, analyzing both method selection and parameter usage across Python and Java, using three representative LLMs (StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot). Based on extensive manual annotation of 3,209 method-level and 3,492 parameter-level misuses, we identify and categorize four recurring misuse types by building on and refining prior API misuse taxonomies. Our evaluation of three widely used LLMs, StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot, reveals persistent challenges in API usage, particularly hallucination and intent misalignment. To address these issues, we propose Dr.Fix, an LLM-based automatic repair approach guided by our taxonomy. Dr.Fix improves repair accuracy compared to baseline prompting and existing repair methods, with gains of up to 38.4 BLEU and 40% in exact match on benchmark datasets. This work offers important insights into the current limitations of LLMs in API usage and provides insights into current limitations and points to directions for improving automated misuse repair in code generation systems. Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo 0001, John C. Grundy, Xiaoning Du 0001 |
IEEE Trans. Software Eng. | 7 |
| 2025 | FailMapper: Automated Generation of Unit Tests Guided by Failure ScenariosabstractThe automation of unit test generation has become a critical task for improving the overall efficiency of software development and testing. Many existing techniques attempt to generate a sufficient number of test cases to achieve high code coverage. However, it has been shown that a high coverage does not necessarily guarantee effective bug discovery. A potential enhancement is to guide the unit test generation based on bug properties. However, this solution is challenged by the large number and diversity of bug types, making it difficult to comprehensively summarize bug properties.We observe that failures, presented as the results of bugs, manifest in a limited number of scenarios. Therefore, instead of bug properties, in this paper, we propose an innovative framework, named FailMapper, which uses failure scenarios to guide the generation of unit tests. We summarize nine failure scenarios and design the corresponding failure-triggering test strategies. This significantly improves the efficacy of generating test cases towards triggering bugs. To systematically explore possible failure scenarios, FailMapper employs the Monte Carlo Tree Search algorithm to search for the faults that may lead to a failure. Experiments demonstrate that, on 50 known bugs in the Defects4J benchmark, FailMapper can detect many more bugs than five typical unit testing approaches, including EvoSuite, Randoop, CoverUp, HITS, and SymPrompt (40 versus at most 12, out of all 50 bugs). Meanwhile, FailMapper detects 12 out of 20 bugs in the GitBug-Java and Bears-benchmark datasets. We reveal 36 potential issues from 2 Apache projects, and 14 of them have been confirmed as bugs, further demonstrating FailMapper’s effectiveness. The experimental results show that our new framework can significantly enhance the overall efficacy of unit testing. Ruiqi Dong, Zehang Deng, Xiaogang Zhu 0001, Xiaoning Du 0001, Huai Liu, Shaohua Wang 0002, Sheng Wen, Yang Xiang 0001 |
ASE | 4 |
| 2025 | Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient ShorthandabstractLarge language models (LLMs) have shown exceptional performance in code generation and understanding tasks, yet their high computational costs hinder broader adoption. One important factor is the inherent verbosity of programming languages, such as unnecessary formatting elements and lengthy boilerplate code. This leads to inflated token counts in both input and generated outputs, which increases inference costs and slows down the generation process. Prior work improves this through simplifying programming language grammar, reducing token usage across both code understanding and generation tasks. However, it is confined to syntactic transformations, leaving significant opportunities for token reduction unrealized at the semantic level.In this work, we propose Token Sugar, a concept that replaces frequent and verbose code patterns with reversible, token-efficient shorthand in the source code. To realize this concept in practice, we designed a systematic solution that mines high-frequency, token-heavy patterns from a code corpus, maps each to a unique shorthand, and integrates them into LLM pretraining via code transformation. With this solution, we obtain 799 (code pattern, shorthand) pairs, which can reduce up to 15.1% token count in the source code and is complementary to existing syntax-focused methods. We further trained three widely used LLMs on Token Sugar-augmented data. Experimental results show that these models not only achieve significant token savings (up to 11.2% reduction) during generation but also maintain near-identical Pass@1 scores compared to baselines trained on unprocessed code. Zhensu Sun, Chengran Yang, Xiaoning Du 0001, Zhou Yang 0003, Li Li 0029, David Lo 0001 |
ASE | 3 |
| 2025 | SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
Guangke Chen, Yedi Zhang, Fu Song, Ting Wang 0004, Xiaoning Du 0001, Yang Liu 0003 |
NDSS | 5 |
| 2025 | GenDetect: Generative Large Language Model Usage in Smart Contract Vulnerability Detection
Peter Ince, Jiangshan Yu, Joseph K. Liu, Xiaoning Du 0001, Xiapu Luo |
ProvSec | 4 |
| 2025 | Is MPC Secure? Leveraging Neural Network Classifiers to Detect Data Leakage Vulnerabilities in MPC ImplementationsabstractDue to the emerging privacy-protection laws and regulations (e.g. GDPR in the EU) in recent years, dozens of multi-party computation (MPC for short) protocols have been proposed and widely applied by companies and institutions. These MPC protocols enable companies and institutions to perform joint analyses and machine learning on their private data while protecting their data's privacy. However, due to the complexity of MPC protocols, their implementations of-ten contain data leakage vulnerabilities, which can critically undermine the intended privacy protection. Additionally, most existing security analyses of MPC protocols rely on theoretical proofs, neglecting to detect possible vulnerabilities in MPC im-plementations. Therefore, detecting data leakage vulnerabilities in MPC implementations is an urgent necessity. In this paper, we propose MPCGuard, a practical frame-work for detecting data leakage vulnerabilities in MPC imple-mentations. Different from traditional memory vulnerabilities, data leakage vulnerabilities in MPC implementations cannot be identified by existing sanitizers. To resolve this challenge, we first establish a leakage identifier in MPCGuard with two neural network classifiers to identify whether an MPC implementation contains data leakage vulnerabilities. To enhance identification effectiveness, the structures of neural network classifiers are designed according to the characteristics of MPC protocols. After identifying a data leakage vulnerability, we employ a delta method to assist in locating the vulnerability. To demonstrate the effectiveness of MPCGuard, we apply MPCGuard to test 29 commonly-used MPC implementations in three main-stream MPC frameworks, i.e. Crypten, TF-Encrypted, and MP-SPDZ. We discover that 12 out of 29 implementations contain data leakage vulnerabilities, some of which can lead to the reconstruction of raw data. Until the moment this paper is written, all vulnerabilities, two of which have been assigned with CVE-IDs, have been confirmed. To the best of our knowledge, these two CVE-IDs are the first CVE-IDs assigned for data leakage vulnerabilities in MPC implementations. Guopeng Lin, Xiaoning Du 0001, Lushan Song, Weili Han, Junming Ma, Wenjing Fang |
SP | 2 |
| 2025 | Detecting and Explaining Python Name Errors
Jiawei Wang 0003, Li Li 0029, Kui Liu 0001, Xiaoning Du 0001 |
Inf. Softw. Technol. | 4 |
| 2025 | Runtime Backdoor Detection for Federated Learning via Representational Dissimilarity AnalysisabstractFederated learning (FL), as a powerful learning paradigm, trains a shared model by aggregating model updates from distributed clients. However, the decoupling of model learning from local data makes FL highly vulnerable to backdoor attacks, where a single compromised client can poison the shared model. While recent progress has been made in backdoor detection, existing methods face challenges with detection accuracy and runtime effectiveness, particularly when dealing with complex model architectures. In this work, we propose a novel approach to detecting malicious clients in an accurate, stable, and efficient manner. Our method utilizes a sampling-based network representation method to quantify dissimilarities between clients, identifying model deviations caused by backdoor injections. We also propose an iterative algorithm to progressively detect and exclude malicious clients as outliers based on these dissimilarity measurements. Evaluations across a range of benchmark tasks demonstrate that our approach outperforms state-of-the-art methods in detection accuracy and defense effectiveness. When deployed for runtime protection, our approach effectively eliminates backdoor injections with marginal overheads. Xiyue Zhang 0001, Xiaoyong Xue, Xiaoning Du 0001, Xiaofei Xie, Yang Liu 0003, Meng Sun 0002 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case PairsabstractAutomated Program Repair (APR) aims to automatically generate patches for rectifying software bugs. Recent strides in Large Language Models (LLM), such as ChatGPT, have yielded encouraging outcomes in APR, especially within the conversation-driven APR framework. Nevertheless, the efficacy of conversation-driven APR is contingent on the quality of the feedback information. In this article, we propose ContrastRepair , a novel conversation-based APR approach that augments conversation-driven APR by providing LLMs with contrastive test pairs. A test pair consists of a failing test and a passing test, which offer contrastive feedback to the LLM. Our key insight is to minimize the difference between the generated passing test and the given failing test, which can better isolate the root causes of bugs. By providing such informative feedback, ContrastRepair enables the LLM to produce effective bug fixes. The implementation of ContrastRepair is based on the state-of-the-art LLM, ChatGPT, and it iteratively interacts with ChatGPT until plausible patches are generated. We evaluate ContrastRepair on multiple benchmark datasets, including Defects4J, QuixBugs, and HumanEval-Java. The results demonstrate that ContrastRepair significantly outperforms existing methods, achieving a new state-of-the-art in program repair. For instance, among Defects4J 1.2 and 2.0, ContrastRepair correctly repairs 143 out of all 337 bug cases, while the best-performing baseline fixes 124 bugs. Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2025 | Don't Complete It! Preventing Unhelpful Code Completion for Productive and Sustainable Neural Code Completion SystemsabstractCurrently, large pre-trained language models are widely applied in neural code completion systems. Though large code models significantly outperform their smaller counterparts, around 70% of displayed code completions from Github Copilot are not accepted by developers. Being reviewed but not accepted, their help to developer productivity is considerably limited and may conversely aggravate the workload of developers, as the code completions are automatically and actively generated in state-of-the-art code completion systems as developers type out once the service is enabled. Even worse, considering the high cost of the large code models, it is a huge waste of computing resources and energy, which severely goes against the sustainable development principle of AI technologies. However, such waste has never been realized, not to mention effectively addressed, in the research community for neural code completion. Hence, preventing such unhelpful code completions from happening in a cost-friendly way is of urgent need. To fill this significant gap, we first investigate the prompts of unhelpful code completions, called “low-return prompts.” We empirically identify four observable patterns in low-return prompts, each lacking necessary information, making it difficult to address through enhancements to the model’s accuracy alone. This demonstrates the feasibility of identifying such low-return prompts based on the prompts themselves. Motivated by this finding, we propose an early-rejection mechanism to turn down low-return prompts by foretelling the code completion qualities. The prompts that are estimated to receive unhelpful code completions will not be sent to the model. Furthermore, we investigated five types of estimators to demonstrate the feasibility of the mechanism. The experimental results show that the estimator can reject 20% of code completion requests with a 97.4% precision. To the best of our knowledge, it is the first systemic approach to address the problem of unhelpful code completions and this work also sheds light on an important research direction of large code models. Zhensu Sun, Xiaoning Du 0001, Fu Song, Shangwen Wang, Mingze Ni, Li Li 0029, David Lo 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | PonziFinder: Attention-Based Edge-Enhanced Ponzi Contract DetectionabstractPonzi contractsare fraudulent investment scams that promise high returns with little risk to investors. However, existing methods for detecting Ponzi contracts have several limitations. For example, they struggle to deal with the class imbalance problem, and their analysis of function call transactions is inadequate, resulting in redundant features. To tackle the challenges of detecting Ponzi contracts, we present PonziFinder, a novel approach that leverages convolutional-based edge-enhanced graph neural network and attention mechanism for the classification of contract transaction graphs. In contrast to previous methods, we not only consider transaction value and timestamp but also analyze transaction input to standardize and sort transactions. We extract node and edge features that capture the unique characteristics of Ponzi contracts. The edge feature, reflecting interaccount correlation, enhances the propagation and updating of node features for effective Ponzi contract detection. To prevent oversmoothing of node embedding caused by the shallow transaction graph and extract important account node information, we introduce an attention-based global layerwise aggregation mechanism (ALGA) for generating the final contract graph representation for classification. Moreover, we optimize the node feature set and use an effective strategy based on undersampling and ensemble learning to address the issue of class imbalance. Experimental results show that PonziFinder can detect all types of Ponzi contracts (100%) with 97% accuracy when there is sufficient transaction data, outperforming other models. The analysis of input values and the ALGA mechanism are experimentally shown to improve accuracy by 4% and 2%, respectively. In summary, PonziFinder is a novel and effective method for detecting Ponzi contracts. Our approach addresses the limitations of existing methods and demonstrates significant improvements in accuracy and efficiency. Bixin Li, Yan Xiao 0002, Xiaoning Du 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | Detect Llama - Finding Vulnerabilities in Smart Contracts Using Large Language Models
Peter Ince, Xiapu Luo, Jiangshan Yu, Joseph K. Liu, Xiaoning Du 0001 |
ACISP (3) | 5 |
| 2024 | When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model InferenceabstractLeveraging recent advancements in large language models, modern neural code completion models have demonstrated the capability to generate highly accurate code suggestions. However, their massive size poses challenges in terms of computational costs and environmental impact, hindering their widespread adoption in practical scenarios. Dynamic inference emerges as a promising solution, as it allocates minimal computation during inference while maintaining the model's performance. In this research, we explore dynamic inference within the context of code completion. Initially, we conducted an empirical investigation on GPT-2, focusing on the inference capabilities of intermediate layers for code completion. We found that 54.4% of tokens can be accurately generated using just the first layer, signifying significant computational savings potential. Moreover, despite using all layers, the model still fails to predict 14.5% of tokens correctly, and the subsequent completions continued from them are rarely considered helpful, with only a 4.2% Acceptance Rate. These findings motivate our exploration of dynamic inference in code completion and inspire us to enhance it with a decision-making mechanism that stops the generation of incorrect code. We thus propose a novel dynamic inference method specifically tailored for code completion models. This method aims not only to produce correct predictions with largely reduced computation but also to prevent incorrect predictions proactively. Our extensive evaluation shows that it can averagely skip 1.7 layers out of 16 layers in the models, leading to an 11.2% speedup with only a marginal 1.1% reduction in ROUGE-L. Zhensu Sun, Xiaoning Du 0001, Fu Song, Shangwen Wang, Li Li 0029 |
ICSE | 2 |
| 2024 | Testing Graph Database Systems with Graph-State Persistence OracleabstractGraph Database Management Systems (GDBMSs) store data in a graph format, facilitating rapid querying of nodes and relationships. This structure is particularly advantageous for applications like social networks and recommendation systems, which often involve frequent writing operations—such as adding new nodes, creating relationships, or modifying existing data—that potentially introduce bugs. However, existing GDBMS testing approaches tend to overlook these writing functionalities, failing to detect bugs arising from such operations. In this paper we present GraspDB, the first metamorphic testing approach specifically designed to identify bugs related to writing operations in graph database systems. GraspDB employs the Graph-State Persistence oracle, which is based on the Labeled Property Graph Isomorphism (LPG-Isomorphism) and Labeled Property Subgraph Isomorphism (LPSG-Isomorphism) relations. We also develop three classes of mutation rules aimed at engaging more diverse writing-related code logic. GraspDB has successfully detected 77 unique, previously unknown bugs across four popular open source graph database engines, among which 58 bugs are confirmed by developers, 43 bugs have been fixed and 31 are related to writing operations. Shuang Liu 0007, Junhao Lan, Xiaoning Du 0001, Jiyuan Li, Wei Lu 0015, Jiajun Jiang, Xiaoyong Du 0001 |
ISSTA | 3 |
| 2024 | FDI: Attack Neural Code Generation Systems through User Feedback ChannelabstractNeural code generation systems have recently attracted increasing attention to improve developer productivity and speed up software development. Typically, these systems maintain a pre-trained neural model and make it available to general users as a service (e.g., through remote APIs) and incorporate a feedback mechanism to extensively collect and utilize the users' reaction to the generated code, i.e., user feedback. However, the security implications of such feedback have not yet been explored. With a systematic study of current feedback mechanisms, we find that feedback makes these systems vulnerable to feedback data injection (FDI) attacks. We discuss the methodology of FDI attacks and present a pre-attack profiling strategy to infer the attack constraints of a targeted system in the black-box setting. We demonstrate two proof-of-concept examples utilizing the FDI attack surface to implement prompt injection attacks and backdoor attacks on practical neural code generation systems. The attacker may stealthily manipulate a neural code generation system to generate code with vulnerabilities, attack payload, and malicious and spam messages. Our findings reveal the security implications of feedback mechanisms in neural code generation systems, paving the way for increasing their security. Zhensu Sun, Xiaoning Du 0001, Xiapu Luo, Fu Song, David Lo 0001, Li Li 0029 |
ISSTA | 2 |
| 2024 | AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationabstractArtificial Intelligence (AI) models have emerged as another important audience for programming languages alongside humans and machines, as we enter the era of large language models (LLMs). LLMs can now perform well in coding competitions and even write programs like developers to solve various tasks, including mathematical problems. However, the grammar and layout of current programs are designed to cater the needs of human developers -- with many grammar tokens and formatting tokens being used to make the code easier for humans to read. While this is helpful, such a design adds unnecessary computational work for LLMs, as each token they either use or produce consumes computational resources. To improve inference efficiency and reduce computational costs, we propose the concept of AI-oriented grammar.This aims to represent code in a way that better suits the working mechanism of AI models. Code written with AI-oriented grammar discards formats and uses a minimum number of tokens to convey code semantics effectively. To demonstrate the feasibility of this concept, we explore and implement the first AI-oriented grammar for Python, named Simple Python (SimPy). SimPy is crafted by revising the original Python grammar through a series of heuristic rules. Programs written in SimPy maintain identical Abstract Syntax Tree (AST) structures to those in standard Python. This allows for not only execution via a modified AST parser, but also seamless transformation between programs written in Python and SimPy, enabling human developers and LLMs to use Python and SimPy, respectively, when they need to collaborate. We also look into methods to help existing LLMs understand and use SimPy effectively. In the experiments, compared with Python, SimPy enables a reduction in token usage by 13.5% and 10.4% for CodeLlama and GPT-4, respectively, when completing the same set of code-related tasks. Additionally, these models can maintain or even improve their performance when using SimPy instead of Python for these tasks. With these promising results, we call for further contributions to the development of AI-oriented program grammar within our community. Zhensu Sun, Xiaoning Du 0001, Zhou Yang 0003, Li Li 0029, David Lo 0001 |
ISSTA | 2 |
| 2024 | Are Latent Vulnerabilities Hidden Gems for Software Vulnerability Prediction? An Empirical StudyabstractCollecting relevant and high-quality data is integral to the development of effective Software Vulnerability (SV) prediction models. Most of the current SV datasets rely on SV-fixing commits to extract vulnerable functions and lines. However, none of these datasets have considered latent SVs existing between the introduction and fix of the collected SVs. There is also little known about the usefulness of these latent SVs for SV prediction. To bridge these gaps, we conduct a large-scale study on the latent vulnerable functions in two commonly used SV datasets and their utilization for function-level and line-level SV predictions. Leveraging the state-of-the-art SZZ algorithm, we identify more than 100k latent vulnerable functions in the studied datasets. We find that these latent functions can increase the number of SVs by 4× on average and correct up to 5k mislabeled functions, yet they have a noise level of around 6%. Despite the noise, we show that the state-of-the-art SV prediction model can significantly benefit from such latent SVs. The improvements are up to 24.5% in the performance (F1-Score) of function-level SV predictions and up to 67% in the effectiveness of localizing vulnerable lines. Overall, our study presents the first promising step toward the use of latent SVs to improve the quality of SV datasets and enhance the performance of SV prediction tasks. Triet Huynh Minh Le, Xiaoning Du 0001, Muhammad Ali Babar 0001 |
MSR | 2 |
| 2023 | CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsabstractCode datasets are of immense value for training neural-network-based code completion models, where companies or organizations have made substantial investments to establish and process these datasets. Unluckily, these datasets, either built for proprietary or public usage, face the high risk of unauthorized exploits, resulting from data leakages, license violations, etc. Even worse, the "black-box" nature of neural models sets a high barrier for externals to audit their training datasets, which further connives these unauthorized usages. Currently, watermarking methods have been proposed to prohibit inappropriate usage of image and natural language datasets. However, due to domain specificity, they are not directly applicable to code datasets, leaving the copyright protection of this emerging and important field of code data still exposed to threats. To fill this gap, we propose a method, named CodeMark, to embed user-defined imperceptible watermarks into code datasets to trace their usage in training neural code completion models. CodeMark is based on adaptive semantic-preserving transformations, which preserve the exact functionality of the code data and keep the changes covert against rule-breakers. We implement CodeMark in a toolkit and conduct an extensive evaluation of code completion models. CodeMark is validated to fulfill all desired properties of practical watermarks, including harmlessness to model accuracy, verifiability, robustness, and imperceptibility. Zhensu Sun, Xiaoning Du 0001, Fu Song, Li Li 0029 |
ESEC/SIGSOFT FSE | 2 |
| 2023 | DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsabstractDeep learning (DL) models are trained on sampled data, where the distribution of training data differs from that of real-world data (i.e., the distribution shift), which reduces the model's robustness. Various testing techniques have been proposed, including distribution-unaware and distribution-aware methods. However, distribution-unaware testing lacks effectiveness by not explicitly considering the distribution of test cases and may generate redundant errors (within same distribution). Distribution-aware testing techniques primarily focus on generating test cases that follow the training distribution, missing out-of-distribution data that may also be valid and should be considered in the testing process. In this paper, we propose a novel distribution-guided approach for generating valid test cases with diverse distributions, which can better evaluate the model's robustness (i.e., generating hard-to-detect errors) and enhance the model's robustness (i.e., enriching training data). Unlike existing testing techniques that optimize individual test cases, DistXplore optimizes test suites that represent specific distributions. To evaluate and enhance the model's robustness, we design two metrics: distribution difference, which maximizes the similarity in distribution between two different classes of data to generate hard-to-detect errors, and distribution diversity, which increase the distribution diversity of generated test cases for enhancing the model's robustness. To evaluate the effectiveness of DistXplore in model evaluation and enhancement, we compare DistXplore with 14 state-of-the-art baselines on 10 models across 4 datasets. The evaluation results show that DisXplore not only detects a larger number of errors (e.g., 2×+ on average). Furthermore, DistXplore achieves a higher improvement in empirical robustness (e.g., 5.2% more accuracy improvement than the baselines on average). Longtian Wang, Xiaofei Xie, Xiaoning Du 0001, Qing Guo 0005, Chao Shen 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2023 | MATH - Finding and Fixing Exploits in AlgorandabstractWith the growth in assets managed on-chain comes more attention from hackers. As Algorand uses its own Algorand Virtual Machine (AVM), there is a need for new vulnerability and exploit detection tools, as those built for Ethereum’s EVM are not suitable. This paper presents the MATH static analysis tool with detectors for the math exploit and byte subtraction vulnerability on the Algorand network using only the deployed smart contract base64 code. We use it to analyse 144,006 stateful smart contracts, find and verify three new instances of the math exploit, and evaluate the tools’ runtime and effectiveness. Peter Ince, Xiapu Luo, Jiangshan Yu, Joseph K. Liu, Xiaoning Du 0001 |
TrustCom | 5 |
| 2023 | FuzzJIT: Oracle-Enhanced Fuzzing for JavaScript Engine JIT Compiler
Junjie Wang 0008, Shuang Liu 0019, Xiaoning Du 0001, Junjie Chen 0003 |
USENIX Security Symposium | 4 |
| 2022 | On the Importance of Building High-quality Training Datasets for Neural Code SearchabstractThe performance of neural code search is significantly influenced by the quality of the training data from which the neural models are derived. A large corpus of high-quality query and code pairs is demanded to establish a precise mapping from the natural language to the programming language. Due to the limited availability, most widely-used code search datasets are established with compromise, such as using code comments as a replacement of queries. Our empirical study on a famous code search dataset reveals that over one-third of its queries contain noises that make them deviate from natural user queries. Models trained through noisy data are faced with severe performance degradation when applied in real-world scenarios. To improve the dataset quality and make the queries of its samples semantically identical to real user queries is critical for the practical usability of neural code search. In this paper, we propose a data cleaning framework consisting of two subsequent filters: a rule-based syntactic filter and a model-based semantic filter. This is the first framework that applies semantic query cleaning to code search datasets. Experimentally, we evaluated the effectiveness of our framework on two widely-used code search models and three manually-annotated code retrieval benchmarks. Training the popular DeepCS model with the filtered dataset from our framework improves its performance by 19.2% MRR and 21.3% [email protected], on average with the three validation benchmarks. Zhensu Sun, Li Li 0091, Xiaoning Du 0001, Li Li 0029 |
ICSE | 4 |
| 2022 | CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningabstractGithub Copilot, trained on billions of lines of public code, has recently become the buzzword in the computer science research and practice community. Although it is designed to help developers implement safe and effective code with powerful intelligence, practitioners and researchers raise concerns about its ethical and security problems, e.g., should the copyleft licensed code be freely leveraged or insecure code be considered for training in the first place? These problems pose a significant impact on Copilot and other similar products that aim to learn knowledge from large-scale open-source code through deep learning models, which are inevitably on the rise with the fast development of artificial intelligence. To mitigate such impacts, we argue that there is a need to invent effective mechanisms for protecting open-source code from being exploited by deep learning models. Here, we design and implement a prototype, CoProtector, which utilizes data poisoning techniques to arm source code repositories for defending against such exploits. Our large-scale experiments empirically show that CoProtector is effective in achieving its purpose, significantly reducing the performance of Copilot-like deep learning models while being able to stably reveal the secretly embedded watermark backdoors. Zhensu Sun, Xiaoning Du 0001, Fu Song, Mingze Ni, Li Li 0029 |
WWW | 2 |
| 2021 | Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) have demonstrated their effectiveness in learning and processing sequential data (e.g., speech and natural language). However, due to the black-box nature of neural networks, understanding the decision logic of RNNs is quite challenging. Some recent progress has been made to approximate the behavior of an RNN by weighted automata. They provide better interpretability, but still suffer from poor scalability. In this paper, we propose a novel approach to extracting weighted automata with the guidance of a target RNN's decision and context information. In particular, we identify the patterns of RNN's step-wise predictive decisions to instruct the formation of automata states. Further, we propose a state composition method to enhance the context-awareness of the extracted model. Our in-depth evaluations on typical RNN tasks, including language model and classification, demonstrate the effectiveness and advantage of our method over the state-of-the-arts. The evaluation results show that our method can achieve accurate approximation of an RNN even on large-scale tasks. Xiyue Zhang 0001, Xiaoning Du 0001, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003, Meng Sun 0002 |
AAAI | 2 |
| 2021 | Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsabstractSpeaker recognition (SR) is widely used in our daily life as a biometric authentication or identification mechanism. The popularity of SR brings in serious security concerns, as demonstrated by recent adversarial attacks. However, the impacts of such threats in the practical black-box setting are still open, since current attacks consider the white-box setting only.In this paper, we conduct the first comprehensive and systematic study of the adversarial attacks on SR systems (SRSs) to understand their security weakness in the practical black-box setting. For this purpose, we propose an adversarial attack, named FAKEBOB, to craft adversarial samples. Specifically, we formulate the adversarial sample generation as an optimization problem, incorporated with the confidence of adversarial samples and maximal distortion to balance between the strength and imperceptibility of adversarial voices. One key contribution is to propose a novel algorithm to estimate the score threshold, a feature in SRSs, and use it in the optimization problem to solve the optimization problem. We demonstrate that FAKEBOB achieves 99% targeted attack success rate on both open-source and commercial systems. We further demonstrate that FAKEBOB is also effective on both open-source and commercial systems when playing over the air in the physical world. Moreover, we have conducted a human study which reveals that it is hard for human to differentiate the speakers of the original and adversarial voices. Last but not least, we show that four promising defense methods for adversarial attack from the speech recognition domain become ineffective on SRSs against FAKEBOB, which calls for more effective defense methods. We highlight that our study peeks into the security implications of adversarial attacks on SRSs, and realistically fosters to improve the security robustness of SRSs. Guangke Chen, Sen Chen 0001, Lingling Fan 0003, Xiaoning Du 0001, Zhe Zhao 0007, Fu Song, Yang Liu 0003 |
SP | 4 |
| 2021 | Trace-Length Independent Runtime Monitoring of Quantitative PoliciesabstractMetric linear-time logic (MTL) has been widely used to specify runtime policies. Traditionally this use of MTL is to capture the qualitative aspects of the monitored systems, but recent developments in its extensions with aggregate operators allow some quantitative policies to be specified. Our interest in MTL-based policy languages is driven by applications in runtime malware or intrusion detection in platforms like Android and autonomous vehicles, which requires the monitoring algorithm to be independent of the length of the system event traces so that its performance does not degrade as the traces grow. We propose a policy language based on a past-time variant of MTL, extended with an aggregate operator called the metric temporal counting quantifier to specify a policy based on the number of times some sub-policies are satisfied in the specified past time interval. We show that a broad class of policies, but not all policies, specified with our language can be monitored in a trace-length independent way, and provide a concrete algorithm to do so. We implement and test our algorithm in both an existing Android monitoring framework and an autonomous vehicle simulation platform, and show that our approach can effectively specify and monitor quantitative policies drawn from real-world studies. Xiaoning Du 0001, Alwen Tiu, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyabstractOver the past decade, deep learning (DL) has been successfully applied to many industrial domain-specific tasks. However, the current state-of-the-art DL software still suffers from quality issues, which raises great concern especially in the context of safety- and security-critical scenarios. Adversarial examples (AEs) represent a typical and important type of defects needed to be urgently addressed, on which a DL software makes incorrect decisions. Such defects occur through either intentional attack or physical-world noise perceived by input sensors, potentially hindering further industry deployment. The intrinsic uncertainty nature of deep learning decisions can be a fundamental reason for its incorrect behavior. Although some testing, adversarial attack and defense techniques have been recently proposed, it still lacks a systematic study to uncover the relationship between AEs and DL uncertainty. Xiyue Zhang 0001, Xiaofei Xie, Lei Ma 0003, Xiaoning Du 0001, Yang Liu 0003, Jianjun Zhao 0001, Meng Sun 0002 |
ICSE | 4 |
| 2020 | Marble: Model-based Robustness Analysis of Stateful Deep Learning SystemsabstractState-of-the-art deep learning (DL) systems are vulnerable to adversarial examples, which hinders their potential adoption in safety-and security-critical scenarios. While some recent progress has been made in analyzing the robustness of feed-forward neural networks, the robustness analysis for stateful DL systems, such as recurrent neural networks (RNNs), still remains largely uncharted. In this paper, we propose Marble, a model-based approach for quantitative robustness analysis of real-world RNN-based DL systems. Marble builds a probabilistic model to compactly characterize the robustness of RNNs through abstraction. Furthermore, we propose an iterative refinement algorithm to derive a precise abstraction, which enables accurate quantification of the robustness measurement. We evaluate the effectiveness of Marble on both LSTM and GRU models trained separately with three popular natural language datasets. The results demonstrate that (1) our refinement algorithm is more efficient in deriving an accurate abstraction than the random strategy, and (2) Marble enables quantitative robustness analysis, in rendering better efficiency, accuracy, and scalability than the state-of-the-art techniques. Xiaoning Du 0001, Yi Li 0008, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 1 |
| 2019 | Leopard: identifying vulnerable code for vulnerability assessment through program metricsabstractIdentifying potentially vulnerable locations in a code base is critical as a pre-step for effective vulnerability assessment; i.e., it can greatly help security experts put their time and effort to where it is needed most. Metric-based and pattern-based methods have been presented for identifying vulnerable code. The former relies on machine learning and cannot work well due to the severe imbalance between non-vulnerable and vulnerable code or lack of features to characterize vulnerabilities. The latter needs the prior knowledge of known vulnerabilities and can only identify similar but not new types of vulnerabilities. In this paper, we propose and implement a generic, lightweight and extensible framework, LEOPARD, to identify potentially vulnerable functions through program metrics. LEOPARD requires no prior knowledge about known vulnerabilities. It has two steps by combining two sets of systematically derived metrics. First, it uses complexity metrics to group the functions in a target application into a set of bins. Then, it uses vulnerability metrics to rank the functions in each bin and identifies the top ones as potentially vulnerable. Our experimental results on 11 real-world projects have demonstrated that, LEOPARD can cover 74.0% of vulnerable functions by identifying 20% of functions as vulnerable and outperform machine learning-based and static analysis-based techniques. We further propose three applications of LEOPARD for manual code review and fuzzing, through which we discovered 22 new bugs in real applications like PHP, radare2 and FFmpeg, and eight of them are new vulnerabilities. Xiaoning Du 0001, Bihuan Chen 0001, Yuekang Li, Jianmin Guo, Yaqin Zhou, Yang Liu 0003, Yu Jiang 0001 |
ICSE | 1 |
| 2019 | A Quantitative Analysis Framework for Recurrent Neural NetworkabstractRecurrent neural network (RNN) has achieved great success in processing sequential inputs for applications such as automatic speech recognition, natural language processing and machine translation. However, quality and reliability issues of RNNs make them vulnerable to adversarial attacks and hinder their deployment in real-world applications. In this paper, we propose a quantitative analysis framework - DeepStellar - to pave the way for effective quality and security analysis of software systems powered by RNNs. DeepStellar is generic to handle various RNN architectures, including LSTM and GRU, scalable to work on industrial-grade RNN models, and extensible to develop customized analyzers and tools. We demonstrated that, with DeepStellar, users are able to design efficient test generation tools, and develop effective adversarial sample detectors. We tested the developed applications on three real RNN models, including speech recognition and image classification. DeepStellar outperforms existing approaches three hundred times in generating defect-triggering tests and achieves 97% accuracy in detecting adversarial attacks. A video demonstration which shows the main features of DeepStellar is available at: https://sites.google.com/view/deepstellar/tool-demo. Xiaoning Du 0001, Xiaofei Xie, Yi Li 0008, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 1 |
| 2019 | Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural NetworksabstractVulnerability identification is crucial to protect the software systems from attacks for cyber security. It is especially important to localize the vulnerable functions among the source code to facilitate the fix. However, it is a challenging and tedious process, and also requires specialized security expertise. Inspired by the work on manually-defined patterns of vulnerabilities from various code representation graphs and the recent advance on graph neural networks, we propose Devign, a general graph neural network based model for graph-level classification through learning on a rich set of code semantic representations. It includes a novel Conv module to efficiently extract useful features in the learned rich node representations for graph-level classification. The model is trained over manually labeled datasets built on 4 diversified large-scale open-source C projects that incorporate high complexity and variety of real source code instead of synthesis code used in previous works. The results of the extensive evaluation on the datasets demonstrate that Devign outperforms the state of the arts significantly with an average of 10.51% higher accuracy and 8.68% F1 score, increases averagely 4.66% accuracy and 6.37% F1 by the Conv module. Yaqin Zhou, Shangqing Liu, Jing Kai Siow, Xiaoning Du 0001, Yang Liu 0003 |
NeurIPS | 4 |
| 2019 | DeepStellar: model-based quantitative analysis of stateful deep learning systemsabstractDeep Learning (DL) has achieved tremendous success in many cutting-edge applications. However, the state-of-the-art DL systems still suffer from quality issues. While some recent progress has been made on the analysis of feed-forward DL systems, little study has been done on the Recurrent Neural Network (RNN)-based stateful DL systems, which are widely used in audio, natural languages and video processing, etc. In this paper, we initiate the very first step towards the quantitative analysis of RNN-based DL systems. We model RNN as an abstract state transition system to characterize its internal behaviors. Based on the abstract model, we design two trace similarity metrics and five coverage criteria which enable the quantitative analysis of RNNs. We further propose two algorithms powered by the quantitative measures for adversarial sample detection and coverage-guided test generation. We evaluate DeepStellar on four RNN-based systems covering image classification and automated speech recognition. The results demonstrate that the abstract model is useful in capturing the internal behaviors of RNNs, and confirm that (1) the similarity metrics could effectively capture the differences between samples even with very small perturbations (achieving 97% accuracy for detecting adversarial samples) and (2) the coverage criteria are useful in revealing erroneous behaviors (generating three times more adversarial samples than random testing and hundreds times more than the unrolling approach). Xiaoning Du 0001, Xiaofei Xie, Yi Li 0008, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2018 | Towards Building a Generic Vulnerability Detection Platform by Combining Scalable Attacking Surface Analysis and Directed Fuzzing
Xiaoning Du 0001 |
ICFEM | 1 |
| 2015 | Trace-Length Independent Runtime Monitoring of Quantitative Policies in LTL
Xiaoning Du 0001, Yang Liu 0003, Alwen Tiu |
FM | 1 |