Xiaodong Zhang 0014

dblp:37/4356-14 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-8380-1019ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SafetyReminder: Reviving Delayed Safety Awareness of Vision-Language Models to Defend Against Jailbreak Attacks
abstract
Vision-Language Models (VLMs) extend Large Language Models (LLMs) with visual perception capabilities, unlocking broad applications across many domains. However, ensuring their safety remains a critical challenge, as adversarial visual inputs can easily bypass built-in safeguards and elicit harmful content. In this paper, we uncover a phenomenon we call delayed safety awareness, where a jailbroken VLM initially produces harmful content but ultimately recognizes the harmfulness at the end of the generation process. We attribute this phenomenon to the fact that the model's safety awareness against jailbreaks cannot be effectively transferred to the intermediate stages of text generation. Motivated by this insight, we introduce SafetyReminder, a simple yet effective defense that optimizes a learnable soft prompt using our proposed Safety-Activation Prompt Tuning (SAPT). This soft prompt is inserted into the generated text to activate the safety awareness of the model, steering it toward refusal when harmful content arises while preserving helpfulness in benign scenarios. We evaluate our method on three established harmful benchmarks and across three types of adversarial attacks. Experimental results demonstrate that our method achieves state-of-the-art defense performance with strong generalization, offering a practical and lightweight solution for safe deployment of VLMs.
Peiyuan Tang, Haojie Xin, Xiaodong Zhang 0014, Jun Sun 0001, Qin Xia, Zijiang Yang 0006
AAAI3
2026 ConFixer: Robustness Semantics Based Configuration Bug Fixing for Automated Driving Systems
abstract
Abstract Automated Driving Systems (ADSs) coordinate multiple modules (e.g., perception, prediction, planning, and control) and expose a large configuration surface that engineers must tune for different platforms and operational conditions. In industrial practice, misconfigured parameters are a common source of unsafe or law-violating behaviors, yet fixing such configuration bugs remains largely overlooked. Debugging is difficult because parameter effects propagate through long, nonlinear pipelines (often involving learning-based components), and manual trial-and-error provides little guidance in high-dimensional spaces while risking regressions in previously passing scenarios. We propose ConFixer , an automated approach for repairing ADS configuration bugs against formal correctness specifications, including traffic laws encoded in signal temporal logic (STL). ConFixer uses robustness semantics to derive gradient-like signals that localize bug-relevant parameters and guide fine-tuning toward compliance. To support adoption in scenario-based validation workflows, ConFixer evaluates candidate fixes on large scenario suites and prioritizes regression-safe solutions. Our evaluation on Baidu Apollo with the LGSVL simulator shows that ConFixer fixes 173 configuration bugs without introducing regressions.
Xiaodong Zhang 0014, Songyang Yan, Zijiang Yang 0004
FM (2)1
2025 DPFuzzer: Discovering Safety Critical Vulnerabilities for Drone Path Planners
abstract
State-of-the-art drone path planners enable drones to autonomously travel through obstacles in GPS-denied, uncharted, cluttered environments. However, our investigation shows that path planners fail to maneuver drones correctly in specific scenarios, leading to incidents such as collisions. To minimize such risks, drone path planners should be tested thoroughly against diverse scenarios before deployment. Existing research for testing drones to uncover safety-critical vulnerabilities is only focused on flight control programs and is limited in the capability to generate diverse obstacle scenarios for testing drone path planners. In this work, we propose DPFuzzer, an automated framework for testing drone path planners. DPFuzzer is an evolutionary algorithm (EA) based testing framework. It aims to uncover vulnerabilities in drone path planners by generating diverse critical scenarios that can trigger vulnerabilities. To better guide the critical scenario generation, we introduce Environmental Risk Factor (ERF), a metric we propose, to abstract potential safety threats of scenarios. We evaluate DPFuzzer on state-of-the-art drone path planners and the experimental result shows that DPFuzzer can effectively find diverse vulnerabilities. Additionally, we demonstrate that these vulnerabilities are exploitable in the real world.
Yue Wang 0063, Chao Yang 0016, Xiaodong Zhang 0014, Yuwanqi Deng, Jianfeng Ma 0001
ICSE3
2025 Measuring and Explaining the Effects of Android App Transformations in Online Malware Detection
abstract
It is well known that antivirus engines are vulnerable to evasion techniques (e.g., obfuscation) that transform malware into its variants.However, it cannot be necessarily attributed to the effectiveness of these evasions, and the limits of engines may also make this unsatisfactory result.In this study, we propose a data-driven approach to measure the effect of app transformations to malware detection, and further explain why the detection result is produced by these engines.First, we develop an interaction model for antivirus engines, illustrating how they respond with different detection results in terms of varying inputs.Six app transformation techniques are implemented in order to generate a large number of Android apps with traceable changes.Then we undertake a onemonth tracking of app detection results from multiple antivirus engines, through which we obtain over 971K detection reports from VirusTotal for 179K apps in total.Last, we conduct a comprehensive analysis of antivirus engines based on these reports from the perspectives of signature-based, static analysis-based, and dynamic analysis-based detection techniques.The results, together with 7 highlighted findings, identify a number of sealed working mechanisms occurring inside antivirus engines and what are the indicators of compromise in apps during malware detection.
Guozhu Meng, Zhixiu Guo, Xiaodong Zhang 0014, Haoyu Wang 0001, Kai Chen 0012, Yang Liu 0003
Internetware3
2025 RSFuzz: A Robustness-Guided Swarm Fuzzing Framework Based on Behavioral Constraints
abstract
Multi-robot swarms play an essential role in complex missions including battlefield reconnaissance, agricultural pest monitoring, as well as disaster search and rescue. Unfortunately, given the complexity of swarm algorithms, logical vulnerabilities are inevitable and often lead to severe safety and security consequences. Although various methods have been presented for detecting logical vulnerabilities through software testing, when they are used in swarm environments, these techniques face significant challenges: 1) Due to the swarm’s vast composable parameter space, it is extremely difficult to generate failure-triggering scenarios, which is crucial to effectively expose logical vulnerabilities; 2) Because of the swarm’s high flexibility and dynamism, it is challenging to model and evaluate the global swarm state, particularly in terms of cooperative behaviors, which makes it difficult to detect logical vulnerabilities.In this work, we propose RSFuzz, a robustness-guided swarm fuzzing framework designed to detect logical vulnerabilities in multi-robot systems. It leverages the robustness of behavioral constraints to quantitatively evaluate the swarm state and guide the generation of failure-triggering scenarios. In addition, RSFuzz identifies and targets key swarm nodes for perturbations, effectively reducing the input space. Upon the RSFuzz framework, we construct two swarm fuzzing schemes, Single Attacker Fuzzing (SA-Fuzzing) and Multiple Attacker Fuzzing (MA-Fuzzing), which employ single and multiple attackers, respectively, during fuzzing to disturb swarm mission execution. We evaluated RSFuzz’s performance with three popular swarm algorithms in simulated environments. The results show that RSFuzz outperforms the state-of-the-art with an average improvement of 17.75% in effectiveness and a 38.4% increase in efficiency. We also validated some detected vulnerabilities in real-world environments. Our code and data are publicly available.
Ruoyu Zhou, Zhiwei Zhang 0004, Haocheng Han, Xiaodong Zhang 0014, Zehan Chen, Jun Sun 0001, Yulong Shen 0001, Dehai Xu
ASE4
2025 Effectively Detecting Software Vulnerabilities via Leveraging Features on Program Slices
abstract
Detecting software vulnerabilities has become increasingly challenging with the growing size and complexity of modern software. Traditional static and dynamic analysis methods often suffer from poor accuracy and reliance on expert knowledge. In recent years, deep learning has shown great promise in this domain due to its ability to automatically learn subtle features from software data. However, existing deep-learning-based methods face two main limitations: 1) difficulty in effectively processing long source code sequences, leading to suboptimal feature representation and 2) insufficient exploration and utilization of common vulnerability features, which hampers further performance improvements. To address these challenges, we propose DV-LVF, a novel deep-learning-based vulnerability detection method that combines program slicing with gated recurrent unit (GRU) embedding techniques to enhance feature representation. Additionally, we introduce a vulnerability dictionary (vulDict) that explicitly captures and leverages common vulnerability patterns to improve detection accuracy. Our evaluation demonstrates that DV-LVF outperforms state-of-the-art methods, achieving accuracies of 98.59% at the function level and 99.27% at the statement level. Notably, DV-LVF successfully identifies 11 previously unknown vulnerabilities across six open-source software projects, including GPAC, Vim, NanoMQ, PJSIP, Libmobi, and Radare2.
Xiaodong Zhang 0014, Zhiwei Zhang 0004, Guiyuan Tang, Jun Sun 0001, Yulong Shen 0001, Jianfeng Ma 0001
IEEE Internet Things J.1
2024 REDriver: Runtime Enforcement for Autonomous Vehicles
abstract
Autonomous driving systems (ADSs) integrate sensing, perception, drive control, and several other critical tasks in autonomous vehicles, motivating research into techniques for assessing their safety. While there are several approaches for testing and analysing them in high-fidelity simulators, ADSs may still encounter additional critical scenarios beyond those covered once they are deployed on real roads. An additional level of confidence can be established by monitoring and enforcing critical properties when the ADS is running. Existing work, however, is only able to monitor simple safety properties (e.g., avoidance of collisions) and is limited to blunt enforcement mechanisms such as hitting the emergency brakes. In this work, we propose REDriver, a general and modular approach to runtime enforcement, in which users can specify a broad range of properties (e.g., national traffic laws) in a specification language based on signal temporal logic (STL). REDriver monitors the planned trajectory of the ADS based on a quantitative semantics of STL, and uses a gradient-driven algorithm to repair the trajectory when a violation of the specification is likely. We implemented REDriver for two versions of Apollo (i.e., a popular ADS), and subjected it to a benchmark of violations of Chinese traffic laws. The results show that REDriver significantly improves Apollo's conformance to the specification with minimal overhead.
Yang Sun 0008, Christopher M. Poskitt, Xiaodong Zhang 0014, Jun Sun 0001
ICSE3
2024 Detecting Vulnerabilities via Explicitly Leveraging Vulnerability Features on Program Slices
Xiaodong Zhang 0014, Zhiwei Zhang 0004, Yulong Shen 0001
TASE2
2022 Understanding Real-world Threats to Deep Learning Models in Android Apps
abstract
Famous for its superior performance, deep learning (DL) has been popularly used within many applications, which also at the same time attracts various threats to the models. One primary threat is from adversarial attacks. Researchers have intensively studied this threat for several years and proposed dozens of approaches to create adversarial examples (AEs). But most of the approaches are only evaluated on limited models and datasets (e.g., MNIST, CIFAR-10). Thus, the effectiveness of attacking real-world DL models is not quite clear. In this paper, we perform the first systematic study of adversarial attacks on real-world DNN models and provide a real-world model dataset named RWM. Particularly, we design a suite of approaches to adapt current AE generation algorithms to the diverse real-world DL models, including automatically extracting DL models from Android apps, capturing the inputs and outputs of the DL models in apps, generating AEs and validating them by observing the apps' execution. For black-box DL models, we design a semantic-based approach to build suitable datasets and use them for training substitute models when performing transfer-based attacks. After analyzing 245 DL models collected from 62,583 real-world apps, we have a unique opportunity to understand the gap between real-world DL models and contemporary AE generation algorithms. To our surprise, the current AE generation algorithms can only directly attack 6.53% of the models. Benefiting from our approach, the success rate upgrades to 47.35%.
Zizhuang Deng, Kai Chen 0012, Guozhu Meng, Xiaodong Zhang 0014
CCS4
2020 Tell You a Definite Answer: Whether Your Data is Tainted During Thread Scheduling
abstract
With the advent of multicore processors, there is a great need to write parallel programs to take advantage of parallel computing resources. However, due to the nondeterminism of parallel execution, the malware behaviors sensitive to thread scheduling are extremely difficult to detect. Dynamic taint analysis is widely used in security problems. By serializing a multithreaded execution and then propagating taint tags along the serialized schedule, existing dynamic taint analysis techniques lead to under-tainting with respect to other possible interleavings under the same input. In this paper, we propose an approach called DSTAM that integrates symbolic analysis and guided execution to systematically detect tainted instances on all possible executions under a given input. Symbolic analysis infers alternative interleavings of an executed trace that cover new tainted instances, and computes thread schedules that guide future executions. Guided execution explores new execution traces that drive future symbolic analysis. We have implemented a prototype as part of an educational tool that teaches secure C programming, where accuracy is more critical than efficiency. To the best of our knowledge, DSTAM is the first algorithm that addresses the challenge of taint analysis for multithreaded program under fixed inputs.
Xiaodong Zhang 0014, Zijiang Yang 0006, Yu Hao 0006, Ting Liu 0002
IEEE Trans. Software Eng.1
2017 Automated Testing of Definition-Use Data Flow for Multithreaded Programs
abstract
With the advent of multicore processors, there is a trend towards multithreading to take advantage of parallel computing resources. Due to greatly increased complexity, programmers need effective testing methodology that can thoroughly test multithreaded programs. There has been significant progress based on symbolic execution that attempts to exhaustively explore all the intra-thread paths and inter-thread interleavings. However, such testing approach faces two insuperable challenges. Firstly, exploring an astronomically large number of paths and interleavings limits its scalability. Secondly, a path itself does not directly help programmers understand program behavior. In this paper, we propose an alternate testing methodology that focuses on definition-use data flow instead of paths/interleavings. Such approach not only leads to orders of magnitude reduction in testing complexity, but also gives programmers direct help on examining the shared variable usage in a multithreaded program.
Xiaodong Zhang 0014, Zijiang Yang 0006, Jialiang Chang, Yu Hao 0006, Ting Liu 0002
ICST1
2016 Frequent Subgraph Based Familial Classification of Android Malware
abstract
The rapid growth of Android malware poses great challenges to anti-malware systems because the sheer number of malware samples overwhelm malware analysis systems. A promising approach for speeding up malware analysis is to classify malware samples into families so that the common features in malwares belonging to the same family can be exploited for malware detection and inspection. However, the accuracy of existing classification solutions is limited because of two reasons. First, since the majority of Android malware is constructed by inserting malicious components into popular apps, the malware's legitimate part may misguide the classification algorithms. Second, the polymorphic variants of Android malware could evade the detection by employing transformation attacks. In this paper, we propose a novel approach that constructs frequent subgraph (fregraph) to represent the common behaviors of malwares in the same family for familial classification of Android malware. Moreover, we propose and develop FalDroid, an automatic system for classifying Android malware according to fregraph, and apply it to 6,565 malware samples from 30 families. The experimental results show that FalDroid can correctly classify 94.5% malwares into their families using around 4.4s per app.
Ming Fan 0002, Jun Liu 0002, Xiapu Luo, Kai Chen 0012, Zhenzhou Tian, Xiaodong Zhang 0014, Ting Liu 0002
ISSRE7
2014 Plagiarism detection for multithreaded software based on thread-aware software birthmarks
abstract
The availability of inexpensive multicore hardware presents a turning point in software development. In order to benefit from the continued exponential throughput advances in new processors, the software applications must be multithreaded programs. As multithreaded programs become increasingly popular, plagiarism of multithreaded programs starts to plague the software industry. Although there has been tremendous progress on software plagiarism detection technology, existing dynamic approaches remain optimized for sequential programs and cannot be applied to multithreaded programs without significant redesign. This paper fills the gap by presenting two dynamic birthmark based approaches. The first approach extracts key instructions while the second approach extracts system calls. Both approaches consider the effect of thread scheduling on computing software birthmarks. We have implemented a prototype based on the Pin instrumentation framework. Our empirical study shows that the proposed approaches can effectively detect plagiarism of multithread programs and exhibit strong resilience to various semantic-preserving code obfuscations.
Zhenzhou Tian, Ting Liu 0002, Ming Fan 0002, Xiaodong Zhang 0014, Zijiang Yang 0006
ICPC5