Xiaogang Zhu 0001

dblp:171/2501 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-0647-4747ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 12 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TWINFUZZ: Dual-Model Fuzzing for Robustness Generalization in Deep Learning
abstract
Deep learning (DL) models are increasingly deployed in safety-critical applications such as face recognition, autonomous driving, and medical diagnosis. Despite their impressive accuracy, they remain vulnerable to adversarial examples - subtle perturbations that can cause incorrect predictions, i.e., the robustness issues. While adversarial training improves robustness against known attacks, it often fails to generalize to unseen or stronger threats, revealing a critical gap in robustness generalization. In this work, we propose a dual-model fuzzing framework to enhance generalized robustness in DL models. Central to our method is a lightweight metric, the Lagrangian Information Bottleneck (LIB), which guides entropy-based mutation toward semantically meaningful and high-risk regions of the input space. The executor uses a resistant model and a more error-prone vulnerable model; their prediction consistency forms the basis of agreement mining, a label-free oracle for isolating decision-boundary samples. To ensure fuzzing effectiveness, we further introduce a task-driven seed selection strategy (e.g., SSIM for vision) that filters out low-quality inputs. We implement a prototype, TWINFUZZ, and evaluate it on six benchmark datasets and nine DL models. Compared with state-of-the-art testing approaches, TWINFUZZ achieves superior improvements in both training-specific and generalized robustness.
Enze Dai, Wentao Mo, Kun Hu 0008, Xiaogang Zhu 0001, Xi Xiao 0001, Sheng Wen, Shaohua Wang 0002, Yang Xiang 0001
AAAI4
2026 IsolatOS: Detecting Double Fetch Bugs in COTS RTOS by Re-enabling Kernel Isolation
Yingjie Cao, Xiaogang Zhu 0001, Dean Sullivan, Lei Xue 0001, Chenxiong Qian, Minrui Yan, Xiapu Luo
NDSS2
2026 Securing the low-altitude economy: a survey
abstract
Abstract The rapid growth of the low-altitude economy, including unmanned aerial vehicles (UAVs) and urban air mobility (UAM), is reshaping industries from transportation to emergency response. Powered by advances in fifth-generation (5G) and 5G-advanced (5.5G) connectivity, artificial intelligence (AI), and new energy systems, these platforms are becoming increasingly autonomous and capable. However, their growing software complexity introduces critical cybersecurity risks. Vulnerabilities in communication protocols, onboard firmware, and AI systems can be exploited to hijack UAVs, disrupt operations, or leak sensitive data. While research has addressed isolated aspects, a unified security perspective is still lacking. This work presents a systematic review of software-level security challenges and defenses in low-altitude UAV/UAM systems. We first categorize major attack surfaces across communication, firmware, and AI layers. Furthermore, we survey defense mechanisms suited to real-time, resource-constrained aerial platforms. Finally, we propose future directions, including quantum-resistant communication protocols, hardware-software cosecurity, and edge-AI-driven architectures. Our work aims to inform researchers, practitioners, and regulators in developing integrated, resilient security strategies for the evolving low-altitude ecosystem.
Minrui Yan, Ruiqi Dong, Qing-Long Han, Zehang Deng, Wanlun Ma, Xiaogang Zhu 0001, Wei Zhou 0044, Sheng Wen, Yang Xiang 0001
Sci. China Inf. Sci.6
2026 Reverse Engineering of Industrial Protocols From Network Traffic
abstract
Reliable protocol knowledge is often difficult to obtain in industrial networks, as industrial communications come with limited documentation, vendor-specific encodings, and opaque payloads. This lack of transparency hinders message interpretation and protocol analysis. To recover this missing protocol knowledge, network-trace-based protocol reverse engineering (PRE) infers message structure, field roles, and interaction logic directly from recorded traces. This enables protocol-aware intrusion detection, process monitoring, and protocol testing and fuzzing without access to device internals. Although PRE has advanced rapidly, existing techniques are developed under diverse objectives and assumptions. As a result, it is often unclear how isolated results relate to an end-to-end reverse-engineering workflow, and how evaluation outcomes should be compared across tasks and protocols. In this article, we cast reverse engineering of industrial protocols from network traces as a task-driven pipeline and articulate a unified task decomposition spanning message type identification, protocol syntax and semantic inference, payload pattern recognition and semantic inference, and protocol state machine reconstruction. For each task, we describe key methodological themes, common evaluation practices, and practical limitations that affect robustness and deployability in industrial settings. We further discuss security, privacy, and ethical risks that accompany increasingly capable PRE, and identify promising research directions toward more systematic, dependable, and deployment-oriented PRE methodologies.
Chuan Sheng, Shan Jiang 0023, Qing-Long Han, Wei Zhou 0044, Wanlun Ma, Xiaogang Zhu 0001, Sheng Wen, Yang Xiang 0001
IEEE Trans. Ind. Informatics6
2025 RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
abstract
Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotation-invariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks.
Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
AAAI4
2025 Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language Models
abstract
Deep learning (DL) libraries are widely used to form the basis of various AI applications in computer vision, natural language processing, and software engineering domains. Despite their popularity, DL libraries are known to have vulnerabilities, such as buffer overflows, use-after-free, and integer overflows, that can be exploited to compromise the security or effectiveness of the underlying libraries. While traditional fuzzing techniques have been used to find bugs in software, they are not well-suited for DL libraries. In general, the complexity of DL libraries and the diversity of their APIs make it challenging to test them thoroughly. To date, mainstream DL libraries like TensorFlow and PyTorch have featured over 1,000 APIs, and the number of APIs is still growing. Fuzzing all these APIs is a daunting task, especially when considering the complexity of the input data and the diversity of the API usage patterns. Recent advances in large language models (LLMs) have illustrated the high potential of LLMs in understanding and synthesizing human-like code. Despite their high potential, we find that emerging LLM-based fuzzers are less optimal for DL library API fuzzing, given their lack of in-depth knowledge on API input edge cases and inefficiency in generating test inputs. In this paper, we propose DFuzz, a LLM-driven DL library fuzzing approach. We have two key insights: (1) With high reasoning ability, LLMs can replace human experts to reason edge cases (likely error-triggering inputs) from checks in an API's code, and transfer the extracted knowledge to test other (new or rarely-tested) APIs. (2) With high generation ability, LLMs can synthesize initial test programs with high accuracy that automates API testing. DFuzz provides LLMs with a novel “white-box view” of DL library APIs, and therefore, can leverage LLMs' reasoning and generation abilities to achieve comprehensive fuzzing. Our experimental results on popular DL libraries demonstrate that DFuzz is able to cover more APIs than SOTA (LLM-based) fuzzers on TensorFlow and PyTorch, respectively. Moreover, DFuzz successfully detected 37 bugs, with 8 already fixed and 19 replicated by the developer but still under investigation.
Shuai Wang 0011, Jitao Han, Xiaogang Zhu 0001, Shaohua Wang 0002, Sheng Wen
ICSE4
2025 FailMapper: Automated Generation of Unit Tests Guided by Failure Scenarios
abstract
The automation of unit test generation has become a critical task for improving the overall efficiency of software development and testing. Many existing techniques attempt to generate a sufficient number of test cases to achieve high code coverage. However, it has been shown that a high coverage does not necessarily guarantee effective bug discovery. A potential enhancement is to guide the unit test generation based on bug properties. However, this solution is challenged by the large number and diversity of bug types, making it difficult to comprehensively summarize bug properties.We observe that failures, presented as the results of bugs, manifest in a limited number of scenarios. Therefore, instead of bug properties, in this paper, we propose an innovative framework, named FailMapper, which uses failure scenarios to guide the generation of unit tests. We summarize nine failure scenarios and design the corresponding failure-triggering test strategies. This significantly improves the efficacy of generating test cases towards triggering bugs. To systematically explore possible failure scenarios, FailMapper employs the Monte Carlo Tree Search algorithm to search for the faults that may lead to a failure. Experiments demonstrate that, on 50 known bugs in the Defects4J benchmark, FailMapper can detect many more bugs than five typical unit testing approaches, including EvoSuite, Randoop, CoverUp, HITS, and SymPrompt (40 versus at most 12, out of all 50 bugs). Meanwhile, FailMapper detects 12 out of 20 bugs in the GitBug-Java and Bears-benchmark datasets. We reveal 36 potential issues from 2 Apache projects, and 14 of them have been confirmed as bugs, further demonstrating FailMapper’s effectiveness. The experimental results show that our new framework can significantly enhance the overall efficacy of unit testing.
Ruiqi Dong, Zehang Deng, Xiaogang Zhu 0001, Xiaoning Du 0001, Huai Liu, Shaohua Wang 0002, Sheng Wen, Yang Xiang 0001
ASE3
2025 WingMuzz: Blackbox Testing of IoT Protocols via Two-dimensional Fuzzing Schedule
abstract
The Internet of Things (IoT) is widely used in various sectors but is often prone to vulnerabilities. With the proprietary nature of IoT devices, their source code and firmware are frequently unavailable for open review, rendering blackbox fuzzing a viable approach. However, the effectiveness of blackbox fuzzing is often challenging due to the lack of feedback, especially the information of code coverage. In this paper, we propose WingMuzz to provide blackbox fuzzing of IoT protocols with effective feedback. The key is to guide blackbox fuzzing by utilizing runtime information from greybox fuzzing on counterpart open-source code. This is based on our observation that IoT protocols and open-source code conform to the same specifications, indicating that inputs exploring different code regions on open-source code may also discover new coverage on IoT protocols. WingMuzz uses a two-dimensional fuzzing schedule to optimize the process of fuzzing IoT protocols. The first dimension involves scheduling open-source implementations, referred to as wingmates, so that similar ones are preferred to guide blackbox fuzzing. The second dimension utilizes coverage-guided greybox fuzzing to test open-source code. This solution can bridge the performance gap between blackbox fuzzing and greybox fuzzing on IoT protocols. We evaluate the performance of WingMuzz across eight IoT protocols and compare it with six widely-used blackbox fuzzers. On average, WingMuzz can discover 42.1%, 26.92%, 25.01%, 34.95%, 23.56% and 11.63% more edges than Boofuzz, Spike, Peach, Snipuzz, Pulsar and ChatAFL, respectively. Additionally, WingMuzz exposes 10 bugs in IoT protocols while other fuzzers expose no more than 3 bugs. It also exposes 2 new protocol vulnerabilities in IoT devices while other fuzzers cannot identify any.
Xiaogang Zhu 0001, Enze Dai, Xiaotao Feng, Shaohua Wang 0002, Xin Xia 0001, Sheng Wen, Kwok-Yan Lam, Yang Xiang 0001
ASE1
2025 Codebreaker: Dynamic Extraction Attacks on Code Language Models
abstract
With the rapid adoption of LLM-based code assistants to enhance programming experiences, concerns over extraction attacks targeting private training data have intensified. These attacks specifically aim to extract Personal Information (PI) embedded within the training data of code generation models (CodeLLMs). Existing methods, using either manual or semi-automated techniques, have successfully extracted sensitive data from these CodeLLMs. However, the limited amount of data currently retrieved by extraction attacks risks significantly underestimating the true extent of training data leakage. In this paper, we propose an automatic PI data extraction attack framework against LLM-based code assistants, named Codebreaker. This framework is built on two core components: (i) the introduction of semantic entropy, which evaluates the likelihood of a prompt triggering the model to respond with training data; and (ii) an automatic dynamic mutation mechanism that seamlessly integrates with Codebreaker, reinforcing the iterative process across the framework and promoting greater interconnection between different PI elements within a single response. This boosts reasoning diversity, model memorization, and finally attack performance. Using six series of open-source CodeLLMs (i.e., CodeParrot, StarCoder2, Code Llama, CodeGemma, DeepSeek-Coder, DeepSeek-V3) and two commercial code assistants (i.e., CodeFuse and GPT), we demonstrate the effectiveness of our proposed framework: (i) Codebreaker outperforms all current state-of-the-art extraction attacks by 6.22% ~ 44.9% (averaging 21.79%); (ii) when PI within a single response originates from the same GitHub repository, our framework - considering multiple interconnections in the response - exceeds others by 3.88% ~ 32.37% (averaging 15.31%). Furthermore, we discuss potential defenses, highlighting the urgent need for stronger measures to prevent PI leakage at the base model level.
Changzhou Han, Zehang Deng, Wanlun Ma, Xiaogang Zhu 0001, Minhui Xue 0001, Tianqing Zhu, Sheng Wen, Yang Xiang 0001
SP4
2025 Blockchain Cross-Chain Bridge Security: Challenges, Solutions, and Future Outlook
abstract
Cross-chain bridges, one of the foundational infrastructures of blockchain, provide the infrastructure and solutions for inter-operability, asset liquidity, data transfer, decentralized finance, and cross-chain governance between blockchain networks. However, because cross-chain bridges often have to handle communication and asset transfers between multiple blockchains, they involve complex protocols and technologies. This complexity increases the likelihood of vulnerabilities and potential attacks. In order to ensure the security and reliability of cross-chain bridges, this article launches a thorough investigation of existing cross-chain bridge projects, clarifying bridging mechanisms, bridge types, and security features. The following part goes into the subject of security and sheds light on the considerable challenges faced by cross-chain bridges. It conducts a thorough analysis of security flaws, covering problems like smart contract vulnerabilities, centralization risks, liquidity issues, and oracle manipulations. Furthermore, this study promotes a compendium of security solutions and best practises, pointing the way toward a cross-chain bridge scenario that is more secure.
Ningran Li, Minfeng Qi, Xiaogang Zhu 0001, Wei Zhou 0044, Sheng Wen, Yang Xiang 0001
Distributed Ledger Technol. Res. Pract.4
2025 One Mutation Fits All: Exploring Universal Library Fuzzing Based on Exogenous Mutation
abstract
Fuzzing is a critical technique for uncovering vulnerabilities in software libraries. However, current approaches often struggle with cross-language compatibility and integration with diverse fuzzing tools. We proposeEXo-Muta, a novel universal library fuzzing framework based on exogenous mutation. We use the term ‘exogenous’ to describe this new mutation process because it operates externally to the fuzzer's core engine, executing within the fuzz driver as an independent component, unlike traditional endogenous mutations tightly integrated within the fuzzer itself. By decoupling the mutation process from specific fuzzers,EXo-Mutaachieves unprecedented adaptability across diverse programming languages and fuzzing tools. It leverages static analysis to extract structured data representations and applies language-independent mutation operators at the code level. This design enables seamless integration with various existing fuzzers, enhancing their performance regardless of the target language. We further utilize large language models (LLM) for efficient cross-language data conversion. In experiments, we evaluatedEXo-Mutaon 20 real-world libraries across C++, Python, Java, and JavaScript, integrating it with multiple stateof- the-art fuzzers, such as AFL++ and libFuzzer, and languagespecific tools, such as Atheris and Jazzer. Results show significant improvements in code coverage across different fuzzers and languages, with up to 58% more edges discovered in C++ projects when integrated with libFuzzer, and consistent outperformance in other scenarios (27% on Python, 9% on Java, 6% on JavaScript).EXo-Mutarepresents a significant advancement in fuzzing technology, offering a universal, language-agnostic approach that substantially improves code coverage across diverse programming languages and fuzzing tools, thereby expanding the reach and effectiveness of library API testing.
Ruiqi Dong, Fanke Tong, Xiaogang Zhu 0001, Xi Xiao 0001, Shaohua Wang 0002, Sheng Wen, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.4
2025 BazzAFL: Moving Fuzzing Campaigns Towards Bugs via Grouping Bug-Oriented Seeds
abstract
As one of the most successful techniques in hunting software bugs, Coverage-guided Greybox Fuzzing (CGF) intends to move fuzzing campaigns towards executions that can trigger bugs. This process can be divided into two steps, including reaching suspicious code regions and exploring their execution states. Many CGFs propose approaches to efficiently reach suspicious code regions and individual execution states, but fail to explore complex execution states. The challenge is how to explore execution states so that fuzzing can detect multiple types of bugs, while maintaining the code coverage. To address this challenge, we proposeBazzAFLto investigate code coverage and multiple types of bugs. The crux ofBazzAFLis to maintain a bunch of seed groups, where each seed saves the best performance on one objective. With the seed group,BazzAFLprioritizes code regions that most likely contain bugs based on multi-objective optimization and adaptively divides energy among the seeds in a group based on Shannon's entropy. Meanwhile, during mutation,BazzAFLtends to mutating the bytes that can change the execution states. With these solutions,BazzAFLgradually moves fuzzing campaigns towards locations and execution states of bugs. Experimental results show thatBazzAFLidentifies at least 62 more bugs on 24 programs compared with other fuzzers.
Xiaogang Zhu 0001, Xi Xiao 0001, Sheng Wen, Minhui Xue 0001, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.2
2025 Network Traffic Fingerprinting for IIoT Device Identification: A Survey
abstract
As the Industrial Internet of Things (IIoT) continues to expand, the need for effective device identification becomes critical for securing industrial environments. Network traffic fingerprinting has emerged as an important technique for IIoT device identification, leveraging the unique communication patterns embedded in network traffic. Despite significant efforts in this area, a comprehensive overview of the relevant research is still missing. To address the lack of comprehensive research, this paper, for the first time, identifies critical knowledge gaps constraining IIoT device identification through network traffic analysis: obscure fingerprint feature space, limited generalizability to unknowns, and scarce data sources. Focusing on these gaps, existing methods are analyzed and summarized in detail across network traffic fingerprinting, IIoT device identification, and public IIoT datasets. Specifically, network traffic fingerprinting methods are categorized into three levels: Packet-level, flow-level, and business-level, and relevant methods are examined in terms of data formats, segmentation units, and extraction or generation techniques. In the context of IIoT device identification, tasks such as device type, model, and instance recognition, as well as abnormal device detection, are extensively investigated using rule-based, traditional machine learning- based, and deep learning-based approaches, with a focus on device fingerprints and application scenarios. Furthermore, main public datasets from the IoT, ICS, and IIoT scenarios are highlighted to support the development of fingerprinting and identification methods. Finally, several future research directions are proposed to guide new advancements in this area.
Chuan Sheng, Wei Zhou 0044, Qing-Long Han, Wanlun Ma, Xiaogang Zhu 0001, Sheng Wen, Yang Xiang 0001
IEEE Trans. Ind. Informatics5
2024 SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
Lintao Wang 0002, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
MMAsia3
2024 ShapFuzz: Efficient Fuzzing via Shapley-Guided Byte Selection
Xiaogang Zhu 0001, Xi Xiao 0001, Minhui Xue 0001, Chao Zhang 0008, Sheng Wen
NDSS2
2024 How COVID-19 impacts telehealth: an empirical study of telehealth services, users and the use of metaverse
abstract
Since the outbreak of the coronavirus 2019 (COVID-19) pandemic, telehealth services are regarded as a good approach to keep health workers and patients safe while simultaneously managing available resources.In this paper, we discuss the impact that COVID-19 has on telehealth services and on telehealth users' opinion of the service.We collected 245 Android telehealth apps, 144 iOS telehealth apps and 86 telehealth websites, and performed a systematic analysis on this dataset.In this analysis, we conducted a comparison analysis and relevant content analysis of the telehealth apps as well as their security risks.Apart from the mobile platforms, we also inspected the telehealth websites' features, particularly those related to the use of metaverse to improve current telehealth solutions.To further understand people's attitude towards telehealth services, we invited users to participate in a user study aimed at revealing what impact COVID-19 has on users' willingness to adopt telehealth services and revealing the gap between the telehealth service and its users.Our result shows that 27.1% new iOS apps and 27.4% new Android apps were released after the COVID-19 announcement, and a surge of updates were noted within 4 weeks after the COVID-19 announcement.We further found that COVID-19 is frequently mentioned in telehealth app reviews in the second and third quarter of 2020, and the most mentioned aspects related to COVID-19 include family, test result and vaccine.According to our user study, COVID-19 has a significant impact on the selection of telehealth services, especially for female participants, people aged 46-55, and students.The investigation also finds out that the use of metaverse will significantly improves the effectiveness of traditional telehealth solutions.
Lihong Tang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Wei Zhou 0044, Xiaogang Zhu 0001, Yang Xiang 0001
Connect. Sci.6
2024 Fuzzing Android Native System Libraries via Dynamic Data Dependency Graph
abstract
Google suggests using only the APIs documented in Android SDK. However, many app developers still choose Java Native Interface (JNI) to access system libraries because of the flexibility and freedom that non-SDK methods provide in implementing complex functions. However, using JNI may have unexpected consequences, including low-level bug-driven crashes. The bugs in system libraries can propagate to Android apps, and further cost much time and energy for developers to debug them. We develop a fuzzing tool, called JDYNUZZ, that exposes the bugs in system JNI to mitigate the aftermath of direct invocation of JNI. To fuzz a system library, one needs to not only prepare appropriate inputs, but also deal with the challenge of maintaining a correct sequence of API calls, both syntactically and semantically. To solve the challenge, the crux of JDYNUZZ is the dynamic refinement of a data dependency graph, which gradually resolves the problem of syntactic and semantic incorrectness when constructing API sequences. JDYNUZZ achieves the dynamic refinement based on the feature of Java reflection, which enables us to dynamically modify API sequences and test different code regions. We evaluate JDYNUZZ on the most recent version of Android Open Source Project (AOSP),i.e., version android-12.0.0 r31. In our experiments, JDYNUZZ discovers 34 new bugs in system JNI libraries, all confirmed by Google.
Xiaogang Zhu 0001, Sheng Wen, Yang Xiang 0001
IEEE Trans. Inf. Forensics Secur.1
2023 Detecting Union Type Confusion in Component Object Model
Xiaogang Zhu 0001, Daojing He, Minhui Xue 0001, Shouling Ji, Mohammad Sayad Haghighi, Sheng Wen, Zhiniang Peng
USENIX Security Symposium2
2023 On the security of fully homomorphic encryption for data privacy in Internet of Things
abstract
Summary To achieve data privacy in Internet of Things (IoT), fully homomorphic encryption (FHE) technique is used to encrypt the data while allowing others to compute on the encrypted data. However, there are many well‐known problems with FHE such as chosen‐ciphertext attack security and circuit privacy problem. In this article, we demonstrate that a famous FHE application named Brakerski/Fan–Vercauteren scheme, a circuit privacy application based on fast private set intersection, and an encoding application that encodes integer or floating point numbers based on Microsoft Simple Encryption Arithmetic Library homomorphic encryption library, are insecure against chosen ciphertext attacks due to insecurity of the underlying fully homomorphic schemes. These results show that using cryptographic primitives even with security proofs causes serious security vulnerabilities on the applications themselves. The results also give evidences that the security of adopted cryptographic primitives in IoT should be proved in appropriate formal security models as well as proof of the scheme itself.
Zhiniang Peng, Wei Zhou 0044, Xiaogang Zhu 0001, Youke Wu, Sheng Wen
Concurr. Comput. Pract. Exp.3
2022 Path Transitions Tell More: Optimizing Fuzzing Schedules via Runtime Program States
abstract
Coverage-guided Greybox Fuzzing (CGF) is one of the most successful and widely-used techniques for bug hunting. Two major approaches are adopted to optimize CGF: (i) to reduce search space of inputs by inferring relationships between input bytes and path constraints; (ii) to formulate fuzzing processes (e.g., path transitions) and build up probability distributions to optimize power schedules, i.e., the number of inputs generated per seed. However, the former is subjective to the inference results which may include extra bytes for a path constraint, thereby limiting the efficiency of path constraints resolution, code coverage discovery, and bugs exposure; the latter formalization, concentrating on power schedules for seeds alone, is inattentive to the schedule for bytes in a seed.
Xi Xiao 0001, Xiaogang Zhu 0001, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen
ICSE3
2022 CSI-Fuzz: Full-Speed Edge Tracing Using Coverage Sensitive Instrumentation
abstract
Coverage-guided fuzzing is one of the most effective solutions for vulnerability discovery. Among coverage-guided fuzzing, full-speed fuzzing, such as UnTracer, traces test cases only when they discover new coverage. Due to the high expense of tracing test cases, full-speed fuzzers improve the efficiency of fuzzing by tracing only coverage-increasing test cases. However, the existing full-speed fuzzer (i.e., UnTracer) is based on basic block coverage, suffering a severe problem called edge collision. Moreover, such fuzzers neglect the path frequency, which affects fuzzing effectiveness. In this article, we propose CSI-Fuzz, a fuzzer utilizing coverage sensitive instrumentation to address the problems of existing full-speed fuzzing. CSI-Fuzz directly instruments at edges, which solves the problem of edge collision. Meanwhile, CSI-Fuzz sets path identifiers to count the frequency of covered paths. Our CSI-Fuzz can be recognized as an add-on and seamlessly applied to existing coverage-guided fuzzers. We accordingly implement CSI-Fuzz based on two widely-adopted fuzzers, AFL and AFLFast, to evaluate its performance. The experiments demonstrate that CSI-Fuzz discovers more edges than AFL, AFLFast, and UnTracer. Additionally, CSI-Fuzz exposes more bugs than the other fuzzers.
Xiaogang Zhu 0001, Xiaotao Feng, Xiaozhu Meng, Sheng Wen, Seyit Ahmet Çamtepe, Yang Xiang 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.1
2021 Snipuzz: Black-box Fuzzing of IoT Firmware via Message Snippet Inference
abstract
The proliferation of Internet of Things (IoT) devices has made people's lives more convenient, but it has also raised many security concerns. Due to the difficulty of obtaining and emulating IoT firmware, in the absence of internal execution information, black-box fuzzing of IoT devices has become a viable option. However, existing black-box fuzzers cannot form effective mutation optimization mechanisms to guide their testing processes, mainly due to the lack of feedback. In addition, because of the prevalent use of various and non-standard communication message formats in IoT devices, it is difficult or even impossible to apply existing grammar-based fuzzing strategies. Therefore, an efficient fuzzing approach with syntax inference is required in the IoT fuzzing domain.
Xiaotao Feng, Ruoxi Sun 0001, Xiaogang Zhu 0001, Minhui Xue 0001, Sheng Wen, Dongxi Liu, Surya Nepal, Yang Xiang 0001
CCS3
2021 Regression Greybox Fuzzing
abstract
What you change is what you fuzz! In an empirical study of all fuzzer-generated bug reports in OSSFuzz, we found that four in every five bugs have been introduced by recent code changes. That is, 77% of 23k bugs are regressions. For a newly added project, there is usually an initial burst of new reports at 2-3 bugs per day. However, after that initial burst, and after weeding out most of the existing bugs, we still get a constant rate of 3-4 bug reports per week. The constant rate can only be explained by an increasing regression rate. Indeed, the probability that a reported bug is a regression (i.e., we could identify the bug-introducing commit) increases from 20% for the first bug to 92% after a few hundred bug reports. In this paper, we introduce regression greybox fuzzing (RGF) a fuzzing approach that focuses on code that has changed more recently or more often. However, for any active software project, it is impractical to fuzz sufficiently each code commit individually. Instead, we propose to fuzz all commits simultaneously, but code present in more (recent) commits with higher priority. We observe that most code is never changed and relatively old. So, we identify means to strengthen the signal from executed code-of-interest. We also extend the concept of power schedules to the bytes of a seed and introduce Ant Colony Optimization to assign more energy to those bytes which promise to generate more interesting inputs. Our large-scale fuzzing experiment demonstrates the validity of our main hypothesis and the efficiency of regression greybox fuzzing. We conducted our experiments in a reproducible manner within Fuzzbench, an extensible fuzzer evaluation platform. Our experiments involved 3+ CPU-years worth of fuzzing campaigns and 20 bugs in 15 open-source C programs available on OSSFuzz.
Xiaogang Zhu 0001, Marcel Böhme
CCS1
2020 SpeedNeuzz: Speed Up Neural Program Approximation with Neighbor Edge Knowledge
abstract
Fuzzing has been a great success in discovering real-world complex programs vulnerabilities. However, fuzzing achieves this effect by blindly generating a large number of test cases, which undoubtedly contains a lot of meaningless mutation inputs. To solve the blindness, machine learning technology is applied to fuzzing in recent work. Some of the machine learning based methods focus on locating and mutating the key bytes in the input, but they do not pay attention to the characteristics in the field of fuzzing when they combine machine learning technology with fuzzing. In this paper, we implement a new fuzzer, called Speed-Neuzz, which uses neural networks to model the branch behaviours of the program based on accurate training data after mitigating the hash collision of AFL. Furthermore, SpeedNeuzz locates and mutates critical bytes in the program input with a gradient-based strategy as well as neighbor edge information. Taking the neighbor edge knowledge into account, we can further reduce the blindness of the mutation based on gradient information so that SpeedNeuzz can generate a large number of quality inputs. Experiments on several real-world programs prove that SpeedNeuzz can achieve higher edge coverage than the state-of-the-art fuzzer NEUZZ under the same time budget.
Xi Xiao 0001, Xiaogang Zhu 0001, Xiao Chen 0002, Sheng Wen, Bin Zhang 0048
TrustCom3
2019 A Feature-Oriented Corpus for Understanding, Evaluating and Improving Fuzz Testing
abstract
Fuzzing is a promising technique for detecting security vulnerabilities. Newly developed fuzzers are typically evaluated in terms of the number of bugs found on vulnerable programs/binaries. However, existing corpora usually do not capture the features that prevent fuzzers from finding bugs, leading to ambiguous conclusions on the pros and cons of the fuzzers evaluated. In this paper, we propose to address the above problem by generating corpora based on search-hampering features. As a proof-of-concept, we designed FEData, a prototype corpus that currently focuses on three search-hampering features to generate vulnerable programs for fuzz testing. Unlike existing corpora that can only answer "how", FEData can also further answer "why" by exposing (or understanding) the reasons for the identified weaknesses in a fuzzer. The "why" information serves as the key to the improvement of fuzzers. Based on the "why" information, our FEData programs enabled us to identify the weakness of AFLFast, called cycle explosion, behind. We further developed an improved version of AFLFast, called AFLFast+, which has overcome the cycle explosion problem. AFLFast+ retains the efficiency of AFLFast in path search while maintaining or even surpassing the bug-finding capability of AFL for the corpus evaluated.
Xiaogang Zhu 0001, Xiaotao Feng, Tengyun Jiao, Sheng Wen, Yang Xiang 0001, Seyit Ahmet Çamtepe, Jingling Xue
AsiaCCS1