VLDB 2026 Research / reviewers in the wild / expert
Shuitao Gan
dblp:153/0205
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-6083-9049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FirmAgent: Leveraging Fuzzing to Assist LLM Agents with IoT Firmware Vulnerability Discovery
Jiangan Ji, Chao Zhang 0008, Shuitao Gan, Lin Jian, Hangtian Liu, Tieming Liu, Zhipeng Jia |
NDSS | 3 |
| 2025 | EAGLEYE: Exposing Hidden Web Interfaces in IoT Devices via Routing Analysis
Hangtian Liu, Shuitao Gan, Chao Zhang 0008, Zicong Gao, Yishun Zeng, Zhiyuan Jiang |
NDSS | 3 |
| 2025 | GLMFuzz: Vulnerability Knowledge Guided Prompting for Efficient Network Protocol FuzzingabstractSecurity vulnerabilities in network protocol implementations have increased rapidly, posing serious threats to network infrastructure. Fuzzing has emerged as a primary method for detecting these vulnerabilities. However, due to the stateful nature of protocol implementations, existing mutationbased protocol fuzzing methods face challenges in generating high-quality initial seeds and developing effective mutation strategies. In this paper, we propose GLMFuzz, which is a vulnerability-knowledge-guided fuzzing method for protocol implementations based on large language models(LLMs). By analyzing historical vulnerability sequences, which contain special states and field information that can trigger boundary conditions and exceptional behaviors in protocol implementations, GLMFuzz constructs a vulnerability knowledge vector database and utilizes LLM to generate diverse initial seeds. During the fuzzing process, GLMFuzz uses vulnerability sequences to guide the LLM in performing directed mutations on the message, producing test cases that cover unusual behaviors. In addition, we design an automatic prompt engineering algorithm to dynamically adjust prompt templates based on fuzzing results. To validate the performance of GLMFuzz, we conduct a comparative evaluation with two state-of-the-art (SOTA) tools, AFLNET and CHATAFL. The experimental results demonstrate that GLMFuzz achieves average improvements of 54.18 % and 16.48 % in state transitions, 14.57 % and 4.97 % in state coverage, 10.61 % and 7.01 % in code coverage compared to AFLNET and CHATAFL, respectively, and identifies a previously unknown vulnerability in protocol implementations, which has been assigned CVE identifier. Shengkai Zhu, Shuitao Gan, Yan Nan |
QRS | 2 |
| 2025 | BridgeRouter: Automated Capability Upgrading of Out-Of-Bounds Write Vulnerabilities to Arbitrary Memory Write Primitives in the Linux KernelabstractMemory corruption vulnerabilities pose a significant threat to the Linux kernel, with out-of-bounds (OOB) vulnerabilities receiving particular attention due to their prevalence. The existing kernel OOB exploitation techniques either require strong capabilities from the vulnerabilities, demand that the vulnerable and victim objects reside in the same memory allocator cache, or rely on extensive page table manipulation. These constraints restrict their applicability and lead to low success rates in completing a full exploitation chain. In this paper, we propose a practical approach that enables arbitrary memory writes from kernel OOB vulnerabilities with limited capabilities. Our method leverages two special kinds of kernel objects to upgrade the capability from an uncontrolled overwrite to a controlled overwrite, ultimately achieving arbitrary memory write. We develop a system to automatically identify and utilize these two kinds of kernel objects. Evaluations on a crafted vulnerability and 14 representative real-world vulnerabilities, along with a comparison against two state-of-the-art works, demonstrate the broad applicability of our approach. Dongchen Xie, Dongnan He, Wei You 0001, Jianjun Huang 0001, Bin Liang 0002, Shuitao Gan, Wenchang Shi |
SP | 6 |
| 2024 | Labrador: Response Guided Directed Fuzzing for Black-box IoT DevicesabstractFuzzing is a popular solution to finding vulnerabilities in software including IoT firmware. However, due to the challenges of emulating or rehosting firmware, some IoT devices (e.g., enterprise-level devices) can only be fuzzed in a black-box manner, which makes fuzzers blind and inefficient due to missing feedbacks (e.g., code coverage or distance). In this paper, we present a novel response guided directed fuzzing solution Labrador, able to test black-box IoT devices efficiently. Specifically, we leverage the network response to infer the execution trace of firmware and deduce the code coverage of testing. Second, we leverage the test case (i.e., request) and its response to estimate the distance to the target sensitive code (i.e., sink). Lastly, we further leverage the distance to guide test case mutation, which efficiently drives directed fuzzing toward candidate vulnerable code. We have implemented a prototype of Labrador and evaluated it on 14 different enterprise-level IoT devices. Results showed that Labrador significantly outperforms state-of-the-art (SOTA) solutions. It finds 44X more vulnerabilities than SNIPUZZ, BOOFUZZ and FIRM-AFL and 8.57X more vulnerabilities than SaTC. In total, it discovered 79 unknown vulnerabilities, of which 61 were assigned with CVEs. Hangtian Liu, Shuitao Gan, Chao Zhang 0008, Zicong Gao, Xiangzhi Wang, Guangming Gao |
SP | 2 |
| 2024 | LLMUZZ: LLM-based seed optimization for black-box device fuzzingabstractAs an increasing number of Internet of Things (IoT) devices are being deployed, the threat from vulnerabilities inside these devices is growing. Fuzzing is a primary method used for discovering vulnerabilities in IoT devices. The quality of the initial seeds and the seed mutation strategy are two crucial components of fuzzing that largely determine the effectiveness of the fuzzing process. However, owing to the diversity of IoT devices and the highly structured nature of inputs, designing universal seed generation and mutation strategies is extremely challenging. In this paper, we propose LLMUZZ, which is a large language model (LLM)-based black-box fuzzing approach for IoT devices. Specifically, we employ prompt engineering techniques in few-shot learning, using HTML form data from frontend files and an example HTTP request as inputs to LLMs to generate initial seeds. Then, we input the requests to be mutated into LLMs to identify the fields requiring mutation, thereby assisting in the seed mutation process. This approach ensures that the mutated seeds remain valid. Additionally, static analysis methods are utilized to discover hidden keywords within the firmware, thereby further expanding the initial seeds. In the experiments, we implement a prototype of LLMUZZ and evaluate it on 8 different IoT devices. A total of 16 previously unknown vulnerabilities are found, for which we have received 4 CVEs; the remaining vulnerabilities still under review, demonstrating that LLMUZZ has a strong capacity for vulnerability discovery. Guangming Gao, Shuitao Gan, Shengkai Zhu |
TrustCom | 2 |
| 2024 | Code is not Natural Language: Unlock the Power of Semantics-Oriented Graph Representation for Binary Code Similarity Detection
Haojie He, Xingwei Lin, Ziang Weng, Ruijie Zhao 0001, Shuitao Gan, Libo Chen 0001, Yuede Ji, Jiashui Wang, Zhi Xue |
USENIX Security Symposium | 5 |
| 2024 | Graphuzz: Data-driven Seed Scheduling for Coverage-guided Greybox FuzzingabstractSeed scheduling is a critical step of greybox fuzzing, which assigns different weights to seed test cases during seed selection, and significantly impacts the efficiency of fuzzing. Existing seed scheduling strategies rely on manually designed models to estimate the potentials of seeds and determine their weights, which fails to capture the rich information of a seed and its execution and thus the estimation of seeds’ potentials is not optimal. In this article, we introduce a new seed scheduling solution, Graphuzz, for coverage-guided greybox fuzzing, which utilizes deep learning models to estimate the potentials of seeds and works in a data-driven way. Specifically, we propose an extended control flow graph called e-CFG to represent the control-flow and data-flow features of a seed's execution, which is suitable for graph neural networks (GNN) to process and estimate seeds’ potential. We evaluate each seed's code coverage increment and use it as the label to train the GNN model. Further, we propose a self-attention mechanism to enhance the GNN model so that it can capture overlooked features. We have implemented a prototype of Graphuzz based on the baseline fuzzer AFLplusplus. The evaluation results show that our model can estimate the potential of seeds and has the robust capability to generalize to different targets. Furthermore, the evaluation using 12 benchmarks from FuzzBench shows that Graphuzz outperforms AFLplusplus and the state-of-the-art seed scheduling solution K-Scheduler and other coverage-guided fuzzers in terms of code coverage, and the evaluation using 8 benchmarks from Magma shows that Graphuzz outperforms the baseline fuzzer AFLplusplus and SOTA solutions in terms of bug detection. Shuitao Gan, Chao Zhang 0008, Zheming Li, Jiangan Ji, Baojian Chen |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Evocatio: Conjuring Bug Capabilities from a Single PoCabstractThe popularity of coverage-guided greybox fuzzers has led to a tsunami of security-critical bugs that developers must prioritize and fix. Knowing the capabilities a bug exposes (e.g., type of vulnerability, number of bytes read/written) enables prioritization of bug fixes. Unfortunately, understanding a bug's capabilities is a time consuming process, requiring (a) an understanding of the bug's root cause, (b) an understanding how an attacker may exploit the bug, and (c) the development of a patch mitigating these threats. This is a mostly-manual process that is qualitative and arbitrary, potentially leading to a misunderstanding of the bug's capabilities. Zhiyuan Jiang, Shuitao Gan, Adrian Herrera, Flavio Toffalini, Lucio Romerio, Chaojing Tang, Manuel Egele, Chao Zhang 0008, Mathias Payer |
CCS | 2 |
| 2022 | Path Sensitive Fuzzing for Native ApplicationsabstractCoverage-guided fuzzing is a widely used and effective solution to find software vulnerabilities. Tracking code coverage and utilizing it to guide fuzzing are crucial to coverage-guided fuzzers. However, tracking full and accurate path coverage is infeasible in practice due to the high instrumentation overhead. Popular fuzzers (e.g., AFL) often usecoarsecoverage information, e.g., edge hit counts stored in a compact bitmap, to achieve highly efficient greybox testing. Such inaccuracy and incompleteness in coverage introduce serious limitations to fuzzers. First, it causespath collisions, which prevent fuzzers from discovering potential paths that lead to new crashes. More importantly, it prevents fuzzers from making wise decisions on fuzzing strategies. In this article, we propose a coverage sensitive fuzzing solution CollAFL. It mitigates path collisions by providing more accurate coverage information, while still preserving low instrumentation overhead. It also utilizes the coverage information to apply three new fuzzing strategies, promoting the speed of discovering new paths and vulnerabilities. We implemented two variants of this solution, namely CollAFL (based on AFL) and CollAFL-bin (based on AFL-dyninst), to test applications with and without source code respectively, and evaluated them on 24 popular applications. The results showed that path collisions are common, i.e., up to 75 percent of edges could collide with others in some applications. But our solutions CollAFL and CollAFL-bin could reduce the edge collision ratio to nearly zero. Moreover, armed with the three fuzzing strategies, they outperform their counterparts (i.e., AFL and AFL-dyninst) in terms of both code coverage and vulnerability discovery. On average, CollAFL covered 20 percent more program paths, and found 320 percent more unique crashes and 260 percent more bugs than AFL in 200 hours. Moreover, CollAFL-bin covered 15 percent more paths, and found 200 percent more unique crashes and 150 percent more vulnerabilities than AFL-dyninst, showing that the proposed solution also works for binary application fuzzing. In total, CollAFL found 157 new security bugs with 95 new CVEs assigned. Shuitao Gan, Chao Zhang 0008, Xiaojun Qin, Xuwen Tu, Zhongyu Pei, Zuoning Chen |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | firm VulSeeker: BERT and Siamese based Vulnerability for Embedded Device Firmware ImagesabstractIn this paper, we propose firmVulSeeker-a vulnerability search tool for embedded firmware images, based on BERT and Siamese network. It first builds a BERT MLM task to observe and learn the semantics of different instructions in their context in a very large unlabeled binary corpus. Then, a finetune mode based on Siamese network is constructed to guide training and matching semantically similar functions using the knowledge learned from the first stage. Finally, it will use a function embedding generated from the fine-tuned model to search in the targeted corpus and find the most similar function which will be confirmed whether it's a real vulnerability manually. We evaluate the accuracy, robustness, scalability and vulnerability search capability of firmVulSeeker. Results show that it can greatly improve the accuracy of matching semantically similar functions, and can successfully find more real vulnerabilities in real-world firmware than other tools. Yingchao Yu, Shuitao Gan, Xiaojun Qin |
ISCC | 2 |
| 2020 | GREYONE: Data Flow Sensitive Fuzzing
Shuitao Gan, Chao Zhang 0008, Peng Chen 0034, Bodong Zhao, Xiaojun Qin, Zuoning Chen |
USENIX Security Symposium | 1 |
| 2018 | CollAFL: Path Sensitive FuzzingabstractCoverage-guided fuzzing is a widely used and effective solution to find software vulnerabilities. Tracking code coverage and utilizing it to guide fuzzing are crucial to coverage-guided fuzzers. However, tracking full and accurate path coverage is infeasible in practice due to the high instrumentation overhead. Popular fuzzers (e.g., AFL) often use coarse coverage information, e.g., edge hit counts stored in a compact bitmap, to achieve highly efficient greybox testing. Such inaccuracy and incompleteness in coverage introduce serious limitations to fuzzers. First, it causes path collisions, which prevent fuzzers from discovering potential paths that lead to new crashes. More importantly, it prevents fuzzers from making wise decisions on fuzzing strategies. In this paper, we propose a coverage sensitive fuzzing solution CollAFL. It mitigates path collisions by providing more accurate coverage information, while still preserving low instrumentation overhead. It also utilizes the coverage information to apply three new fuzzing strategies, promoting the speed of discovering new paths and vulnerabilities. We implemented a prototype of CollAFL based on the popular fuzzer AFL and evaluated it on 24 popular applications. The results showed that path collisions are common, i.e., up to 75% of edges could collide with others in some applications, and CollAFL could reduce the edge collision ratio to nearly zero. Moreover, armed with the three fuzzing strategies, CollAFL outperforms AFL in terms of both code coverage and vulnerability discovery. On average, CollAFL covered 20% more program paths, found 320% more unique crashes and 260% more bugs than AFL in 200 hours. In total, CollAFL found 157 new security bugs with 95 new CVEs assigned. Shuitao Gan, Chao Zhang 0008, Xiaojun Qin, Xuwen Tu, Zhongyu Pei, Zuoning Chen |
IEEE Symposium on Security and Privacy | 1 |
| 2014 | Symbolic Execution of Network Software Based on Unit TestingabstractComplex interactions and the distributed nature of network software make automated testing and debugging before deployment a necessity. Symbolic execution is a systematic program analysis technique that has become increasingly popular in network software testing, due to algorithmic advances and availability of computational power and constraint solving technology. However, A main challenge is to detect determining symbolic values for program variables related to library, loops and cryptograph algorithms which are widely used in network software. In this paper, we propose a unit symbolic analysis, a hybrid technique that enables fully automatic symbolic analysis even for the traditionally challenging code. The novelties of this work are threefold: 1) we flexibly employs static symbolic execution to amplify the effect of dynamic symbolic execution on demand, 2) dynamic executions and regression analysis are performed on the unit tests constructed from the code segments to infer program semantics needed by static analysis, and 3) symbolic analysis is utilized to tackle loop structure and cryptograph algorithm module. We developed the Net Sym framework, consisting of a static component that performs symbolic analysis and partitions a program, a dynamic analysis that synthesizes unit tests and automatically infers symbolic values for program variables, and a protocol that enables static and dynamic analyses to be run interactively and concurrently. Our experimental results show that by handling cryptograph algorithms, loops and library calls that a traditional symbolic analysis cannot process, unit symbolic analysis detects more vulnerabilities in less time. The technique is scalable for real-world programs such as GHttpd, SQL Server and GDI. Shuitao Gan, Xiaojun Qin, Wenbao Han |
NAS | 3 |