Zhanwei Song

dblp:181/4577 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0000-3284-7351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NetID-GPT: Adapting large language models for large-scale internet-connected device identification
Zhi Li 0018, Shunchao Xu, Fengshi Zhang, Zhanwei Song, Dongliang Fang, Yongle Chen, Limin Sun 0001
Comput. Networks5
2025 PNetGPT: Proprietary Protocol Network Traffic Generation with Pre-trained Transformer
abstract
Generative pre-trained transformers are exceedingly effective as generative models and classifiers, widely used in natural language processing and computer vision. This work contributes to the exploration of generative pre-trained transformer-based models in the proprietary protocol network traffic. However, building a pre-trained model for proprietary protocol network traffic is non-trivial due to the heterogeneous unknown formats and the extreme scarcity of proprietary protocol network traffic datasets. In this paper, we present PNetGPT, a pre-trained transformer-based model for generating proprietary protocol network traffic. We have constructed the inaugural dataset of 2 real-world proprietary protocols. After training on this dataset, PNetGPT possesses the capacity to generate high-quality proprietary protocol network traffic to support various applications of proprietary protocols, including reverse analysis, protocol fuzzy testing, intrusion detection, etc. We evaluated PNetGPT with two real proprietary protocols and demonstrated state-of-the-art (SOTA) performance in handling heterogeneous unknown formats. The code and datasets are available at: https://github.com/Snail1502/PNetGPT
Zedong Li, Dongliang Fang, Xin Chen 0123, Zhanwei Song, Zhi Li 0018, Shichao Lv, Limin Sun 0001
ICASSP5
2025 Exp-Arch: A Novel LLM-Powered Approach for Facilitating Exploit Primitive Assessment in the Linux Kernel
abstract
Transforming Linux kernel exploit primitives into full Privilege Escalation (PE) exploits is a critical, expertiseintensive, and time-consuming challenge, especially with constantly evolving kernel mitigations. While previous research has advanced automated kernel exploit development, these efforts often focused on specialized scenarios rather than providing a generalized, end-to-end framework for diverse primitives. This limitation restricts the exploration of a primitive's true exploit potential. This paper introduces Exp-Arch, a novel approach leveraging Large Language Models (LLMs) for the automated, end-to-end generation of PE exploits from kernel primitives. This process includes comprehensive initial assessment and subsequent exploitation. Exp-Arch's LLM-powered workflow systematically performs an in-depth semantic analysis of the input primitive, devises an intelligent strategic plan for the exploitation route, and then automates the synthesis and iterative closed-loop validation of the final PE exploit code. Exp-Arch offers accurate assessment of a primitive's exploitability and significantly accelerates the exploit development lifecycle. We evaluated ExpArch using various primitives from public Linux kernel 1-day vulnerabilities with commercial LLMs. The results show that Exp-Arch effectively converted 73 % (11 out of 15) of test cases into working kernel exploits, demonstrating its effectiveness in primitive evaluation.
Zuxin Chen, Zhi Li 0018, Zhanwei Song, Zhiqiang Shi, Limin Sun 0001
ICPADS3
2025 Exploiting Binary Semantics: Enhancing Function Name Inference in Stripped Binaries via LLMs
abstract
Function name inference in stripped binaries is a crucial task that supports various security applications, including vulnerability detection and malware analysis. Existing methods suffer from limited model capacity and insufficient exploitation of function semantics, which constrains their ability to comprehend binary code and leads to poor generalization on unseen binaries. To address these problems, we propose BinLLM, a novel framework that leverages large language models (LLMs) to exploit the semantic potential of binary code, thereby enhancing function name inference. Specially, BinLLM integrates three key innovations: (1) source code semantics-guided function name refinement, which mitigates the negative effects of low-quality semantic identifiers during training; (2) A context-aware data collection algorithm that seeks richer semantic dependencies to improve model training and inference performance; (3) parameter-efficient fine-tuning on a domain-specific dataset enriched with semantic knowledge to enhance the model's understanding of binary semantics. These components collectively enhance the model's performance in function name inference on unseen binaries. We evaluate BinLLM on a large-scale dataset comprising$2,864,719$functions across four architectures (x86-64, x86-32, ARM, MIPS) and four optimization levels ($\mathrm{O} 0-\mathrm{O} 3$). Experimental results show that BinLLM achieves substantial improvements over state-of-the-art (SOTA) methods, with relative gains of$320.1 \%, 274.8 \%$, and 297.6 % in precision, recall, and F1-score. Ablation studies further validate the effectiveness of each component in enhancing overall performance.
Kailong Wang 0007, Dongliang Fang, Zhongwei Gu, Zhanwei Song, Yongle Chen, Zhiqiang Shi, Limin Sun 0001
IPCCC5
2025 ICSPFuzzer: An Efficient Fuzzing Technique for ICS Protocols
Zhanwei Song, Dongliang Fang, Shunchao Xu, Yaowen Zheng, Hong Li 0004, Shichao Lv, Zhiqiang Shi, Limin Sun 0001
WASA (2)1
2025 Backsolver: Adapting Preceding Execution Paths to Solve Constraints for Concolic Execution
abstract
Concolic execution follows the execution paths of concrete inputs, capable of generating new inputs for unexplored code by solving negated path constraints. However, implicit flows can hinder concolic execution, reducing the code coverage. Implicit flows occur when inputs influence control flow, and the control flow variation affects the values of some variables. During concolic execution, the preceding path selections limit the potential values of these variables. This limitation may result in unsolvable constraints, subsequently restricting the generation of new inputs for unexplored paths. Our insight is that following the same preceding paths is unnecessary, and we can adapt preceding paths to make the latest constraints solvable. We divide states into general states and implicit-flow-solving states (IFSSs). We utilize the general states to perform concolic execution. When solving constraints influenced by implicit flows, we switch to the IFSSs. We use the IFSSs to explore the relevant code region and adapt paths. To mitigate path explosion and construct the relation between inputs and the variables, we merge the IFSSs. State merging does not burden the general states, and we limit the code regions for the IFSSs to minimize the introduced overhead. Finally, we replace the variable symbols in the target constraints with new expressions and attempt to solve the new constraints. We implement our approach in Backsolver and build a test suite to evaluate it. Backsolver successfully identifies all the implicit flows in the test suite and resolves most of them. When evaluated on six real-world binaries, Backsolver resolves the highest number of branches related to implicit flows in total. Besides, Backsolver has the highest code coverage in PlutoSVG and finds a 0-day vulnerability. We reported the vulnerability and obtained a CVE ID.
Yicheng Zeng, Zhanwei Song, Guo Lv, Hongsong Zhu, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.2
2024 DeLink: Source File Information Recovery in Binaries
abstract
Program comprehension can help analysts understand the primary behavior of a binary and enhance the efficiency of reverse engineering analysis. The existing works focus on instruction translation and function name prediction. However, they are limited in understanding the entire program. The recovered source file information can offer insights into the primary behavior of a binary, serving as high-level program summaries. Nevertheless, the files recovered by the function clustering-based approach contain binary functions with discontinuous distributions, resulting in low accuracy. Additionally, there is no existing research related to predicting the names of these recovered files. To this end, we propose a framework for source file information recovery in binaries, DeLink. This framework first leverages a file structure recovery approach based on boundary location to recognize files within a binary. Then, it utilizes an encoder-decoder model to predict the names of these files. The experimental results show that our file structure recovery approach achieves an average improvement of 14% across six evaluation metrics and requires only an average time of 16.74 seconds, outperforming the state-of-the-art work in both recovery quality and efficiency. Additionally, our file name prediction model achieves 70.09% precision and 63.91% recall. Moreover, we demonstrate the effective application of DeLink in malware homology analysis.
Zhe Lang, Zhengzi Xu, Shichao Lv, Zhanwei Song, Zhiqiang Shi, Limin Sun 0001
ISSTA5
2024 A Self-Supervised Targeted Process Anomaly Detection Method Based on the Minimum Set of Observed Events
abstract
In scenarios involving a single targeted application service, it is essential to monitor the security of business-oriented processes. However, there is currently a lack of lightweight, real-time online anomaly detection methods for targeted processes that do not require labeled data. This paper presents a self-supervised anomaly detection model for targeted processes, based on a minimal observation event set and utilizing eBPF and deep learning techniques. The model first selects a minimal set of observed events for the targeted process, which includes critical system calls, process scheduling, resource usage, and I/O operations. Unlike traditional system call sequence features, this model focuses on the rate of change in the frequency of feature selection calls. The detection model employs the VAE-LSTM algorithm, where the VAE module constructs robust short windows and the LSTM module estimates long-term correlations within the sequence. Through self-supervised learning, the model learns the normal behavior of targeted processes, extracts robust behavioral features, and reduces feature dimensions. By performing online detection of these learned features against runtime processes, the model achieves second-level anomaly detection for targeted processes. Finally, by simulating and constructing eight different types of process attack scenarios, the experimental results demonstrate a detection accuracy exceeding 92%, with a performance overhead of less than 10%.
Haojun Xia, Limin Sun 0001, Zhanwei Song, Bibo Tu
TrustCom5
2023 FlowEmbed: Binary function embedding model based on relational control flow graph and byte sequence
abstract
Binary function embedding models are applicable to various downstream tasks within IoT device software systems and have demonstrated advantages in numerous binary analysis tasks, such as vulnerability (homologous) function search and compilation optimization option identification. However, current binary function embedding methods either learn embedding based on code sequence, which lack the program semantics of functions (e.g., control flow, etc.) or based on program structure graphs, which omit global sequential information. As a result, these methods fall short in enabling models to learn the complete semantic of function. In this paper, we introduce FlowEmbed, a novel approach that synergistically integrates control flow and global semantic learning to facilitate exhaustive code comprehension. Initially, FlowEmbed harnesses a distinct relational control flow graph combined with the power of BERT and RGCN models to aptly capture the nuances of control flow semantics. Moreover, by deploying the DPCNN model on a byte sequence constructed from function machine code, FlowEmbed adeptly discerns the inherent global sequential semantics of binary functions. Through rigorous evaluations spanning three IoT-related tasks, FlowEmbed’s efficacy becomes evident, showcasing notable improvements: a 20.6% improvement in compilation optimization option identification, a 1.8% improvement in binary function similarity analysis, and an 11.9% improvement in homologous function search. Collectively, these results underscore FlowEmbed’s superior capability, positioning it as a invaluable asset in a binary analysis application.
Yongpan Wang, Chaopeng Dong, Siyuan Li 0014, Renjie Su, Zhanwei Song, Hong Li 0004
ICPADS6
2022 SIFOL: Solving Implicit Flows in Loops for Concolic Execution
abstract
Concolic execution is widely used for binary analysis and is commonly embedded in hybrid fuzzing to find bugs. However, implicit flows in loops can hinder concolic execution and lead to the reduction of code coverage. The implicit flow variables cannot be symbolized and will block the constraint solver from generating new inputs. We propose a new approach to mitigate the problem. We obtain the implicit flow variables by taint analysis in advance and symbolize them during the concolic execution. Then, when the symbols of the variables are in the path constraints and need to be solved, we backtrack to the corresponding loops and perform static symbolic executions in the loops. During the static symbolic executions, we relate the variables with the input symbols by state merging and solve the constraints to generate inputs for new execution paths. We present SIFOL, a hybrid fuzzer based on Driller, and evaluate it on CB-multios. Results show that SIFOL has 5.4% higher code coverage than Driller and finds 5.9% more crashes. Furthermore, after manually adding implicit flows and checks to the target programs, SIFOL only drops 2.6% on coverage and 5.6% on the crash number, while Driller is severely affected (drops 46.1% on coverage and 47.1% on the crash number).
Yicheng Zeng, Jiaqian Peng, Zhanwei Song, Hongsong Zhu, Limin Sun 0001
IPCCC4
2022 Fuzzing proprietary protocols of programmable controllers to find vulnerabilities that affect physical control
Puzhuo Liu, Yaowen Zheng, Zhanwei Song, Dongliang Fang, Shichao Lv, Limin Sun 0001
J. Syst. Archit.3
2021 ICS3Fuzzer: A Framework for Discovering Protocol Implementation Bugs in ICS Supervisory Software by Fuzzing
abstract
The supervisory software is widely used in industrial control systems (ICSs) to manage field devices such as PLC controllers. Once compromised, it could be misused to control or manipulate these physical devices maliciously, endangering manufacturing process or even human lives. Therefore, extensive security testing of supervisory software is crucial for the safe operation of ICS. However, fuzzing ICS supervisory software is challenging due to the prevalent use of proprietary protocols. Without the knowledge of the program states and packet formats, it is difficult to enter the deep states for effective fuzzing.
Dongliang Fang, Zhanwei Song, Le Guan, Puzhuo Liu, Anni Peng, Yaowen Zheng, Peng Liu 0005, Hongsong Zhu, Limin Sun 0001
ACSAC2
2019 An Efficient Greybox Fuzzing Scheme for Linux-based IoT Programs Through Binary Static Analysis
abstract
With the rapid growth of Linux-based IoT devices such as network cameras and routers, the security becomes a concern and many attacks utilize vulnerabilities to compromise the devices. It is crucial for researchers to find vulnerabilities in IoT systems before attackers. Fuzzing is an effective vulnerability discovery technique for traditional desktop programs, but could not be directly applied to Linux-based IoT programs due to the special execution environment requirement. In our paper, we propose an efficient greybox fuzzing scheme for Linux-based IoT programs which consist of two phases: binary static analysis and IoT program greybox fuzzing. The binary static analysis is to help generate useful inputs for efficient fuzzing. The IoT program greybox fuzzing is to reinforce the IoT firmware kernel greybox fuzzer to support IoT programs. We implement a prototype system and the evaluation results indicate that our system could automatically find vulnerabilities in real-world Linux-based IoT programs efficiently.
Yaowen Zheng, Zhanwei Song, Yuyan Sun, Hongsong Zhu, Limin Sun 0001
IPCCC2
2019 Abnormal detection method of industrial control system based on behavior model
Zhanwei Song, Zenghui Liu
Comput. Secur.1