Peiyu Liu 0003

dblp:85/670-3 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0001-7793-7633ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity Detection
abstract
Binary Code Similarity Detection (BCSD) plays a vital role in various security applications, including vulnerability identification, malware analysis, and code plagiarism detection.With the growing adoption of deep neural networks (DNNs), substantial progress has been made in recognizing and classifying similar code segments.However, DNN-based BCSD methods often exhibit low accuracy and robustness because they struggle to capture fine-grained and high-level program semantics.In contrast, such semantics are typically captured through natural language interpretations of source code by large language models (LLMs).Yet, LLM-based BCSD methods are constrained by their large model sizes and high inference latency.To alleviate these limitations, this paper proposes BinSKD.The key idea is to leverage an LLM-based BCSD method as the teacher model and transfer its knowledge of high-level program semantics to various DNNbased student models.Specifically, to avoid propagating errors from the teacher to the student, we introduce selective distillation, selecting targets with accurate semantics according to their detection retrieval.In addition, to mitigate the noise introduced by a number of negative samples during distillation, we further propose discrepancy-weighted sampling to focus on the samples where the student's prediction notably deviates from the teacher's.Our experiments show that BinSKD yields Recall@1 improvements of 14.5%-91.2%for DNN-based BCSD methods and enables HermesSim to match the teacher's performance with ordersof-magnitude efficiency.
Shize Zhou, Peiyu Liu 0003, Lirong Fu, Wenhai Wang
ACL (1)2
2026 Enhancing ICS equipment security through fuzzing: Automated protocol inference and response-driven exploration
Liyang Hou, Peiyu Liu 0003, Jiaao Sheng, Yangjun Chen, Chunyu Miao, Wenhai Wang
Comput. Secur.2
2026 Agents4PLC: Automating Closed-Loop PLC Code Generation and Verification in Industrial Control Systems Using LLM-Based Agents
abstract
In industrial control systems, the generation and verification of Programmable Logic Controller (PLC) code are crucial for ensuring operational efficiency and safety. While Large Language Models (LLMs) have made strides in automated code generation, they fall short in providing correctness guarantees and specialized support for PLC programming (which has its own programming language and clear logical structures). To address these challenges, this paper introduces Agents4PLC, a novel framework that not only automates PLC code generation but also introduces code-level verification and repair built upon an LLM-based multi-agent system, which together is capable of directly producing operational PLC code without any human interaction. To comprehensively evaluate our framework, we first establish a new benchmark specially designed for the critical area ofverifiable PLC code generation, which includes hundreds of natural language requirements, human-written and verified formal specifications, and finally reference PLC code. Then, we carefully designed a multi-agent workflow combining a set of expert agents responsible for different code generation tasks including planning, coding, validation and debugging towards generating correct PLC code. For each agent, we also incorporate optimization strategies such as Retrieval-Augmented Generation (RAG), advanced prompt engineering techniques, and Chain-of-Thought strategies which are shown to be effective to enhance the ability of these expert ‘agents’. Evaluation against the benchmark demonstrates that Agents4PLC significantly outperforms existing methods, achieving superior results across a series of increasingly rigorous evaluation metrics. This research highlights the potential of LLM agent-based code generation in real-world industrial control systems and the importance of code-level verification in generating correct code with formal guarantees.
Ruinan Zeng, Dongxia Wang 0002, Gengyun Peng, Peiyu Liu 0003, Wenhai Wang, Jingyi Wang 0004
IEEE Trans. Software Eng.7
2025 BinEGA: Enhancing DNN-based Binary Code Similarity Detection through Efficient Graph Alignment
abstract
Binary Code Similarity Detection (BCSD) is essential in various binary code security applications, enabling tasks such as vulnerability identification, malware analysis, and detection of code plagiarism. With the growing adoption of deep neural networks (DNNs) in BCSD, there has been significant progress in the identification and classification of similar code segments. However, DNN-based BCSD approaches often suffer from high false positive rates, because DNNs inevitably map different binary functions with complex structures and semantics to similar low-dimensional embeddings. To alleviate this issue, this paper introduces BinEGA, a novel graph alignment-based approach to enhance the accuracy of DNN-based BCSD approaches. The main idea of BinEGA is to employ a general and low-cost equivalence check through lightweight graph alignment, allowing for the identification and elimination of semantically deviating functions among the top-k candidates retrieved by DNN-based BCSD approaches. During the graph alignment process, we first obtain the node embeddings according to structure and attribute feature. Then we employs pairwise comparison of these node embeddings to filter the false positives because binary code compiled from the same source code always shares similar basic blocks. Our experimental results demonstrate that BinEGA effectively enhances the performance of various edge-cutting DNN-based BCSD approaches across diverse scenarios. For instance, BinEGA significantly enhances RECALL@10 in the cross-optimization scenario for state-of-the-art (SOTA) approaches, with an average improvement of 29.2% for BinaryAI and 33.5% for jTrans. Moreover, BinEGA achieves 88.9 % reduction in execution time compared to other enhancement techniques. In summary, this work provides a robust, generalizable, and efficient solution to improve the reliability of BCSD tools in real-world applications.
Shize Zhou, Lirong Fu, Peiyu Liu 0003, Wenhai Wang
SANER3
2025 Firmrca: Towards Post-Fuzzing Analysis on ARM Embedded Firmware with Efficient Event-Based Fault Localization
abstract
While fuzzing has demonstrated its effectiveness in exposing vulnerabilities within embedded firmware, the discovery of crashing test cases is only the first step in improving the security of these critical systems. The subsequent fault localization process, which aims to precisely identify the root causes of observed crashes, is a crucial yet time-consuming post-fuzzing work. Unfortunately, the automated root cause analysis on embedded firmware crashes remains an underexplored area, which is challenging from several perspectives: (1) the fuzzing campaign towards the embedded firmware lacks adequate debugging mechanisms, making it hard to automatically extract essential runtime information for analysis; (2) the inherent raw binary nature of embedded firmware often leads to over-tainted and noisy suspicious instructions, which provides limited guidance for analysts in manually investigating the root cause and remediating the underlying vulnerability. To address these challenges, we design and implement FirmRCA, a practical fault localization framework tailored specifically for embedded firmware. FirmRCA introduces an event-based footprint collection approach that leverages concrete memory accesses in the crash reproducing process to aid and significantly expedite reverse execution. Next, to solve the complicated memory alias problem, FirmRCA proposes a history-driven method by tracking data propagation through the execution trace, enabling precise identification of deep crash origins. Finally, FirmRCA proposes a novel strategy to highlight key instructions related to the root cause, providing practical guidance in the final investigation. To demonstrate the efficacy of FirmRCA, we evaluate it with both synthetic and real-world targets, including 41 crashing test cases across 17 firmware images. The results show that FIRMRCA can effectively (92.7% success rate) identify the root cause of crashing test cases within the top 10 instructions. Compared to state-of-the-art works, FIRMRCA demonstrates its superiority in 27.8% improvement in full execution trace analysis capability, polynomial-level acceleration in overall efficiency and 73.2% higher success rate within the top 10 instructions in effectiveness.
Boyu Chang, Peiyu Liu 0003, Yuan Tian 0001, Raheem A. Beyah, Shouling Ji
SP4
2025 Waltzz: WebAssembly Runtime Fuzzing with Stack-Invariant Transformation
Jiacheng Xu 0006, Peiyu Liu 0003, Qinge Xie, Yuan Tian 0001, Jianhai Chen, Shouling Ji
USENIX Security Symposium4
2025 SSFuzz: State-Guided Fuzzing With Shared Feedback for Black-Box IoT Devices
abstract
The rapid growth of Internet of Things (IoT) devices has enhanced convenience but introduced significant security risks. Due to limited visibility into device internals, black-box fuzzing has become the primary method for IoT vulnerability detection. However, it is often difficult to recognize the triggered states, which limits the ability to explore different regions of the state space and, as a result, hinders the discovery of vulnerabilities. Additionally, it lacks feedback, preventing the fuzzer from refining its test inputs based on the results and reducing its effectiveness in discovering vulnerabilities. To address these challenges, we propose SSFuzz, an automated black-box fuzzing framework leveraging large language models (LLMs) to extract state nodes from interaction messages, enabling a state-guided approach. Additionally, we design a cross-device feedback-sharing mechanism based on source code similarities, aiming to make more effective use of the limited feedback available. Evaluated against five leading tools on 18 IoT devices, SSFuzz identified 38 previously undisclosed vulnerabilities, significantly outperforming existing methods. SSFuzz discovered 38 previously unknown vulnerabilities, significantly outperforming Snipuzz (five vulnerabilities) and IoTHunter (one vulnerability).
Liyang Hou, Peiyu Liu 0003, Jianchun Ding, Jiaao Sheng, Huan Le, Yangjun Chen, Wenhai Wang
IEEE Internet Things J.2
2025 LuaTaint: A Static Analysis System for Web Configuration Interface Vulnerability of Internet of Things Devices
abstract
The diversity of Web configuration interfaces for Internet of Things (IoT) devices has exacerbated issues, such as inadequate permission controls and insecure interfaces, resulting in various vulnerabilities. Owing to the varying interface configurations across various devices, the existing methods are inadequate for identifying these vulnerabilities precisely and comprehensively. This study addresses these issues by introducing an automated vulnerability detection system, called LuaTaint. It is designed for the commonly used Web configuration interface of IoT devices. LuaTaint combines static taint analysis with a large language model (LLM) to achieve widespread and high-precision detection. The extensive traversal of the static analysis ensures the comprehensiveness of the detection. The system also incorporates rules related to page handler control logic within the taint detection process to enhance its precision and extensibility. Moreover, we leverage the prodigious abilities of LLM for code analysis tasks. By utilizing LLM in the process of pruning false alarms, the precision of LuaTaint is enhanced while significantly reducing its dependence on manual analysis. We develop a prototype of LuaTaint and evaluate it using 2447 IoT firmware samples from 11 renowned vendors. LuaTaint has discovered 111 vulnerabilities. Moreover, LuaTaint exhibits a vulnerability detection precision rate of up to 89.29%.
Jiahui Xiang, Lirong Fu, Peiyu Liu 0003, Huan Le, Liming Zhu 0003, Wenhai Wang
IEEE Internet Things J.4
2024 CINDA: Don't Ignore Instructions When Cloning Memory Access Behavior
abstract
Existing workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error.
Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen
CCGrid5
2024 Exploring ChatGPT's Capabilities on Vulnerability Management
Peiyu Liu 0003, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang
USENIX Security Symposium1
2024 Critical Code Guided Directed Greybox Fuzzing for Commits
Xuhong Zhang 0002, Peiyu Liu 0003, Shouling Ji, Jiacheng Xu 0006, Wenhai Wang
USENIX Security Symposium3
2023 CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code
abstract
Automatically generating function summaries for binaries is an extremely valuable but challenging task, since it involves translating the execution behavior and semantics of the low-level language (assembly code) into human-readable natural language.However, most current works on understanding assembly code are oriented towards generating function names, which involve numerous abbreviations that make them still confusing.To bridge this gap, we focus on generating complete summaries for binary functions, especially for stripped binary (no symbol table and debug information in reality).To fully exploit the semantics of assembly code, we present a control flow graph and pseudo code guided binary code summarization framework called CP-BCS.CP-BCS utilizes a bidirectional instruction-level control flow graph and pseudo code that incorporates expert knowledge to learn the comprehensive binary function execution behavior and logic semantics.We evaluate CP-BCS on 3 different binary optimization levels (O1, O2, and O3) for 3 different computer architectures (X86, X64, and ARM).The evaluation results demonstrate CP-BCS is superior and significantly improves the efficiency of reverse engineering. * Corresponding author.with limited high-level information, making it difficult to read and understand, as shown in Figure 1.Even an experienced reverse engineer needs to spend a significant amount of time determining the functionality of an assembly code snippet.
Lingfei Wu 0001, Tengfei Ma 0001, Xuhong Zhang 0002, Yangkai Du, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
EMNLP6
2023 Static Semantics Reconstruction for Enhancing JavaScript-WebAssembly Multilingual Malware Detection
Xuhong Zhang 0002, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
ESORICS (2)4
2023 How IoT Re-using Threatens Your Sensitive Data: Exploring the User-Data Disposal in Used IoT Devices
abstract
With the rapid technology evolution of the Internet of Things (IoT) and increasing user needs, IoT device re-using becomes more and more common nowadays. For instance, more than 300,000 used IoT devices are selling on Craigslist. During IoT re-using, sensitive data such as credentials and biometrics residing in these devices may face the risk of leakage if a user fails properly dispose of the data. Thus, a critical security concern is raised: do (or can) users properly dispose of the sensitive data in used IoT? To the best of our knowledge, it is still an unexplored problem that desires a systematic study.In this paper, we perform the first in-depth investigation on the user-data disposal of used IoT devices. Our investigation integrates multiple research methods to explore the status quo and the root causes of the user-data leakages with used IoT devices. First, we conduct a user study to investigate the user awareness and understanding of data disposal. Then, we conduct a large-scale analysis on 4,749 IoT firmware images to investigate user-data collection. Finally, we conduct a comprehensive empirical evaluation on 33 IoT devices to investigate the effectiveness of existing data disposal methods.Through the systematical investigation, we discover that IoT devices collect more sensitive data than users expect. Specifically, we detect 121,984 sensitive data collections in the tested firmware. Moreover, users usually do not or even cannot properly dispose of the sensitive data. Worse, due to the inherent characteristics of storage chips, 13.2% of the investigated firmware perform "shallow" deletion, which may allow adversaries to obtain sensitive data after data disposal. Given the large-scale IoT re-using, such leakage would cause a broad impact. We have reported our findings to world-leading companies. We hope our findings raise awareness of the failures of user-data disposal with IoT devices and promote the protection of users’ sensitive data in IoT devices.
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Jingchang Qin, Wenhai Wang, Wenzhi Chen
SP1
2023 A traffic anomaly detection approach based on unsupervised learning for industrial cyber-physical system
Tao Yang 0043, Zhenze Jiang, Peiyu Liu 0003, Qiang Yang 0004, Wenhai Wang
Knowl. Based Syst.3
2022 Focus : Function clone identification on cross-platform
abstract
Automatic identification of function clones on cross-platform aims at determining whether two functions are identical or not without access to the source code, which is a fundamental challenge in vulnerability search, code plagiarism detection, and malware classification. With the rapid development of deep neural network in program analysis, the state-of-the-art neural network-based function clone identification methods propose to represent functions as embeddings by graph neural network (GNN). However, such a novel representation of functions brings in two challenges. (1) The feature engineering that accurately maps the raw data of binary code to machine learning features is complicated. (2) A highly accurate embedding of functions requires a customized GNN to focus on the most critical features to identify binary code. To the best of our knowledge, currently, a comprehensive work that can overcome the above challenges is still missing. In this paper, we propose a novel prototype named as Focus, which is designed to accurately and efficiently identify similar functions. Specifically, inspired by natural language processing techniques which effectively learns text semantic across natural languages, Focus can learn representative semantic features of functions by a customized learning model. To address the second challenge, a multi-head attention mechanism can be employed to capture the critical features of a function. Through extensive experiments, we demonstrate that Focus achieves high accuracy of function clone identification on a broad range of eight architectures. In particular, the identification performance (AUC value) of Focus is 97% and 99% for cross-platform and single-platform, respectively. Furthermore, the evaluation in real world applications shows that our Focus identifies 24 vulnerable functions among the top-30 candidates, which is one time higher than the baseline approaches.
Lirong Fu, Shouling Ji, Changchang Liu, Peiyu Liu 0003, Fuzheng Duan, Zonghui Wang, Whenzhi Chen, Ting Wang 0006
Int. J. Intell. Syst.4
2021 CPscan: Detecting Bugs Caused by Code Pruning in IoT Kernels
abstract
To reduce the development costs, IoT vendors tend to construct IoT kernels by customizing the Linux kernel. Code pruning is common in this customization process. However, due to the intrinsic complexity of the Linux kernel and the lack of long-term effective maintenance, IoT vendors may mistakenly delete necessary security operations in the pruning process, which leads to various bugs such as memory leakage and NULL pointer dereference. Yet detecting bugs caused by code pruning in IoT kernels is difficult. Specifically, (1) a significant structural change makes precisely locating the deleted security operations (DSO ) difficult, and (2) inferring the security impact of a DSO is not trivial since it requires complex semantic understanding, including the developing logic and the context of the corresponding IoT kernel.
Lirong Fu, Shouling Ji, Kangjie Lu, Peiyu Liu 0003, Xuhong Zhang 0002, Yuxuan Duan, Wenzhi Chen
CCS4
2021 IFIZZ: Deep-State and Efficient Fault-Scenario Generation to Test IoT Firmware
abstract
IoT devices are abnormally prone to diverse errors due to harsh environments and limited computational capabilities. As a result, correct error handling is critical in IoT. Implementing correct error handling is non-trivial, thus requiring extensive testing such as fuzzing. However, existing fuzzing cannot effectively test IoT error-handling code. First, errors typically represent corner cases, thus are hard to trigger. Second, testing error-handling code would frequently crash the execution, which prevents fuzzing from testing following deep error paths.In this paper, we propose IFIZZ, a new bug detection system specifically designed for testing error-handling code in Linux-based IoT firmware. IFIZZ first employs an automated binary-based approach to identify realistic runtime errors by analyzing errors and error conditions in closed-source IoT firmware. Then, IFIZZ employs state-aware and bounded error generation to reach deep error paths effectively. We implement and evaluate IFIZZ on 10 popular IoT firmware. The results show that IFIZZ can find many bugs hidden in deep error paths. Specifically, IFIZZ finds 109 critical bugs, 63 of which are even in widely used IoT libraries. IFIZZ also features high code coverage and efficiency, and covers 67.3% more error paths than normal execution. Meanwhile, the depth of error handling covered by IFIZZ is 7.3 times deeper than that covered by the state-of-the-art method. Furthermore, IFIZZ has been practically adopted and deployed in a worldwide leading IoT company. We will open-source IFIZZ to facilitate further research in this area.
Peiyu Liu 0003, Shouling Ji, Xuhong Zhang 0002, Qinming Dai, Kangjie Lu, Lirong Fu, Wenzhi Chen, Peng Cheng 0001, Wenhai Wang, Raheem A. Beyah
ASE1
2020 Understanding the Security Risks of Docker Hub
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wei-Han Lee, Wenzhi Chen, Raheem A. Beyah
ESORICS (1)1
2019 Prudent Practices for Designing Virtual Desktop Experiments
abstract
Virtual desktop technology aims at accessing a remote desktop by endpoint hardware.Great attention has been increasingly paid to virtual desktop since it can increase the utilization of computing resources and provide more flexible accesses.However, researchers have not yet come up with a comprehensive set of rigorous standards of experimental design and implementation in this field.Therefore, it is difficult to conduct prudent experiments, which is correct, real, and transparent.In this paper, we assess the experimental evaluations of recently published papers on desktop virtualization.We observe that most works can be further improved, due to the unsuitable experimental environment and the lack of descriptions of experimental settings.In this paper, in order to help researchers, reviewers, and readers, we propose several guidelines for designing correct, real, and transparent desktop virtualization experiment.
Peiyu Liu 0003, Wenzhi Chen, Zonghui Wang, Lirong Fu
SEKE1