VLDB 2026 Research / reviewers in the wild / expert
Puzhuo Liu
dblp:294/5115
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-8995-5924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-author · 7 since 2021Security and privacy · 6 · 6 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIT: Empowering Binary Analysis through the LLVM ToolchainabstractBinary analysis plays a critical role in software comprehension and security analysis, especially when source code is unavailable or difficult to analyze. Lifting binaries to LLVM IR enables reuse of the rich LLVM toolchain for downstream binary analyses. However, existing binary lifters often fail to produce syntactically valid LLVM IR or to restore sufficient semantics, making downstream analyses unreliable or unfeasible. This paper introduces BIT, a novel binary lifter designed to ensure syntactic compliance as well as semantic adequacy BIT achieves this through a multistage approach that includes anchor variable identification, analysis context collection, and IR refinement. In the evaluation, BIT achieved excellent results across multiple downstream analyses when compared with various lifters In static analysis, the F1 score of bug detection is 0.85, which is better than Plankton’s 0.81; in symbolic execution, it outperforms McSema by 3,049× in path exploration and by 1.36× in test case generation, respectively; in reanalysis, BIT can complete all tasks and is consistent with the advanced work McSema. These results highlight BIT’s ability to bridge the gap between binary-level analysis and the LLVM toolchain. Puzhuo Liu, Peng Di, Jingling Xue, Yu Jiang 0001 |
CGO | 1 |
| 2026 | User-Space Dependency-Aware Rehosting for Linux-Based Firmware Binaries
Cen Zhang, Yaowen Zheng, Puzhuo Liu, Jian Zhang 0087, Yeting Li, Yang Liu 0003, Limin Sun 0001 |
NDSS | 4 |
| 2026 | ADGFUZZ: Assignment Dependency-Guided Fuzzing for Robotic Vehicles
Yaowen Zheng, Puzhuo Liu, Dongliang Fang, Jiaxing Cheng, Dingyi Shi, Limin Sun 0001 |
NDSS | 3 |
| 2026 | Bridge: High-Order Taint Vulnerabilities Detection in Linux-Based IoT Firmware
Jiaqian Peng, Puzhuo Liu, Yicheng Zeng, Yongji Liu, Hongsong Zhu |
SP | 2 |
| 2025 | Automated Flaw Detection for Industrial Robot RESTful Service
Puzhuo Liu, Yaowen Zheng, Dongliang Fang, Shuaizong Si, Zhiwen Pan, Limin Sun 0001 |
VMCAI (2) | 2 |
| 2025 | LLM-Powered Static Binary Taint AnalysisabstractThis article proposes LATTE , the first static binary taint analysis that is powered by a large language model (LLM). LATTE is superior to the state of the art (e.g., Emtaint, Arbiter, Karonte) in three aspects. First, LATTE is fully automated while prior static binary taint analyzers need rely on human expertise to manually customize taint propagation rules and vulnerability inspection rules. Second, LATTE is significantly effective in vulnerability detection, demonstrated by our comprehensive evaluations. For example, LATTE has found 37 new bugs in real-world firmware, which the baselines failed to find. Moreover, 10 of them have been assigned CVE numbers. Lastly, LATTE incurs remarkably low engineering cost, making it a cost-efficient and scalable solution for security researchers and practitioners. We strongly believe that LATTE opens up a new direction to harness the recent advance in LLMs to improve vulnerability analysis for binary programs. Puzhuo Liu, Chengnian Sun, Yaowen Zheng, Xuan Feng 0005, Zhi Li 0018, Peng Di, Yu Jiang 0001, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | T-Rec: Fine-Grained Language-Agnostic Program Reduction Guided by Lexical SyntaxabstractProgram reduction strives to eliminate bug-irrelevant code elements from a bug-triggering program, so that (1) a smaller and more straightforward bug-triggering program can be obtained, (2) and the difference among duplicates (i.e., different programs that trigger the same bug) can be minimized or even eliminated. With such reduction and canonicalization functionality, program reduction facilitates debugging for software, especially language toolchains, such as compilers, interpreters, and debuggers. While many program reduction techniques have been proposed, most of them (especially the language-agnostic ones) overlooked the potential reduction opportunities hidden within tokens. Therefore, their capabilities in terms of reduction and canonicalization are significantly restricted. To fill this gap, we propose \(\mathsf{T}\) - \(\mathsf{Rec}\) , a fine-grained language-agnostic program reduction technique guided by lexical syntax. Instead of treating tokens as atomic and irreducible components, \(\mathsf{T}\) - \(\mathsf{Rec}\) introduces a fine-grained reduction process that leverages the lexical syntax of programming languages to effectively explore the reduction opportunities in tokens. Through comprehensive evaluations with versatile benchmark suites, we demonstrate that \(\mathsf{T}\) - \(\mathsf{Rec}\) significantly improves the reduction and canonicalization capability of two existing language-agnostic program reducers (i.e., Perses and Vulcan). \(\mathsf{T}\) - \(\mathsf{Rec}\) enables Perses and Vulcan to further eliminate 1,294 and 1,315 duplicates in a benchmark suite that contains 3,796 test cases that trigger 46 unique bugs. Additionally, \(\mathsf{T}\) - \(\mathsf{Rec}\) can also reduce up to 65.52% and 53.73% bytes in the results of Perses and Vulcan on our multi-lingual benchmark suite, respectively. Yongqiang Tian 0001, Mengxiao Zhang 0004, Puzhuo Liu, Yu Jiang 0001, Chengnian Sun |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | MSGFuzzer: Message Sequence Guided Industrial Robot Protocol FuzzingabstractIndustrial robots are widely used in industrial control systems (ICS). Once compromised, it could be maliciously controlled by attackers, endangering manufacturing processes or even human lives. Therefore, timely discovery of vulnerabilities in industrial robots is essential. Protocol fuzzing is a popular method for discovering protocol implementation vulnerabilities. However, the intricate workflow of industrial robots imposes strict message sequence constraints on message execution. Moreover, the overhead of sequence constraint satisfaction is exacerbated by the redundant messages in message sequences and the inherent delays in physical domain execution. These challenges make it difficult for fuzzers to penetrate deep code paths for fuzzing effectively. In this paper, we propose MSGFuzzer, a message sequence-guided industrial robot protocol fuzzer. Specifically, we filter the original traffic based on message byte characteristics and gener-ate message sequences. After that, we distinguish the sequence constraints for each message through the feedback mechanism of the industrial robot. To reduce state-guidance time, we construct the minimal message sequence based on the constraint conditions of messages. We evaluated MSGFuzzer on a real industrial robot. The results show that MSGFuzzer discovered 12 unique crashes. Note that this is at least 71.4% more effective than state-of-the-art protocol fuzzers in crash discoveries Yang Zhang 0145, Dongliang Fang, Puzhuo Liu, Laile Xi, Xin Chen 0123, Shuaizong Si, Limin Sun 0001 |
ICST | 3 |
| 2024 | Adversarial Attack against Intrusion Detectors in Cyber-Physical Systems With Minimal PerturbationsabstractCyber-Physical Systems (CPS) are crucial for critical infrastructure sectors such as electricity, water, and transportation. Machine Learning (ML) and Deep Learning (DL)-based Intrusion Detection Systems (IDS) are widely used in CPS for security monitoring. Attack and defense confrontation is an eternal topic, leading to increased research on adversarial attacks against IDS. However, most existing research on CPS adversarial attacks focuses on improving evasion capabilities without ensuring the preservation of malicious attack functionality. To address this problem, we propose a Conditional Wasserstein GAN (CWGAN) based framework to generate adversarial examples that can not only evade IDS detection but also impose constraints on the specified target sensors to preserve the original attack functionality. Evaluation results demonstrate that our approach can effectively preserve the intended malicious functionality by significantly reducing the perturbations to specific target sensors. Specifically, we achieve an average reduction of 96.93% and 90.59%, and a maximum reduction of 93.26% and 95.94% compared to the state-of-the-art JSMA and GAN based methods, respectively, while maintaining largely unchanged evasion capabilities against IDS. Mingqiang Bai, Puzhuo Liu, Fei Lv 0010, Dongliang Fang, Shichao Lv, Limin Sun 0001 |
ISPA | 2 |
| 2024 | Battling against Protocol Fuzzing: Protecting Networked Embedded Devices from Dynamic FuzzersabstractN etworked E mbedded D evices (NEDs) are increasingly targeted by cyberattacks, mainly due to their widespread use in our daily lives. Vulnerabilities in NEDs are the root causes of these cyberattacks. Although deployed NEDs go through thorough code audits, there can still be considerable exploitable vulnerabilities. Existing mitigation measures like code encryption and obfuscation adopted by vendors can resist static analysis on deployed NEDs, but are ineffective against protocol fuzzing. Attackers can easily apply protocol fuzzing to discover vulnerabilities and compromise deployed NEDs. Unfortunately, prior anti-fuzzing techniques are impractical as they significantly slow down NEDs, hampering NED availability. To address this issue, we propose Armor—the first anti-fuzzing technique specifically designed for NEDs. First, we design three adversarial primitives–delay, fake coverage, and forged exception–to break the fundamental mechanisms on which fuzzing relies to effectively find vulnerabilities. Second, based on our observation that inputs from normal users consistent with the protocol specification and certain program paths are rarely executed with normal inputs, we design static and dynamic strategies to decide whether to activate the adversarial primitives. Extensive evaluations show that Armor incurs negligible time overhead and effectively reduces the code coverage (e.g., line coverage by 22%-61%) for fuzzing, significantly outperforming the state of the art. Puzhuo Liu, Yaowen Zheng, Chengnian Sun, Hong Li 0004, Zhi Li 0018, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | FITS: Inferring Intermediate Taint Sources for Effective Vulnerability Analysis of IoT Device FirmwareabstractFinding vulnerabilities in firmware is vital as any firmware vulnerability may lead to cyberattacks to the physical IoT devices. Taint analysis is one promising technique for finding firmware vulnerabilities thanks to its high coverage and scalability. However, sizable closed-source firmware makes it extremely difficult to analyze the complete data-flow paths from taint sources (i.e., interface library functions such as recv) to sinks. Puzhuo Liu, Yaowen Zheng, Chengnian Sun, Dongliang Fang, Mingdong Liu, Limin Sun 0001 |
ASPLOS (4) | 1 |
| 2023 | MESCAL: Malicious Login Detection Based on Heterogeneous Graph Embedding with Supervised Contrastive LearningabstractMalicious logins via stolen credentials have become a primary threat in cybersecurity due to their stealthy nature. Recent malicious login detection methods based on graph learning techniques have made progress due to their ability to capture interconnected relationships among log entries. However, limited malicious samples pose a critical challenge to the detection performance of existing methods. In this paper, we propose MESCAL, a novel approach based on heterogeneous graph embedding with supervised contrastive learning to solve this challenge. Concretely, we construct authentication heterogeneous graphs to represent multiple and interconnected log events. Then, we pretrain a feature extractor with supervised contrastive learning to capture rich semantics on the graphs from limited malicious samples. Based on this, cost-sensitive learning is adopted to distinguish malicious logins on imbalanced data. Extensive evaluations show that the F1 score of MESCAL based on the imbalance dataset is 94.63%, which outperforms state-of-the-art approaches. Weiqing Huang, Yangyang Zong, Zhixin Shi, Puzhuo Liu |
ISCC | 4 |
| 2023 | UCRF: Static analyzing firmware to generate under-constrained seed for fuzzing SOHO router
Jiaqian Peng, Puzhuo Liu, Yaowen Zheng, Limin Sun 0001 |
Comput. Secur. | 3 |
| 2022 | Finding Vulnerabilities in Internal-binary of Firmware with CluesabstractEmbedded devices, represented by Internet of Things devices, bring great convenience to our daily life. Firmware is the core of the embedded device operation. However, vulnerabilities in the firmware can be exploited remotely by hackers through the network. Unfortunately, existing methods are only suitable for finding vulnerabilities in binaries (border-binary) that interact directly with users. When applied to other binaries (internal-binary) that indirectly interact with users, the lack of analysis sources and constraint conditions leads to many false negatives and false positives. In this paper, we propose a new keyword-sensitive data flow analysis approach to address the challenge. Specifically, we leverage crawlers to collect clues related to vulnerability reports from the Internet. Then we use the clues and communication paradigm finders to establish the relationship between different binaries in the firmware sample to form binary dependency graphs. At the same time, based on the functional features, we further dig out the binary relationships that have no Internet clues. Finally, we perform static taint analysis based on binary dependency graphs to determine vulnerabilities. We implemented and evaluated our prototype system FBI. Compared with Karonte, a state-of-the-art tool, FBI found significantly more true positives in Karonte’s data set. Puzhuo Liu, Dongliang Fang, Shichao Lv, Hongsong Zhu, Limin Sun 0001 |
ICC | 1 |
| 2022 | Fuzzing proprietary protocols of programmable controllers to find vulnerabilities that affect physical control
Puzhuo Liu, Yaowen Zheng, Zhanwei Song, Dongliang Fang, Shichao Lv, Limin Sun 0001 |
J. Syst. Archit. | 1 |
| 2021 | DSS: Discrepancy-Aware Seed Selection Method for ICS Protocol Fuzzing
Shuangpeng Bai, Hui Wen 0001, Dongliang Fang, Puzhuo Liu, Limin Sun 0001 |
ACNS (2) | 5 |
| 2021 | ICS3Fuzzer: A Framework for Discovering Protocol Implementation Bugs in ICS Supervisory Software by FuzzingabstractThe supervisory software is widely used in industrial control systems (ICSs) to manage field devices such as PLC controllers. Once compromised, it could be misused to control or manipulate these physical devices maliciously, endangering manufacturing process or even human lives. Therefore, extensive security testing of supervisory software is crucial for the safe operation of ICS. However, fuzzing ICS supervisory software is challenging due to the prevalent use of proprietary protocols. Without the knowledge of the program states and packet formats, it is difficult to enter the deep states for effective fuzzing. Dongliang Fang, Zhanwei Song, Le Guan, Puzhuo Liu, Anni Peng, Yaowen Zheng, Peng Liu 0005, Hongsong Zhu, Limin Sun 0001 |
ACSAC | 4 |