EDBT 2026 Demo / reviewers in the wild / expert
Lirong Fu
dblp:246/8180
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionabstractBinary Code Similarity Detection (BCSD) plays a vital role in various security applications, including vulnerability identification, malware analysis, and code plagiarism detection.With the growing adoption of deep neural networks (DNNs), substantial progress has been made in recognizing and classifying similar code segments.However, DNN-based BCSD methods often exhibit low accuracy and robustness because they struggle to capture fine-grained and high-level program semantics.In contrast, such semantics are typically captured through natural language interpretations of source code by large language models (LLMs).Yet, LLM-based BCSD methods are constrained by their large model sizes and high inference latency.To alleviate these limitations, this paper proposes BinSKD.The key idea is to leverage an LLM-based BCSD method as the teacher model and transfer its knowledge of high-level program semantics to various DNNbased student models.Specifically, to avoid propagating errors from the teacher to the student, we introduce selective distillation, selecting targets with accurate semantics according to their detection retrieval.In addition, to mitigate the noise introduced by a number of negative samples during distillation, we further propose discrepancy-weighted sampling to focus on the samples where the student's prediction notably deviates from the teacher's.Our experiments show that BinSKD yields Recall@1 improvements of 14.5%-91.2%for DNN-based BCSD methods and enables HermesSim to match the teacher's performance with ordersof-magnitude efficiency. Shize Zhou, Peiyu Liu 0003, Lirong Fu, Wenhai Wang |
ACL (1) | 3 |
| 2025 | Facial Authentication Security Evaluation Against Deepfake Attacks in Mobile Apps
Chuer Yu, Siyi Xia, Zonghui Wang, Lirong Fu, Zhiyuan Wan, Yandong Gao, Wenzhi Chen |
ACISP (3) | 6 |
| 2025 | BinEGA: Enhancing DNN-based Binary Code Similarity Detection through Efficient Graph AlignmentabstractBinary Code Similarity Detection (BCSD) is essential in various binary code security applications, enabling tasks such as vulnerability identification, malware analysis, and detection of code plagiarism. With the growing adoption of deep neural networks (DNNs) in BCSD, there has been significant progress in the identification and classification of similar code segments. However, DNN-based BCSD approaches often suffer from high false positive rates, because DNNs inevitably map different binary functions with complex structures and semantics to similar low-dimensional embeddings. To alleviate this issue, this paper introduces BinEGA, a novel graph alignment-based approach to enhance the accuracy of DNN-based BCSD approaches. The main idea of BinEGA is to employ a general and low-cost equivalence check through lightweight graph alignment, allowing for the identification and elimination of semantically deviating functions among the top-k candidates retrieved by DNN-based BCSD approaches. During the graph alignment process, we first obtain the node embeddings according to structure and attribute feature. Then we employs pairwise comparison of these node embeddings to filter the false positives because binary code compiled from the same source code always shares similar basic blocks. Our experimental results demonstrate that BinEGA effectively enhances the performance of various edge-cutting DNN-based BCSD approaches across diverse scenarios. For instance, BinEGA significantly enhances RECALL@10 in the cross-optimization scenario for state-of-the-art (SOTA) approaches, with an average improvement of 29.2% for BinaryAI and 33.5% for jTrans. Moreover, BinEGA achieves 88.9 % reduction in execution time compared to other enhancement techniques. In summary, this work provides a robust, generalizable, and efficient solution to improve the reliability of BCSD tools in real-world applications. Shize Zhou, Lirong Fu, Peiyu Liu 0003, Wenhai Wang |
SANER | 2 |
| 2025 | LuaTaint: A Static Analysis System for Web Configuration Interface Vulnerability of Internet of Things DevicesabstractThe diversity of Web configuration interfaces for Internet of Things (IoT) devices has exacerbated issues, such as inadequate permission controls and insecure interfaces, resulting in various vulnerabilities. Owing to the varying interface configurations across various devices, the existing methods are inadequate for identifying these vulnerabilities precisely and comprehensively. This study addresses these issues by introducing an automated vulnerability detection system, called LuaTaint. It is designed for the commonly used Web configuration interface of IoT devices. LuaTaint combines static taint analysis with a large language model (LLM) to achieve widespread and high-precision detection. The extensive traversal of the static analysis ensures the comprehensiveness of the detection. The system also incorporates rules related to page handler control logic within the taint detection process to enhance its precision and extensibility. Moreover, we leverage the prodigious abilities of LLM for code analysis tasks. By utilizing LLM in the process of pruning false alarms, the precision of LuaTaint is enhanced while significantly reducing its dependence on manual analysis. We develop a prototype of LuaTaint and evaluate it using 2447 IoT firmware samples from 11 renowned vendors. LuaTaint has discovered 111 vulnerabilities. Moreover, LuaTaint exhibits a vulnerability detection precision rate of up to 89.29%. Jiahui Xiang, Lirong Fu, Peiyu Liu 0003, Huan Le, Liming Zhu 0003, Wenhai Wang |
IEEE Internet Things J. | 2 |
| 2024 | Exploring ChatGPT's Capabilities on Vulnerability Management
Peiyu Liu 0003, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang |
USENIX Security Symposium | 3 |
| 2023 | JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application ServicesabstractMany Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead. Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen |
CLUSTER | 8 |
| 2023 | How IoT Re-using Threatens Your Sensitive Data: Exploring the User-Data Disposal in Used IoT DevicesabstractWith the rapid technology evolution of the Internet of Things (IoT) and increasing user needs, IoT device re-using becomes more and more common nowadays. For instance, more than 300,000 used IoT devices are selling on Craigslist. During IoT re-using, sensitive data such as credentials and biometrics residing in these devices may face the risk of leakage if a user fails properly dispose of the data. Thus, a critical security concern is raised: do (or can) users properly dispose of the sensitive data in used IoT? To the best of our knowledge, it is still an unexplored problem that desires a systematic study.In this paper, we perform the first in-depth investigation on the user-data disposal of used IoT devices. Our investigation integrates multiple research methods to explore the status quo and the root causes of the user-data leakages with used IoT devices. First, we conduct a user study to investigate the user awareness and understanding of data disposal. Then, we conduct a large-scale analysis on 4,749 IoT firmware images to investigate user-data collection. Finally, we conduct a comprehensive empirical evaluation on 33 IoT devices to investigate the effectiveness of existing data disposal methods.Through the systematical investigation, we discover that IoT devices collect more sensitive data than users expect. Specifically, we detect 121,984 sensitive data collections in the tested firmware. Moreover, users usually do not or even cannot properly dispose of the sensitive data. Worse, due to the inherent characteristics of storage chips, 13.2% of the investigated firmware perform "shallow" deletion, which may allow adversaries to obtain sensitive data after data disposal. Given the large-scale IoT re-using, such leakage would cause a broad impact. We have reported our findings to world-leading companies. We hope our findings raise awareness of the failures of user-data disposal with IoT devices and promote the protection of users’ sensitive data in IoT devices. Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Jingchang Qin, Wenhai Wang, Wenzhi Chen |
SP | 3 |
| 2022 | Focus : Function clone identification on cross-platformabstractAutomatic identification of function clones on cross-platform aims at determining whether two functions are identical or not without access to the source code, which is a fundamental challenge in vulnerability search, code plagiarism detection, and malware classification. With the rapid development of deep neural network in program analysis, the state-of-the-art neural network-based function clone identification methods propose to represent functions as embeddings by graph neural network (GNN). However, such a novel representation of functions brings in two challenges. (1) The feature engineering that accurately maps the raw data of binary code to machine learning features is complicated. (2) A highly accurate embedding of functions requires a customized GNN to focus on the most critical features to identify binary code. To the best of our knowledge, currently, a comprehensive work that can overcome the above challenges is still missing. In this paper, we propose a novel prototype named as Focus, which is designed to accurately and efficiently identify similar functions. Specifically, inspired by natural language processing techniques which effectively learns text semantic across natural languages, Focus can learn representative semantic features of functions by a customized learning model. To address the second challenge, a multi-head attention mechanism can be employed to capture the critical features of a function. Through extensive experiments, we demonstrate that Focus achieves high accuracy of function clone identification on a broad range of eight architectures. In particular, the identification performance (AUC value) of Focus is 97% and 99% for cross-platform and single-platform, respectively. Furthermore, the evaluation in real world applications shows that our Focus identifies 24 vulnerable functions among the top-30 candidates, which is one time higher than the baseline approaches. Lirong Fu, Shouling Ji, Changchang Liu, Peiyu Liu 0003, Fuzheng Duan, Zonghui Wang, Whenzhi Chen, Ting Wang 0006 |
Int. J. Intell. Syst. | 1 |
| 2021 | CPscan: Detecting Bugs Caused by Code Pruning in IoT KernelsabstractTo reduce the development costs, IoT vendors tend to construct IoT kernels by customizing the Linux kernel. Code pruning is common in this customization process. However, due to the intrinsic complexity of the Linux kernel and the lack of long-term effective maintenance, IoT vendors may mistakenly delete necessary security operations in the pruning process, which leads to various bugs such as memory leakage and NULL pointer dereference. Yet detecting bugs caused by code pruning in IoT kernels is difficult. Specifically, (1) a significant structural change makes precisely locating the deleted security operations (DSO ) difficult, and (2) inferring the security impact of a DSO is not trivial since it requires complex semantic understanding, including the developing logic and the context of the corresponding IoT kernel. Lirong Fu, Shouling Ji, Kangjie Lu, Peiyu Liu 0003, Xuhong Zhang 0002, Yuxuan Duan, Wenzhi Chen |
CCS | 1 |
| 2021 | IFIZZ: Deep-State and Efficient Fault-Scenario Generation to Test IoT FirmwareabstractIoT devices are abnormally prone to diverse errors due to harsh environments and limited computational capabilities. As a result, correct error handling is critical in IoT. Implementing correct error handling is non-trivial, thus requiring extensive testing such as fuzzing. However, existing fuzzing cannot effectively test IoT error-handling code. First, errors typically represent corner cases, thus are hard to trigger. Second, testing error-handling code would frequently crash the execution, which prevents fuzzing from testing following deep error paths.In this paper, we propose IFIZZ, a new bug detection system specifically designed for testing error-handling code in Linux-based IoT firmware. IFIZZ first employs an automated binary-based approach to identify realistic runtime errors by analyzing errors and error conditions in closed-source IoT firmware. Then, IFIZZ employs state-aware and bounded error generation to reach deep error paths effectively. We implement and evaluate IFIZZ on 10 popular IoT firmware. The results show that IFIZZ can find many bugs hidden in deep error paths. Specifically, IFIZZ finds 109 critical bugs, 63 of which are even in widely used IoT libraries. IFIZZ also features high code coverage and efficiency, and covers 67.3% more error paths than normal execution. Meanwhile, the depth of error handling covered by IFIZZ is 7.3 times deeper than that covered by the state-of-the-art method. Furthermore, IFIZZ has been practically adopted and deployed in a worldwide leading IoT company. We will open-source IFIZZ to facilitate further research in this area. Peiyu Liu 0003, Shouling Ji, Xuhong Zhang 0002, Qinming Dai, Kangjie Lu, Lirong Fu, Wenzhi Chen, Peng Cheng 0001, Wenhai Wang, Raheem A. Beyah |
ASE | 6 |
| 2020 | Understanding the Security Risks of Docker Hub
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wei-Han Lee, Wenzhi Chen, Raheem A. Beyah |
ESORICS (1) | 3 |
| 2019 | Prudent Practices for Designing Virtual Desktop ExperimentsabstractVirtual desktop technology aims at accessing a remote desktop by endpoint hardware.Great attention has been increasingly paid to virtual desktop since it can increase the utilization of computing resources and provide more flexible accesses.However, researchers have not yet come up with a comprehensive set of rigorous standards of experimental design and implementation in this field.Therefore, it is difficult to conduct prudent experiments, which is correct, real, and transparent.In this paper, we assess the experimental evaluations of recently published papers on desktop virtualization.We observe that most works can be further improved, due to the unsuitable experimental environment and the lack of descriptions of experimental settings.In this paper, in order to help researchers, reviewers, and readers, we propose several guidelines for designing correct, real, and transparent desktop virtualization experiment. Peiyu Liu 0003, Wenzhi Chen, Zonghui Wang, Lirong Fu |
SEKE | 4 |