EDBT 2026 Demo / reviewers in the wild / expert
Pengbin Feng
dblp:198/7341
· DBLP profile ↗
20ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 12 · 3 first-author · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RomeFuzz: Path-aware Directed Greybox Fuzzing via Dyna-Static Indirect Call Analysis
Pengbin Feng, Chao Yang 0016, Zhizhuang Jia, Jianfeng Ma 0001 |
DSN | 2 |
| 2026 | MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment
Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou |
ICPR (15) | 4 |
| 2025 | CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward
Pengbin Feng, Yanbin Lin, Shuzhang Cai, Zongao Bian, Jinghua Yan, Xingquan Zhu 0001 |
IEEE Big Data | 2 |
| 2025 | LLM-Pot: A High-Interaction Honeypot System Driven by Large Language ModelabstractHoneypots are commonly used tools in network security protection. However, low-interaction honeypots cannot obtain in-depth attack information, while the deployment of high-interaction honeypots is costly. This paper presents LLM-Pot, a novel high-interaction honeypot architecture powered by the Large Language Model (LLM), which explores the direction of intelligent honeypots and addresses the limitations of conventional honeypot solutions. LLM-Pot utilizes LLM to generate dynamic, context-aware responses that accurately simulate the behaviors of real operating systems. To demonstrate the effectiveness of LLM-Pot, this work used offline and online evaluations. The offline evaluation compared LLM-Pot and Cowrie by analyzing their responses to selected commands, and the results demonstrated LLM-Pot’s superior ability in handling complex operations. Online evaluation deployed honeypots in the cloud and captured extensive attack data over two weeks. The evaluation results demonstrate that our LLM-driven approach outperforms traditional honeypots across multiple key metrics, validating LLM-Pot’s superior deception capabilities. Xuan Lyu, Pengbin Feng, Ning Xi 0002, XinDi Ma, Li Yang 0005, Di Lu 0001, Jianfeng Ma 0001 |
GLOBECOM | 2 |
| 2025 | Impact assessment of third-party library vulnerabilities through vulnerability reachability analysis
Zhizhuang Jia, Chao Yang 0016, Pengbin Feng, Xinghua Li 0001, Jianfeng Ma 0001 |
Comput. Secur. | 3 |
| 2025 | Resist Dependency Explosion in Attack Investigation With Splittable Tag Propagation and AggregationabstractAdvanced Persistent Threats (APTs) pose significant security risks to the community. Researchers thereby propose techniques to capture the complex and stealthy scenarios of APT attacks through the use of provenance graphs to model system entities and their dependencies. Particularly, to mitigate the dependency explosion problem in attack investigation using provenance graphs, tag-based and priority-based provenance graphs are frequently utilized for analyzing attacks. These methods use threat tag propagation and threat prioritization to reduce the size of the provenance graph for faster analysis. Unfortunately, these methods can allow more complex and potential attacks to evade detection. To overcome these difficulties, we propose an APT attack investigation system,ProTaging, for APT detection and forensic analysis. By using Tactics, Techniques, and Procedures (TTPs) rules to assign and update the node's threat tag, splittable tag propagation to control the scope of threat information, and threat weight aggregation and prioritized backward analysis during the forensic analysis phase, ProTaging effectively reconstructs attack paths in seconds without dependency explosion. Experimental results on both the simulation dataset, DARPA TC E3, E5 dataset, and DARPA OpTC dataset demonstrate that ProTaging generates smaller dependency graphs (2.5 times smaller) and has fewer false positives (6.7 times fewer) compared to state-of-the-art solutions. Additionally, ProTaging significantly reduces manual investigation effort by approximately 99.9%. Anyuan Sang, Junbo Jia, Li Yang 0005, Pengbin Feng, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | GNNDroid: Graph-Learning Based Malware Detection for Android Apps With Native CodeabstractWith the rapid development of mobile apps, developers tend to implement a variety of functionalities to support users’ demands. Thus, they involve the usage of native libraries to fulfill the luxuriant functionalities and maintain fast system responses, instead of using a unitary programming language (i.e., Java). Nonetheless, such an inter-language programming framework also introduces more security issues because attackers can conceal malicious behaviors at the native level to evade Android security vetting. Existing state-of-the-art detection tools mainly rely on the information extracted in the Java code to infer the potential malicious behaviors implemented in native code. None of them could simultaneously study the correlated behaviors in Java and native code. Therefore, in this paper, we proposed a static semantic-driven malware detection tool,GNNDroid, to distinguish malware by combining the behaviors implemented in both Java and native code. First,GNNDroidseparately analyzes Java and native code to construct Java function call graphs and native function call graphs. It then utilizes a regex-based function recognition approach to explore the correlations between Java code and native code. According to the code correlations,GNNDroidconstructs Multi-Relational Directed Graphs (MRDGs) to extract the comprehensive behaviors. Finally, it executes a Gated Graph Neural Network (GGNN) to analyze the MRDGs and distinguish malicious apps. We assessedGNNDroidby analyzing 40,000 Android apps and compared them with state-of-the-art tools. The result demonstrated thatGNNDroidnot only performs well when analyzing Java+native apps (i.e., apps implemented by both Java and native code), achieving an F1 of 98.57% but also effectively exploits Java only apps (i.e., apps implemented by Java code), achieving an F1 of 96.31%. Ning Xi 0002, Pengbin Feng, Siqi Ma 0001, Jianfeng Ma 0001, Yulong Shen 0001, Yale Yang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Flash: Federated Graph Learning-Based Malicious Bash Script Detection for Industrial Cyber-Physical Systems
Pengbin Feng, Ning Xi 0002, Jiong Jin, Jun Zhang 0010, Jianfeng Ma 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | BinGo: Identifying Security Patches in Binary Code with Graph Representation LearningabstractA timely software update is vital to combat the increasing security vulnerabilities. However, some software vendors may secretly patch their vulnerabilities without creating CVE entries or even describing the security issue in their change log. Thus, it is critical to identify these hidden security patches and defeat potential N-day attacks. Researchers have employed various machine learning techniques to identify security patches in open-source software, leveraging the syntax and semantic features of the software changes and commit messages. However, all these solutions cannot be directly applied to the binary code, whose instructions and program flow may dramatically vary due to different compilation configurations. In this paper, we propose BinGo, a new security patch detection system for binary code. The main idea is to present the binary code as code property graphs to enable a comprehensive understanding of program flow and perform a language model over each basic block of binary code to catch the instruction semantics. BinGo consists of four phases, namely, patch data pre-processing, graph extraction, embedding generation, and graph representation learning. Due to the lack of an existing binary security patch dataset, we construct such a dataset by compiling the pre-patch and post-patch source code of the Linux kernel. Our experimental results show BinGo can achieve up to 80.77% accuracy in identifying security patches between two neighboring versions of binary code. Moreover, BinGo can effectively reduce the false positives and false negatives caused by the different compilers and optimization levels. Shu Wang 0004, Pengbin Feng, Xinda Wang 0001, Qi Li 0002, Kun Sun 0001 |
AsiaCCS | 3 |
| 2024 | DP-CLMI:Differentially Private Contrastive Learning Against Membership Inference Attack
Yiwen Xia, XinDi Ma, Qi Jiang 0001, Ning Xi 0002, Di Lu 0001, Pengbin Feng, Sheng Gao 0002, Jianfeng Ma 0001 |
ICA3PP (5) | 7 |
| 2024 | GlareShell: Graph learning-based PHP webshell detection for web server of industrial internet
Pengbin Feng, Dawei Wei, Qiaoyang Li, Youbing Hu, Ning Xi 0002 |
Comput. Networks | 1 |
| 2024 | DawnGNN: Documentation augmented windows malware detection using graph neural network
Pengbin Feng, Le Gai, Li Yang 0005, Qin Wang 0008, Teng Li 0003, Ning Xi 0002, Jianfeng Ma 0001 |
Comput. Secur. | 1 |
| 2023 | BejaGNN: behavior-based Java malware detection via graph neural network
Pengbin Feng, Li Yang 0005, Di Lu 0001, Ning Xi 0002, Jianfeng Ma 0001 |
J. Supercomput. | 1 |
| 2022 | Consistency is All I Ask: Attacks and Countermeasures on the Network Context of Distributed Honeypots
Pengbin Feng, Jiahao Cao 0001, Tommy Chin, Kun Sun 0001, Qi Li 0002 |
DIMVA | 2 |
| 2022 | BinProv: Binary Code Provenance Identification without DisassemblyabstractProvenance identification, which is essential for binary analysis, aims to uncover the specific compiler and configuration used for generating the executable. Traditionally, the existing solutions extract syntactic, structural, and semantic features from disassembled programs and employ machine learning techniques to identify the compilation provenance of binaries. However, their effectiveness heavily relies on disassembly tools (e.g., IDA Pro) and tedious feature engineering, since it is challenging to obtain accurate assembly code, particularly, from the stripped or obfuscated binaries. In addition, the features in machine learning approaches are manually selected based on the domain knowledge of one specific architecture, which cannot be applied to other architectures. In this paper, we develop an end-to-end provenance identification system BinProv, which leverages a BERT (Bidirectional Encoder Representations from Transformers) based embedding model to learn and represent the context semantics and syntax directly from the binary code. Therefore, BinProv avoids the disassembling step and manual feature selection in provenance identification. Moreover, BinProv can distinguish the compilers and the four optimization levels (O0/O1/O2/O3) by fine-tuning the classifier model with the embedding inputs for specific provenance identification tasks. Experimental results show that BinProv achieves 92.14%, 99.4%, and 99.8% accuracy at byte sequence, function, and binary levels, respectively. We further demonstrate that BinProv works well on obfuscated binary code, suggesting that BinProv is a viable approach to remarkably mitigate the disassembler dependence in future provenance identification tasks. Finally, our case studies show that BinProv can better identify compiler helper functions and improve the performance of binary code similarity detection. Shu Wang 0004, Yunlong Xing, Pengbin Feng, Haining Wang 0001, Qi Li 0002, Songqing Chen, Kun Sun 0001 |
RAID | 4 |
| 2022 | Enhancing malware analysis sandboxes with emulated user behavior
Pengbin Feng, Shu Wang 0004, Kun Sun 0001, Jiahao Cao 0001 |
Comput. Secur. | 2 |
| 2021 | PatchDB: A Large-Scale Security Patch DatasetabstractSecurity patches, embedding both vulnerable code and the corresponding fixes, are of great significance to vulnerability detection and software maintenance. However, the existing patch datasets suffer from insufficient samples and low varieties. In this paper, we construct a large-scale patch dataset called PatchDB that consists of three components, namely, NVD-based dataset, wild-based dataset, and synthetic dataset. The NVD-based dataset is extracted from the patch hyperlinks indexed by the NVD. The wild-based dataset includes security patches that we collect from the commits on GitHub. To improve the efficiency of data collection and reduce the effort on manual verification, we develop a new nearest link search method to help find the most promising security patch candidates. Moreover, we provide a synthetic dataset that uses a new oversampling method to synthesize patches at the source code level by enriching the control flow variants of original patches. We conduct a set of studies to investigate the effectiveness of the proposed algorithms and evaluate the properties of the collected dataset. The experimental results show that PatchDB can help improve the performance of security patch identification. Xinda Wang 0001, Shu Wang 0004, Pengbin Feng, Kun Sun 0001, Sushil Jajodia |
DSN | 3 |
| 2019 | UBER: Combating Sandbox Evasion via User Behavior Emulators
Pengbin Feng, Kun Sun 0001 |
ICICS | 1 |
| 2017 | Data-Oriented Instrumentation against Information Leakages of Android ApplicationsabstractAs one of the most prominent threat, information leakages usually take sensitive data from some private sources and improperly release the data through malicious or misused method invocations and intercommunications. As a countermeasure against this threat, a number of detection approaches have been developed based on static analysis, esp. taint analysis. But we still have not reached a satisfactory solution to the patching and mitigation against this threat. In this paper, we propose an approach to automatically instrument malicious Android applications with cryptographic primitives and data randomization. With the help of an off-the-shelf taint analyzer, we detect the parts of code that might leak private information. In order to mitigate these information leakages, the standard cipher transformations and randomization are used to enforce different security policies according to the positions of related information sinks and intermediate system calls along malicious flow paths. The evaluation on different benchmark suites and real-world applications demonstrates that our approach can avoid false positives and mitigate around 91% information leakages in real applications, with acceptable cost on analysis and instrumentations affordable by desktops. Cong Sun 0001, Pengbin Feng, Teng Li 0003, Jianfeng Ma 0001 |
COMPSAC (2) | 2 |
| 2016 | Measuring the risk value of sensitive dataflow path in Android applicationsabstractAbstract Nowadays, smartphones carry large amounts of user privacy and sensitive data. With the popularity of the Android operating system, the cases of sensitive date leakage in Android applications are on the rise and are causing a great loss to Android users. In order to mitigate this condition, static and dynamic taint analysis are applied to precisely detect sensitive data leakages. These approaches cannot distinguish sensitive data leakages in benign apps from the ones in malicious apps. Recently, the difference on sensitive data flows between benign apps and malicious apps has been found to be significant. In this paper, we further find that there exists great difference between benign and malicious apps on the frequencies of sensitive dataflow paths. This difference can be used to enforce a risk value over every sensitive dataflow path. This risk value can guide the identification of sensitive data leakages in malicious apps. We present RISKPATH, a tool that automatically calculates the risk values for sensitive dataflow paths in Android applications. Applying the result of RISKPATH to MUDFLOW framework, we increase the true positive rate of malware detection by 3.96–6.54% on different datasets with reasonable increase in time and memory consumption. Copyright © 2017 John Wiley & Sons, Ltd. Pengbin Feng, Cong Sun 0001, Jianfeng Ma 0001 |
Secur. Commun. Networks | 1 |