Dongliang Fang

dblp:294/4627 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
26since 2021 · last 2026
0009-0005-7484-1333ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 2 first-author · 11 since 2021Computer networks · 7 · 7 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NS-FirmID: A Neuro-Symbolic Multi-Agent Framework for Reliable Firmware Version Identification at Internet Scale
Fengshi Zhang, Zhi Li 0018, Shunchao Xu, Yongle Chen, Dongliang Fang, Limin Sun 0001
DSN6
2026 Chronos: Large-Scale Online Firmware Version Detection via Inadvertent Chronological Fingerprints
Fengshi Zhang, Zhi Li 0018, Shunchao Xu, Dongliang Fang, Yongle Chen, Limin Sun 0001
INFOCOM6
2026 ADGFUZZ: Assignment Dependency-Guided Fuzzing for Robotic Vehicles
Yaowen Zheng, Puzhuo Liu, Dongliang Fang, Jiaxing Cheng, Dingyi Shi, Limin Sun 0001
NDSS4
2026 NetID-GPT: Adapting large language models for large-scale internet-connected device identification
Zhi Li 0018, Shunchao Xu, Fengshi Zhang, Zhanwei Song, Dongliang Fang, Yongle Chen, Limin Sun 0001
Comput. Networks6
2025 PNetGPT: Proprietary Protocol Network Traffic Generation with Pre-trained Transformer
abstract
Generative pre-trained transformers are exceedingly effective as generative models and classifiers, widely used in natural language processing and computer vision. This work contributes to the exploration of generative pre-trained transformer-based models in the proprietary protocol network traffic. However, building a pre-trained model for proprietary protocol network traffic is non-trivial due to the heterogeneous unknown formats and the extreme scarcity of proprietary protocol network traffic datasets. In this paper, we present PNetGPT, a pre-trained transformer-based model for generating proprietary protocol network traffic. We have constructed the inaugural dataset of 2 real-world proprietary protocols. After training on this dataset, PNetGPT possesses the capacity to generate high-quality proprietary protocol network traffic to support various applications of proprietary protocols, including reverse analysis, protocol fuzzy testing, intrusion detection, etc. We evaluated PNetGPT with two real proprietary protocols and demonstrated state-of-the-art (SOTA) performance in handling heterogeneous unknown formats. The code and datasets are available at: https://github.com/Snail1502/PNetGPT
Zedong Li, Dongliang Fang, Xin Chen 0123, Zhanwei Song, Zhi Li 0018, Shichao Lv, Limin Sun 0001
ICASSP3
2025 Exploiting Binary Semantics: Enhancing Function Name Inference in Stripped Binaries via LLMs
abstract
Function name inference in stripped binaries is a crucial task that supports various security applications, including vulnerability detection and malware analysis. Existing methods suffer from limited model capacity and insufficient exploitation of function semantics, which constrains their ability to comprehend binary code and leads to poor generalization on unseen binaries. To address these problems, we propose BinLLM, a novel framework that leverages large language models (LLMs) to exploit the semantic potential of binary code, thereby enhancing function name inference. Specially, BinLLM integrates three key innovations: (1) source code semantics-guided function name refinement, which mitigates the negative effects of low-quality semantic identifiers during training; (2) A context-aware data collection algorithm that seeks richer semantic dependencies to improve model training and inference performance; (3) parameter-efficient fine-tuning on a domain-specific dataset enriched with semantic knowledge to enhance the model's understanding of binary semantics. These components collectively enhance the model's performance in function name inference on unseen binaries. We evaluate BinLLM on a large-scale dataset comprising$2,864,719$functions across four architectures (x86-64, x86-32, ARM, MIPS) and four optimization levels ($\mathrm{O} 0-\mathrm{O} 3$). Experimental results show that BinLLM achieves substantial improvements over state-of-the-art (SOTA) methods, with relative gains of$320.1 \%, 274.8 \%$, and 297.6 % in precision, recall, and F1-score. Ablation studies further validate the effectiveness of each component in enhancing overall performance.
Kailong Wang 0007, Dongliang Fang, Zhongwei Gu, Zhanwei Song, Yongle Chen, Zhiqiang Shi, Limin Sun 0001
IPCCC3
2025 Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM Verification
abstract
Binary Code Similarity Detection (BCSD), essential for binary-code related tasks like vulnerability detection, has attracted increasing attention in recent years. However, existing methods frequently fall short of achieving both high precision and recall at scale, and their results often lack interpretability due to the neglect of function context and reliance on purely similarity-driven outputs. Our key insights are twofold: 1) Binary functions are not self-contained; they depend on other code and data beyond their content to fulfill their functionalities. 2) Large language models (LLMs) excel not only at analyzing code but also at generating reasonable explanations. Motivated by these insights, we propose a general BCSD framework, Co2F uLL. We first systematically select stable and representative code and data features, along with their corresponding dependencies on the functions, to construct the function context. Then, by fusing function context with content similarities computed by the existing BCSD approach, we substantially narrow down the search space. Ultimately, we employ LLMs with a carefully designed prompt to verify the remaining candidates and produce clear, human-readable explanations. We conduct comprehensive experiments on a large function pool under varying compilation settings and after binary stripping. The results show that Co2F uLL based on HermesSim and DeepSeek-V3 achieves 80.5% precision and 94.4% recall, improving the baseline HermesSim by 142.5% and 42.2%, respectively, providing an accurate and interpretable solution for BCSD.
Chaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li 0008, Dongliang Fang, Yang Xiao 0011, Yongle Chen, Limin Sun 0001
ASE5
2025 Automated Flaw Detection for Industrial Robot RESTful Service
Puzhuo Liu, Yaowen Zheng, Dongliang Fang, Shuaizong Si, Zhiwen Pan, Limin Sun 0001
VMCAI (2)4
2025 Discovering PLC Web Application Vulnerabilities Impacting Physical Control Using LLM-Based Fuzzing
Jiaxing Cheng, Dongliang Fang, Zhongwei Gu, Shichao Lv, Shuaizong Si, Limin Sun 0001
WASA (1)2
2025 ICSPFuzzer: An Efficient Fuzzing Technique for ICS Protocols
Zhanwei Song, Dongliang Fang, Shunchao Xu, Yaowen Zheng, Hong Li 0004, Shichao Lv, Zhiqiang Shi, Limin Sun 0001
WASA (2)2
2025 SFACIF: A safety function attack and anomaly industrial condition identified framework
Kaixiang Liu, Yongfang Xie, Yuqi Chen 0001, Shiwen Xie, Xin Chen 0123, Dongliang Fang, Limin Sun 0001
Comput. Networks6
2025 InvisiGuard: Data Integrity for Microcontroller-Based Devices via Hardware-Triggered Write Monitoring
abstract
This paper considers a strongly connected network of agents, each capable of partially observing and controlling a discrete-time linear time-invariant (LTI) system that is jointly observable and controllable. Additionally, agents collaborate to achieve a shared estimated state, computed as the average of their local state estimates. Recent studies suggest that increasing the number of average consensus steps between state estimation updates allows agents to choose from a wider range of state feedback controllers, thereby potentially enhancing control performance. However, such approaches require that agents know the input matrices of all other nodes, and the selection of control gains is, in general, centralized. Motivated by the limitations of such approaches, we propose a new technique where: (i) estimation and control gain design is fully distributed and finite-time, and (ii) agent coordination involves a finite-time exact average consensus subroutine, allowing arbitrary selection of the convergence rate of the overall asymptotic estimation process despite the estimator's distributed nature. We verify our methodology's effectiveness using illustrative numerical simulations.
Dongliang Fang, Anni Peng, Le Guan, Erik van der Kouwe, Klaus von Gleissenthall, Wenwen Wang 0001, Yuqing Zhang 0001, Limin Sun 0001
IEEE Trans. Dependable Secur. Comput.1
2025 PREXP: Uncovering and Exploiting Security-Sensitive Objects in the Linux Kernel
abstract
Security-Sensitive Objects (SSOs) are often critical components in the exploitation of Linux kernel memory corruption vulnerabilities. While existing research has advanced SSOs identification and classification, there remains a significant gap in systematically understanding how these objects can be effectively exploited in real-world security analysis. To address this challenge, we present PREXP, a novel approach to analyzing SSOs exploitability and automating the transformation of Proof-of-Concept (PoC) into exploitable states. Our approach encompasses three key techniques: (1) capability analysis and attribute modeling of vulnerable object (2) extraction and filtering of target SSOs and (3) automatically augmenting PoCs with SSO-specific code to create exploitation capabilities. To evaluate our approach, we tested our prototype on 30 public CVEs, successfully parsing vulnerable object in 22 cases (73.3%) and achieving accurate SSO matches in 18 (60.0%). PREXP outperformed state-of-the-art tools such as SCAVY and AlphaEXP in structure-matching, and enabled the generation of new Control Flow Hijacking Primitives (CFHPs) for 3 previously unexploited vulnerabilities, demonstrating its practical value in real-world exploit development.
Zuxin Chen, Yaowen Zheng, Hong Li 0004, Siyuan Li 0014, Weijie Wang 0005, Dongliang Fang, Zhiqiang Shi, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.6
2025 EMFuzz: Use Electromagnetic Fuzzing for Automated Attack Surface Assessment of Actuators
abstract
Actuators are essential components in cyber-physical systems, enabling system modules to perform diverse and complex tasks. Unfortunately, the pursuit of higher functional complexity often correlates with a broader attack surface in actuators. Thus, an efficient automated attack surface assessment is crucial to avoid cyber incidents in critical infrastructures. Limited by enormous parameter spaces, current methods rely on heuristic tests to evaluate interference potential but cannot thoroughly investigate the full spectrum of potential hidden interference. The observation that similar interference trigger configurations lead to the same impact has motivated us to use machine learning algorithms for understanding different impact samples around decision boundaries. By leveraging generalized knowledge of responses against specific attack scenarios, we aim to improve the efficiency of automated attack surface assessment of electromagnetic interference on new targets. To this end, we introduce EMFuzz, an automated mechanism to fuzz hardware to quantify varying adverse effects. We evaluate EMFuzz on 16 new servos within real-world scenarios, where it achieves an 86% accuracy in classifying different attack vectors. With the same test time, EMFuzz uncovers over twice the effective attack configurations of the baseline, greatly improving assessment efficiency. To further validate its efficacy, we apply EMFuzz to assess the attack surface of a new actuator from a robot transfer unit, and it can successfully reveal three distinct adverse effects.
Shiquan Dong, Zhi Li 0018, Jianshuo Liu, Hong Li 0004, Dongliang Fang, Shichao Lv, Haining Wang 0001, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.5
2024 MSGFuzzer: Message Sequence Guided Industrial Robot Protocol Fuzzing
abstract
Industrial robots are widely used in industrial control systems (ICS). Once compromised, it could be maliciously controlled by attackers, endangering manufacturing processes or even human lives. Therefore, timely discovery of vulnerabilities in industrial robots is essential. Protocol fuzzing is a popular method for discovering protocol implementation vulnerabilities. However, the intricate workflow of industrial robots imposes strict message sequence constraints on message execution. Moreover, the overhead of sequence constraint satisfaction is exacerbated by the redundant messages in message sequences and the inherent delays in physical domain execution. These challenges make it difficult for fuzzers to penetrate deep code paths for fuzzing effectively. In this paper, we propose MSGFuzzer, a message sequence-guided industrial robot protocol fuzzer. Specifically, we filter the original traffic based on message byte characteristics and gener-ate message sequences. After that, we distinguish the sequence constraints for each message through the feedback mechanism of the industrial robot. To reduce state-guidance time, we construct the minimal message sequence based on the constraint conditions of messages. We evaluated MSGFuzzer on a real industrial robot. The results show that MSGFuzzer discovered 12 unique crashes. Note that this is at least 71.4% more effective than state-of-the-art protocol fuzzers in crash discoveries
Yang Zhang 0145, Dongliang Fang, Puzhuo Liu, Laile Xi, Xin Chen 0123, Shuaizong Si, Limin Sun 0001
ICST2
2024 Adversarial Attack against Intrusion Detectors in Cyber-Physical Systems With Minimal Perturbations
abstract
Cyber-Physical Systems (CPS) are crucial for critical infrastructure sectors such as electricity, water, and transportation. Machine Learning (ML) and Deep Learning (DL)-based Intrusion Detection Systems (IDS) are widely used in CPS for security monitoring. Attack and defense confrontation is an eternal topic, leading to increased research on adversarial attacks against IDS. However, most existing research on CPS adversarial attacks focuses on improving evasion capabilities without ensuring the preservation of malicious attack functionality. To address this problem, we propose a Conditional Wasserstein GAN (CWGAN) based framework to generate adversarial examples that can not only evade IDS detection but also impose constraints on the specified target sensors to preserve the original attack functionality. Evaluation results demonstrate that our approach can effectively preserve the intended malicious functionality by significantly reducing the perturbations to specific target sensors. Specifically, we achieve an average reduction of 96.93% and 90.59%, and a maximum reduction of 93.26% and 95.94% compared to the state-of-the-art JSMA and GAN based methods, respectively, while maintaining largely unchanged evasion capabilities against IDS.
Mingqiang Bai, Puzhuo Liu, Fei Lv 0010, Dongliang Fang, Shichao Lv, Limin Sun 0001
ISPA4
2024 Fast Firmware Fuzz with Input/Output Reposition
Mingfeng Xin, Liting Deng, Hui Wen 0001, Dongliang Fang, Shichao Lv, Limin Sun 0001
SecureComm (3)4
2024 PowerGuard: Using Power Side-Channel Signals to Secure Motion Controllers in ICS
abstract
Motion control systems, extensively utilized in domains like 3D printing, CNC machining, and robotic arm operations, are pivotal in modern manufacturing and automation processes. Consequently, a specific category of attacks, designed to target these systems, can manipulate the movements of controlled objects while replaying false sensor readings to evade existing tools, thereby severely disrupting these essential operations without being detected. To make things worse, the limited computing resources of embedded devices in these systems constrain the implementation of robust security protections and monitoring mechanisms locally. To solve this, we propose a novel side-channel method that leverages current signals emitted by motors to reconstruct trajectories for attack detection. In this paper, we design and implement a two-stage detection framework, dubbed PowerGuard. In the offline learning stage, PowerGuard first captures the current signals emitted by the servo motors and models the correlation between these signals and corresponding movement trajectories. In the real-time monitoring stage, PowerGuard finds outliers that deviate from the desired trajectory described in the benign G-code file. We have evaluated PowerGuard using a typical motion control system that contains CNC machine tools from different vendors (e.g., Siemens 828D, 840D-sl, Fanuc 0i-md, 0i-tf). We conducted extensive experiments to evaluate the reconstruction accuracy and attack detection performance. Experimental results show that PowerGuard can reconstruct movement trajectories with an error of 0.047mm, and detect 93.35% of various trajectory anomalies.
Yuqi Chen 0001, Xin Chen 0123, Zedong Li, Dongliang Fang, Kaixiang Liu, Shichao Lv, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Bitmap-Based Security Monitoring for Deeply Embedded Systems
abstract
Deeply embedded systems powered by microcontrollers are becoming popular with the emergence of Internet-of-Things (IoT) technology. However, these devices primarily run C/C \({+}{+}\) code and are susceptible to memory bugs, which can potentially lead to both control data attacks and non-control data attacks. Existing defense mechanisms (such as control-flow integrity (CFI), dataflow integrity (DFI) and write integrity testing (WIT), etc.) consume a massive amount of resources, making them less practical in real products. To make it lightweight, we design a bitmap-based allowlist mechanism to unify the storage of the runtime data for protecting both control data and non-control data. The memory requirements are constant and small, regardless of the number of deployed defense mechanisms. We store the allowlist in the TrustZone to ensure its integrity and confidentiality. Meanwhile, we perform an offline analysis to detect potential collisions and make corresponding adjustments when it happens. We have implemented our idea on an ARM Cortex-M-based development board. Our evaluation results show a substantial reduction in memory consumption when deploying the proposed CFI and DFI mechanisms, without compromising runtime performance. Specifically, our prototype enforces CFI and DFI at a cost of just 2.09% performance overhead and 32.56% memory overhead on average.
Anni Peng, Dongliang Fang, Le Guan, Erik van der Kouwe, Wenwen Wang 0001, Limin Sun 0001, Yuqing Zhang 0001
ACM Trans. Softw. Eng. Methodol.2
2023 FITS: Inferring Intermediate Taint Sources for Effective Vulnerability Analysis of IoT Device Firmware
abstract
Finding vulnerabilities in firmware is vital as any firmware vulnerability may lead to cyberattacks to the physical IoT devices. Taint analysis is one promising technique for finding firmware vulnerabilities thanks to its high coverage and scalability. However, sizable closed-source firmware makes it extremely difficult to analyze the complete data-flow paths from taint sources (i.e., interface library functions such as recv) to sinks.
Puzhuo Liu, Yaowen Zheng, Chengnian Sun, Dongliang Fang, Mingdong Liu, Limin Sun 0001
ASPLOS (4)5
2023 CEFI: Command Execution Flow Integrity for Embedded Devices
Anni Peng, Dongliang Fang, Wei Zhou 0026, Erik van der Kouwe, Yuqing Zhang 0001
DIMVA2
2022 Finding Vulnerabilities in Internal-binary of Firmware with Clues
abstract
Embedded devices, represented by Internet of Things devices, bring great convenience to our daily life. Firmware is the core of the embedded device operation. However, vulnerabilities in the firmware can be exploited remotely by hackers through the network. Unfortunately, existing methods are only suitable for finding vulnerabilities in binaries (border-binary) that interact directly with users. When applied to other binaries (internal-binary) that indirectly interact with users, the lack of analysis sources and constraint conditions leads to many false negatives and false positives. In this paper, we propose a new keyword-sensitive data flow analysis approach to address the challenge. Specifically, we leverage crawlers to collect clues related to vulnerability reports from the Internet. Then we use the clues and communication paradigm finders to establish the relationship between different binaries in the firmware sample to form binary dependency graphs. At the same time, based on the functional features, we further dig out the binary relationships that have no Internet clues. Finally, we perform static taint analysis based on binary dependency graphs to determine vulnerabilities. We implemented and evaluated our prototype system FBI. Compared with Karonte, a state-of-the-art tool, FBI found significantly more true positives in Karonte’s data set.
Puzhuo Liu, Dongliang Fang, Shichao Lv, Hongsong Zhu, Limin Sun 0001
ICC2
2022 Fuzzing proprietary protocols of programmable controllers to find vulnerabilities that affect physical control
Puzhuo Liu, Yaowen Zheng, Zhanwei Song, Dongliang Fang, Shichao Lv, Limin Sun 0001
J. Syst. Archit.4
2021 DSS: Discrepancy-Aware Seed Selection Method for ICS Protocol Fuzzing
Shuangpeng Bai, Hui Wen 0001, Dongliang Fang, Puzhuo Liu, Limin Sun 0001
ACNS (2)3
2021 ICS3Fuzzer: A Framework for Discovering Protocol Implementation Bugs in ICS Supervisory Software by Fuzzing
abstract
The supervisory software is widely used in industrial control systems (ICSs) to manage field devices such as PLC controllers. Once compromised, it could be misused to control or manipulate these physical devices maliciously, endangering manufacturing process or even human lives. Therefore, extensive security testing of supervisory software is crucial for the safe operation of ICS. However, fuzzing ICS supervisory software is challenging due to the prevalent use of proprietary protocols. Without the knowledge of the program states and packet formats, it is difficult to enter the deep states for effective fuzzing.
Dongliang Fang, Zhanwei Song, Le Guan, Puzhuo Liu, Anni Peng, Yaowen Zheng, Peng Liu 0005, Hongsong Zhu, Limin Sun 0001
ACSAC1
2021 Automatic Inference of Taint Sources to Discover Vulnerabilities in SOHO Router Firmware
Dongliang Fang, Huizhao Wang, Yaowen Zheng, Limin Sun 0001
SEC2