Yangyang Geng

dblp:226/9510 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 5 since 2021Security and privacy · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 A method for function name recovery in binaries via hierarchical semantic propagation
abstract
Abstract The large-scale deployment of IoT devices has led to widespread use of stripped firmware binaries, where function names are removed, leaving many functions anonymous. This obscures program semantics and heightens security risks by hiding vulnerabilities. Recovering meaningful function names is thus crucial for reverse engineering and security analysis. Existing approaches treat naming as an isolated prediction task or rely on static, limited calling contexts, struggling with deep call chains and severe semantic loss. To overcome this, we propose SemFlow , a call-graph-driven, hierarchical semantic propagation method. SemFlow reframes naming as an iterative process that progressively resolves anonymity by leveraging the call graph’s structure: it first recovers bottom-layer functions with richer context and propagates their names upward as semantic anchors to enrich context for mid- and top-layer functions. We evaluate SemFlow on real-world binaries across four architectures (x86-64, x86-32, ARM, MIPS) and four optimization levels (O0–O3). Results show consistent superiority over the state-of-the-art SYMGEN: F1 scores improve by up to 13.4% (x86-64), 20.3% (x86-32), 34.3% (ARM), and 158.5% (MIPS). Notably, on mid-layer anonymous functions—where semantic scarcity is most acute—SemFlow achieves an average F1 gain of 162.7%, demonstrating the effectiveness of hierarchical semantic propagation in mitigating context loss in deep call chains.
Debao Kong, Yangyang Geng, Chaojie Wei, Fang Jing, Liupeng He, Xueman Kong
Cybersecur.2
2026 Functional Semantic Inference Method for Industrial Control Proprietary Protocol
abstract
Industrial control protocols (ICPs), especially proprietary protocols, lack public specifications due to security and commercial considerations, posing profound challenges for security audits and vulnerability detection. Although the function code fields in ICPs define critical operational semantics (e.g., PLC start/stop), existing protocol reverse engineering (PRE) methods struggle to accurately infer the functional implications of their value dependencies. To address this critical gap, we propose FuncSeinfer, an operation-traffic correlation-driven method for inferring functional semantics in proprietary ICPs by associating graphical user interface (GUI) operations of programming software with network traffic. FuncSeinfer first employs an entropy-based dynamic field extraction algorithm to generate candidate fields. Subsequently, it applies heuristic rules and clustering scoring rules to accurately infer function code fields. Finally, it identifies semantics through variance analysis of traffic marked with operational labels. We rigorously evaluated FuncSeinfer on 7 widely used ICPs (S7comm, UMAS-A, UMAS-B, Modbus, PCCC, Melsoft, and Fins) in 9 real-world PLCs from 5 leading manufacturers, achieving 100% accuracy in function code field inference, outperforming state-of-the-art PRE tools such as Netzob, FieldHunter, Netplier, and FSIBP. Furthermore, FuncSeinfer can accurately identify the semantic meaning of 78 function codes, which is an average improvement of 20.2% compared to Wireshark. This study fills the gap in functional semantics within ICP reverse engineering, enabling in-depth security analysis without requiring protocol specifications.
Yahui Yang, Yangyang Geng, Man Zhou 0006, Zhuo Lv, Liupeng He
IEEE Internet Things J.2
2026 An Automated Semantic Analysis Framework for Controller Variables Based on Network Traffic
abstract
Programmable logic controllers (PLCs) play a crucial role in various industrial manufacturing processes. Recent attack events show that attackers have a strong interest in controller variables of PLCs, including the device status and internal program logic. Detecting anomalous messages targeting PLC controller variables, which relies on the analysis of controller variable semantics, has proven to be an effective method for identifying such attacks. However, the proprietary nature of industrial control protocols (ICPs) poses a challenge to extracting the required semantics. In this paper, we propose an automated framework namedSePannerto extract the semantics of controller variables from proprietary ICPs based on network traffic. Specifically, we first collect multiple groups of interaction traffic of PLCs and perform the starting-aligned comparisons on them to locate the semantic fields directly. Then, we identify and investigate a new problem in semantic extraction — interference resulting from misordered messages — and propose a set of filtering criteria to eliminate it effectively. We evaluate SePanner using the S7COMM protocol, and the results indicate that SePanner can successfully extract the semantics of controller variables with 100% accuracy. Additionally, we employ SePanner to analyze 7 proprietary ICPs, successfully extracting the semantics of 63 controller variables and their 134 states. Additionally, we demonstrate the extensive applications of SePanner in multiple ICS security scenarios and present its better performance compared with existing ICP semantic analyzing tools.
Zeyu Yang 0001, Zhenyong Zhang, Yangyang Geng, Ruilong Deng, Peng Cheng 0001, Jiming Chen 0001, Jianying Zhou 0001
IEEE Trans. Dependable Secur. Comput.4
2026 CCPA via Load Redistribution: Sequential Strategies and Vulnerability Analysis in Power Systems
abstract
Coordinated cyber-physical attacks (CCPAs) pose a critical threat to the secure operation of smart power grids. While existing studies often assume simultaneous or sequence-agnostic attack strategies, this paper proposes a sequence-aware CCPA framework that explicitly models the temporal coupling between cyber manipulation and physical sabotage. We enhance the classical load redistribution attack (LRA) by addressing four key limitations: detectability due to infeasible power flows, violation of power balance, omission of post-attack system response, and insensitivity to attack sequencing. Specifically, we formulate two distinct bilevel attack models, namely Cyber-to-Physical (C→P) and Physical-to-Cyber (P→C), and solve them via an exact KKT-based Mixed-Integer Linear Programming (MILP) reformulation and a scalable Benders decomposition (BD) framework. Experiments on IEEE 14-, 57-, and 118-bus systems demonstrate that C→P attacks induce significantly more line overloads than P→C, validating the heightened risk of cyber-initiated cascades. Moreover, our BD approach accurately identifies spatial vulnerability hotspots with high fidelity, even when the physical attack budget is extended fromRp= 1 toRp= 2, confirming the framework’s scalability and practical relevance. The results provide actionable insights for adaptive grid protection against sophisticated, sequential threats.
Huihui Huang, YunKai Song, Ruilong Deng, Yangyang Geng
IEEE Trans. Inf. Forensics Secur.5
2025 Physical semantic inference method for industrial control proprietary protocol data fields
Yahui Yang, Yangyang Geng, Man Zhou 0006
Comput. Secur.2
2025 Unmanned aerial vehicle swarm-assisted reliable federated learning for traffic flow prediction
Man Zhou 0006, Lansheng Han, Yangyang Geng
Future Gener. Comput. Syst.3
2024 Reverse Engineering Industrial Protocols Driven By Control Fields
abstract
Industrial protocols are widely used in Industrial Control Systems (ICSs) to network physical devices, thus playing a crucial role in securing ICSs. However, most commercial industrial protocols are proprietary and owned by their vendors, which impedes the implementation of protections against cyber threats. In this paper, we design REInPro to Reverse Engineer Industrial Protocols. REInPro is inspired by the fact that the structure of industrial protocols can be determined by a particular field referred to control field. By applying a probabilistic model of network traffic behavior, REInPro automatically identifies the control field and groups the associated network traffic into clusters. REInPro then infers critical semantics of industrial protocols by differentiating the features of corresponding protocol fields. We have experimentally implemented and evaluated REInPro using 8 different industrial protocols across 6 Programmable Logic Controllers (PLCs) belonging to 5 original equipment manufacturers. The experimental results show REInPro to reverse-engineer the formats and semantics of industrial protocols with an average correctness/perfection of 0.70/0.58 and 0.96/0.39.
Zeyu Yang 0001, Yangyang Geng, Hengye Zhu, Peng Cheng 0001, Jiming Chen 0001
INFOCOM3
2024 Control Logic Attack Detection and Forensics Through Reverse-Engineering and Verifying PLC Control Applications
abstract
Industrial control systems (ICSs) are prevalent in critical infrastructures, where programmable logic controllers (PLCs) and physical instruments are integrated. However, multiple successful attacks against PLC control logic programs have caused significant damage to ICSs, which has led to an urgent need for detection and forensics of such attacks. Although several off-the-shelf defending mechanisms have been presented in the past, few of them can detect and locate the control logic attacks at run time. In this article, we propose a practical and automatic control logic attack detection and forensics framework (CLADF) to conduct control logic attack detection and forensics in ICSs. Specifically, the core of CLADF includes: 1) a control application extraction module to extract PLC binary control applications by simulating PLC normal upload functionality; 2) a control application reverse engineering module to disassemble binary control applications; and 3) an attack detection and forensics module for verifying the integrity of PLC control applications, recovering the normal control application, and locating the modified control instructions. We extensively evaluated CLADF in five different application scenarios and two real-world Schneider PLCs. For each PLC, we generated three types of 150 mutated control logic attacks. The results demonstrate that CLADF can effectively extract the run-time binary control application in different application scenarios and disassemble these binary control applications into assembly instructions. Moreover, CLADF can accurately detect the attacks and locate the modified subroutines.
Yangyang Geng, Rongkuan Ma, Mufeng Wang, Yuqi Chen 0001
IEEE Internet Things J.1
2023 SePanner: Analyzing Semantics of Controller Variables in Industrial Control Systems based on Network Traffic
abstract
Programmable logic controllers (PLCs), the essential components of critical infrastructure, play a crucial role in various industrial manufacturing processes. Recent attack events show that attackers have a strong interest in tampering with the controller variables, such as the device status and internal program logic. A typical attack strategy is that the attackers just send malicious network traffic of industrial control protocols (ICPs) to change the controller variables of PLCs. To defend against this attack, a lot of countermeasures have been proposed to detect anomalies in network traffic based on the semantic analysis.
Zeyu Yang 0001, Zhenyong Zhang, Yangyang Geng, Ruilong Deng, Peng Cheng 0001, Jiming Chen 0001, Jianying Zhou 0001
ACSAC4
2023 Defending Cyber-Physical Systems Through Reverse-Engineering-Based Memory Sanity Check
abstract
Cyber–physical systems (CPSs) are ubiquitous in critical infrastructures, where programmable logic controllers (PLCs) and physical components intertwine. However, multiple successful attacks targeting safety-related CPSs, in particular the PLCs, manifest their vulnerability toward malicious cyber attacks, which may cause significant damage consequently. Though several kinds of defending techniques exist in the literature, few of them can be practically and widely applied to real-world CPSs equipped with PLCs from leading vendors, primarily due to the lack of specific hardware or unrealistic defense assumptions. In this article, we propose PLC-READER, a practical memory attacks detection and response framework to secure the CPS. The core of PLC-READER includes: 1) a comprehensive semantic analysis solution specifically for PLC’s proprietary protocol based on software reverse engineering and network traffic difference analysis and 2) a fine-grained memory structure analysis solution to identify the critical memory data. Based on the results of such reverse engineering, PLC-READER further performs sanity checks for the PLC’s critical memory by periodically checking the hash values and dynamic checksum values of these memory data. We extensively evaluated PLC-READER against four types of 366 different memory attacks, with some newly developed ones which got six CVE IDs from Schneider and Rockwell, by analyzing three kinds of proprietary protocols and six kinds of memory structures in six kinds of real-world PLCs from three leading manufacturers. The results demonstrate that the PLC-READER can detect all memory attacks with an accuracy of 100% and perform corresponding emergency responses in time.
Yangyang Geng, Yuqi Chen 0001, Rongkuan Ma, Jingyi Wang 0004, Peng Cheng 0001
IEEE Internet Things J.1
2023 AFall: Wi-Fi-Based Device-Free Fall Detection System Using Spatial Angle of Arrival
abstract
Falling is a common health problem for elderly people. Early detection of falls allows earlier rescue measures to be implemented. Most existing Wi-Fi-based fall detection systems employ learning-based methods, which require large amounts of labeled data for prior training. To address this issue, we in this paper present AFall, a robust model-based fall detection system that does not require prior training for a single person based on Wi-Fi Channel State Information (CSI). Different from previous Wi-Fi-based fall detection systems, we model the relationship between human falls and changes of Angle of Arrival (AoA) of Wi-Fi signals reflected from human body by multiple signal classification (MUSIC) algorithm. In particular, we deploy two receivers in orthogonal spatial layouts to capture diversified AoA information. Since AoA reflected from human body is independent of environments and subjects, the performance of AFall can remain stable when the environment changes slightly, which can meet the daily needs of the elderly people. We implement AFall using commodity Wi-Fi devices and evaluate it in five different indoor environments. The experimental results demonstrate that AFall achieves an average accuracy of 84.31% and an average F1 score of 84.56%.
Wei Yang 0011, Yang Xu 0020, Yangyang Geng, Bangzhou Xin, Liusheng Huang
IEEE Trans. Mob. Comput.4
2022 Federated synthetic data generation with differential privacy
Bangzhou Xin, Yangyang Geng, Wei Yang 0011, Shaowei Wang 0003, Liusheng Huang
Neurocomputing2
2022 Achieving Secure and Dynamic Range Queries Over Encrypted Cloud Data
abstract
Cloud computing is motivating data owners to outsource their databases to the cloud. However, for privacy concerns, the sensitive data has to be encrypted before outsourcing, which inevitably posts a challenging task for effective data utilization. Existing work either focuses on keyword searches, or suffers from inadequate security guarantees or inefficiency. In this paper, we concentrate on multi-dimensional range queries over dynamic encrypted cloud data. We first propose a tree-based private range query scheme over dynamic encrypted cloud data (TRQED), which supports faster-than-linear range queries and protects single-dimensional privacy. Then, we discuss the defects of TRQED in terms of privacy-preservation. We modify the framework of the system by adopting a two-server model and put forward a safer range query scheme, called TRQED$^{+}$. By newly designed secure node query (SNQ) and secure point query (SPQ), we propose the perturbation-based oblivious R-tree traversal (ORT) operation to preserve both path pattern and stronger single-dimensional privacy. Finally, we conduct comprehensive experiments on real-world datasets and perform comparisons with existing works to evaluate the performance of the proposed schemes. Experimental results show that our TRQED and TRQED$^+$surpass the state-of-the-art methods in privacy-preservation level and efficiency.
Wei Yang 0011, Yangyang Geng, Xike Xie, Liusheng Huang
IEEE Trans. Knowl. Data Eng.2
2021 Private FLI: Anti-Gradient Leakage Recovery Data Privacy Architecture
abstract
While machine learning brings convenience, it also faces the issue of data privacy. For privacy issues, most researches focus on implementing homomorphic encryption or differential privacy to protect data, while ignoring the potential threats caused by the leakage of model parameters. However, a malicious attacker can still recover sensitive data information through model parameters. On the one hand, traditional methods cannot take both high accuracy and low computation time into account. On the other hand, they cannot resist the reconstruction attack from the model's parameter. In order to address this problem, this paper designs a privacy protection framework named FLI, which is inspired by public key infrastructure. In FLI, all participants and the server are trained and aggregated under one framework based on federated learning, which includes key exchange and shares with the idea of homomorphic encryption. Under the algorithm we design, the malicious adversary cannot recover effective information after obtaining the transformed parameters, while the server can still perform effective parameter aggregation. To evaluate the performance of FLI, we conduct extensive experiments. The experimental results show that the computation time is within an acceptable range while ensuring high accuracy.
Huichao Wang, Wei Yang 0011, Bangzhou Xin, Yangyang Geng, Zhenbo Shi, Liusheng Huang
IJCNN4
2020 Private FL-GAN: Differential Privacy Synthetic Data Generation Based on Federated Learning
abstract
Generative Adversarial Network (GAN) has already made a big splash in the field of generating realistic "fake" data. However, when data is distributed and data-holders are reluctant to share data for privacy reasons, GAN’s training is difficult. To address this issue, we propose private FL-GAN, a differential privacy generative adversarial network model based on federated learning. By strategically combining the Lipschitz limit with the differential privacy sensitivity, the model can generate high-quality synthetic data without sacrificing the privacy of the training data. We theoretically prove that private FL-GAN can provide strict privacy guarantee with differential privacy, and experimentally demonstrate our model can generate satisfactory data.
Bangzhou Xin, Wei Yang 0011, Yangyang Geng, Shaowei Wang 0003, Liusheng Huang
ICASSP3
2020 TransNet: Training Privacy-Preserving Neural Network over Transformed Layer
Qijian He, Wei Yang 0011, Bingren Chen, Yangyang Geng, Liusheng Huang
Proc. VLDB Endow.4
2018 COUSTIC: Combinatorial Double Auction for Crowd Sensing Task Assignment in Device-to-Device Clouds
Yutong Zhai, Liusheng Huang, Long Chen 0006, Yangyang Geng
ICA3PP (1)5