EDBT 2026 Demo / reviewers in the wild / expert
Hui Shu
dblp:05/7742
· DBLP profile ↗
27ranked-venue papers
1as first author
26since 2021 · last 2026
0000-0002-2797-1355ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 12 · 12 since 2021Software engineering, systems software and programming languages · 7 · 7 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TDTCE: generating adversarial instances against NIDS via diffusion modelabstractAbstract The integration of machine learning (ML), particularly deep learning (DL), in network intrusion detection systems (NIDSs) has witnessed exponential growth. Notably, these intelligent systems exhibit significant vulnerability to adversarial manipulations, especially in security-critical environments. A critical limitation of conventional adversarial generation approaches—predominantly relying on generative adversarial networks (GANs)—lies in their inability to produce diverse traffic patterns. This shortcoming allows detectors to quickly identify such instances after retraining. To bridge this gap, we propose an adversarial attack framework, TDTCE (Traffic Diffusion Models with Transformers against NIDS in Constrained Environments), which surpasses previous efforts in the following aspects: (i) Generative diversity. TDTCE utilizes a Transformer-based diffusion model to generate adversarial instances that more effectively deceive NIDSs; (ii) Practical feasibility. By restricting prior knowledge, gray-box and black-box attacks are conducted in real-world scenarios and a mapping function is utilized to adjust generated instances for adherence to traffic protocol standards. A comprehensive evaluation of TDTCE using extensive datasets demonstrates that the proposed framework achieves significant gains in adaptive pattern diversity with improvements of up to 79.4% over GAN-based approaches. The framework achieves a minimum detection rate of 9.16%, with an escape increase rate of 89.64%–16.52% higher than state-of-the-art baselines. ZhongHang Sui, Hui Shu |
Comput. J. | 3 |
| 2026 | SIGMA: An interpretable and efficient clean-label adversarial framework for evaluating the robustness of NIDS
ZhongHang Sui, Hui Shu |
Comput. Networks | 3 |
| 2026 | ChatPRE: Knowledge-aware protocol analysis with LLMs for intelligent segmentation
Guoyu Huo, Yuyao Huang 0003, Hui Shu |
J. Netw. Comput. Appl. | 4 |
| 2026 | GAEDM: Genetic Algorithm-Enhanced Static Analysis for Detection of API Hashing Obfuscation in MalwareabstractMalware authors increasingly exploit API Hashing to create “invisible” system calls, replacing explicit function names with dynamically computed hashes that evade detection systems. This sophisticated obfuscation technique poses three critical challenges: accurately identifying hash functions within obfuscated code, linking computed hashes to their corresponding API calls, and detecting the growing diversity of hash algorithm variants. Existing rule-based approaches fail against these adaptive threats and cannot identify modern hash variants. We propose GAEDM , a novel framework that combines deep learning with program analysis to address these challenges. Our key innovation integrates static taint analysis with a genetic algorithm-enhanced assembly language model that generates diverse training variants, enabling robust detection of previously unseen obfuscation patterns. Experimental evaluation demonstrates that GAEDM achieves 91.9% MRR and 94.6% Recall@k in hash function identification, representing improvements of 18.4% and 8.2% respectively over state-of-the-art methods. GAEDM detects sophisticated obfuscation patterns that completely evade existing approaches, enabling security analysts to uncover previously undetectable threats and significantly advancing malware defense capabilities. Hui Shu, Zihan Sha, Xiaobing Xiong |
ACM Trans. Priv. Secur. | 2 |
| 2026 | SemAder: Evading LLM-Based Binary Code Analysis via Structure-Semantics Joint InductionabstractWith the rapid advancement of artificial intelligence (AI), particularly the widespread adoption of large language models (LLMs) in code comprehension and analysis, their strong semantic parsing capabilities have introduced new threats to software security. Attackers can exploit LLMs to reverse-engineer the deeper semantic logic of code, steal core algorithms, or uncover vulnerabilities, thereby endangering software intellectual property and system security. This work introduces SemAder , a structure–semantics joint induction framework that generates adversarial yet function-preserving binaries to mislead LLMs’ functional judgments in binary analysis, thereby reducing the reliability of LLM-assisted semantic analysis during reverse engineering. SemAder comprises three core components: (1) a control-flow-labeled induced corpus annotated with structural tags and code semantics; (2) a hybrid similarity-driven corpus selection mechanism that favors structural proximity with semantic divergence; and (3) a reinforcement-learning-driven semantic fusion pipeline that incorporates constant externalization and context-aware semantic enhancement to strengthen induction against high-capability LLMs. Experimental results across eight LLM evaluators demonstrate that SemAder consistently shifts model predictions toward the induced target category, achieving an average induction gap of 0.77 and maintaining effectiveness under adversarial prompt variants and multi-agent post-processing workflows. SemAder also misleads the CLAP code classification model (-63.5% original-class confidence) and reduces similarity scores across four binary similarity detectors (Asm2Vec, BinDiff, SAFE, Gemini) to an average of 0.51, with only 12.7% average binary size increase and 8.8% average runtime overhead. Hui Shu, Xiaobing Xiong, Ju Yang |
ACM Trans. Priv. Secur. | 3 |
| 2026 | HyRES: Recovering Data Structures in Binaries via Semantic Enhanced Hybrid ReasoningabstractBinary reverse engineering is pivotal in the realm of cybersecurity, enabling critical applications such as malware analysis, legacy code hardening, and vulnerability detection. However, the challenge of recovering structural information from binaries, especially stripped ones, persists due to the significant loss of variable boundaries, types, names, and dataflow information during compilation. In this article, we introduce Hy brid RE asoning for S tructure Recovery ( HyRES ), an innovative hybrid reasoning technique that energizes static analysis, Large Language Model (LLM), and heuristic methods to recover data structures from stripped binaries. It analyzes the structure layout and proficiently infer its semantics via LLM, and utilizes semantics to perform semantic-enhanced structure aggregation, which overcomes the need for complete dataflow. HyRES outperforms State-of-the-Art (SOTA) solutions in terms of structure pointer identification and layout recovery. Specifically, HyRES achieves 65.1% higher recall and 33.4% higher accuracy than the SOTA, while also being 64.2% faster than existing SOTA solutions. Comprehensive experiments demonstrate HyRES ’s superior performance and practical utility in real-world reverse engineering tasks, marking a significant advancement in binary analysis. Zihan Sha, Hui Shu, Hao Wang 0226, Chao Zhang 0008 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | MALPRE: Malware Protocol Reverse Engineering through Code Slicing and Agentic WorkflowabstractClarifying malware communication protocols is critical for enhancing system security. Existing protocol reverse engineering (PRE) methods lack effective strategies, failing to recover protocol structures or infer precise field semantics. To address these challenges, we propose MALPRE, an execution trace-based PRE framework that integrates precise program analysis with large language models (LLMs) for automated malware protocol format recovery. MALPRE first embeds and hierarchically clusters the code slices to restore message formats. It then introduces a multi-agentic workflow comprising code analyst, malware expert, and protocol puzzler roles to collaboratively infer field semantics. Evaluations on the dataset containing six popular malware frameworks demonstrate that MALPRE outperforms the state-of-the-art methods—including four conventional tools (e.g., BINPRE) and three LLM-based approaches (e.g., DEGPT)—by 23.8% (F1) in field structure recovery and 4.8% (FSS-score) in semantic inference. MALPRE has successfully analyzed APT backdoor communications and emerging botnets, with two extracted traffic rules assigned by Open ET Ruleset. Yuyao Huang 0003, Hui Shu, Guoyu Huo |
ISSRE | 3 |
| 2025 | SPFuzz: Program-State-Aware Fuzzing for Mail ProtocolsabstractMail has become a crucial tool in daily work and communication, making the security testing of mail protocols essential for identifying potential vulnerabilities. Fuzzing, as one of the most widely adopted vulnerability discovery techniques, has evolved into an automated and mature method extensively applied in both software and protocol testing across industry. Recently, several fuzzers targeting network protocols have been proposed; however, they suffer from notable limitations. These include insufficient or imprecise representations of protocol states, often relying on generalized state models that fail to capture the specific internal states of individual protocols.In this paper, we address these limitations by proposing SPFuzz, a state-aware fuzzing approach for mail protocols based on program states. We investigate typical implementations of mail protocol services to determine how their internal program states interact with clients. First, we model the program state using key variables; then, we perform compile-time instrumentation to track these variables and infer runtime state transitions of the mail protocol implementations. Finally, we leverage the state information to prioritize test cases that are more likely to trigger new states and apply state feedback to refine the mutation strategy of a coverage-guided fuzzer. We implemented a prototype of SPFuzz and evaluated it on multiple real-world mail protocol programs. Experimental results demonstrate that SPFuzz significantly outperforms state-of-the-art fuzzers, including AFLNET and StateAFL, in terms of both code and state coverage. Specifically, SPFuzz achieves average improvements of 148.75% in the number of discovered states, 7.09% in state space coverage, 43.98% in map density, and 11.48% in branch coverage. Hangzhou Fei, Xiaobing Xiong, Hui Shu |
TrustCom | 3 |
| 2025 | VarAgent: LLM-Enhanced Variable Name Recovery in Binaries via Reinforcement Learning and Semantic FusionabstractVariable name recovery is a fundamental task in software reverse engineering, crucial for various cybersecurity applications including binary code understanding, protocol format recovery, and malware analysis. The core challenge lies in the semantic loss during compilation to binary code. Although intelligence-driven methods have shown progress, several issues persist: generated names remain vague, semantic inconsistency occurs across different contexts, and inference ordering lacks systematic planning. In this paper, we propose VarAgent, a variable naming agent driven by two domain-enhanced large language models: one focuses on inferring variable names within functions through reinforcement learning from reverse-engineering feedback, while the other propagates and fuses semantics across variables using graph neural networks embedded within the LLM. We also design a confusion-guided planner that mimics expert reverse engineers’ analysis strategies. Experimental results demonstrate that VarAgent achieves average precision and F1 scores of 34.9% and 34.7% on classic and novel datasets, outperforming state-of-the-art tools by 4.0%–15.4% (including ReSym, GenNm, DeGPT, and VarBERT). We further demonstrate its effectiveness in supporting downstream tasks, including protocol field inference, binary semantic search, and malware summarization. Yuyao Huang 0003, Hui Shu |
TrustCom | 2 |
| 2025 | RL-Optimized Lightweight Obfuscation Against Binary Code Similarity DetectionabstractBinary Code Similarity Detection (BCSD) identifies functions with similar functionality across binaries. Reverse engineers leverage BCSD techniques to locate critical functions in software, thereby compromising associated functional modules. This poses significant challenges for software protection. Although code obfuscation can counter BCSD, existing methods typically apply global, undifferentiated obfuscation, leading to unnecessary code size and runtime overhead. This paper proposes ReinObf, a reinforcement learning-optimized method for lightweight and differentiated obfuscation. In the LLVM intermediate representation of functions, we implement a custom pass to extract statistical and structural features from basic blocks and construct Attributed Control Flow Graphs (ACFGs). A hierarchical RL agent encodes ACFGs using graph neural networks, selects critical blocks and optimal obfuscation methods, then iteratively generates optimized obfuscation sequences that maximize evasion effectiveness while minimizing overhead. Experimental validation demonstrates that ReinObf effectively evades BinDiff, jTrans, CLAP, and Gemini, with an average similarity of 0.391. Compared to OLLVM’s combined obfuscation, it reduces code size overhead by 66.3% and runtime overhead by 34.7%, achieving a favorable security-performance tradeoff. Hui Shu |
TrustCom | 2 |
| 2025 | Code obfuscation based on deep integrationabstractAbstract Code obfuscation is essential for software security. However, current obfuscation techniques demonstrate limited resilience against systematic reverse engineering attacks—including taint analysis and code similarity detection. Moreover, these methods often incur considerable resource overheads and recognizable obfuscation features. In this paper, we propose an innovative obfuscation algorithm that integrates the instruction and data flows of two programs at the intermediate representation level. The resulting program maintains the complete functionality of both original programs. This strategy utilizes the static and dynamic features of the parent program to obfuscate the target program. Extracting the target code from the integrated program is a significant challenge, thereby enhancing the target code’s resistance to deobfuscation. We evaluate our algorithm across various metrics: functionality correctness, obfuscation efficiency, protection strength, and resilience to automated reverse engineering techniques. Our comprehensive evaluation demonstrates that our method imposes significantly lower overhead while delivering markedly improved protection effectiveness, marking a significant advancement in software protection. Xiaobing Xiong, Zihan Sha, Hui Shu |
Comput. J. | 3 |
| 2025 | OpTrans: enhancing binary code similarity detection with function inlining re-optimization
Zihan Sha, Chao Zhang 0008, Hao Wang 0003, Hui Shu |
Empir. Softw. Eng. | 7 |
| 2025 | PromeTrans: Bootstrap binary functionality classification with knowledge transferred from pre-trained models
Zihan Sha, Chao Zhang 0008, Hao Wang 0003, Hui Shu |
Empir. Softw. Eng. | 7 |
| 2025 | UniASM: Binary code similarity detection without fine-tuning
Yeming Gu, Hui Shu |
Neurocomputing | 2 |
| 2025 | llasm: Naming Functions in Binaries by Fusing Encoder-only and Decoder-only LLMsabstractPredicting function names in stripped binaries, which requires succinctly summarizing semantics of binary code in natural languages, is a crucial but challenging task. Recently, many machine learning based solutions have been proposed. However, they have poor generalizability, i.e., fail to handle unseen binaries. To advance the state of the art, we present large assembly language Model ( llasm ) , a novel framework which fuses encoder-only and decoder-only LLMs for function name prediction. It refines encoder-only models to preserve more binary information and learn better binary representations. Then it adopts a novel architecture to project the encoding to the input space of a decoder-only natural language model, which enables it to have better capability of inferring general knowledge and better generalizability. We have evaluated llasm in the BinaryCorp and Debin datasets. llasm outperforms the state-of-the-art function name prediction tools by up to 19.9%, 40.7%, and 36.5% in precision, recall, and F1 score, with significantly better generalizability in unseen binaries. Our case studies further demonstrate the practical use cases of llasm in analyzing real-world malware, showing the usefulness of function name prediction. Zihan Sha, Hao Wang 0226, Hui Shu, Chao Zhang 0008 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Formatted Stateful Greybox Fuzzing of TLS ServerabstractThe TLS protocol is one of the most crucial foundations for ensuring internet security. Consequently, vulnerabilities within the TLS protocol have a significant impact on the Internet security. This paper aims to explore more efficient methods of discovering vulnerabilities in the TLS protocol. Fuzzing stands out as one of the most important techniques for vulnerability discovery in the TLS protocol. To tackle the high complexity of the TLS protocol, stateful greybox fuzzers such as AFLnet have been introduced to enable stateful fuzzing of TLS servers. However, these mutation-based fuzzers often encounter chal-lenges in preserving the message format information during the mutation process, which can undermine the testing results. As a result, this paper proposes a novel approach that incorporates a formatted mutation strategy into the stateful greybox fuzzing process, with the aim of achieving more efficient mutation results. The evaluation process involves four mainstream fuzzers, with OpenSSL's TLS server serving as the target. The results demonstrate that the proposed method significantly enhances the quality of generated seeds, code coverage, and state coverage across all four fuzzers. Jiangan Ji, Hui Shu, Zheming Li, Tieming Liu, Chao Zhang 0008 |
ICST | 3 |
| 2024 | SLI-YOLO: A Lightweight Unauthorized Login Detection Model Based on Multiscale Convolutional AttentionabstractBased on camera-based security surveillance is generally used for physical site security supervision, and it can also be used for network security monitoring, such as monitoring and alerting unauthorized login situations of important systems in specific occasions. Considering the limited computational resources of surveillance cameras, lightweight models need to be designed for efficient operation in constrained environments. Therefore, in this paper, we design a lightweight unauthorized login detection method SLI (Software Login Image)-YOLO for the embedded side. Firstly, SLI-YOLO is based on YOLOv5 as the basic framework, based on which a new attention mechanism CPMS is introduced, which fuses multi-scale convolutional channel attention and multi-scale depth-separable convolutional spatial attention to improve the model’s ability to extract feature information. Secondly, the SPPF structure is redesigned to obtain the global view information and mitigate the effects caused by different scale sizes. Finally, by employing MobileViTv3 and Slim-neck as alternatives to the YOLOv5s backbone extraction network and neck layer, the number of parameters in the model has been reduced, achieving model lightweightness. The experimental results show that the number of parameters and computation volume of the detection algorithm is reduced by 69.4 % and 61.8 % compared with YOLOv5s, the recognition accuracy of the login page reaches 94.5 %, and the mAP of the proposed algorithm only decreases by 0.7 % compared with that of YOLOv5s. Meanwhile, when the model is deployed in the embedded platform of Raspberry Pi, the time for each recognition is around 1.2s, which can achieve ideal performance and lightweight in detection. Hui Shu |
IJCNN | 2 |
| 2024 | SBCM: Semantic-Driven Reverse Engineering Framework for Binary Code ModularizationabstractSoftware reverse analysis is a key technology in the field of cyber-security. With the increasing scale and complexity of software, this technology is facing great challenges. Binary code modularization (BCM), as the basic work of software reverse, plays an important role in extracting semantics, narrowing the analysis scope and locating key position. The semantics of strings underlying code is a significant hint, with being processed using natural language processing and artificial intelligence technology help to reverse analysis effectively. However, most of the existing modularization methods ignore these semantics, which limits the in-depth understanding of binary code. This paper proposes a semantic-driven reverse engineering framework for binary code modularization (SBCM). Firstly, the rich string is extracted from the binary file into a large language model for semantics analysis. Then, the semantic information of the string is combined with the control flow graph to construct the function semantic graph (FSG). Subsequently, a function summary is generated based on the FSG. Finally, semantic embedding is generated for Summaries and semantic-driven integrated clustering is carried out to realize binary code modularization. The experiment results show that SBCM improves the F1 value by 12.6% on average compared with the existing methods, which proves its effectiveness and superiority in binary code modularization. Shuang Duan, Hui Shu, Zihan Sha, Yuyao Huang 0003 |
TrustCom | 2 |
| 2024 | Malware2ATT&CK: A sophisticated model for mapping malware to ATT&CK techniques
Huaqi Sun, Hui Shu, Yuntian Zhao, Yuyao Huang 0003 |
Comput. Secur. | 2 |
| 2023 | BinAIV: Semantic-enhanced vulnerability detection for Linux x86 binaries
Yeming Gu, Hui Shu |
Comput. Secur. | 2 |
| 2023 | PeerRemove: An adaptive node removal strategy for P2P botnet based on deep reinforcement learning
Hui Shu |
Comput. Secur. | 2 |
| 2022 | Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning
Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Hui Shu, Jinde Song, Peng Lan, Guohuan Xu, Fei Wu 0001, Shaojie Tang 0001, Fan Wu 0006, Guihai Chen |
OSDI | 12 |
| 2022 | Protocol Reverse-Engineering Methods and Tools: A Survey
Yuyao Huang 0003, Hui Shu, Yan Guang |
Comput. Commun. | 2 |
| 2022 | MinSIB: Minimized static instrumentation for fuzzing binaries
Yeming Gu, Hui Shu, Pan Yang 0024, Rongkuan Ma |
Comput. Secur. | 2 |
| 2022 | DeMal: Module decomposition of malware based on community discovery
Yuyao Huang 0003, Hui Shu |
Comput. Secur. | 2 |
| 2022 | Model of Execution Trace Obfuscation Between ThreadsabstractAdvanced reverse analysis tools have significantly improved the ability of attackers to crack software via dynamic analysis techniques, such as symbol execution and taint analysis. These techniques are widely used in malicious fields such as vulnerability exploitation or theft of intellectual property. In this paper, we present an obfuscation strategy called “execution trace obfuscation,” wherein the program execution trace repeatedly switches between multiple threads. Our technique realizes equivalent code transformation by abstracting the obfuscation problems into pruning, cloning, and coloring problems in graph theory. Based on this, we further propose the cascade encryption of a function that depends on execution trace information with a key derived from the function address calculation process, followed by removing this key from the program. We have implemented a compiler-level system that inputs a source program and automatically generates an obfuscated file. Finally, random test proves the universality of obfuscation algorithm and verify the system’s performance. Results shows that our system can effectively interfere advanced reverse analysis tools. Zihan Sha, Hui Shu, Xiaobing Xiong |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2016 | A study on the equilibrium of regional industrial spatial distribution based on grey prediction modelabstractThis paper discusses the industrial space layout optimization under different TFP evolutionary paths simulated by grey prediction models. We construct a simulation and analysis model about the progress of regional industrial layout adjustment in project level. On the basis of the model, we analyze the economic-ecology balance of the regional industrial space layout. The results show that the short-term evolution trends of the TFP of Chinese industries are more along the path of traditional GM (1, 1) model. The economic and ecological balance of regional industrial space layout is in the reasonable range under this path, while the regional economic-ecology balance is greatly improved under the TFP evolution path provided by the grey Versulst model. Hui Shu, Xiong Ping-ping |
SMC | 1 |