Zihan Sha

dblp:333/5196 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-1020-9006ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GAEDM: Genetic Algorithm-Enhanced Static Analysis for Detection of API Hashing Obfuscation in Malware
abstract
Malware authors increasingly exploit API Hashing to create “invisible” system calls, replacing explicit function names with dynamically computed hashes that evade detection systems. This sophisticated obfuscation technique poses three critical challenges: accurately identifying hash functions within obfuscated code, linking computed hashes to their corresponding API calls, and detecting the growing diversity of hash algorithm variants. Existing rule-based approaches fail against these adaptive threats and cannot identify modern hash variants. We propose GAEDM , a novel framework that combines deep learning with program analysis to address these challenges. Our key innovation integrates static taint analysis with a genetic algorithm-enhanced assembly language model that generates diverse training variants, enabling robust detection of previously unseen obfuscation patterns. Experimental evaluation demonstrates that GAEDM achieves 91.9% MRR and 94.6% Recall@k in hash function identification, representing improvements of 18.4% and 8.2% respectively over state-of-the-art methods. GAEDM detects sophisticated obfuscation patterns that completely evade existing approaches, enabling security analysts to uncover previously undetectable threats and significantly advancing malware defense capabilities.
Hui Shu, Zihan Sha, Xiaobing Xiong
ACM Trans. Priv. Secur.3
2026 HyRES: Recovering Data Structures in Binaries via Semantic Enhanced Hybrid Reasoning
abstract
Binary reverse engineering is pivotal in the realm of cybersecurity, enabling critical applications such as malware analysis, legacy code hardening, and vulnerability detection. However, the challenge of recovering structural information from binaries, especially stripped ones, persists due to the significant loss of variable boundaries, types, names, and dataflow information during compilation. In this article, we introduce Hy brid RE asoning for S tructure Recovery ( HyRES ), an innovative hybrid reasoning technique that energizes static analysis, Large Language Model (LLM), and heuristic methods to recover data structures from stripped binaries. It analyzes the structure layout and proficiently infer its semantics via LLM, and utilizes semantics to perform semantic-enhanced structure aggregation, which overcomes the need for complete dataflow. HyRES outperforms State-of-the-Art (SOTA) solutions in terms of structure pointer identification and layout recovery. Specifically, HyRES achieves 65.1% higher recall and 33.4% higher accuracy than the SOTA, while also being 64.2% faster than existing SOTA solutions. Comprehensive experiments demonstrate HyRES ’s superior performance and practical utility in real-world reverse engineering tasks, marking a significant advancement in binary analysis.
Zihan Sha, Hui Shu, Hao Wang 0226, Chao Zhang 0008
ACM Trans. Softw. Eng. Methodol.1
2025 Code obfuscation based on deep integration
abstract
Abstract Code obfuscation is essential for software security. However, current obfuscation techniques demonstrate limited resilience against systematic reverse engineering attacks—including taint analysis and code similarity detection. Moreover, these methods often incur considerable resource overheads and recognizable obfuscation features. In this paper, we propose an innovative obfuscation algorithm that integrates the instruction and data flows of two programs at the intermediate representation level. The resulting program maintains the complete functionality of both original programs. This strategy utilizes the static and dynamic features of the parent program to obfuscate the target program. Extracting the target code from the integrated program is a significant challenge, thereby enhancing the target code’s resistance to deobfuscation. We evaluate our algorithm across various metrics: functionality correctness, obfuscation efficiency, protection strength, and resilience to automated reverse engineering techniques. Our comprehensive evaluation demonstrates that our method imposes significantly lower overhead while delivering markedly improved protection effectiveness, marking a significant advancement in software protection.
Xiaobing Xiong, Zihan Sha, Hui Shu
Comput. J.2
2025 OpTrans: enhancing binary code similarity detection with function inlining re-optimization
Zihan Sha, Chao Zhang 0008, Hao Wang 0003, Hui Shu
Empir. Softw. Eng.1
2025 PromeTrans: Bootstrap binary functionality classification with knowledge transferred from pre-trained models
Zihan Sha, Chao Zhang 0008, Hao Wang 0003, Hui Shu
Empir. Softw. Eng.1
2025 Generation method of suspense stories based on new-type inference
Wenjuan Bu, Zhuolun Li, Zihan Sha, Yuntian Zhao
J. Intell. Inf. Syst.3
2025 llasm: Naming Functions in Binaries by Fusing Encoder-only and Decoder-only LLMs
abstract
Predicting function names in stripped binaries, which requires succinctly summarizing semantics of binary code in natural languages, is a crucial but challenging task. Recently, many machine learning based solutions have been proposed. However, they have poor generalizability, i.e., fail to handle unseen binaries. To advance the state of the art, we present large assembly language Model ( llasm ) , a novel framework which fuses encoder-only and decoder-only LLMs for function name prediction. It refines encoder-only models to preserve more binary information and learn better binary representations. Then it adopts a novel architecture to project the encoding to the input space of a decoder-only natural language model, which enables it to have better capability of inferring general knowledge and better generalizability. We have evaluated llasm in the BinaryCorp and Debin datasets. llasm outperforms the state-of-the-art function name prediction tools by up to 19.9%, 40.7%, and 36.5% in precision, recall, and F1 score, with significantly better generalizability in unseen binaries. Our case studies further demonstrate the practical use cases of llasm in analyzing real-world malware, showing the usefulness of function name prediction.
Zihan Sha, Hao Wang 0226, Hui Shu, Chao Zhang 0008
ACM Trans. Softw. Eng. Methodol.1
2024 CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
abstract
Binary code representation learning has shown significant performance in binary analysis tasks. But existing solutions often have poor transferability, particularly in few-shot and zero-shot scenarios where few or no training samples are available for the tasks. To address this problem, we present CLAP (Contrastive Language-Assembly Pre-training), which employs natural language supervision to learn better representations of binary code (i.e., assembly code) and get better transferability. At the core, our approach boosts superior transfer learning capabilities by effectively aligning binary code with their semantics explanations (in natural language), resulting a model able to generate better embeddings for binary code. To enable this alignment training, we then propose an efficient dataset engine that could automatically generate a large and diverse dataset comprising of binary code and corresponding natural language explanations. We have generated 195 million pairs of binary code and explanations and trained a prototype of CLAP. The evaluations of CLAP across various downstream tasks in binary analysis all demonstrate exceptional performance. Notably, without any task-specific training, CLAP is often competitive with a fully supervised baseline, showing excellent transferability.
Hao Wang 0226, Chao Zhang 0008, Zihan Sha, Yuchen Zhou 0007, Wenyu Zhu, Wenju Sun, Han Qiu 0001, Xi Xiao 0001
ISSTA4
2024 SBCM: Semantic-Driven Reverse Engineering Framework for Binary Code Modularization
abstract
Software reverse analysis is a key technology in the field of cyber-security. With the increasing scale and complexity of software, this technology is facing great challenges. Binary code modularization (BCM), as the basic work of software reverse, plays an important role in extracting semantics, narrowing the analysis scope and locating key position. The semantics of strings underlying code is a significant hint, with being processed using natural language processing and artificial intelligence technology help to reverse analysis effectively. However, most of the existing modularization methods ignore these semantics, which limits the in-depth understanding of binary code. This paper proposes a semantic-driven reverse engineering framework for binary code modularization (SBCM). Firstly, the rich string is extracted from the binary file into a large language model for semantics analysis. Then, the semantic information of the string is combined with the control flow graph to construct the function semantic graph (FSG). Subsequently, a function summary is generated based on the FSG. Finally, semantic embedding is generated for Summaries and semantic-driven integrated clustering is carried out to realize binary code modularization. The experiment results show that SBCM improves the F1 value by 12.6% on average compared with the existing methods, which proves its effectiveness and superiority in binary code modularization.
Shuang Duan, Hui Shu, Zihan Sha, Yuyao Huang 0003
TrustCom3
2022 Model of Execution Trace Obfuscation Between Threads
abstract
Advanced reverse analysis tools have significantly improved the ability of attackers to crack software via dynamic analysis techniques, such as symbol execution and taint analysis. These techniques are widely used in malicious fields such as vulnerability exploitation or theft of intellectual property. In this paper, we present an obfuscation strategy called “execution trace obfuscation,” wherein the program execution trace repeatedly switches between multiple threads. Our technique realizes equivalent code transformation by abstracting the obfuscation problems into pruning, cloning, and coloring problems in graph theory. Based on this, we further propose the cascade encryption of a function that depends on execution trace information with a key derived from the function address calculation process, followed by removing this key from the program. We have implemented a compiler-level system that inputs a source program and automatically generates an obfuscated file. Finally, random test proves the universality of obfuscation algorithm and verify the system’s performance. Results shows that our system can effectively interfere advanced reverse analysis tools.
Zihan Sha, Hui Shu, Xiaobing Xiong
IEEE Trans. Dependable Secur. Comput.1