Hao Wang 0226

dblp:181/2812-226 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-0536-5039ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HyRES: Recovering Data Structures in Binaries via Semantic Enhanced Hybrid Reasoning
abstract
Binary reverse engineering is pivotal in the realm of cybersecurity, enabling critical applications such as malware analysis, legacy code hardening, and vulnerability detection. However, the challenge of recovering structural information from binaries, especially stripped ones, persists due to the significant loss of variable boundaries, types, names, and dataflow information during compilation. In this article, we introduce Hy brid RE asoning for S tructure Recovery ( HyRES ), an innovative hybrid reasoning technique that energizes static analysis, Large Language Model (LLM), and heuristic methods to recover data structures from stripped binaries. It analyzes the structure layout and proficiently infer its semantics via LLM, and utilizes semantics to perform semantic-enhanced structure aggregation, which overcomes the need for complete dataflow. HyRES outperforms State-of-the-Art (SOTA) solutions in terms of structure pointer identification and layout recovery. Specifically, HyRES achieves 65.1% higher recall and 33.4% higher accuracy than the SOTA, while also being 64.2% faster than existing SOTA solutions. Comprehensive experiments demonstrate HyRES ’s superior performance and practical utility in real-world reverse engineering tasks, marking a significant advancement in binary analysis.
Zihan Sha, Hui Shu, Hao Wang 0226, Chao Zhang 0008
ACM Trans. Softw. Eng. Methodol.3
2025 SmartTrans: Advanced Similarity Analysis for Detecting Vulnerabilities in Ethereum Smart Contracts
abstract
In the ever-evolving landscape of Ethereum smart contracts, the specter of vulnerabilities intensified by code reuse presents a significant challenge to the security of the blockchain. Recent studies employ deep learning for similarity analysis to identify these vulnerabilities, yet their effectiveness wanes as the volume of analyzed code increases. This article introducesSmartTrans, an advanced similarity analysis model designed to efficiently and accurately retrieve similar vulnerabilities within Ethereum bytecodes. Leveraging a novel jump-aware Transformer-based model, our approach captures the semantics and control flow of bytecodes. It not only refines the representation of functions by integrating program analysis with natural language processing techniques but also innovates a contract-level similarity detection scheme tailored for the expansive scale of contracts. Our experiments show thatSmartTransoutperforms state-of-the-art techniques at both function and contract levels, proving its capability to detect n-day vulnerabilities across Ethereum bytecodes accurately. Vulnerabilities recalling experiments show thatSmartTransachieves 95.43% and 99.37% accuracy at two levels. Furthermore, we stand out as the first work to retrieve N-day vulnerabilities across the Ethereum bytecode corpus, unveiling 4,988 vulnerable contracts. Our methodology secures an accuracy of 88.60%, which is 1.30 times higher than the best baseline.
Hao Wang 0226, Yuchen Zhou 0007, Taiyu Wong, Jialai Wang, Chao Zhang 0008
IEEE Trans. Dependable Secur. Comput.2
2025 llasm: Naming Functions in Binaries by Fusing Encoder-only and Decoder-only LLMs
abstract
Predicting function names in stripped binaries, which requires succinctly summarizing semantics of binary code in natural languages, is a crucial but challenging task. Recently, many machine learning based solutions have been proposed. However, they have poor generalizability, i.e., fail to handle unseen binaries. To advance the state of the art, we present large assembly language Model ( llasm ) , a novel framework which fuses encoder-only and decoder-only LLMs for function name prediction. It refines encoder-only models to preserve more binary information and learn better binary representations. Then it adopts a novel architecture to project the encoding to the input space of a decoder-only natural language model, which enables it to have better capability of inferring general knowledge and better generalizability. We have evaluated llasm in the BinaryCorp and Debin datasets. llasm outperforms the state-of-the-art function name prediction tools by up to 19.9%, 40.7%, and 36.5% in precision, recall, and F1 score, with significantly better generalizability in unseen binaries. Our case studies further demonstrate the practical use cases of llasm in analyzing real-world malware, showing the usefulness of function name prediction.
Zihan Sha, Hao Wang 0226, Hui Shu, Chao Zhang 0008
ACM Trans. Softw. Eng. Methodol.2
2024 CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
abstract
Binary code representation learning has shown significant performance in binary analysis tasks. But existing solutions often have poor transferability, particularly in few-shot and zero-shot scenarios where few or no training samples are available for the tasks. To address this problem, we present CLAP (Contrastive Language-Assembly Pre-training), which employs natural language supervision to learn better representations of binary code (i.e., assembly code) and get better transferability. At the core, our approach boosts superior transfer learning capabilities by effectively aligning binary code with their semantics explanations (in natural language), resulting a model able to generate better embeddings for binary code. To enable this alignment training, we then propose an efficient dataset engine that could automatically generate a large and diverse dataset comprising of binary code and corresponding natural language explanations. We have generated 195 million pairs of binary code and explanations and trained a prototype of CLAP. The evaluations of CLAP across various downstream tasks in binary analysis all demonstrate exceptional performance. Notably, without any task-specific training, CLAP is often competitive with a fully supervised baseline, showing excellent transferability.
Hao Wang 0226, Chao Zhang 0008, Zihan Sha, Yuchen Zhou 0007, Wenyu Zhu, Wenju Sun, Han Qiu 0001, Xi Xiao 0001
ISSTA1
2024 CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity Detection
abstract
Binary code similarity detection (BCSD) is a fundamental technique for various applications. Many BCSD solutions have been proposed recently, which mostly are embedding-based, but have shown limited accuracy and efficiency especially when the volume of target binaries to search is large. To address this issue, we propose a cost-effective BCSD framework, CEBin, which fuses embedding-based and comparison-based approaches to significantly improve accuracy while minimizing overheads. Specifically, CEBin utilizes a refined embedding-based approach to extract features of target code, which efficiently narrows down the scope of candidate similar code and boosts performance. Then, it utilizes a comparison-based approach that performs a pairwise comparison on the candidates to capture more nuanced and complex relationships, which greatly improves the accuracy of similarity detection. By bridging the gap between embedding-based and comparison-based approaches, CEBin is able to provide an effective and efficient solution for detecting similar code (including vulnerable ones) in large-scale software ecosystems. Experimental results on three well-known datasets demonstrate the superiority of CEBin over existing state-of-the-art (SOTA) baselines. To further evaluate the usefulness of BCSD in real world, we construct a large-scale benchmark of vulnerability, offering the first precise evaluation scheme to assess BCSD methods for the 1-day vulnerability detection task. CEBin could identify the similar function from millions of candidate functions in just a few seconds and achieves an impressive recall rate of 85.46% on this more practical but challenging task, which are several order of magnitudes faster and 4.07× better than the best SOTA baseline.
Hao Wang 0226, Chao Zhang 0008, Yuchen Zhou 0007, Han Qiu 0001, Xi Xiao 0001
ISSTA1
2022 jTrans: jump-aware transformer for binary code similarity detection
abstract
Binary code similarity detection (BCSD) has important applications in various fields such as vulnerabilities detection, software component analysis, and reverse engineering. Recent studies have shown that deep neural networks (DNNs) can comprehend instructions or control-flow graphs (CFG) of binary code and support BCSD. In this study, we propose a novel Transformer-based approach, namely jTrans, to learn representations of binary code. It is the first solution that embeds control flow information of binary code into Transformer-based language models, by using a novel jump-aware representation of the analyzed binaries and a newly-designed pre-training task. Additionally, we release to the community a newly-created large dataset of binaries, BinaryCorp, which is the most diverse to date. Evaluation results show that jTrans outperforms state-of-the-art (SOTA) approaches on this more challenging dataset by 30.5% (i.e., from 32.0% to 62.5%). In a real-world task of known vulnerability searching, jTrans achieves a recall that is 2X higher than existing SOTA baselines.
Hao Wang 0226, Wenjie Qu 0001, Gilad Katz, Wenyu Zhu, Han Qiu 0001, Jianwei Zhuge, Chao Zhang 0008
ISSTA1