VLDB 2026 Research / reviewers in the wild / expert
Shaoyin Cheng
dblp:21/2249
· DBLP profile ↗
22ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-3992-9509ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An empirical study on the effectiveness of large language models for binary code understanding
Xiuwei Shang, Zhenkan Fu, Shaoyin Cheng, Gangyang Li, Weiming Zhang 0001, Nenghai Yu |
Empir. Softw. Eng. | 3 |
| 2026 | FoC: Figure Out the Cryptographic Functions in Stripped Binaries with LLMsabstractAnalyzing the behavior of cryptographic functions in stripped binaries is a challenging but essential task, which is crucial in software security fields such as malware analysis and legacy code inspection. However, the inherent high logical complexity of cryptographic algorithms makes their analysis more difficult than that of ordinary code, and the general absence of symbolic information in binaries exacerbates this challenge. Existing methods for cryptographic algorithm identification frequently rely on data or structural pattern matching, which limits their generality and effectiveness while requiring substantial manual effort. In response to these challenges, we present F igure o ut the C ryptographic functions (FoC), a novel framework that leverages Large Language Models (LLMs) to identify and analyze cryptographic functions in stripped binaries. In FoC, we first build an LLM-based generative model ( FoC-BinLLM ) to summarize the semantics of cryptographic functions in natural language form, which is intuitively readable to analysts. Subsequently, based on the semantic insights provided by FoC-BinLLM, we further develop a binary code similarity detection model ( FoC-Sim ), which allows analysts to effectively retrieve similar implementations of unknown cryptographic functions from a library of known cryptographic functions. The predictions of generative model like FoC-BinLLM are inherently difficult to reflect minor alterations in binary code, such as those introduced by vulnerability patches. In contrast, the change-sensitive representations generated by FoC-Sim compensate for the shortcomings to some extent. To support the development and evaluation of these models, and to facilitate further research in this domain, we also construct a comprehensive cryptographic binary dataset and introduce an automatic method to create semantic labels for extensive binary functions. Our evaluation results are promising. FoC-BinLLM outperforms ChatGPT by 14.61% on the ROUGE-L score, demonstrating superior capability in summarizing the semantics of cryptographic functions. FoC-Sim also surpasses previous best methods with a 52% higher Recall@1 in retrieving similar cryptographic functions. Beyond these metrics, our method has proven its practical utility in real-world scenarios, including cryptographic-related virus analysis and 1-day vulnerability detection. Xiuwei Shang, Shaoyin Cheng, Shikai Guo, Weiming Zhang 0001, Nenghai Yu |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent SystemabstractLi Hu, Guoqiang Chen, Xiuwei Shang, Shaoyin Cheng, Benlong Wu, LiGangyang LiGangyang, Xu Zhu, Weiming Zhang, Nenghai Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiuwei Shang, Shaoyin Cheng, Benlong Wu, LiGangyang LiGangyang, Weiming Zhang 0001, Nenghai Yu |
ACL (1) | 4 |
| 2025 | WelkIR: Flow-Sensitive Pre-trained Embeddings from Compiler IR for Vulnerability Detection
Xiuwei Shang, Shaoyin Cheng, Weiming Zhang 0001, Nenghai Yu |
ESORICS (3) | 4 |
| 2025 | BinMetric: A Comprehensive Binary Code Analysis Benchmark for Large Language ModelsabstractBinary analysis is crucial for software security, offering insights into compiled programs without source code. As large language models (LLMs) excel in language tasks, their potential for complex decoding binary data structures is growing. However, the lack of standardized benchmarks hinders their evaluation and progress in this domain. To bridge this gap, we introduce BinMetric, a first comprehensive benchmark designed specifically to evaluate LLMs performance on binary analysis tasks. BinMetric comprises 1,000 questions derived from 20 real-world open-source projects across 6 practical binary analysis tasks, including decompilation, code summarization, etc., which reflect actual reverse engineering scenarios. Our empirical study on this benchmark investigates various state-of-the-art LLMs, revealing their strengths and limitations. The findings indicate that while LLMs show strong potential, challenges still exist, particularly in the areas of precise binary lifting and assembly synthesis. In summary, BinMetric makes a significant step forward in measuring binary analysis capabilities of LLMs, establishing a new benchmark leaderboard, and our study offers valuable insights for advancing LLMs in software security. Xiuwei Shang, Shaoyin Cheng, Benlong Wu, Gangyang Li, Weiming Zhang 0001, Nenghai Yu |
IJCAI | 3 |
| 2025 | PseudoFix: Refactoring Distorted Structures in Decompiled C PseudocodeabstractDecompilation can convert binary programs into clear C-style pseudocode, which is of great value in a wide range of security applications. Existing research primarily focuses on recovering symbolic information in pseudocode, such as function names, variable names, and data types, but neglecting structural information. We observe that even when symbolic information is fully preserved, severe and complex structure distortions remain in the pseudocode, greatly impairing code readability and comprehension. In this work, we first systematically investigate structure distortions in decompiled pseudocode, revealing their variation patterns through quantitative analysis. Using open coding, we derive a taxonomy comprising six top-level categories of structure distortions. Building upon this taxonomy, we propose PseudoFix, a novel framework that combines large language models (LLMs) with retrieval-based in-context learning. PseudoFix employs semantic retrieval to select the most relevant few-shot examples that provide structure distortion knowledge, and combines this with the well-structured coding patterns learned by LLMs from vast source code repositories, to efficiently refactor distorted pseudocode. Comprehensive evaluations demonstrate that PseudoFix significantly improves pseudocode readability, achieving up to a 34% reduction in Halstead Complexity Effort and a 105% increase in BLEU-4 score. Notably, it significantly outperforms state-of-the-art approaches in both temporary variable elimination and goto statement removal tasks. Additionally, human evaluations yield consistently positive feedback from users across readability, consistency, and reasonability. Gangyang Li, Xiuwei Shang, Shaoyin Cheng, Weiming Zhang 0001, Nenghai Yu |
ASE | 3 |
| 2025 | The Ghost Navigator: Revisiting the Hidden Vulnerability of Localization in Autonomous Driving
Shaoyin Cheng, Linqing Hu, Jie Zhang 0073, Chengyu Shi, Xingshuo Han, Tianwei Zhang 0004, Yueqiang Cheng, Weiming Zhang 0001 |
USENIX Security Symposium | 2 |
| 2024 | How Far Have We Gone in Binary Code Understanding Using Large Language ModelsabstractBinary code analysis plays a pivotal role in various software security applications, such as software maintenance, malware detection, software vulnerability discovery, patch analysis, etc. However, unlike source code, understanding binary code is challenging for reverse engineers due to the absence of semantic information. Therefore, automated tools are needed to assist human players in interpreting binary code. In recent years, two groups of technologies have shown promising prospects: (1) Deep learning-based technologies have demonstrated competitive results in tasks related to binary code understanding, furthermore, (2) Large Language Models (LLMs) have been extensively pre-trained at the source-code level for tasks such as code understanding and generation. This makes participants wonder about the ability of LLMs in binary code understanding. In this work, we propose a benchmark to evaluate the effectiveness of LLMs in real-world reverse engineering scenarios. The benchmark covers two key binary code understanding tasks, including function name recovery and binary code summarization. We gain valuable insights into their capabilities and limitations through extensive evaluations of popular LLMs using our benchmark. Our evaluations reveal that existing LLMs can understand binary code to a certain extent, thereby improving the efficiency of binary code analysis. Our results highlight the great potential of the LLMs in advancing the field of binary code understanding. Xiuwei Shang, Shaoyin Cheng, Gangyang Li, Weiming Zhang 0001, Nenghai Yu |
ICSME | 2 |
| 2024 | A Deep Reinforcement Learning Approach for Adaptive GPS Spoofing against Multi-Sensor Fusion Localization SystemabstractRecent studies have revealed that Multi-Sensor Fusion (MSF) algorithms used in autonomous driving localization systems are still vulnerable to sensor spoofing attacks. However, the attack strategies explored in current research have largely remained unchanged, typically involving two fixed stages, as demonstrated in FusionRipper[1]. To address this limitation and investigate new strategies as well as higher attack thresholds, this paper presents an innovative sensor spoofing attack method that utilizes deep reinforcement learning, specifically targeting the widely used Error State Kalman Filter-based MSF localization algorithm. This technique allows for persistent spoofing attacks along a trajectory by training an agent to modify GPS sensor data inputs. In a scenario where the vehicle travels at a constant speed in a straight line, we demonstrate the effectiveness of this approach, achieving a higher attack offset upper limit than the current state-of-the-art methods. This new perspective on injecting false sensor data into the fusion algorithm not only establishes a higher attack threshold but also poses a significant threat to the security of autonomous driving systems. Linqing Hu, Shaoyin Cheng, Weiming Zhang 0001, Nenghai Yu |
MSN | 3 |
| 2023 | HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceabstractDecompilation is a widely used process for reverse engineers to significantly enhance code readability by lifting assembly code to a higher-level C-like language, pseudo-code. Nevertheless, the process of compilation and stripping irreversibly discards high-level semantic information that is crucial to code comprehension, such as comments, identifier names, and types. Existing approaches typically recover only one type of information, making them suboptimal for semantic inference. In this paper, we treat pseudo-code as a special programming language, then present a unified pre-trained model, HexT5, that is trained on vast amounts of natural language comments, source identifiers, and pseudo-code using novel pseudo-code-based pre-training objectives. We fine-tune HexT5 on various downstream tasks, including code summarization, variable name recovery, function name recovery, and similarity detection. Comprehensive experiments show that HexT5 achieves state-of-the-art performance on four downstream tasks, and it demonstrates the robust effectiveness and generalizability of HexT5 for binary-related tasks. Jiaqi Xiong, Kejiang Chen, Han Gao 0014, Shaoyin Cheng, Weiming Zhang 0001 |
ASE | 5 |
| 2023 | Investigating Neural-based Function Name Reassignment from the Perspective of Binary Code RepresentationabstractBuilding a model to reassign descriptive names for binary functions is considerable assistance for reverse engineering. Existing methods proposed for this issue are based on the low-level representation of binary code (e.g., assembly code), and especially the recent approaches employed neural-based models on instruction sequences. However, their performance is still unsatisfactory. Meanwhile, modern decompilers provide lifted representations of binary code, and their effectiveness has not been adequately studied. This paper further explores the issue of function name reassignment from the perspective of binary code representation. Specifically, we present a general and flexible NEural-based function name Reassignment framework NER, which leverages a decompiler to obtain a specific representation and applies the corresponding serialization strategy on it. NER then uses an alternative neural network to make predictions. Three levels of representation are investigated, including assembly code, Intermediate Representation (IR), and pseudo-code. We observe the binary code representations are significant for the final performance. It demonstrates that the pseudo-code is the most effective one. Based on these findings, we leverage the framework to implement a reassignment model NER-pc, which has 25% and 10% F1 score improvements against the state-of-the-art methods. Besides, more experiments are conducted to verify the design of NER and the effectiveness of NER-pc. Han Gao 0014, Jie Zhang 0073, Yanru He, Shaoyin Cheng, Weiming Zhang 0001 |
PST | 5 |
| 2021 | A lightweight framework for function name reassignment based on large-scale stripped binariesabstractSoftware in the wild is usually released as stripped binaries that contain no debug information (e.g., function names). This paper studies the issue of reassigning descriptive names for functions to help facilitate reverse engineering. Since the essence of this issue is a data-driven prediction task, persuasive research should be based on sufficiently large-scale and diverse data. However, prior studies can only be based on small-scale datasets because their techniques suffer from heavyweight binary analysis, making them powerless in the face of big-size and large-scale binaries. Han Gao 0014, Shaoyin Cheng, Yinxing Xue, Weiming Zhang 0001 |
ISSTA | 2 |
| 2021 | GDroid: Android malware detection and classification with graph convolutional network
Han Gao 0014, Shaoyin Cheng, Weiming Zhang 0001 |
Comput. Secur. | 2 |
| 2019 | Image steganography using texture features and GANsabstractAs steganography is the main practice of hidden writing, many deep neural networks are proposed to conceal secret information into images, whose invisibility and security are unsatisfactory. In this paper, we present an encoder-decoder framework with an adversarial discriminator to conceal messages or images into natural images. The message is embedded into QR code first which significantly improves the fault-tolerance. Considering the mean squared error (MSE) is not conducive to perfectly learn the invisible perturbations of cover images, we introduce a texture-based loss that is helpful to hide information into the complex texture regions of an image, improving the invisibility of hidden information. In addition, we design a truncated layer to cope with stego image distortions caused by data type conversion and a moment layer to train our model with varisized images. Finally, our experiments demonstrate that the proposed model improves the security and visual quality of stego images. Jinjing Huang, Shaoyin Cheng, Songhao Lou, Fan Jiang 0005 |
IJCNN | 2 |
| 2017 | FFFuzzer: Filter Your Fuzz to Get Accuracy, Efficiency and Schedulability
Fan Jiang 0005, Cen Zhang, Shaoyin Cheng |
ACISP (2) | 3 |
| 2016 | AALRSMF: An Adaptive Learning Rate Schedule for Matrix Factorization
Feng Wei 0001, Hao Guo 0016, Shaoyin Cheng, Fan Jiang 0005 |
APWeb (2) | 3 |
| 2015 | ADKAM: A-Diversity K-Anonymity Model via Microaggregation
Shaoyin Cheng, Fan Jiang 0005 |
ISPEC | 2 |
| 2013 | Novel user influence measurement based on user interaction in microblogabstractWith the development of science and technology, various social networks have emerged in recent years and microblog is a prevailing one. This paper focuses on how to identify the most influential users quantitatively in microblog and proposes a new ranking method which employs the fact that a follower's contribution to the influences of his/her followees varies and depends greatly on the interactions between them. We consider bidirectional interactions from perspectives of followees and followers, and measure the interactive degree by four factors comprised of retweeting strength, commenting intensity, mentioning density and a special indicator to the potential interactions called keyword similarity. The experimental results show that our method based on user interaction is better in calculating the user influence. Xiang Li 0067, Shaoyin Cheng, Fan Jiang 0005 |
ASONAM | 2 |
| 2013 | DroidFuzzer: Fuzzing the Android Apps with Intent-Filter TagabstractThe Android system is getting more and more popular on the mobile devices. Thus, lots of apps have sprung up to facilitate people's daily life. However, many of the apps are released without sufficient testing work, so the users encounter a sudden app crash now and then. This will undoubtedly impact the user's experience and even lead to economic loss. Because current testing tools on Android apps mainly focus on the motion event on the screen, like click event, bugs concerned with data handling module in an app is neglected. In this paper, we propose an automated testing method to fuzz testing the Android apps. The test targets are the Activities which accept outside MIME data. These Activities are picked out by analyzing the Intent-filter tag in the AndroidManifest.xml file. An automated fuzzing tool, DroidFuzzer, is implemented based on the method. Finally, experiments are conducted to prove the effectiveness of it. Shaoyin Cheng, Lanbo Zhang, Fan Jiang 0005 |
MoMM | 2 |
| 2011 | An efficient SVM-based method for multi-class network traffic classificationabstractMulti-class network traffic classification is a fundamental function for network services and management. Support vector machine (SVM) based network traffic classification has recently attracted increasing interest, for its high accuracy and low training sample size requirement. However, to better fit applications with delay requirements, it is desirable to reduce the high computation cost of existing SVM-based traffic classifiers. In this paper, we propose a novel scheme for SVM-based traffic classification (called fuzzy tournament). Experiment results based on real network traffic traces show that our proposed scheme can reduce computation cost by as much as 7.65 times; in the mean time, misclassification ratio is consistently reduced by up to 2.35 times as well. Ning Jing, Shaoyin Cheng, Qunfeng Dong |
IPCCC | 3 |
| 2011 | LoongChecker: Practical Summary-Based Semi-simulation to Detect Vulnerability in Binary CodeabstractThe automatic detection of security vulnerabilities in binary code is challenging and lacks efficient tools. This paper presents a novel semi-simulation approach to statically detect potential vulnerabilities in binary code. The semi-simulation approach simulates address related instructions accurately using value set analysis, and only traces data dependence on other instructions using data dependence analysis. We have implemented this approach on a tool called LoongChecker, and evaluate it on three real world programs, and detect three known vulnerabilities and two zero-day vulnerabilities. The results show our approach is practical and can be applied to large real world software. Shaoyin Cheng, Jinding Wang, Fan Jiang 0005 |
TrustCom | 1 |
| 2010 | PDVDS: A Pattern-Driven Software Vulnerability Detection SystemabstractThe automatic detection of security vulnerabilities in binary program is challenging and lacks efficient tools. Current research and tools are mostly restricted to a specific platform and environment, which induces the trouble to detect all kinds of vulnerabilities with unified approach. Moreover, Existing methods need many manual operations and rely on the experience of researchers. This paper presents a cross-platform system for automatically software vulnerability detection based on uniform intermediate representation. It supports many platforms, including x86, PowerPC and ARM. The system lifts underlying instructions to intermediate representation from several platforms. Platform-independent analysis method is implemented based on intermediate representation by static analysis. It also uses a vulnerability pattern driver extracted from experience and knowledge to drive the automatic vulnerability detection during the analysis. The system called PDVDS has been realized. We have evaluated its effectiveness through validating many known vulnerabilities and detecting three zero-day vulnerabilities. Shaoyin Cheng, Jinding Wang, Fan Jiang 0005 |
EUC | 1 |