VLDB 2026 Research / reviewers in the wild / expert
Ruigang Liang
dblp:232/3062
· DBLP profile ↗
21ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-8751-9918ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 16 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation AttacksabstractPeizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang, Xingshuo Han, XiaoFeng Wang, Wei Dong, Yuling Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peizhuo Lv, Ruihua Zhou, Ruigang Liang, Xingshuo Han |
ACL (1) | 4 |
| 2025 | TypeForge: Synthesizing and Selecting Best-Fit Composite Data Types for Stripped BinariesabstractStatic binary analysis is a widely used approach for ensuring the security of closed-source software. However, the absence of type information in stripped binaries, particularly for composite data types, poses significant challenges for both static analyzers and reverse engineering experts in achieving efficient and accurate analysis. Existing methods often struggle with inaccuracies and scalability limitations when dealing with such data types. To address these problems, we present Typeforge, a novel approach inspired by the workflow of reverse engineering experts, which uses a two-stage synthesis-selection strategy to automate the recovery of composite data types from stripped binaries. We design a new graph structure, the Type Flow Graph (TFG) to represent type information within stripped binaries. In the first stage, TFG-based Type Synthesis focuses on efficiently and accurately building constraints and synthesizing possible composite type declarations from the stripped binaries. In the second stage, we propose an LLM-assisted double-elimination framework to select the best-fit type declaration from the candidates by assessing the readability of the decompiled code. Our comparison with state-of-the-art approaches demonstrates that TYPEFORGE achieves F1 scores of 81.7% and 88.2% in Composite Data Type Identification and Layout Recovery, respectively, substantially outperforming existing methods. Additionally, TYPEFORGE achieves an F1 score of 72.1% in Relationship Recovery, a particularly challenging task for previous approaches. Furthermore, TYPEFORGE has significantly lower time overhead, requiring only about 3.8% of the time taken by OSPREY, the best-performing existing approach, making it a promising solution for various real-world reverse engineering tasks. Yanzhong Wang, Ruigang Liang, Peiwei Hu, Kai Chen 0012 |
SP | 2 |
| 2025 | EGRTE: adversarially training a self-explaining smoothed classifier for certified robustnessabstractAbstract Deep learning has transformed fields such as computer vision, natural language processing, and audio analysis through its powerful pattern recognition and predictive capabilities. However, the robustness of these models remains a major concern, as they are highly vulnerable to adversarial attacks-subtle, intentional perturbations that lead to incorrect predictions. While recent defenses like adversarial training and defensive distillation aim to improve robustness, they have notable drawbacks, including overfitting and degraded performance under strong attacks. Certified defenses, such as robust training and Randomized Smoothing, offer theoretical guarantees within a specific perturbation radius, yet struggle to reflect real-world robustness due to efficiency bottlenecks and the unpredictable nature of actual adversarial attacks. These challenges reveal a critical gap between current defenses and real-world attack scenarios, highlighting the need for more practical and resilient solutions. To address the challenges of defense-attack gaps and the inefficiency in robust training, we introduce the Explanation-Guided Robust Training Enhancer (EGRTE). EGRTE combines a self-explaining mechanism, which guides adversarial training to focus on generalized features for improved robustness and accuracy, with a masking mechanism that transforms noised data for easier model learning. This approach not only mitigates noise effects, including adversarial perturbations, but also eliminates the need for time-intensive gradient calculations, greatly enhancing training efficiency. Comprehensive experiments on several datasets show EGRTE’s superior certified accuracy and robustness against adversarial attacks, with a 6.24-fold efficiency increase over comparable methods, positioning EGRTE as a highly effective solution for robust and efficient deep learning. Zijin Lin, Jinwen He, Yue Zhao 0018, Ruigang Liang, Zhendong Wu |
Cybersecur. | 4 |
| 2025 | LLM4TDG: test-driven generation of large language models based on enhanced constraint reasoningabstractAbstract With the evolution of modern software development paradigms, component reuse, and low-code approaches have emerged as mainstream in software development. However, developers often lack an in-depth understanding of reused code. The inability of components to operate autonomously leads to insufficient testing of software functionalities and security, further exacerbating the contradiction between the increasing complexity of software architectures and the demand for accurate and efficient software automation testing. This, in turn, increases the frequency of software supply chain security incidents. This paper proposes a test-driven generation framework, LLM4TDG, based on large language models (LLMs). By formally defining the constraint dependency graph and converting it into context constraints, LLMs’ ability to understand natural language descriptions such as test requirements and documents is enhanced. Constraint reasoning and backtracking mechanisms are then used to generate test drivers that satisfy the defined constraints automatically. Using the EvalPlus dataset, we evaluate the comprehensive capabilities of LLM4TDG in test case generation using four general-domain LLMs and five code-generation-domain LLMs. The experimental results indicate that our approach significantly enhances LLMs’ ability to comprehend constraints in testing objectives, achieving a 47.62% increase in constraint understanding across 147 testing tasks. Employing LLM4TDG significantly improves the average pass@k metric of all LLMs by 10.41%. The pass@k metric for CodeQwen-chat has improved by up to 18.66%. The metric surpasses the state-of-the-art GPT-4, with a performance of 92.16% on HUMANEVAL and 87.14% on HUMANEVAL+, which enhances the error correction and functional correctness in test-driven code generation. Meanwhile, Our experiments were conducted on a dataset of Python third-party libraries containing malicious behavior in the context of security testing tasks, validating the effectiveness of our method in real-world applications and its generalization capabilities. Jingqiang Liu, Ruigang Liang, Xiaoxi Zhu, Qixu Liu |
Cybersecur. | 2 |
| 2024 | Attention-Based Decompilation Through Neural Machine Translation
Ruigang Liang, Ying Cao 0006, Peiwei Hu, Kai Chen 0012 |
Inscrypt (1) | 1 |
| 2024 | Optimizing Decompiler Output by Eliminating Redundant Data Flow in Self-Recursive InliningabstractDecompilation, which aims to lift a binary to a high-level language such as C, is one of the most common approaches software security analysts use for analyzing binary code. Recovering decompiled code with high readability is essential, as humans must understand the code's functionality correctly. However, some compilation optimization strategies will introduce obfuscation into the binary code, thereby reducing the readability of decompiled code. Among them, the function inlining related optimization strategies combine functions, causing the original function's code volume and complexity to multiply. Especially with self-recursive inlining optimization, it transforms initially simple functions into ones with significantly increased code volume and complex logic, greatly hindering the understanding of security engineers. In this paper, we present Erase, the first approach to reverse the self-recursive inlining optimization technique. We compare Erase with state-of-the-art decompilers Ghidra and Hex-Rays to evaluate ERASE's improvement for the functions affected by self-recursive inlining. Experimental results show that Erase's output is 78.4% and 88.9% more compact (fewer lines of code) than Ghidra and Hex-Rays, respectively. Moreover, reverse engineers spend 88.5% less time analyzing ERASE's output than analyzing Ghidra and 90.4% less time than analyzing Hex-Rays, and the accuracy of analyzing Erase's output is 2.75 times higher than both Ghidra and Hex-Rays. Ying Cao 0006, Ruigang Liang, Peiwei Hu, Kai Chen 0012 |
ICSME | 3 |
| 2024 | Evaluating the Effectiveness of DecompilersabstractIn software security tasks like malware analysis and vulnerability mining, reverse engineering is pivotal, with C decompilers playing a crucial role in understanding program semantics. However, reverse engineers still predominantly rely on assembly code rather than decompiled code when analyzing complex binaries. This practice underlines the limitations of current decompiled code, which hinders its effectiveness in reverse engineering. Identifying and analyzing the problems of existing decompilers and making targeted improvements can effectively enhance the efficiency of software analysis. In this study, we systematically evaluate current mainstream decompilers’ semantic consistency and readability. Semantic evaluation results show that the state-of-the-art decompiler Hex-Rays has about 55% accuracy at almost all optimization, which contradicts the common belief among many reverse engineers that decompilers are usually accurate. Readability evaluation indicates that despite years of efforts to improve the readability of the decompiled code, decompilers’ template-based approach still predominantly yields code akin to binary structures rather than human coding patterns. Additionally, our human study indicates that to enhance decompilers’ accuracy and readability, introducing human or compiler-aware strategies like a speculate-verify-correct approach to obtain recompilable decompiled code and iteratively refine it to more closely resemble the original binary, potentially offers a more effective optimization method than relying on static analysis and rule expansion. Ying Cao 0006, Ruigang Liang, Kai Chen 0012 |
ISSTA | 3 |
| 2024 | DeGPT: Optimizing Decompiler Output with LLM
Peiwei Hu, Ruigang Liang, Kai Chen 0012 |
NDSS | 2 |
| 2024 | SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised Learning
Peizhuo Lv, Shenchen Zhu, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, Guozhu Meng |
NDSS | 6 |
| 2024 | MEA-Defender: A Robust Watermark against Model Extraction AttackabstractRecently, numerous highly-valuable Deep Neural Networks (DNNs) have been trained using deep learning algorithms. To protect the Intellectual Property (IP) of the original owners over such DNN models, backdoor-based watermarks have been extensively studied. However, most of such watermarks fail upon model extraction attack, which utilizes input samples to query the target model and obtains the corresponding outputs, thus training a substitute model using such input-output pairs. In this paper, we propose a novel watermark to protect IP of DNN models against model extraction, named MEA-Defender. In particular, we obtain the watermark by combining two samples from two source classes in the input domain and design a watermark loss function that makes the output domain of the watermark within that of the main task samples. Since both the input domain and the output domain of our watermark are indispensable parts of those of the main task samples, the watermark will be extracted into the stolen model along with the main task during model extraction. We conduct extensive experiments on four model extraction attacks, using five datasets and six models trained based on supervised learning and self-supervised learning algorithms. The experimental results demonstrate that MEA-Defender is highly robust against different model extraction attacks, and various watermark removal/detection approaches. Peizhuo Lv, Hualong Ma, Kai Chen 0012, Jiachen Zhou 0001, Shengzhi Zhang, Ruigang Liang, Shenchen Zhu |
SP | 6 |
| 2023 | Invisible Backdoor Attacks Using Data Poisoning in Frequency DomainabstractBackdoor attacks have become a significant threat to deep neural networks (DNNs), whereby poisoned models perform well on benign samples but produce incorrect outputs when given specific inputs with a trigger. These attacks are usually implemented through data poisoning by injecting poisoned samples (samples patched with a trigger and mislabelled to the target label) into the dataset, and the models trained with that dataset will be infected with the backdoor. However, most current backdoor attacks lack stealthiness and robustness because of the fixed trigger patterns and mislabelling, which humans or some backdoor defense approach can easily detect. To address this issue, we propose a frequency-domain-based backdoor attack method that implements backdoor implantation without mislabeling the poisoned samples or accessing the training process. We evaluated our approach on four benchmark datasets and two popular scenarios: no-label self-supervised and clean-label supervised learning. The experimental results demonstrate that our approach achieved a high attack success rate (above 90%) on all tasks without significant performance degradation on main tasks and robust against mainstream defense approaches. Chang Yue, Peizhuo Lv, Ruigang Liang, Kai Chen 0012 |
ECAI | 3 |
| 2023 | DBIA: Data-Free Backdoor Attack Against Transformer NetworksabstractRecently, transformer architecture has demonstrated its significance in both Natural Language Processing (NLP) and Computer Vision (CV) tasks. Although other network models are known to be vulnerable to the backdoor attack, which embeds triggers in the models and controls the models’ behavior when the triggers are presented, little is known about how such an attack performs on the transformer models. In this paper, we propose DBIA, a novel Data-free1Backdoor Attack against the CV-oriented transformer networks, leveraging the inherent attention mechanism of transformers to generate triggers and injecting the backdoor using a poisoned substitute dataset. We conducted extensive experiments using three benchmark transformers, i.e., ViT, DeiT, and Swin Transformer, on four mainstream image classification tasks, i.e., ImageNet, CIFAR-10, GTSRB, and Youtube Face. The evaluation results demonstrate that, with fewer resources, our approach can embed backdoors with a high success rate and a low impact on the performance of the victim transformers. Peizhuo Lv, Hualong Ma, Jiachen Zhou 0001, Ruigang Liang, Kai Chen 0012, Shengzhi Zhang, Yunfei Yang 0001 |
ICME | 4 |
| 2023 | AURC: Detecting Errors in Program Code and Documentation
Peiwei Hu, Ruigang Liang, Ying Cao 0006, Kai Chen 0012 |
USENIX Security Symposium | 2 |
| 2023 | A Data-free Backdoor Injection Approach in Neural Networks
Peizhuo Lv, Chang Yue, Ruigang Liang, Yunfei Yang 0001, Shengzhi Zhang, Hualong Ma, Kai Chen 0012 |
USENIX Security Symposium | 3 |
| 2023 | SkillSim: voice apps similarity detectionabstractAbstract Virtual personal assistants (VPAs), such as Amazon Alexa and Google Assistant, are software agents designed to perform tasks or provide services to individuals in response to user commands. VPAs extend their functions through third-party voice apps, thereby attracting more users to use VPA-equipped products. Previous studies demonstrate vulnerabilities in the certification, installation, and usage of these third-party voice apps. However, these studies focus on individual apps. To the best of our knowledge, there is no prior research that explores the correlations among voice apps.Voice apps represent a new type of applications that interact with users mainly through a voice user interface instead of a graphical user interface, requiring a distinct approach to analysis. In this study, we present a novel voice app similarity analysis approach to analyze voice apps in the market from a new perspective. Our approach, called SkillSim, detects similarities among voice apps (i.e. skills) based on two dimensions: text similarity and structure similarity. SkillSim measures 30,000 voice apps in the Amazon skill market and reveals that more than 25.9% have at least one other skill with a text similarity greater than 70%. Our analysis identifies several factors that contribute to a high number of similar skills, including the assistant development platforms and their limited templates. Additionally, we observe interesting phenomena, such as developers or platforms creating multiple similar skills with different accounts for purposes such as advertising. Furthermore, we also find that some assistant development platforms develop multiple similar but non-compliant skills, such as requesting user privacy in a non-compliance way, which poses a security risk. Based on the similarity analysis results, we have a deeper understanding of voice apps in the mainstream market. Zhixiu Guo, Ruigang Liang, Guozhu Meng, Kai Chen 0012 |
Cybersecur. | 2 |
| 2023 | A Robustness-Assured White-Box Watermark in Neural NetworksabstractRecently, stealing highly-valuable and large-scale deep neural network (DNN) models becomes pervasive. The stolen models may be re-commercialized, e.g., deployed in embedded devices, released in model markets, utilized in competitions, etc, which infringes the Intellectual Property (IP) of the original owner. Detecting IP infringement of the stolen models is quite challenging, even with the white-box access to them in the above scenarios, since they may have experienced fine-tuning, pruning, functionality-equivalent adjustment to destruct any embedded watermark. Furthermore, the adversaries may also attempt to extract the embedded watermark or forge a similar watermark to falsely claim ownership. In this article, we propose a novel DNN watermarking solution, named$HufuNet$, to detect IP infringement of DNN models against the above mentioned attacks. Furthermore, HufuNet is the first one theoretically proved to guarantee robustness against fine-tuning attacks. We evaluate HufuNet rigorously on four benchmark datasets with five popular DNN models, including convolutional neural network (CNN) and recurrent neural network (RNN). The experiments and analysis demonstrate that HufuNet is highly robust against model fine-tuning/pruning, transfer learning, kernels cutoff/supplement, functionality-equivalent attacks and fraudulent ownership claims, thus highly promising to protect large-scale DNN models in the real world. Peizhuo Lv, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Hualong Ma, Yue Zhao 0018, Yingjiu Li |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | Boosting Neural Networks to Decompile Optimized BinariesabstractDecompilation aims to transform a low-level program language (LPL) (eg., binary file) into its functionally-equivalent high-level program language (HPL) (e.g., C/C++). It is a core technology in software security, especially in vulnerability discovery and malware analysis. In recent years, with the successful application of neural machine translation (NMT) models in natural language processing (NLP), researchers have tried to build neural decompilers by borrowing the idea of NMT. They formulate the decompilation process as a translation problem between LPL and HPL, aiming to reduce the human cost required to develop decompilation tools and improve their generalizability. However, state-of-the-art learning-based decompilers do not cope well with compiler-optimized binaries. Since real-world binaries are mostly compiler-optimized, decompilers that do not consider optimized binaries have limited practical significance. In this paper, we propose a novel learning-based approach named NeurDP, that targets compiler-optimized binaries. NeurDP uses a graph neural network (GNN) model to convert LPL to an intermediate representation (IR), which bridges the gap between source code and optimized binary. We also design an Optimized Translation Unit (OTU) to split functions into smaller code fragments for better translation performance. Evaluation results on datasets containing various types of statements show that NeurDP can decompile optimized binaries with 45.21% higher accuracy than state-of-the-art neural decompilation frameworks. Ying Cao 0006, Ruigang Liang, Kai Chen 0012, Peiwei Hu |
ACSAC | 2 |
| 2021 | Neutron: an attention-based neural decompilerabstractAbstract Decompilation aims to analyze and transform low-level program language (PL) codes such as binary code or assembly code to obtain an equivalent high-level PL. Decompilation plays a vital role in the cyberspace security fields such as software vulnerability discovery and analysis, malicious code detection and analysis, and software engineering fields such as source code analysis, optimization, and cross-language cross-operating system migration. Unfortunately, the existing decompilers mainly rely on experts to write rules, which leads to bottlenecks such as low scalability, development difficulties, and long cycles. The generated high-level PL codes often violate the code writing specifications. Further, their readability is still relatively low. The problems mentioned above hinder the efficiency of advanced applications (e.g., vulnerability discovery) based on decompiled high-level PL codes.In this paper, we propose a decompilation approach based on the attention-based neural machine translation (NMT) mechanism, which converts low-level PL into high-level PL while acquiring legibility and keeping functionally similar. To compensate for the information asymmetry between the low-level and high-level PL, a translation method based on basic operations of low-level PL is designed. This method improves the generalization of the NMT model and captures the translation rules between PLs more accurately and efficiently. Besides, we implement a neural decompilation framework called Neutron. The evaluation of two practical applications shows that Neutron’s average program accuracy is 96.96%, which is better than the traditional NMT model. Ruigang Liang, Ying Cao 0006, Peiwei Hu, Kai Chen 0012 |
Cybersecur. | 1 |
| 2021 | Evaluation indicators for open-source software: a reviewabstractAbstract In recent years, the widespread applications of open-source software (OSS) have brought great convenience for software developers. However, it is always facing unavoidable security risks, such as open-source code defects and security vulnerabilities. To find out the OSS risks in time, we carry out an empirical study to identify the indicators for evaluating the OSS. To achieve a comprehensive understanding of the OSS assessment, we collect 56 papers from prestigious academic venues (such as IEEE Xplore, ACM Digital Library, DBLP, and Google Scholar) in the past 21 years. During the process of the investigation, we first identify the main concerns for selecting OSS and distill five types of commonly used indicators to assess OSS. We then conduct a comparative analysis to discuss how these indicators are used in each surveyed study and their differences. Moreover, we further undertake a correlation analysis between these indicators and uncover 13 confirmed conclusions and four cases with controversy occurring in these studies. Finally, we discuss several possible applications of these conclusions, which are insightful for the research on OSS and software supply chain. Ruigang Liang, Xiang Chen 0005 |
Cybersecur. | 2 |
| 2020 | FuzzGuard: Filtering out Unreachable Inputs in Directed Grey-box Fuzzing through Deep Learning
Peiyuan Zong, Dawei Wang 0021, Zizhuang Deng, Ruigang Liang, Kai Chen 0012 |
USENIX Security Symposium | 5 |
| 2019 | Seeing isn't Believing: Towards More Robust Adversarial Attack Against Real World Object DetectorsabstractRecently Adversarial Examples (AEs) that deceive deep learning models have been a topic of intense research interest. Compared with the AEs in the digital space, the physical adversarial attack is considered as a more severe threat to the applications like face recognition in authentication, objection detection in autonomous driving cars, etc. In particular, deceiving the object detectors practically, is more challenging since the relative position between the object and the detector may keep changing. Existing works attacking object detectors are still very limited in various scenarios, e.g., varying distance and angles, etc. In this paper, we presented systematic solutions to build robust and practical AEs against real world object detectors. Particularly, for Hiding Attack (HA), we proposed thefeature-interference reinforcement (FIR) method and theenhanced realistic constraints generation (ERG) to enhance robustness, and for Appearing Attack (AA), we proposed thenested-AE, which combines two AEs together to attack object detectors in both long and short distance. We also designed diverse styles of AEs to make AA more surreptitious. Evaluation results show that our AEs can attack the state-of-the-art real-time object detectors (i.e., YOLO V3 and faster-RCNN) at the success rate up to 92.4% with varying distance from 1m to 25m and angles from -60º to 60º. Our AEs are also demonstrated to be highly transferable, capable of attacking another three state-of-the-art black-box models with high success rate. Yue Zhao 0018, Ruigang Liang, Qintao Shen, Shengzhi Zhang, Kai Chen 0012 |
CCS | 3 |