EDBT 2026 Demo / reviewers in the wild / expert
Yang Xiao 0011
dblp:181/1848-11
· DBLP profile ↗
29ranked-venue papers
2as first author
27since 2021 · last 2026
0009-0005-8009-2252ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 13 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 13 · 1 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LifeFuzz: Lifecycle-Guided Fuzzing for Windows Driver Cross-Handler VulnerabilitiesabstractThird-party Windows drivers expose a critical attack surface. However, vulnerabilities that require cross-handler I/O Control (IOCTL) sequences remain hard to find, despite their prevalence, and often lead to privilege escalation. Static analysis suffers from high false positives, path explosion, and complex resource modeling. Meanwhile, dynamic fuzzers often exercise handlers in isolation or combine them randomly, leaving implicit state dependencies unchecked. To address the gap, we present LifeFuzz, a lifecycle-guided fuzzing framework that models global-variable lifecycles to construct dependency-respecting IOCTL sequences. Specifically, it identifies variable operations across handlers, preserves seeds that affect driver state, and then combines them into meaningful sequences. Consequently, LifeFuzz explores deep paths unreachable for existing fuzzers. We evaluate LifeFuzz on 26 Windows WDM drivers. It discovers 86 vulnerabilities, including 32 cross-handler cases, with six assigned CVE IDs. Moreover, it finds 357% more cross-handler vulnerabilities than msFuzz and achieves 19.2% higher average coverage. Overall, 37% of discovered vulnerabilities require cross-handler interactions, thereby validating lifecycle-aware, cross-handler fuzzing for driver security. Chendong Yu, Yuekang Li, Yang Xiao 0011, Jie Lu 0009, Yeting Li, Defang Bo, Wei Huo 0005 |
EuroSys | 3 |
| 2026 | Through the Authentication Maze: Detecting Authentication Bypass Vulnerabilities in Firmware Binaries
Nanyu Zhong, Yuekang Li, Yanyan Zou 0002, Jiaxu Zhao 0004, Jinwei Dong, Yang Xiao 0011, Bingwei Peng, Yeting Li, Wei Huo 0005 |
NDSS | 6 |
| 2026 | SPVR: syntax-to-prompt vulnerability repair based on large language models
Ruoke Wang, Zongjie Li, Cuiyun Gao 0001, Chaozheng Wang, Yang Xiao 0011, Xuan Wang 0002 |
Autom. Softw. Eng. | 5 |
| 2025 | Vulnerability-Affected Versions Identification: How Far Are We?abstractIdentifying which software versions are affected by a vulnerability is critical for patching, risk mitigation. Despite a growing body of tools, their real-world effectiveness remains unclear due to narrow evaluation scopes—often limited to early SZZ variants, outdated techniques, and small or coarse-grained datasets. In this paper, we present the first comprehensive empirical study of vulnerability-affected versions identification. We curate a high-quality benchmark of 1,128 real-world C/C++ vulnerabilities and systematically evaluate 12 representative tools from both tracing and matching paradigms across four dimensions: effectiveness at both vulnerability and version levels, root causes of false positives and negatives, sensitivity to patch characteristics, and ensemble potential. Our findings reveal fundamental limitations: no tool exceeds 45.0% accuracy, with key challenges stemming from heuristic dependence, limited semantic reasoning, and rigid matching logic. Patch structures such as add-only and cross-file changes further hinder performance. Although ensemble strategies can improve results by up to 10.1%, overall accuracy remains below 60.0%, highlighting the need for fundamentally new approaches. Moreover, our study offers actionable insights to guide tool development, combination strategies, and future research in this critical area. Finally, we release the replicated code and benchmark on our website to encourage future contributions. Xingchu Chen, Jialun Cao, Yang Xiao 0011, Xinyue Cai, Yeting Li, Tianqi Sun, Haiming Chen 0001, Wei Huo 0005 |
ASE | 4 |
| 2025 | Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationabstractBinary Code Similarity Detection (BCSD), essential for binary-code related tasks like vulnerability detection, has attracted increasing attention in recent years. However, existing methods frequently fall short of achieving both high precision and recall at scale, and their results often lack interpretability due to the neglect of function context and reliance on purely similarity-driven outputs. Our key insights are twofold: 1) Binary functions are not self-contained; they depend on other code and data beyond their content to fulfill their functionalities. 2) Large language models (LLMs) excel not only at analyzing code but also at generating reasonable explanations. Motivated by these insights, we propose a general BCSD framework, Co2F uLL. We first systematically select stable and representative code and data features, along with their corresponding dependencies on the functions, to construct the function context. Then, by fusing function context with content similarities computed by the existing BCSD approach, we substantially narrow down the search space. Ultimately, we employ LLMs with a carefully designed prompt to verify the remaining candidates and produce clear, human-readable explanations. We conduct comprehensive experiments on a large function pool under varying compilation settings and after binary stripping. The results show that Co2F uLL based on HermesSim and DeepSeek-V3 achieves 80.5% precision and 94.4% recall, improving the baseline HermesSim by 142.5% and 42.2%, respectively, providing an accurate and interpretable solution for BCSD. Chaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li 0008, Dongliang Fang, Yang Xiao 0011, Yongle Chen, Limin Sun 0001 |
ASE | 6 |
| 2025 | A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and PractitionersabstractBinary Function Similarity Detection (BFSD) is a foundational technique in software security, underpinning a wide range of applications including vulnerability detection, malware analysis. Recent advances in AI-based BFSD tools have led to significant performance improvements. However, existing evaluations of these tools suffer from three key limitations: a lack of in-depth analysis of performance-influencing factors, an absence of realistic application analysis, and reliance on small-scale or low-quality datasets.In this paper, we present the first large-scale empirical study of AI-based BFSD tools to address these gaps. We construct two high-quality and diverse datasets: BinAtlas, comprising 12,453 binaries and over 7 million functions for capability evaluation; and BinAres, containing 12,291 binaries and 54 real-world 1-day vulnerabilities for evaluating vulnerability detection performance in practical IoT firmware settings. Using these datasets, we evaluate nine representative BFSD tools, analyze the challenges and limitations of existing BFSD tools, and investigate the consistency among BFSD tools. We also propose an actionable strategy for combining BFSD tools to enhance overall performance (an improvement of 13.4%). Our study not only advances the practical adoption of BFSD tools but also provides valuable resources and insights to guide future research in scalable and automated binary similarity detection. Yang Xiao 0011, Yuekang Li, Zhengzi Xu, Sihao Qiu, Keyu Qi, Yeting Li, Xingchu Chen, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
ASE | 3 |
| 2025 | From Constraints to Cracks: Constraint Semantic Inconsistencies as Vulnerability Beacons for Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Yang Xiao 0011, Naijia Jiang, Yeting Li, Nanyu Zhong, Bingwei Peng, Kunpeng Jian, Wei Huo 0005 |
USENIX Security Symposium | 4 |
| 2025 | VULCANBOOST: Boosting ReDoS Fixes through Symbolic Representation and Feature Normalization
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Haiming Chen 0001, Xinyi Wang 0013, Hengyu Yang, Huina Chao, Cen Zhang, Yang Xiao 0011, Yanyan Zou 0002, Feng Li 0045, Wei Huo 0005 |
USENIX Security Symposium | 9 |
| 2025 | ZIPPER: Static Taint Analysis for PHP Applications with Precision and Efficiency
Xinyi Wang 0013, Yeting Li, Jie Lu 0009, Shizhe Cui, Chenghang Shi, Qin Mai, Yunpei Zhang, Yang Xiao 0011, Feng Li 0045, Wei Huo 0005 |
USENIX Security Symposium | 8 |
| 2024 | LibvDiff: Library Version Difference Guided OSS Version Identification in BinariesabstractOpen-source software (OSS) has been extensively employed to expedite software development, inevitably exposing downstream software to the peril of potential vulnerabilities. Precisely identifying the version of OSS not only facilitates the detection of vulnerabilities associated with it but also enables timely alerts upon the release of 1-day vulnerabilities. However, current methods for identifying OSS versions rely heavily on version strings or constant features, which may not be present in compiled OSS binaries or may not be representative when only function code changes are made. As a result, these methods are often imprecise in identifying the version of OSS binaries being used. Chaopeng Dong, Siyuan Li 0014, Shouguo Yang, Yang Xiao 0011, Yongpan Wang, Hong Li 0004, Zhi Li 0018, Limin Sun 0001 |
ICSE | 4 |
| 2024 | SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionabstractRecently, there has been a growing interest in automatic software vulnerability detection. Pre-trained model-based approaches have demonstrated superior performance than other Deep Learning (DL)-based approaches in detecting vulnerabilities. However, the existing pre-trained model-based approaches generally employ code sequences as input during prediction, and may ignore vulnerability-related structural information, as reflected in the following two aspects. First, they tend to fail to infer the semantics of the code statements with complex logic such as those containing multiple operators and pointers. Second, they are hard to comprehend various code execution sequences, which is essential for precise vulnerability detection. To mitigate the challenges, we propose a Structured Natural Language Comment tree-based vulnerAbiLity dEtection framework based on the pre-trained models, named . The proposed Structured Natural Language Comment Tree (SCT) integrates the semantics of code statements with code execution sequences based on the Abstract Syntax Trees (ASTs).Specifically, comprises three main modules: (1) Comment Tree Construction, which aims at enhancing the model’s ability to infer the semantics of code statements by first incorporating Large Language Models (LLMs) for comment generation and then adding the comment node to ASTs. (2) Structured Natural Language Comment Tree Construction, which aims at explicitly involving code execution sequence by combining the code syntax templates with the comment tree. (3) SCT-Enhanced Representation, which finally incorporates the constructed SCTs for well capturing vulnerability patterns. Experimental results demonstrate that outperforms the best-performing baseline, including the pre-trained model and LLMs, with improvements of 2.96%, 13.47%, and 3.75% in terms of F1 score on the FFMPeg+Qemu, Reveal, and SVulD datasets, respectively. Furthermore, can be applied to different pre-trained models, such as CodeBERT and UniXcoder, yielding the F1 score performance enhancements ranging from 1.37% to 10.87%. Xin-Cheng Wen, Cuiyun Gao 0001, Shuzheng Gao, Yang Xiao 0011, Michael R. Lyu |
ISSTA | 4 |
| 2024 | File Hijacking Vulnerability: The Elephant in the Room
Chendong Yu, Yang Xiao 0011, Jie Lu 0009, Yuekang Li, Yeting Li, Lian Li 0002, Jian Wang 0067, Defang Bo, Wei Huo 0005 |
NDSS | 2 |
| 2024 | Leveraging Semantic Relations in Code and Data to Enhance Taint Analysis of Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Zhaohui Liang, Yang Xiao 0011, Yeting Li, Bingwei Peng, Nanyu Zhong, Xinyi Wang 0013, Wei Huo 0005 |
USENIX Security Symposium | 5 |
| 2024 | Asteria-Pro: Enhancing Deep Learning-based Binary Code Similarity Detection by Incorporating Domain KnowledgeabstractWidespread code reuse allows vulnerabilities to proliferate among a vast variety of firmware. There is an urgent need to detect these vulnerable codes effectively and efficiently. By measuring code similarities, AI-based binary code similarity detection is applied to detecting vulnerable code at scale. Existing studies have proposed various function features to capture the commonality for similarity detection. Nevertheless, the significant code syntactic variability induced by the diversity of IoT hardware architectures diminishes the accuracy of binary code similarity detection. In our earlier study and the tool Asteria , we adopted a Tree-LSTM network to summarize function semantics as function commonality, and the evaluation result indicates an advanced performance. However, it still has utility concerns due to excessive time costs and inadequate precision while searching for large-scale firmware bugs. To this end, we propose a novel deep learning-enhancement architecture by incorporating domain knowledge-based pre-filtration and re-ranking modules, and we develop a prototype named Asteria-Pro based on Asteria . The pre-filtration module eliminates dissimilar functions, thus reducing the subsequent deep learning-model calculations. The re-ranking module boosts the rankings of vulnerable functions among candidates generated by the deep learning model. Our evaluation indicates that the pre-filtration module cuts the calculation time by 96.9%, and the re-ranking module improves MRR and Recall by 23.71% and 36.4%, respectively. By incorporating these modules, Asteria-Pro outperforms existing state-of-the-art approaches in the bug search task by a significant margin. Furthermore, our evaluation shows that embedding baseline methods with pre-filtration and re-ranking modules significantly improves their precision. We conduct a large-scale real-world firmware bug search, and Asteria-Pro manages to detect 1,482 vulnerable functions with a high precision 91.65%. Shouguo Yang, Chaopeng Dong, Yang Xiao 0011, Yiran Cheng, Zhiqiang Shi, Zhi Li 0018, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2023 | Enhancing OSS Patch Backporting with SemanticsabstractKeeping open-source software (OSS) up to date is one potential solution to prevent known vulnerabilities. However, it requires frequent and costly testing and may introduce compatibility issues. Consequently, developers often choose to backport security patches to the vulnerable versions instead. Manual backporting is time-consuming, especially for large OSS such as the Linux kernel. Therefore, automating this process is urgently needed to save considerable time. Existing automated approaches for backporting patches involve either automatic patch generation or automatic patch migration. However, these methods are often ineffective and error-prone since they failed to locate the precise patch locations or generate the correct patch, operating only on the syntactic level. Su Yang 0003, Yang Xiao 0011, Zhengzi Xu, Chengyi Sun, Yuqing Zhang 0001 |
CCS | 2 |
| 2023 | An Enhanced Vulnerability Detection in Software Using a Heterogeneous Encoding EnsembleabstractDetecting vulnerabilities in source code is essential to prevent cybersecurity attacks. Deep learning-based vulnerability detection is an active research topic in software security. However, existing deep learning-based vulnerability detectors (VD) are limited to using either serialization-based or graph-based methods, which do not combine serialized global and structured local information at the same time. As a result, a single method cannot perform well for semantic information that exists in complex source code, leading to low detection accuracy. In this paper, we present EL-VDetect, a stacked ensemble learning approach for vulnerability detection that eliminates these issues. EL-VDetect enhances feature selection techniques to represent the best relevant vulnerability features with the slice code and subgraphs, reducing redundant information of vulnerabilities. Our model combines serialization-based and graph-based neural networks to successfully capture the global and local context information of source code, effectively understands code semantics, and focuses on vulnerable nodes based on the attention mechanism to accurately detect vulnerabilities. To evaluate EL-VDetect's effectiveness, we crawl a real-world dataset from CVEDetails, consisting of functions for eight applications. A comprehensive performance analysis of the real-world dataset shows that EL-VDetect achieves 90.72% accuracy, outperforming baseline deep learning models by 1.75-26.21 %. Our proposed model can better identify vulnerabilities in software than other existing vulnerability detection models. Hao Sun 0028, Yongji Liu, Zhenquan Ding, Yang Xiao 0011, Zhiyu Hao, Hongsong Zhu |
ISCC | 4 |
| 2023 | ACETest: Automated Constraint Extraction for Testing Deep Learning OperatorsabstractDeep learning (DL) applications are prevalent nowadays as they can help with multiple tasks. DL libraries are essential for building DL applications. Furthermore, DL operators are the important building blocks of the DL libraries, that compute the multi-dimensional data (tensors). Therefore, bugs in DL operators can have great impacts. Testing is a practical approach for detecting bugs in DL operators. In order to test DL operators effectively, it is essential that the test cases pass the input validity check and are able to reach the core function logic of the operators. Hence, extracting the input validation constraints is required for generating high-quality test cases. Existing techniques rely on either human effort or documentation of DL library APIs to extract the constraints. They cannot extract complex constraints and the extracted constraints may differ from the actual code implementation. To address the challenge, we propose ACETest, a technique to automatically extract input validation constraints from the code to build valid yet diverse test cases which can effectively unveil bugs in the core function logic of DL operators. For this purpose, ACETest can automatically identify the input validation code in DL operators, extract the related constraints and generate test cases according to the constraints. The experimental results on popular DL libraries, TensorFlow and PyTorch, demonstrate that ACETest can extract constraints with higher quality than state-of-the-art (SOTA) techniques. Moreover, ACETest is capable of extracting 96.4% more constraints and detecting 1.95 to 55 times more bugs than SOTA techniques. In total, we have used ACETest to detect 108 previously unknown bugs on TensorFlow and PyTorch, with 87 of them confirmed by the developers. Lastly, five of the bugs were assigned with CVE IDs due to their security impacts. Yang Xiao 0011, Yuekang Li, Yeting Li, Dongsong Yu, Chendong Yu, Hui Su, Wei Huo 0005 |
ISSTA | 2 |
| 2023 | Software Vulnerability Detection Using an Enhanced Generalization Strategy
Hao Sun 0028, Zhe Bu, Yang Xiao 0011, Chengsheng Zhou, Zhiyu Hao, Hongsong Zhu |
SETTA | 3 |
| 2023 | Learning Program Semantics for Vulnerability Detection via Vulnerability-Specific Inter-procedural SlicingabstractLearning-based approaches that learn code representations for software vulnerability detection have been proven to produce inspiring results. However, they still fail to capture complete and precise vulnerability semantics for code representations. To address the limitations, in this work, we propose a learning-based approach namely SnapVuln, which first utilizes multiple vulnerability-specific inter-procedural slicing algorithms to capture vulnerability semantics of various types and then employs a Gated Graph Neural Network (GGNN) with an attention mechanism to learn vulnerability semantics. We compare SnapVuln with state-of-the-art learning-based approaches on two public datasets, and confirm that SnapVuln outperforms them. We further perform an ablation study and demonstrate that the completeness and precision of vulnerability semantics captured by SnapVuln contribute to the performance improvement. Bozhi Wu, Shangqing Liu, Yang Xiao 0011, Jun Sun 0001, Shangwei Lin 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2023 | Effective ReDoS Detection by Principled Vulnerability Modeling and Exploit GenerationabstractRegular expression Denial-of-Service (ReDoS) is one kind of algorithmic complexity attack. For a vulnerable regex, attackers can craft certain strings to trigger the super-linear worst-case matching time, which causes denial-of-service to regex engines. Various ReDoS detection approaches have been proposed recently. Among them, hybrid approaches which absorb the advantages of both static and dynamic approaches have shown their performance superiority. However, two key challenges still hinder the effectiveness of the detection: 1) Existing modelings summarize localized vulnerability patterns based on partial features of the vulnerable regex; 2) Existing attack string generation strategies are ineffective since they neglected the fact that non-vulnerable parts of the regex may unexpectedly invalidate the attack string (we name this kind of invalidation as disturbance.)Rengar is our hybrid ReDoS detector with new vulnerability modeling and disturbance free attack string generator. It has the following key features: 1) Benefited by summarizing patterns from full features of the vulnerable regex, its modeling is a more precise interpretation of the root cause of ReDoS vulnerability. The modeling is more descriptive and precise than the union of existing modelings while keeping conciseness; 2) For each vulnerable regex, its generator automatically checks all potential disturbances and composes generation constraints to avoid possible disturbances.Compared with nine state-of-the-art tools, Rengar detects not only all vulnerable regexes they found but also 3 – 197 times more vulnerable regexes. Besides, it saves 57.41% – 99.83% average detection time compared with tools containing a dynamic validation process. Using Rengar, we have identified 69 zero-day vulnerabilities (21 CVEs) affecting popular projects which have more than dozens of millions weekly download count. Xinyi Wang 0013, Cen Zhang, Yeting Li, Zhiwu Xu 0001, Shuailin Huang, Yi Liu 0069, Yican Yao, Yang Xiao 0011, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
SP | 8 |
| 2023 | Detecting API Post-Handling Bugs Using Code and Description in Patches
Miaoqian Lin, Kai Chen 0012, Yang Xiao 0011 |
USENIX Security Symposium | 3 |
| 2023 | Towards Practical Binary Code Similarity Detection: Vulnerability Verification via Patch Semantic AnalysisabstractVulnerability is a major threat to software security. It has been proven that binary code similarity detection approaches are efficient to search for recurring vulnerabilities introduced by code sharing in binary software. However, these approaches suffer from high false-positive rates (FPRs) since they usually take the patched functions as vulnerable, and they usually do not work well when binaries are compiled with different compilation settings. To this end, we propose an approach, named Robin , to confirm recurring vulnerabilities by filtering out patched functions. Robin is powered by a lightweight symbolic execution to solve the set of function inputs that can lead to the vulnerability-related code. It then executes the target functions with the same inputs to capture the vulnerable or patched behaviors for patched function filtration. Experimental results show that Robin achieves high accuracy for patch detection across different compilers and compiler optimization levels respectively on 287 real-world vulnerabilities of 10 different software. Based on accurate patch detection, Robin significantly reduces the false-positive rate of state-of-the-art vulnerability detection tools (by 94.3% on average), making them more practical. Robin additionally detects 12 new potentially vulnerable functions. Shouguo Yang, Zhengzi Xu, Yang Xiao 0011, Zhe Lang, Yang Liu 0003, Zhiqiang Shi, Hong Li 0004, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | VERJava: Vulnerable Version Identification for Java OSS with a Two-Stage AnalysisabstractThe software version information affected by the CVEs (Common Vulnerabilities and Exposures) provided by the National Vulnerability Database (NVD) is not always accurate. This could seriously mislead the repair priority for software users, and greatly hinder the work of security researchers. Bao et al. improved the well-known Sliwerski-Zimmermann-Zeller (SZZ) algorithm for vulnerabilities (called V-SZZ) to precisely refine vulnerable software versions. But V-SZZ only focuses on those CVEs of which patches only have deleted lines.In this study, we target Java Open Source Software (OSS) by virtue of its pervasiveness and ubiquitousness. Due to Java’s object-oriented characteristic, a single security patch often involves modifications of multiple functions. Existing patch code similarity analysis does not consider patch existence from the point of view of an entire patch, which would generate too many false positives for Java CVEs. In this work, we address these limitations by introducing a two-stage approach named VERJava, to systematically assess vulnerable versions for a target vulnerability in Java OSS. Specifically, vulnerable versions are calculated respectively at a function level and an entire patch level, then the results are synthesized to decide the final vulnerable versions. For evaluation, we manually annotated the vulnerable versions of 167 real CVEs from seven popular Java open source projects. The result shows that VERJava achieves the precision of 90.7% on average, significantly outperforming the state-of-the-art work V-SZZ. Furthermore, our study reveals some interesting findings that have not yet been discussed. Yang Xiao 0011, Feng Li 0045, He Su, Hongyun Huang, Wei Huo 0005 |
ICSME | 3 |
| 2022 | RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Jialun Cao, Yuekang Li, Rongchen Li, Haiming Chen 0001, Shing-Chi Cheung, Yang Liu 0003, Yang Xiao 0011 |
USENIX Security Symposium | 10 |
| 2022 | Unleashing the power of pseudo-code for binary code similarity analysisabstractAbstract Code similarity analysis has become more popular due to its significant applicantions, including vulnerability detection, malware detection, and patch analysis. Since the source code of the software is difficult to obtain under most circumstances, binary-level code similarity analysis (BCSA) has been paid much attention to. In recent years, many BCSA studies incorporating AI techniques focus on deriving semantic information from binary functions with code representations such as assembly code, intermediate representations, and control flow graphs to measure the similarity. However, due to the impacts of different compilers, architectures, and obfuscations, binaries compiled from the same source code may vary considerably, which becomes the major obstacle for these works to obtain robust features. In this paper, we propose a solution, named UPPC (Unleashing the Power of Pseudo-code), which leverages the pseudo-code of binary function as input, to address the binary code similarity analysis challenge, since pseudo-code has higher abstraction and is platform-independent compared to binary instructions. UPPC selectively inlines the functions to capture the full function semantics across different compiler optimization levels and uses a deep pyramidal convolutional neural network to obtain the semantic embedding of the function. We evaluated UPPC on a data set containing vulnerabilities and a data set including different architectures (X86, ARM), different optimization options (O0-O3), different compilers (GCC, Clang), and four obfuscation strategies. The experimental results show that the accuracy of UPPC in function search is 33.2% higher than that of existing methods. Zhengzi Xu, Yang Xiao 0011, Yinxing Xue |
Cybersecur. | 3 |
| 2021 | VIVA: Binary Level Vulnerability Identification via Partial SignatureabstractBinary level code clone detection techniques have been used to identify 1-day vulnerabilities in software. It collects functions with known vulnerabilities and searches for similar functions in the target system. However, existing approaches are limited to detect the same vulnerabilities in different binaries. They can hardly find new recurring vulnerabilities, which share similar logic. Moreover, they only focus on improving the accuracy of binary function matching algorithms while overlooking the presence of security patches, which results in high false-positive rates and requires significant effort to verify the results.To this end, we propose VIVA, a binary level vulnerability and patch semantic summarization and matching tool for accurate recurring vulnerability detection. It uses novel binary program slicing techniques with the aid of pseudo-code trace refinement to generate partial vulnerability and patch signatures, which capture the semantics. It matches the signatures with pre-filtering to efficiently detect 1-day and recurring vulnerabilities. The experimental results show that VIVA outperforms other source code and binary matching tools with a precision of 100% for 1-day vulnerabilities and 87.6% for recurring vulnerabilities and good performance (28.58s per signature search in 4M functions). It detects 92 new vulnerabilities in different series and different versions of real-world projects, with 11 exist without fixing in the latest version. Yang Xiao 0011, Zhengzi Xu, Chendong Yu, Longquan Liu, Zimu Yuan, Yang Liu 0003, Aihua Piao, Wei Huo 0005 |
SANER | 1 |
| 2021 | B2SMatcher: fine-Grained version identification of open-Source software in binary filesabstractAbstract Codes of Open Source Software (OSS) are widely reused during software development nowadays. However, reusing some specific versions of OSS introduces 1-day vulnerabilities of which details are publicly available, which may be exploited and lead to serious security issues. Existing state-of-the-art OSS reuse detection work can not identify the specific versions of reused OSS well. The features they selected are not distinguishable enough for version detection and the matching scores are only based on similarity.This paper presents B2SMatcher, a fine-grained version identification tool for OSS in commercial off-the-shelf (COTS) software. We first discuss five kinds of version-sensitive code features that are trackable in both binary and source code. We categorize these features into program-level features and function-level features and propose a two-stage version identification approach based on the two levels of code features. B2SMatcher also identifies different types of OSS version reuse based on matching scores and matched feature instances. In order to extract source code features as accurately as possible, B2SMatcher innovatively uses machine learning methods to obtain the source files involved in the compilation and uses function abstraction and normalization methods to eliminate the comparison costs on redundant functions across versions. We have evaluated B2SMatcher using 6351 candidate OSS versions and 585 binaries. The result shows that B2SMatcher achieves a high precision up to 89.2% and outperforms state-of-the-art tools. Finally, we show how B2SMatcher can be used to evaluate real-world software and find some security risks in practice. Gu Ban, Yang Xiao 0011, Xinhua Li, Zimu Yuan, Wei Huo 0005 |
Cybersecur. | 3 |
| 2020 | MVP: Detecting Vulnerabilities using Patch-Enhanced Vulnerability Signatures
Yang Xiao 0011, Bihuan Chen 0001, Chendong Yu, Zhengzi Xu, Zimu Yuan, Feng Li 0045, Binghong Liu, Yang Liu 0003, Wei Huo 0005, Wenchang Shi |
USENIX Security Symposium | 1 |
| 2019 | Open-Source License Violations of Binary Software at Large ScaleabstractOpen-source licenses are widely used in open-source projects. However, developers using or modifying the source code of open-source projects do not always strictly follow the licenses. GPL and AGPL, two of the most popular copyleft licenses, are most likely to be violated, because they require developers to open-source the entire project if any code under GPL/AGPL protection is included whether modified or not. There are few license violation detectors focusing on binary software, owning to the challenge of mapping binary code to source code efficiently and accurately at large scale. In this paper, we propose a scalable and fully-automated system to check open-source license violation of binary software at large scale. We match source code to binary code by analyzing file attributes of executable files and code features that are not affected by compilation and could vary between projects. Moreover, to break the barrier of large-scale analysis, we introduce an automatic extractor to parse executable files from installation packages that are broadly available in software download sites. In empirical experiments of binary-to-source mapping, we have got a remarkable high accuracy of 99.5% and recall of 95.6% without significant loss of precision. Besides, 2270 pairs of binary-to-source mapping relationships are discovered, with 110 license violations of GPL and AGPL licenses related to 7.2% of the 1000 real-world binary software projects. Muyue Feng, Weixuan Mao, Zimu Yuan, Yang Xiao 0011, Gu Ban, Jiahuan Xu, He Su, Binghong Liu, Wei Huo 0005 |
SANER | 4 |