Siyuan Li 0014

dblp:63/9705-14 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0004-4096-1209ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BinEnhance-Pro: Enhancing Binary Code Search by Distinguishing Similar but Non-Homologous Functions
Yongpan Wang, Siyuan Li 0014, Xiaojie Zhu, Xiaodong Gu 0002, Yongle Chen
IEEE Trans. Dependable Secur. Comput.4
2025 LibRI: A Module Analysis Framework for Identifying Complex Reuse Relationship in Binaries
abstract
With the rapid advancement of collaborative software development, the increased reuse of third-party libraries (TPLs) has introduced new security challenges. Detecting reuse relationships between binary programs and TPLs is vital for software maintenance, vulnerability tracing, and component analysis. However, most existing detection methods are confined to identifying code reuse between binaries, often misclassifying nested and pseudo-propagation reuse as direct reuse. While some methods attempt to analyze these intricate relationships, they depend on pre-collected source code structures or unstable constants, which compromises their generality and accuracy. To tackle these challenges, we introduce LibRI, a framework for identifying complex dependency relationships in C/C++ binaries. LibRI modularizes and matches binaries to analyze module-matching scenarios across multiple TPLs, accurately determining true reuse relationships and identifying original module sources. This facilitates the construction of a detailed reuse relationship graph. Experimental results show that LibRI achieves an accuracy of 0.966 in detecting actual direct reuse, significantly outperforming existing methods. In addition, LibRI is able to build a vulnerability propagation graph of TPLs, identify the propagation paths of TPL vulnerabilities, demonstrating its potential in vulnerability tracking.
Wenyan Yu, Siyuan Li 0014, Mingjiang Huang, Rongrong Xi, Hongsong Zhu
CSCWD2
2025 VN-GT: Optimizing Virtual Network Deployment via Game Theory
abstract
The static and homogeneous nature of traditional networks presents a significant challenge for our defense efforts. These characteristics enable an experienced attacker to quickly determine our network topology and gather detailed information about the internal hosts through systematic scanning techniques. Implementing a virtual network view can mitigate this by simulating a virtual topology, thereby consuming the attacker’s resources and time. However, deploying a virtual network view reduces network throughput and increase latency. Additionally, an improperly configured virtual network view can waste resources and degrade Quality of Service (QoS). Most existing studies have focused solely on the defender’s perspective, resulting in overly idealistic solutions that are ineffective in real-world scenarios. To address this, we propose VN-GT, a game-theoretic based model that optimizes virtual network deployment by considering both attackers and defenders. We provide a detailed example scenario, analyze the game’s equilibrium, and validate the effectiveness of our method through a real attack and defense experiment.
Weijie Wang 0005, Yan Wang 0081, Guokun Xu, Zuxin Chen, Siyuan Li 0014, Min Yu 0001, Weiqing Huang, Degang Sun
ICASSP5
2025 TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, making the scope of affected software in CVE reports imprecise. Traditional methods primarily focus on identifying reused vulnerability code within target software, yet they cannot verify if these vulnerabilities can be triggered in new software contexts. This limitation often results in false positives. In this paper, we introduce TransferFuzz, a novel vulnerability verification framework, to verify whether vulnerabilities propagated through code reuse can be triggered in new software. Innovatively, we collected runtime information during the execution or fuzzing of the basic binary (the vulnerable binary detailed in CVE reports). This process allowed us to extract historical traces, which proved instrumental in guiding the fuzzing process for the target binary (the new binary that reused the vulnerable function). TransferFuzz introduces a unique Key Bytes Guided Mutation strategy and a Nested Simulated Annealing algorithm, which transfers these historical traces to implement trace-guided fuzzing on the target binary, facilitating the accurate and efficient verification of the propagated vulnerability. Our evaluation, conducted on widely recognized datasets, shows that TransferFuzz can quickly validate vulnerabilities previously unverifiable with existing techniques. Its verification speed is 2.5 to 26.2 times faster than existing methods. Moreover, TransferFuzz has proven its effectiveness by expanding the impacted software scope for 15 vulnerabilities listed in CVE reports, increasing the number of affected binaries from 15 to 53. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/TransferFuzz.
Siyuan Li 0014, Yuekang Li, Zuxin Chen, Chaopeng Dong, Yongpan Wang, Hong Li 0004, Yongle Chen, Hongsong Zhu
ICSE1
2025 Lares: LLM-driven Code Slice Semantic Search for Patch Presence Testing
abstract
In modern software ecosystems, 1-day vulnerabilities pose significant security risks due to extensive code reuse. Identifying vulnerable functions in target binaries alone is insufficient; it is also crucial to determine whether these functions have been patched. Existing methods, however, suffer from limited usability and accuracy. They often depend on the compilation process to extract features, requiring substantial manual effort and failing for certain software. Moreover, they cannot reliably differentiate between code changes caused by patches or compilation variations.To overcome these limitations, we propose Lares, a scalable and accurate method for patch presence testing. Lares introduces Code Slice Semantic Search, which directly extracts features from the patch source code and identifies semantically equivalent code slices in the pseudocode of the target binary. By eliminating the need for the compilation process, Lares improves usability, while leveraging large language models (LLMs) for code analysis and SMT solvers for logical reasoning to enhance accuracy. Experimental results show that Lares achieves superior precision, recall, and usability. Furthermore, it is the first work to evaluate patch presence testing across optimization levels, architectures, and compilers. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/Lares.
Siyuan Li 0014, Yaowen Zheng, Hong Li 0004, Jingdong Guo, Chaopeng Dong, Chunpeng Yan, Weijie Wang 0005, Yimo Ren, Limin Sun 0001, Hongsong Zhu
ASE1
2025 BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search
Yongpan Wang, Hong Li 0004, Xiaojie Zhu, Siyuan Li 0014, Chaopeng Dong, Shouguo Yang, Kangyuan Qin
NDSS4
2025 PREXP: Uncovering and Exploiting Security-Sensitive Objects in the Linux Kernel
abstract
Security-Sensitive Objects (SSOs) are often critical components in the exploitation of Linux kernel memory corruption vulnerabilities. While existing research has advanced SSOs identification and classification, there remains a significant gap in systematically understanding how these objects can be effectively exploited in real-world security analysis. To address this challenge, we present PREXP, a novel approach to analyzing SSOs exploitability and automating the transformation of Proof-of-Concept (PoC) into exploitable states. Our approach encompasses three key techniques: (1) capability analysis and attribute modeling of vulnerable object (2) extraction and filtering of target SSOs and (3) automatically augmenting PoCs with SSO-specific code to create exploitation capabilities. To evaluate our approach, we tested our prototype on 30 public CVEs, successfully parsing vulnerable object in 22 cases (73.3%) and achieving accurate SSO matches in 18 (60.0%). PREXP outperformed state-of-the-art tools such as SCAVY and AlphaEXP in structure-matching, and enabled the generation of new Control Flow Hijacking Primitives (CFHPs) for 3 previously unexploited vulnerabilities, demonstrating its practical value in real-world exploit development.
Zuxin Chen, Yaowen Zheng, Hong Li 0004, Siyuan Li 0014, Weijie Wang 0005, Dongliang Fang, Zhiqiang Shi, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.4
2025 TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, leading to imprecise scopes of affected software in CVE reports. Traditional methods focus primarily on detecting reused vulnerability code in target software but lack the ability to confirm whether these vulnerabilities can be triggered in new software contexts. In previous work, we introduced the TransferFuzz framework to address this gap by using historical trace-based fuzzing. However, its effectiveness is constrained by the need for manual intervention and reliance on source code instrumentation. To overcome these limitations, we propose TransferFuzz-Pro, a novel framework that integrates Large Language Model (LLM)-driven code debugging technology. By leveraging LLM for automated, human-like debugging and Proof-of-Concept (PoC) generation, combined with binary-level instrumentation, TransferFuzz-Pro extends verification capabilities to a wider range of targets. Our evaluation shows that TransferFuzz-Pro is significantly faster and can automatically validate vulnerabilities that were previously unverifiable using conventional methods. Notably, it expands the number of affected software instances for 15 CVE-listed vulnerabilities from 15 to 53 and successfully generates PoCs for various Linux distributions. These results demonstrate that TransferFuzz-Pro effectively verifies vulnerabilities introduced by code reuse in target software and automatically generation PoCs.
Siyuan Li 0014, Kaiyu Xie, Yuekang Li, Hong Li 0004, Yimo Ren, Limin Sun 0001, Hongsong Zhu
IEEE Trans. Software Eng.1
2024 LibvDiff: Library Version Difference Guided OSS Version Identification in Binaries
abstract
Open-source software (OSS) has been extensively employed to expedite software development, inevitably exposing downstream software to the peril of potential vulnerabilities. Precisely identifying the version of OSS not only facilitates the detection of vulnerabilities associated with it but also enables timely alerts upon the release of 1-day vulnerabilities. However, current methods for identifying OSS versions rely heavily on version strings or constant features, which may not be present in compiled OSS binaries or may not be representative when only function code changes are made. As a result, these methods are often imprecise in identifying the version of OSS binaries being used.
Chaopeng Dong, Siyuan Li 0014, Shouguo Yang, Yang Xiao 0011, Yongpan Wang, Hong Li 0004, Zhi Li 0018, Limin Sun 0001
ICSE2
2024 Malware Classification Method Based on Dynamic Features with Sensitive Behaviors
abstract
Traditional malware classification methods often just scratch the surface by analyzing the sequence of system commands (API calls) used by malware during its operation. These approaches miss out on deeper, complex behaviors that could significantly enhance accuracy in identifying different malware types. To address this, we introduce SenBeMC, a method that delves deeper into the behaviors exhibited by malware. SenBeMC combine API call information vectors with behavioral information to enhance the deep semantic information of input features, enriching the hierarchical structure of feature representation. SenBeMC stands out by employing soft thresholding and attention mechanisms to sift through the noise — extraneous information that can mask the malware's true nature, and a BiLSTM model that excels in understanding the sequence and timing of actions, crucial for spotting sophisticated threats. Experimental evaluations on real-world datasets affirm that SenBeMC effectively improves feature representation and accuracy of malware classification when compared to other contemporary state-of-the-art models.
Yamin Xie, Siyuan Li 0014, Zhengcai Chen, Haichao Du, Xiaoqi Jia, Yuejin Du
SMC2
2024 VDTriplet: Vulnerability detection with graph semantics using triplet model
Hao Sun 0028, Lei Cui 0003, Zhenquan Ding, Siyuan Li 0014, Zhiyu Hao, Hongsong Zhu
Comput. Secur.5
2024 LibAM: An Area Matching Framework for Detecting Third-Party Libraries in Binaries
abstract
Third-party libraries (TPLs) are extensively utilized by developers to expedite the software development process and incorporate external functionalities. Nevertheless, insecure TPL reuse can lead to significant security risks. Existing methods, which involve extracting strings or conducting function matching, are employed to determine the presence of TPL code in the target binary. However, these methods often yield unsatisfactory results due to the recurrence of strings and the presence of numerous similar non-homologous functions. Furthermore, the variation in C/C++ binaries across different optimization options and architectures exacerbates the problem. Additionally, existing approaches struggle to identify specific pieces of reused code in the target binary, complicating the detection of complex reuse relationships and impeding downstream tasks. And, we call this issue the poor interpretability of TPL detection results. In this article, we observe that TPL reuse typically involves not just isolated functions but also areas encompassing several adjacent functions on the Function Call Graph (FCG). We introduce LibAM, a novel Area Matching framework that connects isolated functions into function areas on FCG and detects TPLs by comparing the similarity of these function areas, significantly mitigating the impact of different optimization options and architectures. Furthermore, LibAM is the first approach capable of detecting the exact reuse areas on FCG and offering substantial benefits for downstream tasks. To validate our approach, we compile the first TPL detection dataset for C/C++ binaries across various optimization options and architectures. Experimental results demonstrate that LibAM outperforms all existing TPL detection methods and provides interpretable evidence for TPL detection results by identifying exact reuse areas. We also evaluate LibAM’s scalability on large-scale, real-world binaries in IoT firmware and generate a list of potential vulnerabilities for these devices. Our experiments indicate that the Area Matching framework performs exceptionally well in the TPL detection task and holds promise for other binary similarity analysis tasks. Last but not least, by analyzing the detection results of IoT firmware, we make several interesting findings, for instance, different target binaries always tend to reuse the same code area of TPL. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/LibAM .
Siyuan Li 0014, Yongpan Wang, Chaopeng Dong, Shouguo Yang, Hong Li 0004, Hao Sun 0028, Zhe Lang, Zuxin Chen, Weijie Wang 0005, Hongsong Zhu, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.1
2023 A Privacy-Preserving Online Deep Learning Algorithm Based on Differential Privacy
abstract
Deep Reinforcement Learning (DRL) combines the perceptual capabilities of deep learning with the decision-making capabilities of Reinforcement Learning RL, which can achieve enhanced decision-making. However, the environmental state data contains the privacy of the users. There exists consequently a potential risk of environmental state information being leaked during RL training. Some data desensitization and anonymization technologies are currently being used to protect data privacy. There may still be a risk of privacy disclosure with these desensitization techniques. Meanwhile, policymakers need the environmental state to make decisions, which will cause the disclosure of raw environmental data. To address the privacy issues in DRL, we propose a differential privacy-based online DRL algorithm. The algorithm will add Gaussian noise to the gradients of the deep network according to the privacy budget. More important, we prove tighter bounds for the privacy budget. Furthermore, we train an autocoder to protect the raw environmental state data. In this work, we prove the privacy budget formulation for differential privacy-based online deep RL. Experiments show that the proposed algorithm can improve privacy protection while still having relatively excellent decisionmaking performance.
Jun Li 0085, Fengshi Zhang, Yonghe Guo, Siyuan Li 0014, Guanjun Wu, Dahui Li, Hongsong Zhu
CSCWD4
2023 FlowEmbed: Binary function embedding model based on relational control flow graph and byte sequence
abstract
Binary function embedding models are applicable to various downstream tasks within IoT device software systems and have demonstrated advantages in numerous binary analysis tasks, such as vulnerability (homologous) function search and compilation optimization option identification. However, current binary function embedding methods either learn embedding based on code sequence, which lack the program semantics of functions (e.g., control flow, etc.) or based on program structure graphs, which omit global sequential information. As a result, these methods fall short in enabling models to learn the complete semantic of function. In this paper, we introduce FlowEmbed, a novel approach that synergistically integrates control flow and global semantic learning to facilitate exhaustive code comprehension. Initially, FlowEmbed harnesses a distinct relational control flow graph combined with the power of BERT and RGCN models to aptly capture the nuances of control flow semantics. Moreover, by deploying the DPCNN model on a byte sequence constructed from function machine code, FlowEmbed adeptly discerns the inherent global sequential semantics of binary functions. Through rigorous evaluations spanning three IoT-related tasks, FlowEmbed’s efficacy becomes evident, showcasing notable improvements: a 20.6% improvement in compilation optimization option identification, a 1.8% improvement in binary function similarity analysis, and an 11.9% improvement in homologous function search. Collectively, these results underscore FlowEmbed’s superior capability, positioning it as a invaluable asset in a binary analysis application.
Yongpan Wang, Chaopeng Dong, Siyuan Li 0014, Renjie Su, Zhanwei Song, Hong Li 0004
ICPADS3
2023 Denoising Network of Dynamic Features for Enhanced Malware Classification
abstract
Malware classification based on dynamic feature analysis works by running malware in controlled and isolated environments to observe how it behaves. This technology widely uses the sequence of run-time API calls to classify. Malware often adopts evasion techniques such as obfuscation, encryption, and code injection to obfuscate classification results by introducing noise into the API sequence. The existing methods lack explicit means of filtering noise components in the data, which affects the accuracy of malware detection. To address this issue, we propose DenoMC, a malware classification method with an explicit denoising module. Firstly, we employ dynamic analysis and embedding techniques to encode the API sequence. Then, we introduce a soft thresholding mechanism in the residual network to achieve active filtering of noise components in API sequences. Finally, a BiLSTM model is adopted to enhance the temporal correlation among sequence of API calls and improve classification performance. Experiments conducted on real datasets demonstrate that DenoMC significantly improves malware classification accuracy compared to other state-of-art models. In addition, we validate the effectiveness of each module in DenoMC through extensive ablation studies.
Siyuan Li 0014, Hui Wen 0001, Liting Deng, Zhi Li 0018, Limin Sun 0001
IPCCC1
2023 VN-SMT: An SMT-based Construction Method on Virtual Network to Defend Insider Reconnaissance
abstract
Due to networks’ static and homomorphic nature, experienced attackers can quickly get the target network’s topology and internal host information by scanning. The virtual network view prevents network reconnaissance by simulating a virtual network topology for the network hosts, to consume the attacker’s attack resources and time. However, deploying a virtual network view will reduce network throughput and increase network latency, and an unreasonable virtual network view configuration will waste resources and reduce Quality of Services(QoS). We, therefore, propose a method VN-SMT that can rationally configure virtual network view. This method generates an optimal virtual network view base on existing host configuration, risk constraints, and budget constraints. We define metrics for deception, concealment, and resource consumption to measure the effectiveness of virtual network views. We conduct simulations to verify the effectiveness and feasibility of VN-SMT.
Weijie Wang 0005, Yan Wang 0081, Guokun Xu, Qiujian Lv, Zuxin Chen, Siyuan Li 0014
WCNC6