VLDB 2026 Research / reviewers in the wild / expert
Xin Liu 0050
dblp:76/1820-50
· DBLP profile ↗
21ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-3685-4852ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RlDecompiler: Enhancing LLM-based Decompilation via Reinforcement Learning with a Multi-Faceted Reward FunctionabstractDecompiling binary code into human-readable, high-level source code is a core challenge in reverse engineering. While traditional methods often rely on brittle, pattern-based heuristics, the advent of Large Language Models (LLMs) offers a more flexible and robust approach. However, current LLM-based decompilation efforts are often limited by their training methodologies, which typically treat the task as a simple sequence-to-sequence translation and struggle to enforce the functional correctness of the output. To address these issues, this paper proposes an innovative framework for training LLMs to perform high-fidelity decompilation. A core contribution of our work is a novel data processing pipeline that enriches the model’s input. This pipeline integrates Ghidra-based static analysis to directly embed crucial context, such as static resources (strings, floating-point numbers) and relabeled basic blocks—from the binary into an LLM-friendly prompt. Building on this enriched input, we employ reinforcement learning fine-tuning guided by a multi-faceted reward function that comprehensively evaluates syntactic correctness, AST similarity, compilability, and functional correctness via test cases. Using this framework, we trained the RlDecompiler family of models (1.3B and 3B). Experimental results demonstrate that RlDecompiler achieves state-of-the-art performance, and its generated code quality is also higher than that of the baseline models. The RlDecompiler 1.3B and 3B models achieve rerunnable rates of 27.96% and 40.70%, respectively, outperforming existing baselines. The code is available at https://github.com/ri-char/rldecompile. Yuchi Su, Weina Niu, Jiacheng Gong, Song Li 0006, Xin Liu 0050, Xiaosong Zhang 0001 |
ICPC | 6 |
| 2026 | Towards a comprehensive framework for verifying open-source software license compatibility
Ziang Liu 0006, Xin Liu 0050, Yingli Zhang, Song Li 0006, Weina Niu, Qingguo Zhou, Rui Zhou 0005, Xiaokang Zhou |
Empir. Softw. Eng. | 2 |
| 2026 | MultiSiFer: Detecting Multiple-Speaker Fake Voice Without Speaker-Irrelative FeaturesabstractVoice synthesis technologies have advanced rapidly, raising serious concerns about content security and trust. While many fake voice detectors achieve strong performance in controlled settings, they often overfit to speaker-irrelative features (SiFs), exhibit poor robustness, and fail in multi-speaker scenarios. To address these limitations, we propose MultiSiFer, a novel fake voice detector grounded in a new design philosophy: rather than merely distinguishing synthetic from human voices, it explicitly prioritizes learning essential human voice characteristics. MultiSiFer leverages a pre-trained speech representation model to enhance this learning and is the first detector trained on a newly curated multi-speaker fake voice dataset, enabling effective generalization across speakers. Experiments show that MultiSiFer outperforms existing methods in both standard and multi-speaker settings, achieving 10.84% average equal error rate (EER). Xin Liu 0050, Xuan Hai, Ziyao Yu, Qingyuan Fei, Qingguo Zhou |
IEEE Internet Things J. | 1 |
| 2025 | LZHV: Accelerating LZ77 Compression Algorithm with Hash VerificationabstractThe longest match strategy in LZ77, a major bottleneck in the compression process, is accelerated in enhanced algorithms such as LZ4 and ZSTD by using a hash table. However, it may results in numerous random memory accesses, which modern CPUs handle inefficiently, thus reducing the compression speed. In this paper, we introduce the LZHV algorithm, which significantly reduces unnecessary memory accesses by over 99% through the implementation of hash verification within the hash table. By integrating LZHV into LZ4 at its default compression level and ZSTD at levels 3 and 4, we achieve a compression speed improvement of over 10% across various platforms. Guodong Ye, Xin Liu 0050, Rui Zhou 0005, Qingguo Zhou |
DCC | 3 |
| 2025 | LLM-SZZ: Novel Vulnerability-Inducing Commit Identification Driven by Large Language Model and CVE DescriptionabstractThe SZZ method and its variants are widely employed to identify vulnerability-affected ranges by analyzing vulnerability-fixing commits to trace back vulnerability-inducing commits. However, these methods generally suffer from low precision due to several key factors: 1) Current static method-based variants often incorrectly consider too many irrelevant lines and files in a commit. While methods that extract file references from vulnerability discussions can help narrow down relevant files, obtaining bug discussions for every CVE is often difficult. 2) Learning-based approaches focus exclusively on code to capture semantic relationships for identifying root cause lines. However, these models utilize limited information and demonstrate insufficient capacity for effective capture. 3) The reliance on line mapping algorithms results in inadequate tracing capabilities for complex vulnerabilities, especially when vulnerability-inducing commits are obscured in earlier software versions. To address these issues, this paper innovatively incorporates semantic information from descriptive text and the nature of CVEs derived from vulnerability-fixing commit diffs. By leveraging large language models (LLMs), this approach aims to capture the true root cause lines of vulnerabilities more accurately and enhance the tracing capabilities of the SZZ method, thereby achieving precise localization of the vulnerability impact range. Experimental results indicate that our proposed LLM-SZZ method outperforms existing state-of-the-art approaches, achieving over a 18 % increase in precision across datasets in various programming languages, demonstrating a significant performance advantage. Siqi Fan 0005, Xin Liu 0050, Yingli Zhang, Yuan Tan 0003, Luxing Yin, Zhaorun Chen, Song Li 0006, Rui Zhou 0005 |
ICSME | 2 |
| 2025 | ISGraphVD: Precise Vulnerability Detection for IoT Supply Chains Based on Identifier Sensitive GraphabstractOpen-source software (OSS) is widely reused in Internet of Things (IoT) devices, leading to widespread N-Day vulnerabilities when outdated components remain unpatched. Existing methods typically encode features of different Common Vulnerabilities and Exposures (CVEs) within a shared representation space. However, the model’s limited capacity, combined with the new vulnerability features, can disrupt previously learned patterns. Minimal code modifications in tiny-patch vulnerabilities are often overshadowed by variations introduced by different compilation settings, making it more difficult to distinguish vulnerable functions from their patched counterparts. This paper introduces ISGraphVD, a novel graph-based and function-level vulnerability detection approach that supports cross-compilation settings and enhances detection accuracy. By modeling each CVE independently through a one-model-per-CVE strategy, ISGraphVD reduces feature interference and improves detection accuracy across diverse CVEs. To better detect tinypatch vulnerability, we propose ISGraph, a fine-grained graph representation that models variable dependencies within and across basic blocks by integrating control flow analysis. Then, ISGraphVD utilizes a Graph Matching Network (GMN) with a cross-graph attention mechanism to identify critical vulnerability patterns. Experiments on IoT OSS projects show that ISGraphVD outperforms state-of-the-art methods, achieving a 6.3 percentage-point (pp) accuracy improvement over the strongest baseline, and real-world tests further validate its effectiveness in IoT supply chains. Yingli Zhang, Xin Liu 0050, Ziang Liu 0006, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ISSRE | 2 |
| 2025 | SiFMimicEvader: Evading Fake Voice Detection with Adversarial Neural Mimicry AttacksabstractThe application of deep learning in voice cloning has significantly enhanced the quality of cloned voices. While advanced voice cloning technologies are widely applied across various domains, they also pose serious security challenges such as producing natural Deepfakes. In response, numerous studies have focused on detecting fake voices, with many reporting outstanding performance. However, is the issue truly resolved? This paper introduces Adversarial Neural Mimicry Attack (ANMA) which leverages a specialized model to predict the behavior of other similar models, transforming black-box attacks into white-box scenarios indirectly. Based on ANMA and Speaker-irrelative Features (SiFs), we propose a novel black-box attack framework called SiFMimicEvader, designed to evade fake voice detectors with high success rates and minimal query requirements. The framework utilizes speech representation models as the breakthrough to predict the behaviors of fake voice detectors and employs a series of SiFs editing operations as perturbations to deceive these detectors. Experimental results demonstrate the effectiveness of SiFMimicEvader, achieving an average attack success rate exceeding 50% across various detectors, significantly outperforming other attack methods, while also showing great performance in audio quality and query scale, indicating its high availability in real-world scenarios. Xuan Hai, Xin Liu 0050, Ziyao Yu, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ACM Multimedia | 2 |
| 2025 | CFLBD: Distance-Informed Dynamic Clustering via Bhattacharyya Metrics for Federated Learning
Xiaowen Duan, Rui Zhou 0005, Xin Liu 0050, Qingguo Zhou |
NPC (1) | 5 |
| 2024 | SQLStateGuard: Statement-Level SQL Injection Defense Based on Learning-Driven MiddlewareabstractSQL injection is a significant and persistent threat to web services. Most existing protections against SQL injections rely on traffic-level anomaly detection, which often results in high false-positive rates and can be easily bypassed by attackers. This paper introduces SQLStateGuard, the world's first middleware-driven statement-level SQL injection defense approach, to address these issues. The SQLStateGuard uses a custom SQL middleware based on the idea of Runtime Application Self-Protection to capture raw SQL statements. These statements are then analyzed by SQLSG-Net, a database-oriented detection network based on gated linear units. If SQLSG-Net detects malicious SQL statements, the SQL middleware will block them. Experiments show that the detection accuracy of SQLStateGuard exceeds 99%, outperforming existing approaches, and it can identify the type of a specific SQL injection. Additionally, SQLStateGuard has no fingerprint and does not respond to SQL syntax errors, making it more challenging for attackers to gather information. This paper also presents a novel dataset generation process for SQLStateGuard and shares two statement-level SQL injection datasets with the research community, including over 145,000 malicious SQL statements categorized by the type of SQL injection. Xin Liu 0050, Song Li 0006, Weina Niu, Jun Shen 0001, Qingguo Zhou, Xiaokang Zhou |
SoCC | 1 |
| 2024 | Ghost-in-Wave: How Speaker-Irrelative Features Interfere DeepFake Voice DetectorsabstractRecent speech synthesis technology can generate high-quality speech indistinguishable from human speech, thus introducing various security and privacy risks. Numerous recent studies have focused on fake voice detection to address these risks, with many claiming to achieve ideal performance. However, is this really the case? A recent research work introduced Speaker-Irrelative-Features (SiFs), unrelated to the information in speech files but capable of influencing fake detectors. This means that existing detectors may rely on SiFs to a certain extent to distinguish real and fake speech. In this paper, we introduce an evaluation framework to evaluate the influence of SiFs in existing fake voice detectors in depth. We evaluate three SiFs which include background noise, the mute parts before and after voice, and the sampling rate on ASVspoof2019 and FoR. Our results confirm the substantial influence of SiFs on fake voice detection performance, and we delve into the analysis of the underlying mechanisms. Xuan Hai, Xin Liu 0050, Zhaorun Chen, Yuan Tan 0003, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ICME | 2 |
| 2024 | LiScopeLens: An Open-Source License Incompatibility Analysis Tool Based on Scope Representation of License TermsabstractOpen-source software has emerged as a pivotal force in the advancement of information technology. Robust open-source compliance governance is essential for the sustainable and healthy growth of both open-source software and its communities. License incompatibility analysis, in particular, represents a critical challenge hindering the progress of open-source software. Traditional methods of incompatibility analysis often fail to account for diverse usage scenarios or are tailored to a limited subset of scenarios. This limitation obstructing their ability to handle the intricate compatibility arising from varied programming language interactions, leading to a high false positives. Our study embarks from an examination of license exceptions, delving into the incompatibility analysis challenges through extensive empirical research on these exceptions. We discovered that the majority of exceptions are, in fact, detectable. Leveraging this empirical insight, our research further develops the license compatibility analysis model by introducing a new, refined legal terminology representation alongside a novel method for license compatibility reasoning. This approach begins with modeling different scenarios to represent license compatibility variably. Furthermore, based on these modeling outcomes, we have designed and implemented LiScopeLens, a tool capable of discerning dependency behaviors for granular compatibility assessment, starting with binary dependencies. Our experimental findings affirm that LiScopeLens proficiently determines the license compatibility status of open-source software across various usage scenarios, demonstrating its significant practical utility. Ziang Liu 0006, Xin Liu 0050, Yingli Zhang, Song Li 0006, Weina Niu, Qingguo Zhou, Rui Zhou 0005, Xiaokang Zhou |
ISSRE | 2 |
| 2024 | What's the Real: A Novel Design Philosophy for Robust AI-Synthesized Voice DetectionabstractVoice is one of the most widely used media for information transmission in human society. While high-quality synthetic voices are extensively utilized in various applications, they pose significant risks to content security and trust building. Numerous studies have concentrated on AI-synthesized voice detection to mitigate these risks, with many claiming to achieve promising performance. However, recent research has demonstrated that fake voice detectors suffer from serious overfitting to speaker-irrelative features (SiFs) and cannot be used in real-world scenarios. In this paper, we analyze the limitations of existing fake voice detectors and propose a new design philosophy, guiding the detection model to prioritize learning human voice features rather than the difference between the human voice and the synthetic voice. Based on this philosophy, we propose a novel AI-synthesized voice detection framework named SiFSafer, which uses pre-trained speech representation models to enhance the learning of feature distribution in human voices and the adapter fine-tuning to optimize the performance. The evaluation shows that the average EERs of existing fake voice detectors in the ASVspoof datasets can exceed 20% if the SiFs like silence segments are removed, while SiFSafer achieves an EER of less than 8%, indicating that SiFSafer is robust to SiFs and strongly resistant to existing attacks. Xuan Hai, Xin Liu 0050, Yuan Tan 0003, Song Li 0006, Weina Niu, Rui Zhou 0005, Xiaokang Zhou |
ACM Multimedia | 2 |
| 2023 | Hidden-in-Wave: A Novel Idea to Camouflage AI-Synthesized Voices Based on Speaker-Irrelative FeaturesabstractVoice is an essential medium for human communication and collaboration, and its trustworthiness is of great importance to humans. Synthesizing fake voices and detecting synthesized voices are two sides of a coin. Both sides have made great strides with the recently prospering deep learning techniques. Attackers started using AI techniques to synthesize, even clone, human voices. Researchers also proposed a series of AI-synthesized voice detection approaches and achieved promising results in laboratory environments.In this paper, we introduced the concept of speaker-irrelative features (SiFs) and a novel detection-bypass idea to camouflage AI-synthesized voices: replacing SiFs of AI-synthesized voices with crafted ones. We implemented a proof-of-concept framework named SiF-DeepVC based on our detection-bypass idea. Experiments show that the existing detection systems would consider the voices output by SiF-DeepVC more human-like than human voices, proving our detection-bypass idea is effective and SiFs are noteworthy in camouflaging AI-synthesized voices. Xin Liu 0050, Yuan Tan 0003, Xuan Hai, Qingguo Zhou |
ISSRE | 1 |
| 2023 | SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesabstractVoice is a vital medium for transmitting information. The advancement of speech synthesis technology has resulted in high-quality synthesized voices indistinguishable from human ears. These fake voices have been widely used in natural Deepfake production and other malicious activities, raising serious concerns regarding security and privacy. To deal with this situation, there have been many studies working on detecting fake voices and reporting excellent performance. However, is the story really over? In this paper, we propose SiFDetectCracker, a black-box adversarial attack framework based on Speaker-Irrelative Features (SiFs) against fake voice detection. We select background noise and mute parts before and after the speaker's voice as the primary attack features. By modifying these features in synthesized speech, the fake speech detector will make a misjudgment. Experiments show that SiFDetectCracker achieved a success rate of more than 80% in bypassing existing state-of-the-art fake voice detection systems. We also conducted several experiments to evaluate our attack approach's transferability and activation factor. Xuan Hai, Xin Liu 0050, Yuan Tan 0003, Qingguo Zhou |
ACM Multimedia | 2 |
| 2023 | Code classification with graph neural networks: Have you ever struggled to make it work?
Xin Liu 0050, Qingguo Zhou, Jianwei Zhuge, Chunming Wu 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Novel supply chain vulnerability detection based on heterogeneous-graph-driven hash similarity in IoT
Guodong Ye, Xin Liu 0050, Siqi Fan 0005, Yuan Tan 0003, Qingguo Zhou, Rui Zhou 0005, Xiaokang Zhou |
Future Gener. Comput. Syst. | 2 |
| 2023 | An Efficient Smart Contract Vulnerability Detector Based on Semantic Contract Graphs Using Approximate Graph MatchingabstractThe Internet of Things (IoT) has become a focus of information infrastructure development in recent years. The smart blockchain can provide various solutions for trust, security, and privacy (TSP) challenges to protect IoT data, and smart contracts are the foundation of blockchain intelligence, and greatly enhance the ability of smart blockchain to solve TSP problems. So, the security of smart contracts must be addressed. We propose an efficient smart contract vulnerability detector to improve the safety of smart contracts. It comprises a graph extraction method and a complete vulnerability detection process. The graph extraction method consists of vulnerability pattern extraction and a graph generation process. The vulnerability detection process first uses the approximate graph matching algorithm to select representative SCGraphs from the data set to build vulnerability SCGraph libraries. Second, determine whether the contract contains vulnerabilities by calculating the similarity between the SCGraphs generated from the contracts to be detected and the SCGraphs in the vulnerability library. Experiments show that our approach achieves an inspiring high detection rate and is the fastest among existing vulnerability detection tools, which indicates that it can provide good vulnerability detection for smart contracts. Yingli Zhang, Xin Liu 0050, Guodong Ye, Qun Jin, Jianhua Ma 0002, Qingguo Zhou |
IEEE Internet Things J. | 3 |
| 2022 | PG-VulNet: Detect Supply Chain Vulnerabilities in IoT Devices using Pseudo-code and GraphsabstractBackground: With the boosting development of IoT technology, the supply chains of IoT devices become more powerful and sophisticated, and the security issues introduced by code reuse are becoming more prominent. Therefore, the detection and management of vulnerabilities through code similarity detection technology is of great significance for protecting the security of IoT devices. Aim: We aim to propose a more accurate, parallel-friendly, and realistic software supply chain vulnerability detection solution for IoT devices. Method: This paper presents PG-VulNet, standing for Vulnerability-detection Network based on Pseudo-code Graphs. It is a ”multi-model” cross-architecture vulnerability detection solution based on pseudo-code and Graph Matching Network (GMN). PG-VulNet extracts both behavioral and structural features of pseudo-code to build customized feature graphs and then uses GMN to detect supply chain vulnerabilities based on these graphs. Results: The experiments show that PG-VulNet achieves an average detection accuracy of 99.14%, significantly higher than existing approaches like Gemini, VulSeeker, FIT, and Asteria. In addition to this, PG-VulNet also excels in detection overhead and false alarms. In the real-world evaluation, PG-VulNet detected 690 known vulnerabilities in 1,611 firmwares. Conclusions: PG-VulNet can effectively detect the vulnerabilities introduced by software supply chain in IoT firmwares and is well suited for large-scale detection. Compared with existing approaches, PG-VulNet has significant advantages. Xin Liu 0050, Yixiong Wu, Shangru Song, Qingguo Zhou, Jianwei Zhuge |
ESEM | 1 |
| 2022 | TCN enhanced novel malicious traffic detection for IoT devicesabstractWith the development of IoT technology, more and more IoT devices are connected to the network. Due to the hardware constraints of IoT devices themselves, it is difficult for developers to embed security software into them. Therefore, it is better to protect IoT devices at the traffic level. The effect of malicious traffic detection based on neural networks is promising. Still, the slow computation brings some difficulties to deploying AI-based detection systems on edge servers. Time Convolutional Network (TCN) is a high-speed neural network suitable for massively parallel computation. In this paper, we propose Multi-class S-TCN, an improved network supporting multiple classifications based on TCN for the practical needs of IoT scenarios. Besides, we implement a complete IoT traffic security detection procedure based on deep packet inspection and protocol analysis. The proposed Multi-class S-TCN significantly improves the detection speed without degrading the detection effect. Experiments show that this work has better detection performance and faster detection speed compared to existing approaches, proving the effectiveness of the proposed detection flow and Multi-class S-TCN in IoT scenarios. Xin Liu 0050, Ziang Liu 0006, Yingli Zhang, Dong Lv, Qingguo Zhou |
Connect. Sci. | 1 |
| 2021 | MECGuard: GRU enhanced attack detection in Mobile Edge Computing environment
Xin Liu 0050, Xiaokang Zhou, Qingguo Zhou |
Comput. Commun. | 1 |
| 2019 | Anomaly detection in ad-hoc networks based on deep learning model: A plug and play device
Xin Liu 0050, Binbin Yong, Rui Zhou 0005, Qingguo Zhou |
Ad Hoc Networks | 2 |