Xiaokang Yin 0002

dblp:439/1829-2 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-1617-4561ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RPKClust: region-partitioned keywords inference for binary protocol reverse
abstract
Abstract Protocol reverse engineering is a critical technology for analyzing unknown binary protocols. Message clustering serves as a fundamental and widely adopted step, playing a pivotal role in inferring both protocol format and state machine. Currently, most methods use multiple sequence alignment as a core technique for message clustering, where the degree of difference between messages is calculated. This may lead to the loss of valuable information and incur relatively high costs. To address this issue, we propose a novel binary protocol message clustering method, named RPKClust, based on region-based keyword positioning. By leveraging the characteristics of field offsets in messages, this method divides protocol messages into the fixed-offset region and the non-fixed-offset region. RPKClust adopts different keyword candidate generation strategies in these two regions. Subsequently, keyword fields are inferred through two-stage probability constraints, thus completing the clustering of protocol messages. We evaluated eight widely used protocols, and the results show that RPKClust outperforms the state-of-the-art methods (i.e. Netplier, MDIplier, ProInfer, NEMETYL). Its clustering results achieve a homogeneity of 0.959, a completeness of 0.941, and a V-measure of 0.949, and it significantly reduces the overhead. Furthermore, we validated the effectiveness of RPKClust on two specialized protocols and further verified its significant role in state machine inference.
Qichao Yang, Xiaokang Yin 0002, Fangfang Zhao, Shengli Liu 0003
Comput. J.2
2026 C2Detector: Interaction-enhanced semantic-aware detection method for C2 channels
Youqiang Luo, Ruijie Cai, Xiaokang Yin 0002, Jingman Zhou, Fangfang Zhao, Zhenjie Xie, Shengli Liu 0003
Comput. Networks3
2026 FieldWeaver: A visual-language approach to binary protocol format inference
Qichao Yang, Fangfang Zhao, Xiaokang Yin 0002, Ruijie Cai, Shengli Liu 0003
Comput. Networks3
2026 kAPR: A coverage-guided, context-aware agent for automated repair of Linux kernel bugs
Bingzheng Li, Xiaokang Yin 0002, Yao Zhang 0019, Shengli Liu 0003, Shouling Ji
Inf. Softw. Technol.2
2026 Nonstandard Sinks Matter: A Comprehensive and Efficient Taint Analysis Framework for Vulnerability Detection in Embedded Firmware
abstract
The discovery of vulnerabilities in embedded firmware has received significant attention from security researchers. However, current vulnerability detection methods still suffer from false negatives and inefficiency, which limit detection effectiveness and require substantial analysis time. To alleviate the above problems, we propose a bidirectional path and data flow analysis method, named BPDA, that effectively compensates for the limitations in detecting firmware vulnerabilities at nonstandard sink points. Our key insight is that, some vulnerabilities arise in nonstandard library sinks, and not all user inputs can reach each corresponding sink. Guided by these insights, we design a more comprehensive sink identification algorithm and leverage accurate backward data flow tracking to eliminate the non-vulnerable paths. After that, we execute forward taint analysis and generate the final Proof of Concepts (PoCs). To evaluate the effectiveness of BPDA, we evaluated it on 84 firmware samples (including both Linux and VxWorks firmware) from 8 major brands, comparing it with state-of-the-art methods (i.e., SaTC and Mango). BPDA discovered 163 real vulnerabilities, including 34 0-day vulnerabilities, of which 32 have been confirmed by CVE/CNVD. Besides, results show that BPDA completed its analysis in just 6% of the time required by SaTC, and remarkably identified 21 vulnerabilities that SaTC and Mango had not detected. It also resolved the issue of Mango failing to analyze specific firmware. In addition, we also performed an ablation study to verify the effectiveness of optimization methods in taint analysis. These results demonstrate the superiority of BPDA in terms of effectiveness and efficiency in detecting embedded firmware vulnerabilities.
Enzhou Song, Jinyuan Zhai, Ruijie Cai, Qichao Yang, Xiaokang Yin 0002, Shengli Liu 0003
IEEE Trans. Dependable Secur. Comput.8
2025 Precise Discovery of More Taint-Style Vulnerabilities in Embedded Firmware
abstract
The proliferation of taint-style vulnerabilities in embedded devices poses a significant threat to cybersecurity. However, discovering these vulnerabilities is challenging due to their vast number and variety. While current solutions for discovering vulnerabilities in embedded firmware have achieved some success, they suffer from imprecision, are time-consuming, and fail to consider sensitive sinks and constraints. To address these challenges, we propose a novel taint-style vulnerability discovery method called SinkTaint. SinkTaint incorporates backtracking and constraint analysis to achieve high precision and employs a global taint keyword identification strategy to identify implicit taint keywords. It identifies additional sinks using static analysis and performs backtracking analysis to eliminate sanitized sinks, while retrieving the parameter's length for risky sinks. Furthermore, SinkTaint employs dual-label labeling strategies for taint keywords and data, propagating taint labels based on function return values. Finally, SinkTaint employs symbolic execution-based taint analysis to discover taint-style vulnerabilities. We evaluate SinkTaint on datasets released by SaTC and 10 known overflow vulnerabilities. Compared to state-of-the-art methods, including Karonte, SaTC, and EmTaint, SinkTaint demonstrated superior performance, discovering more vulnerabilities with an increase in vulnerability discovery effectiveness by 472%. To date, SinkTaint has identified 21 high-risk taint-style vulnerabilities that were previously undisclosed.
Xiaokang Yin 0002, Ruijie Cai, Xiaoya Zhu, Qichao Yang, Enzhou Song, Shengli Liu 0003
IEEE Trans. Dependable Secur. Comput.1
2023 ConFunc: Enhanced Binary Function-Level Representation through Contrastive Learning
abstract
Binary code similarity detection (BCSD) has numerous applications, including malware detection, vulnerability search, plagiarism detection, and patch identification. Recent studies have demonstrated that with the rapid progress of machine learning (ML) techniques, various BCSD approaches based on machine learning have exhibited stronger performance than traditional methods. However, current ML-based BCSD approaches tend to ignore the issue of training samples, and most ML-based BCSD approaches are based on supervised learning, which is suffered from the labelling difficulties. To mitigate these issues, we propose ConFunc: a function-level binary code similarity detection framework based on contrastive learning. Performance evaluation shows that ConFunc enhances the Mean Reciprocal Rank (MRR) and Recall rates (Recall@1) of baseline models by fully harnessing the potential of the data. Additionally, ConFunc demonstrates stronger performance in scenarios with scarce data, achieving the baseline model’s performance on the entire dataset using only 10% of the complete dataset. In real-world patch identification and vulnerability search tasks, ConFunc consistently outperforms other baseline models in MRR and Recall@10.
Xiaokang Yin 0002, Xiao Li 0032, Xiaoya Zhu, Shengli Liu 0003
TrustCom2