VLDB 2026 Research / reviewers in the wild / expert
Yanyang Zhao
dblp:78/5804
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-1663-8901ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Camveil: Unveiling Security Camera Vulnerabilities Through Multi-Protocol Coordinated Fuzzing
Fuchen Ma, Yuqiao Yang, Yuanliang Chen, Yanyang Zhao, Ting Chen 0002, Yu Jiang 0001 |
SP | 4 |
| 2026 | Scalable hierarchical protocol format inference via feature-heuristic message delimiter
Yanyang Zhao, Zhengxiong Luo 0002, Ronghua Shi, Yu Jiang 0001, Heyuan Shi |
Empir. Softw. Eng. | 3 |
| 2026 | Multimodal Path Semantics and Contrastive Domain Adaptation for Cross-Domain Defect Pattern Prediction
Xiaoqin Ma, Xiangxiang Huang, Yanyang Zhao |
Softw. Qual. J. | 5 |
| 2025 | CMFuzz: Parallel Fuzzing of IoT Protocols by Configuration Model Identification and SchedulingabstractIoT protocols are essential for the communication among diverse devices. In real-world scenarios, IoT protocols utilize flexible configurations to meet various use cases. These configurations can significantly impact the protocols’ execution paths, with many bugs emerging only under specific configurations. Fuzzing has become a prominent technique for uncovering vulnerabilities in IoT protocol implementations. However, traditional fuzzing approaches are typically conducted using fixed or default configurations, overlooking potential issues that might arise in different settings. This limitation can lead to missing critical bugs that appear only under alternative configurations.In this paper, we propose CMFuzz, a parallel fuzzing framework designed to improve fuzzing effectiveness of IoT protocols through configuration identification and scheduling. CMFuzz first constructs a generalized protocol configuration model by systematically extracting configuration items from protocol implementations. Then, based on this model, CMFUZZ defines the relations among configuration items and introduces a relation-aware allocation mechanism to distribute them across parallel fuzzing instances. For evaluation, We implement CMFuzz on top of the widely-used protocol fuzzer Peach and conduct experiments on six popular IoT protocols. Compared to the original parallel mode of Peach and state-of-the-art parallel protocol fuzzer SPFuzz, CMFuzz covers an average of 34.4% and 28.5% more branches within 24 hours. Additionally, CMFuzz has detected 14 previously-unknown bugs in these real-world IoT protocols. Fuchen Ma, Yuanliang Chen, Feifan Wu, Yanyang Zhao, Heyuan Shi, Yu Jiang 0001 |
DAC | 6 |
| 2025 | Understanding and Detecting SQL Function Bugs: Using Simple Boundary Arguments to Trigger Hundreds of DBMS BugsabstractBuilt-in SQL functions are crucial in Database Management Systems (DBMSs), supporting various operations and computations across multiple data types. They are essential for querying, data transformation, and aggregation. Despite their importance, the bugs in SQL functions have caused widespread problems in the real world, from system failures to arbitrary code execution. However, the understanding of the bug characteristics is limited. More importantly, conventional function testing methods struggle to generate semantically correct SQL test cases, while DBMS testing efforts are hard to measure built-in SQL functions. Jingzhou Fu, Jie Liang 0006, Zhiyong Wu 0010, Yanyang Zhao, Shanshan Li 0001, Yu Jiang 0001 |
EuroSys | 4 |
| 2025 | DualFuzz: Detecting Vulnerability in Wi-Fi NICs through Dual-Directional FuzzingabstractWi-Fi Network Interface Cards (NICs) are vital for enabling wireless connectivity across a wide range of devices. Ensuring their security is critical, as vulnerabilities can expose entire networks to threats. Fuzzing is a promising technique for detecting such flaws. However, existing Wi-Fi fuzzers typically test transmission and reception separately, overlooking their interactions and resulting in inefficient testing.In this work, we present DualFuzz, a dual-directional fuzzing framework designed to simultaneously test both transmission and reception processes in Wi-Fi NICs. First, DualFuzz automatically identifies interaction behaviors within Wi-Fi NICs and constructs a Transmission-Reception Model (TRModel) to characterize Wi-Fi frames that influence these interactions. Leveraging this model, DualFuzz utilizes latency guided fuzzing to efficiently coordinate exploring transmission and reception interaction logics. Finally, we propose liveness and equivalence detectors that enable real-time monitoring to identify abnormal states and uncover potential vulnerabilities in Wi-Fi NICs. We implemented and evaluated DualFuzz on eight widely used Wi-Fi NICs, incorporating chipsets from various manufacturers (e.g., Intel and Realtek). Compared to state-of-the-art Wi-Fi fuzzers like OwFuzz, wpaspy, and Greyhound, DualFuzz detects 75%, 163%, and 250% more vulnerabilities, respectively. In total, it uncovered 21 previously unknown vulnerabilities, 7 of which have been assigned CVEs. Yuanliang Chen, Fuchen Ma, Yanyang Zhao, Yuanyi Li, Yu Jiang 0001 |
ASE | 3 |
| 2025 | Protocol syntax recovery via knowledge transfer
Yanyang Zhao, Zhengxiong Luo 0002, Feifan Wu, Heyuan Shi, Yu Jiang 0001 |
Comput. Networks | 1 |
| 2024 | SPFuzz: Stateful Path based Parallel Fuzzing for Protocols in Autonomous VehiclesabstractProtocols in autonomous vehicles are essential for efficient in-vehicle network communication. To ensure their security, many research efforts have been paid to the fuzz testing of their implementations. However, those fuzzing optimizations often struggle to manage the protocols' complex state, resulting in low efficiency in branch covering and vulnerability detection. Junze Yu, Zhengxiong Luo 0002, Fangshangyuan Xia, Yanyang Zhao, Heyuan Shi, Yu Jiang 0001 |
DAC | 4 |
| 2024 | MDIplier: Protocol Format Recovery via Hierarchical InferenceabstractNetwork protocol reverse engineering is crucial for a wide range of security applications. Many existing techniques accomplish this task by analyzing network traces. However, these methods globally cluster messages and analyze each cluster separately, which causes the loss of valuable field information. To address this problem, we present MDIplier, a protocol reverse engineering tool that leverages the hierarchical structure of protocol messages and performs tailored analysis at each message layer. MDIplier performs an iterative inference process. During each iteration, it identifies the message delimiter for layer separation and infers the format for each layer separately, optimizing the use of available field information. Our evaluation of eight widely used protocols shows that MDIplier outperforms state-of-the-art methods. It identifies fields with a perfection score 4.6×, 1.4×, 5.8×, and 1.8× higher than that of Netzob, Netplier, FieldHunter, and BinaryInferno, respectively. Furthermore, the experiments on proprietary protocols used in three IoT devices demonstrate the effectiveness of MDIplier in real-world scenarios. Zhengxiong Luo 0002, Yanyang Zhao, Ronghua Shi, Yu Jiang 0001, Heyuan Shi |
ISSRE | 3 |
| 2024 | Logos: Log Guided Fuzzing for Protocol ImplementationsabstractNetwork protocols are extensively used in a variety of network devices, making the security of their implementations crucial. Protocol fuzzing has shown promise in uncovering vulnerabilities in these implementations. However traditional methods often require instrumentation of the target implementation to provide guidance, which is intrusive, adds overhead, and can hinder black-box testing. This paper presents Logos, a protocol fuzzer that utilizes non-intrusive runtime log information for fuzzing guidance. Logos first standardizes the unstructured logs and embeds them into a high-dimensional vector space for semantic representation.Then, Logos filters the semantic representation and dynamically maintains a semantic coverage to chart the explored space for customized guidance.We evaluate Logos on eight widely used implementations of well-known protocols. Results show that, compared to existing intrusive or expert knowledge-driven protocol fuzzers, Logos achieves 26.75%-106.19% higher branch coverage within 24 hours. Furthermore, Logos exposed 12 security-critical vulnerabilities in these prominent protocol implementations, with 9 CVEs assigned. Feifan Wu, Zhengxiong Luo 0002, Yanyang Zhao, Qingpeng Du, Junze Yu, Ruikang Peng, Heyuan Shi, Yu Jiang 0001 |
ISSTA | 3 |
| 2024 | DynPRE: Protocol Reverse Engineering via Dynamic Inference
Zhengxiong Luo 0002, Yanyang Zhao, Feifan Wu, Junze Yu, Heyuan Shi, Yu Jiang 0001 |
NDSS | 3 |
| 2024 | Parallel Fuzzing of IoT Messaging Protocols Through Collaborative Packet GenerationabstractInternet of Things (IoT) messaging protocols play an important role in facilitating communications between users and IoT devices. Mainstream IoT platforms employ brokers, server-side implementations of IoT messaging protocols, to enable and mediate this user-device communication. Due to the complex nature of managing communications among devices with diverse roles and functionalities, comprehensive testing of the protocol brokers necessitates collaborative parallel fuzzing. However, being unaware of the relationship between test packets generated by different parties, existing parallel fuzzing methods fail to explore the brokers’ diverse processing logic effectively. This article introduces MPFuzz, a parallel fuzzing tool designed to secure IoT messaging protocols through collaborative packet generation. The approach leverages the critical role of certain fields within IoT messaging protocols that specify the logic for message forwarding and processing by protocol brokers. MPFuzzemploys an information synchronization mechanism to synchronize these key fields across different fuzzing instances and introduces a semantic-aware refinement module that optimizes generated test packets by utilizing the shared information and field semantics. This strategy facilitates a collaborative refinement of test packets across otherwise isolated fuzzing instances, thereby boosting the efficiency of parallel fuzzing. We evaluated MPFuzzon six widely used IoT messaging protocol implementations. Compared to two state-of-the-art protocol fuzzers with parallel capabilities, Peach and AFLNet, as well as two representative parallel fuzzers, SPFuzz and AFLTeam, MPFuzzachieves (6.1%,$174.5\times $), (20.2%,$607.2\times $), (1.9%,$4.1\times $), and (17.4%,$570.2\times $) higher branch coverage and fuzzing speed under the same computing resource. Furthermore, MPFuzzexposed seven previously unknown vulnerabilities in these extensively tested projects, all of which have been assigned with CVE identifiers. Zhengxiong Luo 0002, Junze Yu, Qingpeng Du, Yanyang Zhao, Feifan Wu, Heyuan Shi, Wanli Chang 0001, Yu Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Eliminating the high false-positive rate in defect prediction through BayesNet with adjustable weightabstractAbstract In defect prediction, a high false‐positive rate (FPR) caused by class imbalance not only increases the workload of testing and development but also consumes unnecessary costs. Many defect models against class imbalance have been proposed to improve the accuracy of defect prediction, but their ability to reduce FPR is unclear. To solve these problems, we first proposed a BayesNet with adjustable weights, called WBN, to reduce the FPR in software defect prediction, which is an algorithm independent of data preprocessing techniques. The mechanism of our WBN is to change the sampling probability of the misclassified instances when training the defect model, making the BayesNet model focus more on false alarm instances. And then, we investigate the FPR of five mainstream defect models for solving class imbalance and select them as comparison models to test the validity of our methods. The experimental result on eight open‐source projects shows that a) our WBN, in in‐version defect prediction (IVDP) and cross‐version defect prediction (CVDP), effectively reduces FPR with means of 0.384 and 0.322, respectively; b) compared with improved subclass discriminant analysis (ISDA) that is the lowest FPR in all control models, our WBN not only reduced the FPR but maintained recall whose mean value was 0.797, whereas ISDA did not, with an average recall of only 0.397; c) our WBN, in CVDP, not only reduces FPR, but also has significant superiority over five control defect models and baseline. Besides, we also found that the class imbalance difference between the test set and the training set has an impact on CVDP performance, recommending that practitioners choose the best dataset for CVDP from the defect data of the historical version through special technology. Yanyang Zhao, Dalin Zhang 0003, Yunzhan Gong |
Expert Syst. J. Knowl. Eng. | 1 |
| 2022 | ST-TLF: Cross-version defect prediction framework based transfer learningabstractCross-version defect prediction (CVDP) is a practical scenario in which defect prediction models are derived from defect data of historical versions to predict potential defects in the current version. Prior research employed defect data of the latest historical version as the training set using the empirical recommended method, ignoring the concept drift between versions, which undermines the accuracy of CVDP. We customized a Selected Training set and Transfer Learning Framework (ST-TLF) with two objectives: a) to obtain the best training set for the version at hand, proposing an approach to select the training set from the historical data; b) to eliminate the concept drift, designing a transfer strategy for CVDP. To evaluate the performance of ST-TLF, we investigated three research problems, covering the generalization of ST-TLF for multiple classifiers, the accuracy of our training set matching methods, and the performance of ST-TLF in CVDP compared against state-of-the-art approaches. The results reflect that (a) the eight classifiers we examined are all boosted under our ST-TLF, where SVM improves 49.74% considering MCC, as is similar to others; (b) when performing the best training set matching, the accuracy of the method proposed by us is 82.4%, while the experience recommended method is only 41.2%; (c) comparing the 12 control methods, our ST-TLF (with BayesNet), against the best contrast method P15-NB, improves the average MCC by 18.84%. Our framework ST-TLF with various classifiers can work well in CVDP. The training set selection method we proposed can effectively match the best training set for the current version, breaking through the limitation of relying on experience recommendation, which has been ignored in other studies. Also, ST-TLF can efficiently elevate the CVDP performance compared with random forest and 12 control methods. Yanyang Zhao, Yuwei Zhang 0003, Dalin Zhang 0003, Yunzhan Gong, Dahai Jin |
Inf. Softw. Technol. | 1 |