VLDB 2026 Research / reviewers in the wild / expert
Yixuan Cheng
dblp:223/7877
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CrossCode2Vec: A unified representation across source and binary functions for code similarity detectionabstractCode similarity detection identifies code by analyzing similarities in syntax, semantics, and structure, which includes types of tasks: source-to-source, binary-to-binary, and source-to-binary. Due to encoding and representation disparities between source and binary code, existing methods have mainly focused on individual tasks, without providing a universal solution. Additionally, current source-to-binary tasks only achieve one-to-one matching between source code and binary functions, neglecting the one-to-many relationship inherent between source code and its cross-compiled binaries. In this paper, we propose CrossCode2Vec, a unified framework for representing code in both source and binary functions, which aims to bridge the gap in original coding features and provide a standardized similarity measurement across three code similarity detection tasks. For source code and its corresponding compiled binary, we first design an enhanced Abstract Path Context data preprocessing method, construct an abstract syntax tree (AST) from both source code functions and decompiled binary functions, and implement the function embedding followed by the pre-trained Word2vec model. Then we propose a task-specific data sampling strategy. We establish a one-to-one correspondence between source and binary functions through symbol tables and create a one-to-many relationship between source functions and their cross-compiled binaries based on sampling rules. Finally, we employ a hierarchical LSTM-attention network to facilitate the representation and similarity measurement of functions. We conduct both extrinsic and intrinsic evaluations to confirm the effectiveness of CrossCode2Vec in code representation and code similarity tasks, validating its superiority in model architecture and data processing methods. CrossCode2Vec demonstrates stable and exceptional performance across multiple experiments, reinforcing its ability to bridge the gap between source and binary code representations while effectively measuring their similarities. Gaoqing Yu, Jiuyang Lyu, Wenqing Fan, Yixuan Cheng, Aina Sui |
Neurocomputing | 6 |
| 2025 | CrossSimEmb: transformer-based embedding model for cross-compilation binary code similarity detection
Gaoqing Yu, Jiuyang Lyu, Yixuan Cheng, Aina Sui |
J. Supercomput. | 5 |
| 2023 | MSLFuzzer: black-box fuzzing of SOHO router devices via message segment list inferenceabstractAbstract The popularity of small office and home office routers has brought convenience, but it also caused many security issues due to vulnerabilities. Black-box fuzzing through network protocols to discover vulnerabilities becomes a viable option. The main drawbacks of state-of-the-art black-box fuzzers can be summarized as follows. First, the feedback process neglects to discover the missing fields in the raw message. Secondly, the guidance of the raw message content in the mutation process is aimless. Finally, the randomized validity of the test case structure can cause most fuzzing tests to end up with an invalid response of the tested device. To address these challenges, we propose a novel black-box fuzzing framework called MSLFuzzer. MSLFuzzer infers the raw message structure according to the response from a tested device and generates a message segment list. Furthermore, MSLFuzzer performs semantic, sequence, and stability analyses on each message segment to enhance the complementation of missing fields in the raw message and guide the mutation process. We construct a dataset of 35 real-world vulnerabilities and evaluate MSLFuzzer. The evaluation results show that MSLFuzzer can find more vulnerabilities and elicit more types of responses from fuzzing targets. Additionally, MSLFuzzer successfully discovered 10 previously unknown vulnerabilities. Yixuan Cheng, Wenqing Fan, Gaoqing Yu |
Cybersecur. | 1 |
| 2022 | FIoTFuzzer: Response-Based Black-Box fuzzing for IoT DevicesabstractTo prevent IoT devices from being exploited, it is particularly important to detect vulnerabilities as many as possible during the device development process. The black-box fuzzing test is widely used in vulnerability detection for IoT devices for several reasons. First of all, the source code of the firmware is rarely provided in public, device response messages are a valuable source of device status. In legacy black-box fuzzing tests, there was a lack of checks on network protocols, message formats and encodings. Byte-to-byte mutation without these checks produced a large amount of garbage input data, which could not reach the deep-level function code. The efficiency and accuracy of fuzzing testing were negatively impacted accordingly. Secondly, communication protocol specification of firmware is rarely provided in public too, and it is difficult for existing grammar-based fuzzing strategies to distinguish the meaning of each field of the message. To solve the above issues, this paper proposes a response-based black-box fuzzing method, named FIoTFuzzer. We set up a message adapter to identify the protocol, format, encoding and other information of original communication packets. To improve the syntax inference capability, FIoTFuzzer divides the message segment based on the response, avoiding blind mutation of the content. This method of using mutation strategy based on message segment under the premise of format specification can reach deep functional components of smart devices. This fuzzing method has lightweight dependencies and does not require reverse engineering. Our tests were evaluated on 12 IoT devices, which included routers, smart bulbs and IP cameras. The results show that: (1) FIoFuzzer is able to detect real-world vulnerabilities in IoT devices; (2) In our benchmark comparison tests with Boofuzz and Sulley, FIoTFuzzer detected 9 vulnerabilities while Boofuzz detected only 5 and Sulley detected only 4 among these 9 vulnerabilities. Wenqing Fan, Yixuan Cheng |
ICIS | 4 |
| 2021 | Improved Fluctuation Derived Block Selection Strategy in Pixel Value Ordering Based Reversible Data Hiding
Linna Zhou, Guang Tang, Yuchen Wen, Yixuan Cheng |
IWDW | 5 |