VLDB 2026 Research / reviewers in the wild / expert
Taiyan Wang
dblp:373/4047
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0007-0533-4404ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Software defect detection using large language models: a literature reviewabstractAbstract As software systems grow in complexity, the importance of efficient defect detection escalates, becoming vital to maintain software quality. In recent years, artificial intelligence technology has boomed. In particular, with the proposal of Large Language Models (LLMs), researchers have found the huge potential of LLMs to enhance the performance of software defect detection. This review aims to elucidate the relationship between LLMs and software defect detection. We categorize and summarize existing research based on the distinct applications of LLMs in dynamic and static detection scenarios. Dynamic detection methods are categorized based on the different phases in which they employ LLMs, such as using them for test case generation, providing feedback guidance, and conducting output assessment. Static detection methods are classified according to whether they analyze the source code or the binary of the software under test. Furthermore, we investigate the prompt engineering and model fine-tuning strategies adopted within these studies. Finally, we summarize the emerging trend of integrating LLMs into software defect detection, identify challenges to be addressed and prospect for some potential research directions. Yu Chen 0053, Yi Shen 0012, Taiyan Wang, Shiwen Ou, Yuwei Li 0002, Zulie Pan |
Frontiers Comput. Sci. | 3 |
| 2026 | Unreachable Features? Exposing the Security Risks of Invisible Interfaces in Embedded Web Services of IoT DevicesabstractIoT devices, now integral to our daily routines, offer unparalleled convenience but also face mounting security threats. Embedded web services, prevalent in public networks, pose a major risk to these devices. While research has focused on detecting vulnerabilities in IoT embedded web services, it has overlooked the presence of invisible interfaces, which have emerged as significant security threats. In this paper, we propose InvRadar, a novel framework for detecting vulnerabilities in invisible interfaces of embedded web services in IoT devices. Specifically, InvRadar identifies invisible interfaces by analyzing the differences between the front-end visible interface keywords and the back-end interface keywords through a correlation analysis method. Subsequently, InvRadar uses a static taint analysis method to detect the vulnerabilities that can be triggered by the invisible interfaces. To validate the performance of InvRadar, we conduct extensive experiments and compare InvRadar with the state-of-the-art methods. In testing 13 device firmware, InvRadar identifies 1,793 invisible interfaces and detects 124 vulnerabilities, including 53 newly discovered ones, with 34 receiving new CVE/CNVD IDs. Additionally, InvRadar outperforms the state-of-the-art methods in interface keyword extraction, border binary and data ingestion function identification. Yuanchao Chen, Yuwei Li 0002, Yi Shen 0012, Yu Chen 0053, Yang Li 0215, Taiyan Wang, Yuliang Lu, Zulie Pan, Shouling Ji |
IEEE Internet Things J. | 6 |
| 2025 | Fusing Multimodal Binary Code Representations for Enhanced Similarity DetectionabstractAs software reuse has become increasingly prevalent in the modern era, binary code similarity detection plays a critical role in program analysis. Numerous machine learning methods have been introduced to this field, with the primary challenge lying in the effective representation of binary code. While existing approaches leverage multimodal features for embedding, they still adopt relatively simple fusion methods to combine different modalities. Therefore, their performance can be limited by inadequate modality interaction during the feature fusion step and an over-reliance on a single modality in the final embedding step. To address these issues, we propose FuseBinRepr, a method that fuses text modality and graph modality representation techniques using a fusion model architecture and three specialized learning tasks to enhance binary code similarity detection. We design a fusion architecture integrating both text and graph embeddings via cross-attention and self-attention mechanisms. For model training, we developed three tasks: text-graph alignment (TGA), graph masking recovery (GMR), and contrastive learning (CL), to capture and align high-level semantics across modalities. Through evaluation, FuseBinRepr demonstrates improvements in Mean Reciprocal Rank (MRR) and Recall compared to state-of-the-art baseline methods, achieving increases of up to 40.8 % and 42 %, respectively. Ablation studies confirm the robustness of our model design, as well as the effectiveness of pretraining tasks. In the downstream software vulnerability detection task, FuseBinRepr achieves the best MRR and Recall performance in ranking CVE binary functions. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054 |
SRDS | 1 |
| 2025 | A survey of binary code representation technologyabstractBinary analysis, as an important foundational technology, provides support for numerous applications in the fields of software engineering and security research. With the continuous expansion of software scale and the complex evolution of software architecture, binary analysis technology is facing new challenges. To break through existing bottlenecks, researchers have applied artificial intelligence (AI) technology to the understanding and analysis of binary code. The core lies in characterizing binary code, i.e., how to use intelligent methods to generate representation vectors containing semantic information for binary code, and apply them to multiple downstream tasks of binary analysis. In this paper, we provide a comprehensive survey of recent advances in binary code representation technology, and introduce the workflow of existing research in two parts, i.e., binary code feature selection methods and binary code feature embedding methods. The feature selection section includes mainly two parts: definition and classification of features, and feature construction. First, the abstract definition and classification of features are systematically explained, and second, the process of constructing specific representations of features is introduced in detail. In the feature embedding section, based on the different intelligent semantic understanding models used, the embedding methods are classified into four categories based on the usage of text-embedding models and graph-embedding models. Finally, we summarize the overall development of existing research and provide prospects for some potential research directions related to binary code representation technology. Taiyan Wang, Qingsong Xie, Zulie Pan, Min Zhang 0054 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2025 | Whiskey: Large-Scale Identification of Mobile Mini-App Session Key Leakage With LLMsabstractMini-apps, which run on super-apps, have attracted a large number of users due to their lightweight nature and the convenience of supporting the authorized use of super-app user information. Super-apps employ encryption to protect the transmission of sensitive identity information authorized by users to the mini-app, using the session key as the key. However, we have identified a risk of session key leakage, which could be exploited to maliciously manipulate sensitive user identity information, thereby posing a significant threat to user data security. To reveal this damage, we explore potential business scenarios of session key leakage in detail. Nevertheless, the diversity in design among various mini-apps makes automated testing of these business scenarios at a large scale challenging. This diversity is reflected in the inconsistent naming of identical types of controls and the disparate execution orders of controls within the same business scenarios across different mini-apps. To overcome these challenges, we propose Whiskey, which can adaptively and intelligently optimize dynamic testing strategies for mini-apps with diverse designs using large language models to detect session key leakage at scale. We evaluated Whiskey on 157,063 WeChat mini-apps and 10,000 TikTok mini-apps, and found that 15,712 of WeChat mini-apps and 678 of TikTok mini-apps had session key leakage vulnerabilities. Further analysis showed that this leakage could lead to account takeover and promotion abuse attacks. We responsibly reported the detection results to Tencent and the mini-app vendors. At the time of submission, 17 reported issues had been assigned CNVD IDs. Yu Chen 0053, Yuanchao Chen, Taiyan Wang, Shouling Ji, Hong Shan, Zulie Pan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Enhancing Black-box Compiler Option Fuzzing with LLM through Command FeedbackabstractSince the compiler acts as a core component in software building, it is essential to ensure its availability and reliability through software testing and security analysis. Most research has focused on compiler robustness when compiling various test cases, while the reliability of compiler options lacks attention, especially since each option can activate a specific compiler function. Although some researchers have made efforts in testing it, the insufficient utilization of compiler command feedback messages leads to the poor efficiency, which hinders more diverse and in-depth testing.In this paper, we propose a novel solution to enhance black-box compiler option fuzzing by utilizing command feedback, such as error messages, standard output and compiled files, to guide the error fixing and option pruning via prompting large language models for suggestions. We have implemented the prototype and evaluated it on 4 versions of LLVM. Experiments show that our method significantly improves the detection of crashes, reduces false negatives, and even increase the success rate of compilation when compared to the baseline. To date, our method has identified hundreds of unique bugs, and 9 of them are previously unknown. Among these, 8 have been assigned CVE numbers, and 1 has been fixed following our report. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054, Huimin Ma 0004, Jinghua Zheng |
ISSRE | 1 |
| 2023 | Optir-SBERT: Cross-Architecture Binary Code Similarity Detection Based on Optimized LLVM IR
Yintong Yan, Taiyan Wang, Zulie Pan |
ICDF2C (2) | 3 |