VLDB 2026 Research / reviewers in the wild / expert
Zhiyao Wang
dblp:61/6464
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From noisy feedback to evidence-aware issue specifications: an agent-governed retrieval-augmented generation approachabstractPost-release user feedback is a major control signal for maintenance and evolution in modern software development, yet it is noisy, fragmented, and difficult to translate into developer-usable issue specifications. Large Language Models (LLMs) can assist this transformation, but they often hallucinate or over-commit when evidence is weak, conflicting, or incomplete, limiting their robustness in automated software engineering workflows. We propose AGR (Agent-Governed Retrieval-Augmented Generation), a framework that regulates evidence acquisition and generation decisions via agentic control. AGR first applies an agentic triage step to filter low-signal or off-topic feedback, then retrieves evidence from a three-category hierarchy comprising official documentation, historical bug reports, and targeted web sources. It further performs confidence-weighted fusion across authoritative categories and uses an agentic decision module to verify relevance and sufficiency, trigger additional retrieval or online search when needed, reuse prior reports via memory, and abstain when evidence-supported grounding cannot be established. We evaluate AGR on two open-source software ecosystems, Firefox and VS Code. Results show that AGR achieves strong decision accuracy in triage and evidence verification, and produces more actionable and engineering-useful issue specifications than both raw feedback and a strong LLM baseline, while reducing unsupported details. Zhiyao Wang, Jialong Li 0001, Xiujing Guo, Tatsuhiro Tsuchiya |
Autom. Softw. Eng. | 1 |
| 2025 | RAG4Test: Retrieving GUI States for Multilingual Bug Report and Test Case Generation via LLMsabstractThe scalability and efficiency of software testing are persistently hindered by a reliance on manual practices for creating test cases and bug reports. Moreover, the valuable insights from post-release user feedback are often lost due to the lack of an automated pipeline connecting them to regression testing. To overcome these challenges, we present a novel framework that synergizes the structural representation of software with the generative power of Large Language Models (LLMs). We first represent the application’s Graphical User Interface (GUI) as a directed graph, capturing its components and navigational logic. This queryable graph provides essential, structured context for a Retrieval-Augmented Generation (RAG) model, which then autonomously generates and populates high-quality test cases and bug reports. A key innovation of our work is a fully automated pipeline that processes unstructured user feedback and error reports, transforming them into standardized test cases. We selected some reviews from a popular mobile App for preliminary experiments and verified the feasibility and efficiency improvement of this method. Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya |
APSEC | 1 |
| 2025 | Graph-Centric Approaches for Coverage Optimization in Software Requirement TestingabstractIn software testing, traceability links between software requirements and test cases are crucial for managing test coverage and detecting defects effectively. Accurate traceability enables comprehensive coverage analysis, identification of untested requirements, and targeted defect detection. The manual effort required to establish and update traceability links often leads to high labor costs and a greater risk of human error. Furthermore, as requirements evolve during the development lifecycle, the effort needed to maintain accurate links increases, driving up maintenance costs and further complicating test management.This study proposes a graph-based approach to establish and maintain traceability throughout the software testing process. By comparing various models for their ability to identify semantic relationships between software requirements and test cases, we selected the most effective method to create accurate traceability links. These links are further utilized through graph queries, enabling efficient analysis of test coverage, identification of untested requirements, and discovery of high-similarity requirement clusters, thereby enhancing the overall testing process. Automatically generating test cases based on query results, our approach seamlessly integrates into the software testing lifecycle, enhancing both coverage and efficiency. In a case study involving real-world industrial data, we effectively identified previously untested requirements and generated a substantial number of high-quality test cases. The results validate the applicability and effectiveness of our approach, demonstrating its potential to improve test traceability and reliability in practical software development environments. Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya |
COMPSAC | 1 |
| 2025 | SVFR: A Unified Framework for Generalized Video Face RestorationabstractFace Restoration (FR) is a crucial area within image and video processing, focusing on reconstructing high-quality portraits from degraded inputs. Despite advancements in image FR, video FR remains relatively under-explored, primarily due to challenges related to temporal consistency, motion artifacts, and the limited availability of high-quality video data. Moreover, traditional face restoration typically prioritizes enhancing resolution and may not give as much consideration to related tasks such as facial colorization and inpainting. In this paper, we propose a novel approach for the Generalized Video Face Restoration (GVFR) task, which integrates video blind face restoration (BFR), inpainting, and colorization tasks that we empirically show to benefit each other. We present a unified framework, termed as stable video face restoration (SVFR), which leverages the generative and motion priors of Stable Video Diffusion (SVD) and incorporates task-specific information through a unified face restoration framework. A learnable task embedding is introduced to enhance task identification. Meanwhile, a novel Unified Latent Regularization (ULR) is employed to encourage the shared feature representation learning among different subtasks. To further enhance the restoration quality and temporal stability, we introduce the facial prior learning and the self-referred refinement as auxiliary strategies. The proposed framework effectively combines the complementary strengths of these tasks, enhancing temporal coherence and achieving superior restoration quality. This work advances the state-of-the-art in video FR and establishes a new paradigm for generalized video face restoration. Code and video demo are available at https://github.com/wangzhiyaoo/SVFR.git. Zhiyao Wang, Xu Chen 0024, Chengming Xu 0001, Xiaobin Hu, Jiangning Zhang, Chengjie Wang 0001, Yiyi Zhou, Rongrong Ji |
CVPR | 1 |
| 2025 | Retrieval-Augmented Generation for Software Requirement-Based Test Case GenerationabstractTesters often need to manually write black-box test cases based on software artifacts such as requirement documents. In agile development, this process is often time-consuming and is further complicated by frequent requirement changes, leading to continuous maintenance overhead. Automating this process is therefore essential. Given the strong natural language understanding and generation capabilities of large language models (LLMs), combined with Retrieval-Augmented Generation (RAG), we propose a RAG-based framework for automated test case generation. Before generation, we embed software artifacts to construct a vectorbased knowledge database. At runtime, software requirements are used as queries to retrieve relevant context, which is integrated into a prompt and passed to the LLM for test case generation. This approach addresses several shortcomings of LLMs, including limited context length, attention dilution over large inputs, and the tendency to hallucinate or over-look key domain-specific constraints. By providing query-specific external knowledge, RAG enhances both accuracy and efficiency. We deploy the framework with different models locally and conduct experiments on two open-source datasets. Compared with the manually written benchmark test cases, our method achieves full requirement coverage with fewer test cases, improved efficiency, reduced error potential, and realized better readability. Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya |
QRS | 1 |
| 2023 | A Product Fuzzy Convolutional Network for Detecting Driving FatigueabstractExisting driving fatigue detection methods rarely consider how to effectively fuse the advantages of the electroencephalogram (EEG) and electrocardiogram (ECG) signals to enhance detection performance under noise conditions. To address the issues, this article proposes a new type of the deep learning (DL) framework based on EEG and ECG called the product fuzzy convolutional network (PFCN). It should be noted that this article first investigates how to fuse EEG and ECG signals to deal with driving fatigue detection under noise conditions in both simulated and real-field driving environments. Specifically, the PFCN includes three subnetworks. The first uses a fuzzy neural network (FNN) with feedback and a product layer, effectively capturing the particularity and temporal variation of high-dimensional EEG signals and reducing the time-space complexity. The second subnetwork uses a 1-D convolution to convert the ECG data into feature sequences, providing high accuracy and low computational complexity in ECG data classification. The third subnetwork proposes a fusion-separation mechanism to effectively fuse the extracted ECG and EEG features, suppressing the noise interference and ensuring higher detection accuracy. To evaluate the performance of PFCN, a series of experiments has been set up in both simulated and real-field driving environments. The results indicate that the proposed PFCN model has better robustness and detection accuracy compared with several mainstream fatigue detection models. Guanglong Du, Shuaiying Long, Chunquan Li 0001, Zhiyao Wang, Peter Xiaoping Liu |
IEEE Trans. Cybern. | 4 |
| 2021 | A TSK-Type Convolutional Recurrent Fuzzy Network for Predicting Driving FatigueabstractDriver fatigue monitoring is very important for driving safety, and many intricate factors while driving make fatigue monitoring harder. To effectively predict driving fatigue, this article proposes a new deep learning framework called TSK-type convolution recurrent fuzzy network (TCRFN) based on the spatial and temporal characteristics of electroencephalogram (EEG) signals. In TCRFN, the convolution block is first introduced to extract spatial dependencies from EEG signals. Furthermore, since EEG noise has a strong spatial dependence, this convolutional neural networks is used to reduce the impact of noise. Additionally, a new local feedback method in fuzzy neural network is proposed to process the EEG signals, which can better capture the temporal dependencies from EEG signals. Finally, a logarithmic spatial firing layer function is used in the proposed TCRFN. The activation performance of this function is smoother, which allows more feature numbers and provides better prediction. The experimental results show that the proposed TCRFN model has better antinoise performance and prediction accuracy compared with other widely used and state-of-the-art models. Guanglong Du, Zhiyao Wang, Chunquan Li 0001, Peter Xiaoping Liu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | A Convolution Bidirectional Long Short-Term Memory Neural Network for Driver Emotion RecognitionabstractReal-time recognition of driver emotions can greatly improve traffic safety. With the rapid development of communication technology, it becomes possible to process large amounts of video data and identify the driver's emotions in real time. To effectively recognize driver's emotions, this paper proposes a new deep learning framework called Convolution Bidirectional Long Short-term Memory Neural Network (CBLNN). This method predicts the driver's emotion based on the geometric features extracted from facial skin information and the heart rate extracted from changes in RGB components. The facial geometry features obtained by using Convolutional Neural Network (CNN) are intermediate variables for the heart rate analysis of Bidirectional Long Short Term Memory (Bi-LSTM). Subsequently, the output of Bi-LSTM is used as input to the CNN module to extract the hear rate features. CBLNN uses Multi-modal factorized bilinear pooling (MFB) to fuse the extracted information and classifies it into five common emotions: happiness, anger, sadness, fear and neutrality. Our emotion recognition method was tested, proving that it can be used to quickly and steadily recognize emotions in real time. Guanglong Du, Zhiyao Wang, Boyu Gao 0003, Shahid Mumtaz, Khamael M. Abualnaja, Cuifeng Du |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Secure Distribution of Big Data Based on BitTorrentabstractRecently, big data becomes more and more widespread on the Internet, as various P2P protocols, especially BitTorrent, make great contribution. Accompanied with BitTorrent spreading, however, malicious activities, divulging sensitive data and other security problems arise. Relative researches and analyses indicate that existing means of protecting P2P network are sophisticated but intricate when distributing big data, leading inefficiency of implementation. In this paper, a scheme to distribute big data securely and efficiently on BitTorrent network is proposed, which can be implemented in server, authorizing peers' admittances and actions, hence protecting sensitive data in the network. To achieve this goal, identity verification and cipher system are embedded into BitTorrent protocol, enabling the server to regulate and keep trace of peers' behaviors and sensitive data. The experimental results show the functional effectiveness of this scheme, as well as acceptable overhead on server. Chunjie Xu, Jingchao Qin, Guangjun Qin, Mingfa Zhu, Zhiyao Wang, Mingquan Li, Dongyu Tan |
DASC | 7 |