VLDB 2026 Research / reviewers in the wild / expert
Yuzhou Liu 0001
dblp:173/6435-1
· DBLP profile ↗
22ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0003-2765-4074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature SelectionabstractMulti-view multi-label feature selection aims to identify informative features from heterogeneous views, where each sample is associated with multiple interdependent labels. This problem is particularly important in machine learning involving high-dimensional, multimodal data such as social media, bioinformatics or recommendation systems. Existing Multi-View Multi-Label Feature Selection (MVMLFS) methods mainly focus on analyzing statistical information of data, but seldom consider semantic information. In this paper, we aim to use these two types of information jointly and propose a method that combines Large Language Models (LLMs) semantic reasoning with Graph Neural Networks (GNNs) structural modeling for MVMLFS. Specifically, the method consists of three main components. (1) LLM is first used as an evaluation agent to assess the latent semantic relevance among feature, view, and label descriptions. (2) A semantic-aware heterogeneous graph with two levels is designed to represent relations among features, views and labels: one is a semantic graph representing semantic relations, and the other is a statistical graph. (3) A lightweight Graph Attention Network (GAT) is applied to learn node embedding in the heterogeneous graph as feature saliency scores for ranking and selection. Experimental results on multiple benchmark datasets demonstrate the superiority of our method over state-of-the-art baselines, and it is still effective when applied to small-scale datasets, showcasing its robustness, flexibility, and generalization ability. Yuzhou Liu 0001, Wanfu Gao |
AAAI | 2 |
| 2026 | Redundancy-optimized Multi-head Attention Networks for Multi-view Multi-label Feature SelectionabstractMulti-view multi-label data offers richer perspectives for artificial intelligence, but simultaneously presents significant challenges for feature selection due to the inherent complexity of interrelations among features, views and labels. Attention mechanisms provide an effective way for analyzing these intricate relationships. They can compute importance weights for information by aggregating correlations between Query and Key matrices to focus on pertinent Values. However, existing attention-based feature selection methods predominantly focus on intra-view relationships, neglecting the complementarity of inter-view features and the critical feature-label correlations. Moreover, they often fail to account for feature redundancy, potentially leading to suboptimal feature subsets. To overcome these limitations, we propose a novel method based on Redundancy-optimized Multi-head Attention Networks for Multi-view Multi-label Feature Selection (RMAN-MMFS). Specifically, we employ each individual attention head to model intra-view feature relationships and use the cross-attention mechanisms between different heads to capture inter-view feature complementarity. Furthermore, we design static and dynamic feature redundancy terms: the static term mitigates redundancy within each view, while the dynamic term explicitly models redundancy between unselected and selected features across the entire selection process, thereby promoting feature compactness. Comprehensive evaluations on six real-world datasets, comparing against six multi-view multi-label feature selection methods, demonstrate the superior performance of the proposed method. Yuzhou Liu 0001, Wanfu Gao |
AAAI | 1 |
| 2026 | PBSketch: Finding Periodic Burst Items in Data StreamsabstractDetecting periodic burst (PB) items in data streams is crucial for applications like rate limiting but remains unexplored. % While combining existing sketch algorithms offers a baseline, it suffers from significant inaccuracy and inefficiency. In this paper, we propose PBSketch, the first dedicated sketch algorithm designed for detecting PB items in real time. Its key techniques mainly include: 1) a two-stage hierarchical structure that efficiently maintains potential burst items and discards those without potential; 2) a fine-grained PB selection mechanism during window processing, coupled with the Window Smoothing Processing optimization to amortize performance overhead and eliminate processing spikes. % We provide its error bounds through rigorous theoretical analysis. Our extensive experiments show that PBSketch outperforms the baseline solution in accuracy and speed. By deploying it on an FPGA platform, the throughput is further significantly improved. Moreover, it effectively optimizes a practical application of rate limiting, clearly improving performance with almost negligible overhead. Zhuochen Fan, Zhongxian Liang, Zirui Liu 0002, Dayu Wang, Dong Wen 0004, Wenjun Li 0004, Tong Yang 0003, Yuzhou Liu 0001, Weizhe Zhang |
KDD (1) | 8 |
| 2026 | Meta-learning-based multi-view multi-label feature selection with streaming labels
Yuzhou Liu 0001, Pengyuan Gao, Wanfu Gao |
Knowl. Based Syst. | 1 |
| 2026 | Multimodal information fusion for software vulnerability detection based on both source and binary codes
Yuzhou Liu 0001, Shuang Jiang, Hongxu Tian, Peng Zhang 0053 |
Sci. Comput. Program. | 1 |
| 2025 | API comparison based on the non-functional information mined from Stack Overflow
Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu, Peng Zhang 0053 |
Sci. Comput. Program. | 2 |
| 2025 | DAOR: Distinguish Similar Machine Learning APIs Based on Official Documents and ReviewsabstractABSTRACT Background In recent years, machine learning (ML) APIs have emerged as a valuable resource for addressing complex problems, such as image recognition. However, developers should use ML APIs carefully, as they have their own characteristics different from traditional ones: an ML API has its own training data set, a concrete target task, and its output is often the probability. As a result, developers may use an inappropriate API, and the program can still run without reporting errors, especially as there are many similar ML APIs provided by different platforms. Methods This paper proposes an approach called DAOR to help developers use ML APIs properly in their tasks. First, a comparative analysis of ML APIs is conducted, leveraging information from documentation and user reviews to identify comparable APIs. This involves extracting differences from the documentation, categorized into inputs, functions, and outputs, and summarizing key information from user reviews using GPT‐driven prompts. Finally, a visualization framework is designed to summarize and show the results. Evaluation and Results To evaluate the approach, a series of experiments is conducted based on the ML APIs from two famous platforms, Amazon Web Service AI and IBM Watson. The results show that useful information for distinguishing similar ML APIs can be gained, and it is helpful for developers to use the ML APIs correctly. Shuang Jiang, Junxin Yang, Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu |
Softw. Pract. Exp. | 3 |
| 2025 | Multimodal Fusion for Android Malware Detection Based on Large Pre-Trained ModelsabstractMalware detection is a critical issue in software engineering as it directly threatens user information security. Existing approaches often focus on individual modality (either source code or binary code) for the detection, but it ignores to effectively exploit the complementary information between them. This limits the detection performance, especially in complex and evasive malware scenarios. In this paper, we take Android applications written in Java as objects, and provide a novel fine-grained multimodal fusion method with large pre-trained models to combine the features from source and binary codes for the malware detection. For the source code modality, we employ the graphical user interface (GUI) as a framework to segment the source code into snippets, and use a pre-trained programming language model to extract feature representations. For the binary code modality, we convert binary code into grayscale images and fine-tune a pre-trained vision model to extract features indirectly. We then implement cross-modal attention and devise a contrastive loss to align features across modalities, supplementing this with supervised classification loss to refine the multimodal fusion process specifically for malware detection. Our experiments, conducted using the Data-MD and Data-MC benchmarks, demonstrate that our approach achieves a precision of 0.977 and a recall of 0.984 in detecting malware. This underscores the advantages of using large pre-trained models for feature representation and the fusion of information across different modalities for effective malware detection. Lei Liu 0040, Yuzhou Liu 0001, Yu Zhao 0010, Peng Zhang 0053, Huaxiao Liu |
IEEE Trans. Software Eng. | 3 |
| 2023 | AGAA: An Android GUI Accessibility Adapter for Low Vision UsersabstractThe graphical user interface (GUI) is crucial for users to interact with mobile devices. However, accessibility issues in the GUI, such as undersized text and redundant information, lead to understanding and operating obstacles for billions of low vision users in our society. To alleviate this situation, academia and industry have proposed various accessibility-related methods. Still, their over-dependence on specific detection rules and their inability to automatically repair the GUI source code limit them in helping developers resolve these issues. In this paper, we propose a novel method, named AGAA, for capturing and repairing undersized text and redundant information issues in the GUI. The evaluation on 12 real-world apps and the user study on 36 low vision users demonstrate that AGAA is effective in resolving these issues and is useful in improving the mobile device experience for low vision users, respectively. Yifang Xu, Zhuopeng Li, Huaxiao Liu, Yuzhou Liu 0001 |
COMPSAC | 4 |
| 2023 | Describing the APIs comprehensively: Obtaining the holistic representations from multiple modalities data for different tasks
Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu |
Inf. Softw. Technol. | 4 |
| 2023 | A lightweight API recommendation method for App development based on multi-objective evolutionary algorithm
Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu |
Sci. Comput. Program. | 3 |
| 2022 | Consistent or not? An investigation of using Pull Request Template in GitHub
Huaxiao Liu, Chunyang Chen 0001, Yuzhou Liu 0001, Shuotong Bai |
Inf. Softw. Technol. | 4 |
| 2021 | Supporting features updating of apps by analyzing similar products in App stores
Huaxiao Liu, Yuzhou Liu 0001, Shanquan Gao |
Inf. Sci. | 3 |
| 2021 | API recommendation for the development of Android App features based on the knowledge mined from App stores
Shanquan Gao, Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu |
Sci. Comput. Program. | 3 |
| 2021 | App recommendation based on both quality and securityabstractAbstract With the rapid prevalence of smartphones and the dramatic proliferation of mobile applications, people tend to do everything at their fingertips, including some sensitive activities, such as bank transfers. This makes security become one important factor when recommending apps to users. However, most existing methods recommend apps only on the basis of the apps' functionalities. Even when some methods take security into account, they usually roughly group apps with functionalities and identify the products using extra permissions as risky, but this ignores a common phenomenon that these permissions may be used only to achieve the corresponding functionalities. In this paper, we propose an app recommendation method considering both functionalities and security. For functionalities, we summarized them from app descriptions and further evaluated their completion quality in different products by analyzing their related reviews. For security, we cluster apps with similar functionalities and quality and analyze the permissions of apps in a more comparable range. In this way, our method recommends apps with higher completion quality of functionalities and security degree to users according to their demands. We conducted experiments on apps collected from six categories of Google Play, and the results show that our method has a good recommendation effect. Shanquan Gao, Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu, Peixun Liu |
J. Softw. Evol. Process. | 3 |
| 2021 | Application programming interface recommendation according to the knowledge indexed by app feature mined from app storesabstractAbstract Application programming interfaces (APIs) play an important role in the increasingly competitive mobile application development industry, as they can greatly improve the efficiency of app development. However, finding proper APIs is often time‐consuming for the gap between the knowledge of APIs and app features. To solve this problem, we give an approach to summarize the wisdom of developers contained in the products in app stores and establish the system of API knowledge indexed by app features for the API recommendation. First, we extract features from the app descriptions and define the feature framework. Second, we parse the APK files of apps to gain the methods in code and APIs called by them and further introduce such API knowledge into the feature framework by utilizing method names as bridges. Finally, according to features in developers' queries, we locate corresponding feature nodes in the API knowledge system and recommend related API knowledge to developers. We conduct experiments based on 38,952 apps from five categories on Google Play, and the experimental results show that our approach has a good recommendation effect for the queries on app features. Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu |
J. Softw. Evol. Process. | 3 |
| 2020 | Combining goal model with reviews for supporting the evolution of appsabstractTo support the iterative development process of Apps, the goal model is not only established to describe the requirements at the early stage but also used for identifying the updating strategy in every iteration. In this process, reviews from users provide valuable information for developers to analyse the model with users sentiments. In this study, the authors combine the goal model with reviews for supporting the evolution of Apps. First, the authors introduce the reviews into the goal model as a new factor by comparing keywords. Second, the users sentiments in reviews are mined, and two kinds of information are gained by analysing the model to help developers make decisions on which goals to be improved in next version: one kind of information is about users sentiments on the goals to evaluate whether users like them; another kind is the impact of updating one goal to others. To validate the proposed approach, they conducted experiments and a survey based on the Apps in Google Play. The results show that the proposed approach can establish relationships between goals and reviews reasonably and further provide useful information for optimising the evolution strategy of the App. Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu, Shanquan Gao |
IET Softw. | 1 |
| 2020 | Updating the goal model with user reviews for the evolution of an appabstractAbstract Goal model is an important model in requirements engineering, and it can describe features and their relationships for supporting the development of apps. Since an app evolves continually, the goal model also needs to be updated with new requirements to guide the whole process. As the feedback of users, reviews provide an abundant resource of user requirements for updating the goal model. In this paper, we propose an approach to help developers (a) analyze reviews to gain the information of user requirements by training a classifier and defining keyword‐based linguistic rules as well as grammar‐based rules and (b) update the goal model with the extracted information, including improving existing goals and extending the model with new goals. In addition, we design a framework to represent results so that they can be understood by developers easily. According to our experiments based on the data in Google Play, the F‐measure of classifier on reviews can reach 75.76%, and the average precision for extracting requirements‐related information from reviews is 84.04%, then we can map the information to goals with the F‐measure of 70.21%. Furthermore, the survey on 22 developers shows that the information provided by us is useful for updating the goal model. Shanquan Gao, Lei Liu 0040, Yuzhou Liu 0001, Huaxiao Liu |
J. Softw. Evol. Process. | 3 |
| 2019 | App store mining for iterative domain analysis: Combine app descriptions with user reviewsabstractSummary Compared with traditional software, the domain analysis of apps is conducted not only in the early stage of software development to gain knowledge of a particular domain but also runs throughout each iteration of apps to help developers understand evolution trends of the domain for maintaining their competitiveness. In this paper, we propose an approach to analyze app descriptions combined with reviews in App stores automatically and construct a feature‐based domain state model (FDSM) in the form of state machine to support the domain analysis of apps. In FDSM, the domain knowledge up to a certain moment together is defined as a state. Initial state summarizes the high‐level knowledge by gaining topics of app descriptions, whereas each transition is generated based on the information gained within one period of time and describes the change from the current state to the next one. Furthermore, user opinions in reviews are introduced into the model to quantify the value of information for helping developers get key domain knowledge efficiently. To validate the proposed approach, we conducted a series of experiments based on Google Play. The results show that FDSM can provide valuable information for supporting domain analysis, especially in the evolution process of apps. Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu, Xinglong Yin |
Softw. Pract. Exp. | 1 |
| 2018 | Analyzing reviews guided by App descriptions for the software development and evolutionabstractAbstract Reviews in App stores are a massive and fast‐growing data resource for developers to understand user experiences and their needs. Studies show that users often express their sentiments on App features in reviews, and this information is important for the development and evolution of Apps. To help developers gain such information efficiently, this paper proposes a method using App descriptions, another typical data in App stores, to guide the analysis of reviews. Firstly, we extract App features from descriptions, then summarize them to gain topics of App features as high‐level information; the results are formalized as a topic‐based domain model (TBDM). Secondly, we train classifiers of reviews based on the model to establish the relationships between user sentiments and App features. Finally, a quantified method is given to analyze the model based on developer preferences for recommending and summarizing reviews. To evaluate our approach, experiments were conducted using the App descriptions and reviews collected from Google Play. The results indicate that the approach can classify reviews to their related App features effectively (average F measure is 86.13%), and provides useful information for overall analyzing App features in a domain and identifying (dis)advantages of an App. Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu |
J. Softw. Evol. Process. | 1 |
| 2017 | The verification of program relationships in the context of software cybernetics
Huaxiao Liu, Yuzhou Liu 0001, Lei Liu 0040 |
J. Syst. Softw. | 2 |
| 2017 | Mining domain knowledge from app descriptions
Yuzhou Liu 0001, Lei Liu 0040, Huaxiao Liu |
J. Syst. Softw. | 1 |