VLDB 2026 Research / reviewers in the wild / expert
Feiqiao Mao
dblp:212/8932
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-1512-5503ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CSMS: Boosting Class Code Summarization via Transformer Refinement and Method Summary FusionabstractCode comprehension remains fundamental to software development. Although recent advances in deep learning and Code Large Language Models (CodeLLMs) have significantly improved method-level code summarization, class-level understanding remains challenging due to: (1) limited dedicated research focus, (2) excessive code sequence of a class, and (3) overreliance on source code and Syntax Tree (AST) while neglecting other valuable multimodal information. To address these challenges, we present a framework (called CSMS) with two key innovations. First, it pioneers method-level summary fusion, leveraging CodeT5 to extract semantically rich features from member methods without additional input. Second, its optimized Transformer architecture combines a parameter-efficient TTSEncoder with a multiperspective TAJDecoder, enabling comprehensive feature utilization while handling long sequences. Experimental results on ClassSum and HRCE datasets demonstrate CSMS’s superiority, achieving state-of-the-art BLEU scores while effectively handling long class sequences. The framework’s ability to leverage method-level context while maintaining computational efficiency represents a significant advance in class comprehension. Feiqiao Mao, Xingyang Du, Shaocheng Feng, Jiafeng Guo |
APSEC | 1 |
| 2025 | MTL-CR: A Multitask Learning Approach for Code RepresentationabstractCode representation plays a fundamental role in enabling a wide range of code intelligence tasks, such as code translation, bug fixing, and code completion. Existing approaches typically adopt task-specific modeling paradigms, training separate models for individual downstream tasks. However, such methods often fail to capture semantic and structural commonalities across tasks, resulting in limited generalization and transferability. Therefore, we propose MTL-CR (Multitask Learning for Code Representation), a unified multitask learning framework that jointly optimizes multiple code-related tasks to learn more generalizable and robust code representations. To balance optimization dynamics among tasks, we further introduce a dynamic task weighting strategy based on taskspecific learning speeds. Experimental results demonstrate that MTL-CR consistently outperforms single task baselines and static multitask methods in various downstream tasks and evaluation metrics. For example, in the code translation task, it improves Accuracy and BLEU by 6.29% and 2.33%. Dongxu Yu, Feiqiao Mao, Xingyang Du |
APSEC | 2 |
| 2025 | PathGPS: An Integrated Deep Learning and Machine Learning Framework for Multi-Pathogen Diagnosis Via Gene Pair SignaturesabstractPathogen identification is crucial for the diagnosis of infectious diseases. Clinical practice demands efficient, rapid, and accurate diagnostic methods. While host gene expression-based diagnostic models are widely used for singlepathogen identification, the vast diversity of infectious diseases makes individual testing highly impractical. Here, we designed a multi-pathogen diagnostic strategy, call PathGPS (pathogen gene pair signature), targeting three major infectious diseases prioritized by the World Health Organization (WHO). PathGPS strategy introduces deep learning models based on relative expression levels of gene pairs, which extract diseasespecific biomarkers, to capture pathogen-specific information. Then, it integrates these features through machine learning for joint pathogen classification. Our model achieved an average AUC of over 0.95 in distinguishing pathogens from healthy samples and demonstrated superior accuracy in differentiating the three infectious diseases compared to existing diagnostic approaches. This study identifies a set of biomarkers for three critical infectious diseases, providing a novel method for precise multi-pathogen diagnosis. Huaijin Wen, Feiqiao Mao |
BIBM | 2 |
| 2025 | Bmco-o: a smart code smell detection method based on co-occurrences
Feiqiao Mao, Kaihang Zhong |
Autom. Softw. Eng. | 1 |
| 2025 | MPDA: a data augmentation approach to improve deep learning for software vulnerability detection
Feiqiao Mao, Yingxiang Yuan, Xingyang Du, Zhihua Du |
Empir. Softw. Eng. | 1 |
| 2024 | An XGBoost-assisted evolutionary algorithm for expensive multiobjective optimization problems
Feiqiao Mao, Kaihang Zhong, Jiyu Zeng, Zhengping Liang |
Inf. Sci. | 1 |
| 2022 | BenchSubset: A framework for selecting benchmark subsets based on consensus clusteringabstractThe redundancy in the benchmark suite will increase the time for computer system performance evaluation and simulation. The most typical method to solve this problem is to select subsets based on clustering. However, it is a challenge to validate benchmark subsetting results for unlabeled benchmark suites when using the clustering method, and existing research has not considered this problem. Also, there is no quantitative evaluation method for subsetting which can reflect the universal and the diversity characteristics of the benchmark suite at the same time. To solve the above problems, we propose BenchSubset, a framework for selecting benchmark subsets based on consensus clustering, which includes Group Principal Components Analysis, consensus clustering, and a new evaluation method considering the universal and the diversity characteristics of the benchmark suite. We conducted SPEC CPU2017 subsetting experiments on Huawei's Taishan 200, then verified the effectiveness of BenchSubset in selecting a benchmark subset. Compared with the mainstream principal components analysis with hierarchical clustering (PCA-H) method, the benchmark subset selected by BenchSubset performs better in representing the universal and the diversity characteristics of SPEC CPU2017. Hongping Zhan, Weiwei Lin 0001, Feiqiao Mao, Minxian Xu, Guangxin Wu, Guokai Wu, Jianzhuo Li |
Int. J. Intell. Syst. | 3 |