Feiqiao Mao

dblp:212/8932 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-1512-5503ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CSMS: Boosting Class Code Summarization via Transformer Refinement and Method Summary Fusion
abstract
Code comprehension remains fundamental to software development. Although recent advances in deep learning and Code Large Language Models (CodeLLMs) have significantly improved method-level code summarization, class-level understanding remains challenging due to: (1) limited dedicated research focus, (2) excessive code sequence of a class, and (3) overreliance on source code and Syntax Tree (AST) while neglecting other valuable multimodal information. To address these challenges, we present a framework (called CSMS) with two key innovations. First, it pioneers method-level summary fusion, leveraging CodeT5 to extract semantically rich features from member methods without additional input. Second, its optimized Transformer architecture combines a parameter-efficient TTSEncoder with a multiperspective TAJDecoder, enabling comprehensive feature utilization while handling long sequences. Experimental results on ClassSum and HRCE datasets demonstrate CSMS’s superiority, achieving state-of-the-art BLEU scores while effectively handling long class sequences. The framework’s ability to leverage method-level context while maintaining computational efficiency represents a significant advance in class comprehension.
Feiqiao Mao, Xingyang Du, Shaocheng Feng, Jiafeng Guo
APSEC1
2025 MTL-CR: A Multitask Learning Approach for Code Representation
abstract
Code representation plays a fundamental role in enabling a wide range of code intelligence tasks, such as code translation, bug fixing, and code completion. Existing approaches typically adopt task-specific modeling paradigms, training separate models for individual downstream tasks. However, such methods often fail to capture semantic and structural commonalities across tasks, resulting in limited generalization and transferability. Therefore, we propose MTL-CR (Multitask Learning for Code Representation), a unified multitask learning framework that jointly optimizes multiple code-related tasks to learn more generalizable and robust code representations. To balance optimization dynamics among tasks, we further introduce a dynamic task weighting strategy based on taskspecific learning speeds. Experimental results demonstrate that MTL-CR consistently outperforms single task baselines and static multitask methods in various downstream tasks and evaluation metrics. For example, in the code translation task, it improves Accuracy and BLEU by 6.29% and 2.33%.
Dongxu Yu, Feiqiao Mao, Xingyang Du
APSEC2
2025 PathGPS: An Integrated Deep Learning and Machine Learning Framework for Multi-Pathogen Diagnosis Via Gene Pair Signatures
abstract
Pathogen identification is crucial for the diagnosis of infectious diseases. Clinical practice demands efficient, rapid, and accurate diagnostic methods. While host gene expression-based diagnostic models are widely used for singlepathogen identification, the vast diversity of infectious diseases makes individual testing highly impractical. Here, we designed a multi-pathogen diagnostic strategy, call PathGPS (pathogen gene pair signature), targeting three major infectious diseases prioritized by the World Health Organization (WHO). PathGPS strategy introduces deep learning models based on relative expression levels of gene pairs, which extract diseasespecific biomarkers, to capture pathogen-specific information. Then, it integrates these features through machine learning for joint pathogen classification. Our model achieved an average AUC of over 0.95 in distinguishing pathogens from healthy samples and demonstrated superior accuracy in differentiating the three infectious diseases compared to existing diagnostic approaches. This study identifies a set of biomarkers for three critical infectious diseases, providing a novel method for precise multi-pathogen diagnosis.
Huaijin Wen, Feiqiao Mao
BIBM2
2025 Bmco-o: a smart code smell detection method based on co-occurrences
Feiqiao Mao, Kaihang Zhong
Autom. Softw. Eng.1
2025 MPDA: a data augmentation approach to improve deep learning for software vulnerability detection
Feiqiao Mao, Yingxiang Yuan, Xingyang Du, Zhihua Du
Empir. Softw. Eng.1
2024 An XGBoost-assisted evolutionary algorithm for expensive multiobjective optimization problems
Feiqiao Mao, Kaihang Zhong, Jiyu Zeng, Zhengping Liang
Inf. Sci.1
2022 BenchSubset: A framework for selecting benchmark subsets based on consensus clustering
abstract
The redundancy in the benchmark suite will increase the time for computer system performance evaluation and simulation. The most typical method to solve this problem is to select subsets based on clustering. However, it is a challenge to validate benchmark subsetting results for unlabeled benchmark suites when using the clustering method, and existing research has not considered this problem. Also, there is no quantitative evaluation method for subsetting which can reflect the universal and the diversity characteristics of the benchmark suite at the same time. To solve the above problems, we propose BenchSubset, a framework for selecting benchmark subsets based on consensus clustering, which includes Group Principal Components Analysis, consensus clustering, and a new evaluation method considering the universal and the diversity characteristics of the benchmark suite. We conducted SPEC CPU2017 subsetting experiments on Huawei's Taishan 200, then verified the effectiveness of BenchSubset in selecting a benchmark subset. Compared with the mainstream principal components analysis with hierarchical clustering (PCA-H) method, the benchmark subset selected by BenchSubset performs better in representing the universal and the diversity characteristics of SPEC CPU2017.
Hongping Zhan, Weiwei Lin 0001, Feiqiao Mao, Minxian Xu, Guangxin Wu, Guokai Wu, Jianzhuo Li
Int. J. Intell. Syst.3