Pu Xiong

dblp:307/3999 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0003-4546-4335ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Fast Heterogeneous Multiproblem Surrogates for Transfer Evolutionary Multiobjective Optimization
abstract
Transfer evolutionary multiobjective optimization leverages the relevant knowledge from other source problems (distinct but possibly related) to assist the optimization of the target problem of interest. Multi-problem surrogates stack multiple source surrogates to reduce the number of function evaluations of the target expensive problem. The current multi-problem surrogates only considers several source problems and the source and target problems are assumed to be homogeneous. In order to address the above issues, this paper proposes fast heterogeneous multi-problem surrogates for transfer evolutionary multiobjective optimization with a large number of surrogates. First, an iterative surrogate selection strategy is designed to select the highly relevant surrogates from the large-scale surrogate pool to avoid negative transfer. Second, heterogeneous multi-problem surrogates are established to align the features of the source and target models. Finally, an adaptive k-fold cross-validation method is proposed to obtain the predicted values of the target model with low computational costs. Experiments on the multiobjective optimization benchmark problems and multiobjective neural architecture search problems have demonstrated that the proposed method is able to avoid negative transfer in the large-scale scenarios and reduce the computational costs.
Hao Li 0009, Pu Xiong, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001
IEEE Trans. Evol. Comput.2
2023 Exploring the Impact of Code Clones on Deep Learning Software
abstract
Deep learning (DL) is a really active topic in recent years. Code cloning is a common code implementation that could negatively impact software maintenance. For DL software, developers rely heavily on frameworks to implement DL features. Meanwhile, to guarantee efficiency, developers often reuse the steps and configuration settings for building DL models. These may bring code copy-pastes or reuses inducing code clones. However, there is little work exploring code clones’ impact on DL software. In this article, we conduct an empirical study and show that: (1) code clones are prevalent in DL projects, about 16.3% of code fragments encounter clones, which is almost twice larger than the traditional projects; (2) 75.6% of DL projects contain co-changed clones, meaning changes are propagated among cloned fragments, which can bring maintenance difficulties; (3) Percentage of the clones and Number of clone lines are associated with the emergence of co-changes; (4) the prevalence of Code clones varies in DL projects with different frameworks, but the difference is not significant; (5) Type 1 co-changed clones often spread over different folders, but Types 2 and 3 co-changed clones mainly occur within the same files or folders; (6) 57.1% of all co-changed clones are involved in bugs.
Ran Mo, Yao Zhang 0028, Yushuo Wang, Pu Xiong, Zengyang Li
ACM Trans. Softw. Eng. Methodol.5
2022 Exploring and understanding cross-service code clones in microservice projects
abstract
Microservice is an architecture style that decomposes complex software into loosely coupled services, which could be developed, maintained, and deployed independently. In recent years, the microservice architecture has been drawing more and more attention from both industrial and academic communities. Many companies, such as Google, Netflix, Amazon, and IBM have applied microservice architecture in their projects. Researchers have also studied microservices in different directions, such as microservices extraction, fault localization, and code quality analysis. The recent work has presented cross-service code clones are prevalent in microservice projects and have caused considerable co-modifications among different services, which undermines the independence of microservices. But there is no systematic study to reveal the underlying reasons for the emergence of such clones. In this paper, we first build a dataset consisting of 2,722 pairs of cross-service clones from 22 open-source microservice projects. Then we manually inspect the implementations of files and methods involved in cross-service clones to understand why the clones are introduced. In the file-level analysis, we categorize files into three types: DPFile (Data-processing File), DRFile (Data-related File), and DIFile (Data-irrelevant File), and have presented that DRFiles are more likely to encounter cross-service clones. For each type of files, we further classify them into specific cases. Each case describes the characteristics of involved files and why the clones happen. In the method-level analysis, we dig information from the code of involved methods. On this basis, we propose a catalog containing 4 categories with 10 subcategories of method-level implementations that result in cross-service clones. We believe our analyses have provided the fundamental knowledge of cross-service clones, which can help developers better manage and resolve such clones in microservice projects.
Ran Mo, Yao Zhang 0028, Pu Xiong
ICPC5
2021 Formal Definition and Automatic Generation of Semantic Metrics: An Empirical Study on Bug Prediction
abstract
Bug prediction is helpful for facilitating bug fixes and improving the efficiency in software development and maintenance. In the past decades, researchers have proposed numerous studies on bug prediction by using code metrics. However, most of the existing studies use syntax-based metrics, there exists little work building bug prediction models with semantic metrics from source code. In this paper, we propose a new model, semantic dependency graph (SDG), to represent semantic relationships among source files. Based on the SDG, we formally define a suite of semantic metrics reflecting semantic characteristics of a project’s source files. Moreover, we create a tool to automate the generation of our proposed SDG-based metrics. Through our experimental studies, we have demonstrated that the SDG-based semantic metrics are effective for building bug prediction models, and the SDG-based metrics outperform traditional syntactic metrics on bug prediction. In addition, models using the SDG-based metrics could achieve a better prediction performance than two state-of-the-art models that learn semantic features automatically. Finally, we have also presented that our approach is applicable in practice in terms of execution time and space.
Ran Mo, Pu Xiong, Zengyang Li, Qiong Feng
SCAM3
2021 Predicting and Monitoring Bug-Proneness at the Feature Level
Shaozhi Wei, Ran Mo, Pu Xiong, Zengyang Li
SETTA3