Weifeng Pan 0001

dblp:75/6501 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0001-6355-1385ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 4 · 3 first-authorArtificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Refactoring techniques for software vulnerabilities
Obieda Ananbeh, Wala Alnozami, Dae-Kyoo Kim, Ming Hua 0003, Weifeng Pan 0001
J. Syst. Softw.5
2025 Toward the Fractal Dimension of Classes
abstract
The fractal property has been regarded as a fundamental property of complex networks, characterizing the self-similarity of a network. Such a property is usually numerically characterized by the fractal dimension metric, and it not only helps the understanding of the relationship between the structure and function of complex networks but also finds a wide range of applications in complex systems. The existing literature shows that class-level software networks (i.e., class dependency networks) are complex networks with the fractal property. However, the fractal property at the feature (i.e., methods and fields) level has never been investigated, although it is useful for measuring class complexity and predicting bugs in classes. Furthermore, existing studies on the fractal property of software systems were all performed on un-weighted software networks and have not been used in any practical quality assurance tasks such as bug prediction. Generally, considering the weights on edges can give us more accurate representations of the software structure and thus help us obtain more accurate results. The illustration of an approach’s practical use can promote its adoption in practice. In this article, we examine the fractal property of classes by proposing a new metric. Specifically, we build a Feature-Level Software Network (FLSN) for each class to represent the methods/fields and their couplings (including coupling frequencies) within the class and propose a new metric, Fractal Dimension for Classes (FDC) , to numerically describe the fractal property of classes using FLSNs, which captures class complexity. We evaluate FDC theoretically against Weyuker’s nine properties, and the results show that FDC adheres to eight of the nine properties. Empirical experiments performed on a set of 12 large open source Java systems show that (i) for most classes (larger than \(96\%\) ), there exists the fractal property in their FLSNs, (ii) FDC is capable of capturing additional aspects of class complexity that have not been addressed by existing complexity metrics, (iii) FDC significantly correlates with both the existing class-level complexity metrics and the number of bugs in classes, and (iv) FDC , when used together with existing class-level complexity metrics, can significantly improve bug prediction in classes in three scenarios (i.e., bug-count , bug-classification , and effort-aware ) of the cross-project context, but in the within-project context, it cannot.
Weifeng Pan 0001, Ming Hua 0003, Dae-Kyoo Kim, Zijiang Yang 0006, Yutao Ma
ACM Trans. Softw. Eng. Methodol.1
2024 EASE: An Effort-aware Extension of Unsupervised Key Class Identification Approaches
abstract
Key class identification approaches aim at identifying the most important classes to help developers, especially newcomers, start the software comprehension process. So far, many supervised and unsupervised approaches have been proposed; however, they have not considered the effort to comprehend classes. In this article, we identify the challenge of “ effort-aware key class identification ”; to partially tackle it, we propose an approach, EASE , which is implemented through a modification to existing unsupervised key class identification approaches to take into consideration the effort to comprehend classes. First, EASE chooses a set of network metrics that has a wide range of applications in the existing unsupervised approaches and also possesses good discriminatory power . Second, EASE normalizes the network metric values of classes to quantify the probability of any class to be a key class and utilizes Cognitive Complexity to estimate the effort required to comprehend classes. Third, EASE proposes a metric, RKCP , to measure the relative key-class proneness of classes and further uses it to sort classes in descending order. Finally, an effort threshold is utilized, and the top-ranked classes within the threshold are identified as the cost-effective key classes. Empirical results on a set of 18 software systems show that (i) the proposed effort-aware variants perform significantly better in almost all (≈98.33%) the cases, (ii) they are superior to most of the baseline approaches with only several exceptions, and (iii) they are scalable to large-scale software systems. Based on these findings, we suggest that (i) we should resort to effort-aware key class identification techniques in budget-limited scenarios; and (ii) when using different techniques, we should carefully choose the weighting mechanism to obtain the best performance.
Weifeng Pan 0001, Marouane Kessentini, Ming Hua 0003, Zijiang Yang 0006
ACM Trans. Softw. Eng. Methodol.1
2023 Identifying Key Classes for Initial Software Comprehension: Can We Do It Better?
abstract
Key classes are excellent starting points for developers, especially newcomers, to comprehend an unknown software system. Though many unsupervised key class identification approaches have been proposed in the literature by representing software as class dependency networks (aka software networks) and using some network metrics (e.g., h-index, a-index, and coreness), they are never aware of the field where the nodes exist and the effect of the field on the importance of the nodes in it. According to the classic field theory in physics, every material particle is in a field through which they exert an impact on other particles in the field via non-contact interactions (e.g., electromagnetic force, gravity, and nuclear force). Similarly, every node in a software network might also exist in a field, which might affect the importance of class nodes in it. In this paper, we propose an approach, iFit, to identify key classes in object-oriented software systems. First, we represent software as a CSNWD(Weighted Directed Class-level Software Network) to capture the topological structure of software, including classes, their couplings, and the direction and strength of couplings. Second, we assume that the nodes in the CSNWDexist in a gravitation-like field and propose a new metric, CG (Cumulative Gravitation-like importance), to measure the importance of classes. CG is inspired by Newton's gravitational formula and uses the PageRank value computed by a biased-PageRank algorithm as the masses of classes. Finally, classes in the system are sorted in descending order according to their CG values, and a cutoff is utilized, that is, the top-ranked classes are recommended as key classes. The experiments were performed on a data set composed of six open-source Java systems from the literature. The results show that iFit is superior to the baseline approaches on 93.75% of the total cases, and is scalable to large-scale software systems. Besides, we find that iFit is neutral to the weighting mechanisms used to assign the weights for different coupling types in the CSNWD, that is, when applying iFit to identify key classes, we can use any one of the weighting mechanisms.
Weifeng Pan 0001, Ming Hua 0003, Dae-Kyoo Kim, Zijiang Yang 0006
ICSE1
2023 Pride: Prioritizing Documentation Effort Based on a PageRank-Like Algorithm and Simple Filtering Rules
abstract
Code documentation can be helpful in many software quality assurance tasks. However, due to resource constraints (e.g., time, human resources, and budget), programmers often cannot document their work completely and timely. In the literature, two approaches (one is supervised and the other is unsupervised) have been proposed to prioritize documentation effort to ensure the most important classes to be documented first. However, both of them contain several limitations. The supervised approach overly relies on a difficult-to-obtain labeled data set and has high computation cost. The unsupervised one depends on a graph representation of the software structure, which is inaccurate since it neglects many important couplings between classes. In this paper, we propose an improved approach, named Pride, to prioritize documentation effort. First, Pride uses a weighted directed class coupling network to precisely describe classes and their couplings. Second, we propose a PageRank-like algorithm to quantify the importance of classes in the whole class coupling network. Third, we use a set of software metrics to quantify source code complexity and further propose a simple but easy-to-operate filtering rule. Fourth, we sort all the classes according to their importance in descending order and use the filtering rule to filter out unimportant classes. Finally, a threshold$k$is utilized, and the top-$k$% ranked classes are the identified important classes to be documented first. Empirical results on a set of nine software systems show that, according to the average ranking of the Friedman test, Pride is superior to the existing approaches in the whole data set.
Weifeng Pan 0001, Ming Hua 0003, Dae-Kyoo Kim, Zijiang Yang 0006
IEEE Trans. Software Eng.1
2022 Comments on "Using $k$k-Core Decomposition on Class Dependency Networks to Improve Bug Prediction Model's Practical Performance"
abstract
In a very recent paper by Qu et al. (IEEE Transactions on Software Engineering, vol. 47 no. 2, pp. 348-366, Feb. 1 2021, doi:10.1109/TSE.2019.2892959), the authors propose an effective equation, top-core, to improve the performance of effort-aware bug prediction models. A distinctive feature of top-core is that it takes into account the coreness of a class in a Class Dependency Network (CDN) when calculating the relative risk of a class to be buggy. In this comment, we show that Qu et al.'s paper contains three shortcomings that may influence the performance of top-core or even have the potential to lead to erroneous results. First, we show that the CDN that they use to calculate the coreness of classes is not very accurate, neglecting many important types of dependency relations between classes such as method call relation, access relation, and instantiates relation. Second, they trained a Logistic Regression model using the scikit-learn framework to predict the probability of a specific class to be buggy. It is actually an L2 regularized Logistic Regression model, which is dependent on the scale of the features. But they neglected to normalize the features, making the obtained results erroneous. Finally, the number of execution times (viz. 10 times in the paper of Qu et al.) they used to reduce the bias caused by the randomness (viz. random split of instances and the process to handle class-imbalance problem) in the experiments is too small to ensure that the obtained results converge to stable values; but they failed to signify the precision level of their results for comparison. In this comment, we provide solutions to the problems by using i) an improved CDN (ICDN) to represent the structure of software systems, ii) the z-score method to normalize the features, and iii) an adaptive mechanism to determine the number of execution times. In the experiments, we find that Qu et al.'s approach based on the Logistic Regression model does not perform significantly better than the state-of-the-art approach Ree, which is inconsistent with the conclusion in Qu et al.'s work. We also observe that replacing CDN with ICDN does improve the performance of Qu et al.'s approach.
Weifeng Pan 0001, Ming Hua 0003, Zijiang Yang 0006, Tian Wang 0007
IEEE Trans. Software Eng.1
2021 ElementRank: Ranking Java Software Classes and Packages using a Multilayer Complex Network-Based Approach
abstract
Software comprehension is an important part of software maintenance. To understand a piece of large and complex software, the first problem to be solved is where to start the understanding process. Choosing to start the comprehension process from the important software elements has proven to be a practical way. Research on complex networks opens new opportunities for identifying important elements, and many approaches have been proposed. However, the software networks that existing approaches use neglect the multilayer nature of software systems. That is, nodes in the network can have different types of relationships at the same time, and each type of relationship forms a specific layer. Worse still, they mainly focus on identifying important classes, and little work has been done on quantifying package importance. In this paper, we propose an ElementRank approach to provide a ranked list of classes (or packages) for maintainers to start the comprehension process. The top-ranked classes (or packages) can be seen as the starting points for the software comprehension process at the class (or package) level. First, we introduce two kinds of multilayer software networks to describe the topological structure of software at the class level and package level, respectively. Second, we propose a weighted PageRank algorithm to calculate the weighted PageRank value of classes (or packages) in each layer of the corresponding multilayer software network. Then, we use AHP (Analytic Hierarchy Process) to weigh each layer in the corresponding multilayer software network, and further aggregate the weighted PageRank value to obtain the global weighted PageRank value for each class (or package). Finally, all the classes (or packages) are ranked according to their global weighted PageRank values in a descending order, and the top-ranked classes (or packages) can serve as the starting points for the software comprehension process at the class (or package) level. ElementRank is validated theoretically using the widely accepted Weyuker’s criteria. Theoretical results show that the global weighted PageRank value for classes (or packages) satisfies most of Weyuker’s properties. Furthermore, ElementRank is evaluated empirically using a set of twelve open source software systems. Through a set of experiments, we show the rank correlation between the results of ElementRank and that of the approaches in the related work, and the benefits of ElementRank are also illustrated in comparison with other approaches in the related work. Empirical results also show that ElementRank can be applied to large software systems.
Weifeng Pan 0001, Ming Hua 0003, Carl K. Chang, Zijiang Yang 0006, Dae-Kyoo Kim
IEEE Trans. Software Eng.1
2018 Structure-aware Mashup service Clustering for cloud-based Internet of Things using genetic algorithm based clustering algorithm
Weifeng Pan 0001, Chunlai Chai
Future Gener. Comput. Syst.1
2018 Analyzing the structure of Java software systems by weighted K-core decomposition
Weifeng Pan 0001, Bing Li 0010, Jing Liu 0033, Yutao Ma, Bo Hu 0013
Future Gener. Comput. Syst.1
2018 Identifying key classes in object-oriented software using generalized k-core decomposition
Weifeng Pan 0001, Beibei Song, Kangshun Li
Future Gener. Comput. Syst.1
2013 BIGSIR: A Bipartite Graph Based Service Recommendation Method
abstract
Cloud computing is an Internet-based computing. It relies on sharing computing resources which are delivered as services on the Internet. Web service is one of the most important types of services that can be used in cloud computing. But many of them may be similar in some functional or nonfunctional properties, making how to recommend a suitable web service a problem facing many developers. Researchers have taken the QoS attributes into consideration. However, their research is on the premise that all the recommended web services are compatible, i.e., the recommended web services can be composed with existing web services. It may not always be true. In this paper, we only take the compatibility of web services into consideration, and present a BIpartite Graph based Service Recommendation (BIGSIR) method to address the service compatibility problem. BIGSIR uses the historical usage data of web services to recommend web services to developers. Different from existing web service recommendation approaches, BIGSIR adopts a bipartite graph to visual the web services and the relationship between them. Based on the graph model, an effective recommendation algorithm is introduced to recommend the suitable web services. Our approach is evaluated on a dataset constructed from myExperiment, a search engine that contains about 1, 851 web services and 2, 000 workflows. Experimental results demonstrate that apart from some isolated web services or workflows, BIGSIR can obtain promising results. And we also explore the factors that will influence the performance of BIGSIR. This work not only provides a new dataset, but also highlights a new perspective for service recommendation, i.e. services as a bipartite network.
Bo Jiang 0009, Xiao-xiao Zhang, Weifeng Pan 0001, Bo Hu 0013
SERVICES3
2011 Multi-granularity dynamic analysis of complex software networks
abstract
Software systems represent one of the most complex man-made systems. In this paper, we analyze the evolution of Object-Oriented (OO) software using complex network theory from a multi-granularity perspective. First, the software net works are constructed for a multi-version software system at different levels of granularity. Then, some parameters used in complex network theory are introduced to study the topological characteristics of these software networks. By investigating the parameters' values in consecutive software networks, we have a better understanding about software evolution. A case study on an open source OO project, Azureus, is conducted as an example to illustrate our approach. It uncovers some underlying dynamic characteristics of OO systems. These results provide a different dimension to our understanding of software system dynamics and also are very useful for the design and development of OO software systems.
Bing Li 0010, Weifeng Pan 0001, Jinhu Lü 0001
ISCAS2
2010 Measuring Structural Quality of Object-Oriented Softwares via Bug Propagation Analysis on Weighted Software Networks
Weifeng Pan 0001, Bing Li 0010, Yutao Ma, Yeyi Qin, Xiao-Yan Zhou
J. Comput. Sci. Technol.1
2009 A Novel Method for Mining SaaS Software Tag via Community Detection in Software Services Network
Bing Li 0010, Weifeng Pan 0001, Tao Peng 0007
CloudCom3
2009 Requirements Discovery Based on RGPS Using Evolutionary Algorithm
Tao Peng 0007, Bing Li 0010, Weifeng Pan 0001, Zaiwen Feng
SEKE3
2009 Class structure refactoring of object-oriented softwares using community detection in dependency networks
Weifeng Pan 0001, Bing Li 0010, Yutao Ma, Jing Liu 0033, Yeyi Qin
Frontiers Comput. Sci. China1
2008 A sequence cipher producing method based on two-layer ranking Multi-Objective Evolutionary Algorithm
abstract
Aiming at designing a high safe and high efficiency cryptosystem, the period of the sequence cipher can not be too long, and the cipher sequence produced should approach random numbers. But the key sequence produced by traditional methods sometimes does not have randomness, which makes insecurity the system using this key sequence. Considering this, in this paper, we take two criteria usually used to evaluate the randomness of a key sequence as two objectives of Multi-Objective Evolutionary Algorithm (MOEA), and a new sequence cipher producing method based on two-layer MOEA is proposed (called TLEASCP). Because of TLEASCP is based on the randomness of crossover operator and mutation operator of the high efficient MOEA, the key sequences produced by TLEASCP have the merits of high randomness, chaos and long period.
Kangshun Li, Weifeng Pan 0001, Wensheng Zhang 0003, Zhangxin Chen
IEEE Congress on Evolutionary Computation2
2008 Automatic modeling of a novel gene expression programming based on statistical analysisand critical velocity
abstract
The basic principle of GEP is briefly introduced. And considering the defects of classic GEP such as lack of variety, the problem of convergence and blind searching without learning mechanism, a novel GEP based on statistical analysis and stagnancy velocity is proposed (called AMACGEP). It mainly has the following characteristics: First, improve the initial population by statistic analysis of repeated bodies. Second, introduce the concept of stagnancy velocity to adjust the searching space, evolution velocity, the diversity of individuals and the accuracy of prediction. Third, introduce dynamic mutation operator to improve the diversity of individuals and the velocity of convergence. Compared with other methods like traditional methods, methods of neural network, classic GEP and other improved GEPs in automatic modeling of complex function, the simulation results show that the AMACGEP set up by this paper is better.
Kangshun Li, Weifeng Pan 0001, Wensheng Zhang 0003, Zhangxin Chen
IEEE Congress on Evolutionary Computation2