VLDB 2026 Research / reviewers in the wild / expert
Xiangjuan Yao
dblp:02/86
· DBLP profile ↗
33ranked-venue papers
4as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A dynamic granularity-based multi-objective evolutionary algorithm for coal mine integrated energy system dispatch optimization
Xiaoyu Zhong 0001, Xiangjuan Yao, Kangjia Qiao, Dun-Wei Gong |
Expert Syst. Appl. | 2 |
| 2026 | Mutation testing based on non-cooperative Stackelberg game
Xiangjuan Yao, Changqing Wei, Dun-Wei Gong |
Inf. Softw. Technol. | 2 |
| 2026 | An Indicator-Based Evolutionary Algorithm for Large-Scale Constrained Multiobjective OptimizationabstractMost existing constrained multi-objective evolutionary algorithms (CMOEAs) experience a dramatic performance degradation when solving large-scale constrained multi-objective optimization problems (LSCMOPs), since they converge very slowly and easily get trapped in local optima due to the loss of diversity. To enhance the efficiency of tackling LSCMOPs, this paper proposes an indicator-based evolutionary algorithm, referred to as ILCMO. In ILCMO, two complementary indicators are proposed to assess the contribution of each individual to feasibility, convergence, and diversity. The first is a feasibility-oriented indicator designed to drive the population towards the feasible regions. The second is an infeasibility-assisted dynamic indicator, which comprises two relaxed constraint boundaries. Theoretical studies demonstrate that this dynamic indicator can effectively guide the population to focus on evenly searching the infeasible regions around feasible solutions to enhance local diversity. In addition, a variable grouping-based differential evolution (VGDE) strategy, which includes a group-based intra-learning operator and a group-based inter-learning operator, is devised to improve the quality of reproduction in large-scale search spaces. The effectiveness of the proposed algorithm is validated through comprehensive experiments on four benchmarks and a microgrid dispatch problem against seven state-of-the-art algorithms. Xiaoyu Zhong 0001, Xiangjuan Yao, Kangjia Qiao, Dun-Wei Gong, Yaochu Jin |
IEEE Trans. Evol. Comput. | 2 |
| 2025 | A Network-Assisted Evolutionary Multitask Framework for Multi-objective Optimization Problems with Unknown Constraints
Yong Zhang 0016, Ruizhao Zheng, Ali Wagdy Mohamed, Mingcheng Zuo, Xiangjuan Yao |
ICIC (17) | 8 |
| 2025 | Optimizing test data generation using SI_CNNpro-enhanced MGA for mutation testing
Xiangying Dang, Juxin Hu, Dun-Wei Gong, Guosheng Hao, Xiangjuan Yao, Bingsen Huang |
J. Syst. Softw. | 6 |
| 2025 | Prioritization Method for Crowdsourced Test Report by Integrating Text and Image InformationabstractABSTRACT Crowdsourcing testing has the advantages of efficiency, speed, and reliability, but an excessive number of test reports makes it a challenge for report reviewers to select high‐quality test reports in a limited time. Test reports submitted by crowd workers often tend to be short textual descriptions with a large number of screenshots attached. Most traditional processing methods of test reports target reports that only contain text information, which cannot meet the defect detection requirements of crowdsourced test reports. In view of this, this paper proposes a prioritization method of crowdsourced test reports that integrates text and image information. First, we extract the text and image information from the test reports, based on which the defect detection abilities of the test reports are measured and the similarities between test reports are calculated. Then, a multi‐stage prioritization method of the test reports is presented based on the defect detection levels and similarities of the test reports. In the first stage, based on the defect detection levels and the similarities, the test report set is sorted and clustered to obtain the sorting results of partial reports and the similar set for each sorted report; in the second stage, the similar test report set is sorted with the criteria of minimizing the similarity and maximizing the defect detection level; the sorting results of the two stages are combined to form the final priorities of test reports. To validate our approach, we conducted experiments on five crowdsourced test datasets. The results and the analysis show that our approach can detect all faults faster in a limited time. By comprehensively utilizing text and image information to prioritize test reports, better sorting results can be obtained than state‐of‐the‐art methods. Huijie Tu, Xiangjuan Yao, Dun-Wei Gong |
J. Softw. Evol. Process. | 2 |
| 2025 | Community Detection of Directed Network for Software Ecosystems Based on a Two-Step Information Dissemination ModelabstractABSTRACT A software ecosystem is a complex system that allows developers to cooperate with each other. Community is a universal and important topological property of networks. Detecting the communities of the software ecosystem is of great significance for analyzing its structural characteristics, discovering its hidden patterns, and predicting its behavior. Traditional community detection algorithms of complex networks are mostly for undirected networks. For the social network, the direction of information dissemination between developers cannot be ignored. In addition, the existing algorithms of community detection usually only consider direct influence between individuals while neglecting indirect relationships. To solve these problems, this paper presents a community detection method based on a two‐step information dissemination model for the software ecosystem. First, a two‐step information dissemination model is established to calculate the information gain of nodes. Second, a ranking method of developers' comprehensive influence is given through their influence vectors and information gains. Finally, communities are detected by taking the influential nodes as the cluster centers and the probability of information dissemination as the clustering direction. The proposed method is applied to community detection of typical software ecosystems in GitHub. The experimental results show that our method has good performance in the identification of community structure. Huijie Tu, Xiangjuan Yao, Tingting Hou, Dun-Wei Gong, Mengyi Yang |
J. Softw. Evol. Process. | 2 |
| 2024 | Multi-objective optimization and integrated indicator-driven two-stage project recommendation in time-dependent software ecosystem
Xiangjuan Yao, Dun-Wei Gong, Huijie Tu |
Inf. Softw. Technol. | 2 |
| 2024 | Test data generation for covering mutation-based path using MGA for MPI program
Xiangying Dang, Jinyong Wang, Dun-Wei Gong, Xiangjuan Yao, Changqing Wei |
J. Syst. Softw. | 4 |
| 2024 | Set evolution based test data generation for killing stubborn mutants
Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu, Xiangying Dang |
J. Syst. Softw. | 2 |
| 2024 | Parallel program testing based on critical communication and branch transformation
Tian Tian 0010, Anshi Wang, Xiuting Yang, Dun-Wei Gong, Tie Hou, Xiangjuan Yao |
J. Supercomput. | 6 |
| 2024 | Test Data Generation for Mutation Testing Based on Markov Chain Usage Model and Estimation of Distribution AlgorithmabstractMutation testing, a mainstream fault-based software testing technique, can mimic a wide variety of software faults by seeding them into the target program and resulting in the so-called mutants. Test data generated in mutation testing should be able to kill as many mutants as possible, hence guaranteeing a high fault-detection effectiveness of testing. Nevertheless, the test data generation can be very expensive, because mutation testing normally involves an extremely large number of mutants and some mutants are hard to kill. It is thus a critical yet challenging job to find an efficient way to generate a small set of test data that are able to kill multiple mutants at the same time as well as reveal those hard-to-detect faults. In this paper, we propose a new approach for test data generation in mutation testing, through the novel applications of the Markov chain usage model and the estimation of distribution algorithm. We first utilize the Markov chain usage model to reduce the so-called mutant branches in weak mutation testing and generate a minimal set of extended paths. Then, we regard the problem of generating test data as the problem of covering extended paths and use an estimation of distribution algorithm based on probability model to solve the problem. Finally, we develop a framework, TAMMEA, to implement the new approach of generating test data for mutation testing. The empirical studies based on fifteen object programs show that TAMMEA can kill more mutants using fewer test data compared with baseline techniques. In addition, the computation overhead of TAMMEA is lower than that of the baseline technique based on the traditional genetic algorithm, and comparable to that of the random method. It is clear that the new approach improves both the effectiveness and efficiency of mutation testing, thus promoting its practicability. Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu |
IEEE Trans. Software Eng. | 2 |
| 2023 | Integrating DSGEO into test case generation for path coverage of MPI programs
Baicai Sun, Dun-Wei Gong, Xiangjuan Yao |
Inf. Softw. Technol. | 3 |
| 2023 | Overlapping community detection in software ecosystem based on pheromone guided personalized PageRank algorithm
Xiangjuan Yao, Dun-Wei Gong, Huijie Tu |
Inf. Softw. Technol. | 2 |
| 2023 | Solution of Large-Scale Many-Objective Optimization Problems Based on Dimension Reduction and Solving Knowledge-Guided Evolutionary AlgorithmabstractThere are lots of many-objective optimization problems (MaOPs) in real-world applications, which often have many decision variables. Although a variety of methods have been proposed to solve MaOPs, with the increasing number of decision variables or objective functions, the performance of these algorithms deteriorates appreciably. In view of this, this article proposes a method to solve large-scale MaOPs (LSMaOPs) based on dimension reduction and a solving knowledge-guided evolutionary algorithm (KGEA). First, a dimension reduction method of objective functions is proposed. By clustering and aggregating the objective functions based on their correlation, the dimension of the original LSMaOP is effectively reduced. In addition, the correlations between the reduced objective functions are relatively low, so they can better represent different preferences. Then, we propose a solving KGEA to solve the transformed LSMaOP. In order to get a better set of initial solutions, a population initialization method by mirror partitioning the decision space is given, in which we dynamically modify the sampling probability according to the performance of solutions contained in each subdomain. At the same time, the algorithm will continuously supplement new excellent individuals using the solving knowledge obtained in the evolution of the population. To examine the performance of the proposed method, we carried out a number of comparative experiments. The experimental results demonstrated that the proposed algorithm can effectively tackle LSMaOPs. Xiangjuan Yao, Qian Zhao 0024, Dun-Wei Gong, Song Zhu |
IEEE Trans. Evol. Comput. | 1 |
| 2023 | Decomposition-Based Multiobjective Optimization Algorithms With Adaptively Adjusting Weight Vectors and NeighborhoodsabstractThe decomposition-based multiobjective optimization algorithm (MOEA/D) is an effective method of solving a multiobjective optimization problem (MOP). The main idea of MOEA/D is that the objectives are weighted through different vectors to form different subproblems, and an optimal solution set is obtained by co-evolution in a certain neighborhood. However, with the increase of objectives, the number of nondominated solutions increases exponentially, resulting in the deteriorated capability of searching for optimal solutions. In addition, for an optimization problem with the complex Pareto front (PF), the selection pressure of nondominated solutions is insufficient. To make evolution more efficient, an MOEA/D with adaptively adjusting weight vectors and neighborhoods (MOEA/D-AAWNs) is developed in this article. First, the evolutionary direction of each subproblem is analyzed and the Sparsity function (Spa) is proposed to measure the population density on the PF. By using Spa, a method of generating uniform vectors is presented to improve the diversity of solutions. Besides, a method of adaptively adjusting neighborhoods is given. It adjusts neighborhoods according to the number of iterations and the Spa value of its corresponding subproblem. In this way, the computational resource can be effectively allocated, leading to the improvement in evolutionary efficiency. The proposed algorithm is applied to solve a series of benchmark optimization instances, and the experimental results show that the proposed algorithm outperforms comparison algorithms in runtime, convergence, and diversity. Qian Zhao 0024, Yinan Guo 0001, Xiangjuan Yao, Dun-Wei Gong |
IEEE Trans. Evol. Comput. | 3 |
| 2023 | Evolutionary Generation of Test Suites for Multi-Path Coverage of MPI Programs With Non-DeterminismabstractWhen a large number of target paths in a sequential program need to be covered, we can divide similar target paths into the same group, and generate a test suite covering the same group of target paths at the same time, so as to reduce the testing cost. However, different communication edges may be run under a same test input when executing a Message-PassingInterface (MPI) program with non-determinism, which cause different code fragments may be traversed, indicating the difficulty of generating a test suite to cover each group of target paths. This paper proposes an approach to evolutionary generation of test suites for multi-path coverage of MPI programs with non-determinism, which can significantly reduce the testing cost and difficulty. We first design an indicator for evaluating each traversal set of communication edges, which is used to form a relation matrix between each target path and each traversal set of communication edges, so as to divide all the target paths into a certain amount of groups. Then, we construct an optimization model for test suite generation associated with each group. Finally, an evolutionary optimization algorithm is extended to solve each model, and used to generate a test suite covering each group of target paths. The proposed approach is utilized and compared with several state-of-the-art approaches to seven benchmark MPI programs, as well as the experimental results illustrate that the proposed approach can efficiently generate a test suite, thus supporting the superiority of the proposed approach. Baicai Sun, Dun-Wei Gong, Feng Pan 0008, Xiangjuan Yao, Tian Tian 0010 |
IEEE Trans. Software Eng. | 4 |
| 2022 | Parallel multi-objective evolutionary optimization based dynamic community detection in software ecosystem
Xiangjuan Yao, Huijie Tu, Dun-Wei Gong |
Knowl. Based Syst. | 2 |
| 2022 | Enhancement of Mutation Testing via Fuzzy Clustering and Multi-Population Genetic AlgorithmabstractMutation testing, a fundamental software testing technique, which is a typical way to evaluate the adequacy of a test suite. In mutation testing, a set of mutants are generated by seeding the different classes of faults into a program under test. Test data shall be generated in the way that as many mutants can be killed as possible. Thanks to numerous tools to implement mutation testing for different languages, a huge amount of mutants are normally generated even for small-sized programs. However, a large number of mutants not only leads to a high cost of mutation testing, but also make the corresponding test data generation a non-trivial task. In this paper, we make use of intelligent technologies to improve the effectiveness and efficiency of mutation testing from two perspectives. A machine learning technique, namely fuzzy clustering, is applied to categorize mutants into different clusters. Then, a multi-population genetic algorithm via individual sharing is employed to generate test data for killing the mutants in different clusters in parallel when the problem of test data generation as an optimization one. A comprehensive framework, termed as$\mathbf {FUZGENMUT}$, is thus developed to implement the proposed techniques. The experiments based on nine programs of various sizes show that fuzzy clustering can help to reduce the cost of mutation testing effectively, and that the multi-population genetic algorithm improves the efficiency of test data generation while delivering the high mutant-killing capability. The results clearly indicate that the huge potential of using intelligent technologies to enhance the efficacy and thus the practicality of mutation testing. Xiangying Dang, Dun-Wei Gong, Xiangjuan Yao, Tian Tian 0010, Huai Liu |
IEEE Trans. Software Eng. | 3 |
| 2022 | Integrating an Ensemble Surrogate Model's Estimation into Test Data GenerationabstractFor the path coverage testing of a Message-Passing Interface (MPI) program, test data generation based on an evolutionary optimization algorithm (EOA) has been widely known. However, during the use of the above technique, it is necessary to evaluate the fitness of each evolutionary individual by executing the program, which is generally computationally expensive. In order to reduce the computational cost, this article proposes a method of integrating an ensemble surrogate model’s estimation into the process of generating test data. The proposed method first produces a number of test inputs using an EOA, and forms a training set together with their real fitness. Then, this article trains an ensemble surrogate model (ESM) based on the training set, which is employed to estimate the fitness of each individual. Finally, a small number of individuals with good estimations are selected to further execute the program, so as to have their real fitness for the subsequent evolution. This article applies the proposed method to seven benchmark MPI programs, which is compared with several state-of-the-art approaches. The experimental results show that the proposed method can generate test data with significantly low computational cost. Baicai Sun, Dun-Wei Gong, Tian Tian 0010, Xiangjuan Yao |
IEEE Trans. Software Eng. | 4 |
| 2022 | Orderly Generation of Test Data via Sorting Mutant Branches Based on Their Dominance Degrees for Weak Mutation TestingabstractCompared with traditional structural test criteria, test data generated based on mutation testing are proved more effective at detecting faults. However, not all test data have the same potence in detecting software faults. If test data are prioritized while generating for mutation testing, the defect detectability of the test suite can be further strengthened. In view of this, we propose a method of test data generation for weak mutation testing via sorting mutant branches based on their dominance degrees. First, the problem of weak mutation testing is transformed into that of covering mutant branches for a transformed program. Then, the dominance relation of mutant branches in the transformed program is analyzed to obtain the non-dominated mutant branches and their dominance degrees. Following that, we prioritize all non-dominated mutant branches in descending order by virtue of their dominance degrees. Finally, the test data are generated in an orderly manner by selecting the mutant branches sequentially. The experimental results on 15 programs show that compared with other methods, the proposed test data generation method can not only improve the error detectability of the test suite, but also has higher efficiency. Xiangjuan Yao, Gongjie Zhang, Feng Pan 0008, Dun-Wei Gong, Changqing Wei |
IEEE Trans. Software Eng. | 1 |
| 2021 | Community detection in software ecosystem by comprehensively evaluating developer cooperation intensity
Tingting Hou, Xiangjuan Yao, Dun-Wei Gong |
Inf. Softw. Technol. | 2 |
| 2021 | Spectral clustering based mutant reduction for mutation testing
Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu |
Inf. Softw. Technol. | 2 |
| 2021 | Test Data Generation for Path Coverage of MPI Programs Using SAEOabstractMessage-passing interface (MPI) programs, a typical kind of parallel programs, have been commonly used in various applications. However, it generally takes exhaustive computation to run these programs when generating test data to test them. In this article, we propose a method of test data generation for path coverage of MPI programs using surrogate-assisted evolutionary optimization, which can efficiently generate test data with high quality. We first divide a sample set of a program into a number of clusters according to the multi-mode characteristic of the coverage problem, with each cluster training a surrogate model. Then, we estimate the fitness of each individual using one or more surrogate models when generating test data through evolving a population. Finally, a small number of representative individuals are selected to execute the program, with the purpose of obtaining their real fitness, to guide the subsequent evolution of the population. We apply the proposed method to seven benchmark MPI programs and compare it with several state-of-the-art approaches. The experimental results show that the proposed method can generate test data with reduced computation, thus improving the testing efficiency. Dun-Wei Gong, Baicai Sun, Xiangjuan Yao, Tian Tian 0010 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2020 | New Task Oriented Recommendation method Based on Hungarian algorithm in Crowdsourcing PlatformabstractAs a distributed problem-solving model based on human-machine integration, crowdsourcing has attracted wide attention in industry and academia with the development of Internet technology. There are many prominent problems on the crowdsourcing platform, for example, the task can't be noticed by users and the users are not competent for the task, resulting in huge waste of time and economy. If we can make coping recommendations according to the characteristics of the tasks, the operational efficiency of crowdsourcing platforms will be greatly improved. Therefore, this paper proposed a new task oriented recommendation method based on Hungarian algorithm in crowdsourcing platform. Aiming at the problem that the new users and tasks on the crowdsourcing platform have low matching degree, we establish the multi-objective optimization model of task recommendation with the aim of maximizing quality, time and cost efficiency. Then, the model is solved by the Hungarian algorithm through appropriate transformation. The experimental results show that the proposed method can improve the recommendations accuracy of new tasks, and therefor effectively improve the operational efficiency of crowdsourcing platforms. Zhimin Shi, Dun-Wei Gong, Xiangjuan Yao, Mengyi Yang |
SERVICES | 3 |
| 2020 | Efficiently Generating Test Data to Kill Stubborn Mutants by Dynamically Reducing the Search DomainabstractMutation testing is a fault-oriented software testing technique, and a test suite generated based on the criterion of mutation testing generally has a high capability in detecting faults. A mutant that is hard killed is called a stubborn one. The traditional methods of test data generation often fail to generate test data that kill stubborn mutants. To improve the efficiency of killing stubborn mutants, in this article, we propose a method of generating test data by dynamically reducing the search domain under the criterion of strong mutation testing. To fulfill this task, we first present a method of measuring the stubbornness of a mutant based on the reachability condition of a mutated statement. Then, we formulate the problem of generating test data to kill the mutant as an optimization one with a unique constraint. Finally, we generate test data using a coevolutionary genetic algorithm. Given the fact that the domain of test data that kills a stubborn mutant is generally small, we adopt a method of dynamically reducing the search domain to improve the efficiency of the algorithm. We apply the proposed method to test eight benchmark and industrial programs. The experimental results demonstrate that the proposed method has capabilities in seeking stubborn mutants and efficiently generating test data to kill stubborn mutants. Xiangying Dang, Xiangjuan Yao, Dun-Wei Gong, Tian Tian 0010 |
IEEE Trans. Reliab. | 2 |
| 2017 | Mutant reduction based on dominance relation for weak mutation testing
Dun-Wei Gong, Gongjie Zhang, Xiangjuan Yao, Fan-Lin Meng |
Inf. Softw. Technol. | 3 |
| 2017 | Online algorithms for scheduling on batch processing machines with interval graph compatibilities between jobs
Ji Tian, Ruyan Fu, Xiangjuan Yao |
Theor. Comput. Sci. | 4 |
| 2014 | A study of equivalent and stubborn mutation operators using human analysis of equivalenceabstractThough mutation testing has been widely studied for more than thirty years, the prevalence and properties of equivalent mutants remain largely unknown. We report on the causes and prevalence of equivalent mutants and their relationship to stubborn mutants (those that remain undetected by a high quality test suite, yet are non-equivalent). Our results, based on manual analysis of 1,230 mutants from 18 programs, reveal a highly uneven distribution of equivalence and stubbornness. For example, the ABS class and half UOI class generate many equivalent and almost no stubborn mutants, while the LCR class generates many stubborn and few equivalent mutants. We conclude that previous test effectiveness studies based on fault seeding could be skewed, while developers of mutation testing tools should prioritise those operators that we found generate disproportionately many stubborn (and few equivalent) mutants. Xiangjuan Yao, Mark Harman, Yue Jia 0001 |
ICSE | 1 |
| 2012 | Test data reduction based on dominance relations of target statementsabstractTraditional methods of generating test data may result in redundancy of test data, which brings many troubles to software testing. In order to solve the redundancy of test data, this study proposed a novel approach of generating test data by reducing target statements based on dominant relations. First, basic concepts and principles concerning dominance are listed. Then, an approach is proposed to reduce target statements according to their dominant relations. Finally, test suite covering the reduced set of target statements is generated by a genetic algorithm. The generated test suite can also cover all original target statements, which is guaranteed by the proposed strategy. We applied the method to nine benchmark programs, and compared with traditional and greedy methods. The experimental results show that our method can not only reduce redundancy, but also improve the efficiency of generating test data. Xiangjuan Yao, Dun-Wei Gong, Yongjin Luo, Ming Li 0013 |
IEEE Congress on Evolutionary Computation | 1 |
| 2012 | Grouping target paths for evolutionary generation of test data in parallel
Dun-Wei Gong, Tian Tian 0010, Xiangjuan Yao |
J. Syst. Softw. | 3 |
| 2012 | Testability transformation based on equivalence of target statements
Dun-Wei Gong, Xiangjuan Yao |
Neural Comput. Appl. | 2 |
| 2011 | Evolutionary generation of test data for many paths coverage based on grouping
Dun-Wei Gong, Wanqiu Zhang, Xiangjuan Yao |
J. Syst. Softw. | 3 |