Cuie Yang

dblp:190/2501 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-1997-1854ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HyReaL: Clustering Attributed Graph via Hyper-complex Space Representation Learning
Yang Lu 0009, Mengke Li 0001, Cuie Yang, Yiqun Zhang 0006, Yiu-Ming Cheung
DASFAA (2)4
2026 A Unified AutoEncoder-Based Representation Learning Framework for Fault Detection With Heterogeneous Feature Dimensions
abstract
Industrial data collected from similar processes under varying production specifications or monitoring configurations often exhibit structural heterogeneity, particularly in the form of varying feature dimensions. This presents a fundamental challenge for conventional fault detection models, which typically require fixed-length inputs and assume consistent feature spaces. As a result, practitioners are often forced to either discard valuable heterogeneous data or develop separate models for each product configuration, both of which compromise scalability and efficiency. To address this, we propose a novel heterogeneous AutoEncoder (HAE) framework that enables unified representation learning and fault detection on heterogeneous tabular data with arbitrary feature lengths. HAE adopts a Transformer-based encoder–decoder architecture guided by a binary padding indicator mask, which explicitly distinguishes between observed and padded attributes. To further improve discriminability and reduce intersource variation, a triplet-center loss is introduced to align latent representations and enhance interclass separability. We evaluate our method on two real-world datasets: a heavy-plate production process from the steel industry and a cross-project software defect prediction task. Experimental results show that HAE consistently outperforms a wide range of traditional, statistical, and generative imputation baselines or feature truncation method. Ablation studies further verify the effectiveness of each proposed component. HAE provides a scalable and generalizable solution for heterogeneous fault detection, without requiring strong assumptions on data missingness patterns or feature correspondences.
Shiqi Ren, Jinliang Ding, Cuie Yang, Tongkang Zhang, Yongchao Zhang 0004, Jun Zhao 0004
IEEE Trans. Ind. Informatics3
2025 A Scalable Test Problem Generator for Sequential Transfer Optimization
abstract
Despite the increasing interest in sequential transfer optimization (STO), a comprehensive benchmark suite for systematically comparing various STO algorithms remains underexplored. Existing test problems, which are often manually configured and lack scalability, can result in biased and nongeneralizable algorithm performance. In light of the above, we first introduce four concepts for characterizing STO problems (STOPs) in this study and present an important feature, namely similarity distribution, to quantitatively delineate the relationship between the optimal solutions of source and target tasks. Subsequently, we present general design guidelines for STOPs and introduce a problem generator that demonstrates strong scalability. Specifically, the similarity distribution of a problem can be easily customized through a novel inverse generation strategy, allowing for a continuous spectrum that captures the diverse similarity relationships present in real-world scenarios. Lastly, a benchmark suite comprising 12 STOPs, characterized by a range of customized similarity relationships, has been developed using the proposed generator and will serve as a platform for examining various STO algorithms. For instance, biased transferability representation, irregular mapping learning behaviors, and performance improvements unrelated to search experience are significant empirical findings that previous benchmarks failed to reveal, yet can be effectively identified through our test problems. The source code of the proposed problem generator is available at https://github.com/XmingHsueh/STOP-G.
Xiaoming Xue 0001, Cuie Yang, Liang Feng 0001, Kai Zhang 0029, Linqi Song, Kay Chen Tan
IEEE Trans. Cybern.2
2025 GOIO: Generative Oversampling Approach to Class Imbalance and Overlap of Tabular Data
abstract
Class imbalance, which is common in real-world classification tasks, often leads to biased models favoring majority classes. Data oversampling is a widely used strategy to address this issue. However, traditional oversampling methods often generate incorrect or redundant instances when class overlap occurs, increasing decision boundary complexity. To this end, we propose a novel Generative Oversampling approach to addressing Class Imbalance and Overlap (GOIO) in the classification of tabular data. GOIO combines a Metric-Learning-based Variational Autoencoder (MLVAE) and a Conditional Latent Diffusion Model (CLDM) to handle class imbalance and overlap effectively. The MLVAE employs a triplet-center loss to the adverse effects of class overlap by transforming the data distribution into a more separable latent feature space. Following this, the CLDM is trained with class-center feature prompting and classifier-free guidance strategy to capture class-specific latent distributions accurately. Minority class samples are synthesized in the latent space using the CLDM and then reconstructed into the data space via the MLVAE decoder. Comprehensive experiments on 18 real-world and five synthetic datasets demonstrate that GOIO outperforms the state-of-the-art oversampling methods in F1-score, MCC, and Accuracy. Ablation studies further validate the effectiveness of the proposed contributions in addressing class imbalance and overlap.
Shiqi Ren, Jinliang Ding, Cuie Yang, Yiu-Ming Cheung
IEEE Trans. Knowl. Data Eng.3
2024 A Surrogate-Assisted Coevolutionary Algorithm for Expensive Constrained Multiobjective Optimization
abstract
Surrogate-assisted evolutionary algorithms (SAEAs) are often used for expensive constrained multi-objective optimization problems (CMOPs). However, they exhibit poor convergence and diversity on complex CMOPs because they may consume more evaluations for approximating the constrained Pareto Front (CPF), i.e., the infeasible regions far from the CPF. To address this issue, we propose a surrogate-assisted coevolution-ary algorithm, which improves the convergence and diversity performance by searching on auxiliary problems together with the original CMOP. Firstly, we construct a surrogate model for each objective and constraint function. The primary population searches on both objective and constrained surrogate models, which approximate for the original constrained problem. The secondary population searches on an auxiliary problem, which is a multi-objective optimization problem (MOP) approximated by objective surrogate models. The two populations exchange information to enhance the convergence performance of the primary population. To improve the feasibility, the auxiliary problem is adopted with a certain probability. A novel infill criterion is proposed to guide the selection of high-quality solutions for updating the surrogate model, which relies on the feasible ratio of evaluated solutions. The performance of the proposed algorithm is evaluated on multi-objective and many-objective benchmarks to demonstrate its competitiveness.
Haofeng Wu, Qingda Chen, Cuie Yang, Jinliang Ding, Yaochu Jin
CEC3
2024 Multiobjective Sequential Transfer Optimization: Benchmark Problems and Preliminary Results
abstract
In cases of frequent problem-solving of multiobjective optimization tasks from a domain due to changing conditions or problem features, a growing number of individual tasks will be solved and stored in a database, providing an opportunity for a target task at hand to achieve better optimization performance through knowledge transfer from the previously-solved tasks, which is also known as sequential transfer optimization. Despite a variety of transfer algorithms that have been developed over the years, the research on the design of benchmark problems for evaluating such algorithms received far less attention. Oftentimes, the source and target tasks in existing test problems are manually assembled or extended from specific practical problems, limiting their ability to represent the diverse yet complex source-target similarity relationships in real-world problems. In light of this, we propose design methods to generate multiobjective sequential transfer optimization problems (MSTOPs) systematically in this work, wherein the Pareto manifolds of individual tasks and the manifold-based similarity between the tasks can be customized with ease, enabling a broad spectrum of representation of the diverse similarity relationships between the source-target Pareto manifolds of MSTOPs. Lastly, a benchmark suite with 12 test problems is developed using the proposed methods, which would serve as an arena for electing superior multiobjective sequential transfer optimization algorithms. The source code is available at https://github.com/XmingHsueh/MSTOP.
Xiaoming Xue 0001, Liang Feng 0001, Cuie Yang, Songbai Liu, Linqi Song, Kay Chen Tan
CEC3
2024 Distributed Knowledge Transfer for Evolutionary Multitask Multimodal Optimization
abstract
Evolutionary multitasking Optimization (EMTO) is a paradigm that optimizes multiple tasks simultaneously to improve the overall performance of all tasks by seamlessly transferring useful knowledge among them. Although EMTO has received significant interest, rare studies consider handling tasks that are multimodal optimization problems (MMOPs) with multiple global optimal solutions. Due to the multiple different modalities of each task, a major challenge of solving multiple MMOPs is how to extract and transfer knowledge across modalities of different tasks. To this end, this paper designs a distributed knowledge transfer based evolutionary multitask multimodal optimization (EMTMO-DKT) approach for solving multiple MMOPs simultaneously by discovering and utilizing local knowledge across modalities of different tasks. Specifically, we first divide the population of each task into multiple subpopulations, where each subpopulation explores a modality. Then, we propose an evolution path based similarity measurement to measure the local similarities between subpopulations of different tasks. Since the modalities can be locally similar across tasks, we develop a subpopulation cross matching strategy according to the obtained similarities to pair subpopulations of different tasks. In this stage, the successfully paired subpopulations are allowed to transfer knowledge. Finally, the knowledge transfer probability self-adjusting strategy is applied to each subpopulation to balance knowledge transfer and self-evolution, so as to improve search efficiency. In this paper, a set of multitask multimodal optimization test problems are constructed to assess the efficacy of compared algorithms. Experimental results on both the benchmark functions and the real-world optimization problem demonstrate that the proposed algorithm can quickly locate more global optima in comparison with state-of-the-art EMTO and multimodal optimization algorithms.
Kailai Gao, Cuie Yang, Jinliang Ding, Kay Chen Tan, Tianyou Chai
IEEE Trans. Evol. Comput.2
2024 Solution Transfer in Evolutionary Optimization: An Empirical Study on Sequential Transfer
abstract
Knowledge transfer from optimized problems has emerged as a promising technique for enhancing evolutionary search. However, most studies in this domain primarily concentrate on devising knowledge transfer mechanisms for specific problem domains, often lacking the examination of the fundamental aspects of knowledge transfer, i.e., what, when and how to transfer across diverse scenarios. This not only restricts the generality of these algorithms but also hinders their practical applicability. In light of this, this paper: 1) reviews a vast array of techniques associated with the crucial aspects of solution transfer and 2) conducts a series of experiments to explore the underlying transfer mechanisms that enhance the evolutionary search. In particular, we first define solution transferability in the context of evolutionary search, which provides a new perspective in understanding what, when and how to transfer in enhancing evolutionary search. Next, through comprehensive experiments, we find that the approximation and evaluation of solution transferability is of great importance in designing what, when and how to transfer towards enhanced evolutionary search. Furthermore, our empirical study also discusses the counterintuitive performance improvements unrelated to the search experience of source tasks. The source code for reproducing our experiments is available at https://github.com/XmingHsueh/STO-EC.
Xiaoming Xue 0001, Cuie Yang, Liang Feng 0001, Kai Zhang 0029, Linqi Song, Kay Chen Tan
IEEE Trans. Evol. Comput.2
2024 A Co-Training Framework for Heterogeneous Heuristic Domain Adaptation
abstract
The purpose of this article is to address unsupervised domain adaptation (UDA) where a labeled source domain and an unlabeled target domain are given. Recent advanced UDA methods attempt to remove domain-specific properties by separating domain-specific information from domain-invariant representations, which heavily rely on the designed neural network structures. Meanwhile, they do not consider class discriminate representations when learning domain-invariant representations. To this end, this article proposes a co-training framework for heterogeneous heuristic domain adaptation (CO-HHDA) to address the above issues. First, a heterogeneous heuristic network is introduced to model domain-specific characters. It allows structures of heuristic network to be different between domains to avoid underfitting or overfitting. Specially, we initialize a small structure that is shared between domains and increase a subnetwork for the domain which preserves rich specific information. Second, we propose a co-training scheme to train two classifiers, a source classifier and a target classifier, to enhance class discriminate representations. The two classifiers are designed based on domain-invariant representations, where the source classifier learns from the labeled source data, and the target classifier is trained from the generated target pseudolabeled data. The two classifiers teach each other in the training process with high-quality pseudolabeled data. Meanwhile, an adaptive threshold is presented to select reliable pseudolabels in each classifier. Empirical results on three commonly used benchmark datasets demonstrate that the proposed CO-HHDA outperforms the state-of-the-art domain adaptation methods.
Cuie Yang, Bing Xue 0001, Kay Chen Tan, Mengjie Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 A Data Stream Ensemble Assisted Multifactorial Evolutionary Algorithm for Offline Data-Driven Dynamic Optimization
abstract
Existing work on offline data-driven optimization mainly focuses on problems in static environments, and little attention has been paid to problems in dynamic environments. Offline data-driven optimization in dynamic environments is a challenging problem because the distribution of collected data varies over time, requiring surrogate models and optimal solutions tracking with time. This paper proposes a knowledge-transfer-based data-driven optimization algorithm to address these issues. First, an ensemble learning method is adopted to train surrogate models to leverage the knowledge of data in historical environments as well as adapt to new environments. Specifically, given data in a new environment, a model is constructed with the new data, and the preserved models of historical environments are further trained with the new data. Then, these models are considered to be base learners and combined as an ensemble surrogate model. After that, all base learners and the ensemble surrogate model are simultaneously optimized in a multitask environment for finding optimal solutions for real fitness functions. In this way, the optimization tasks in the previous environments can be used to accelerate the tracking of the optimum in the current environment. Since the ensemble model is the most accurate surrogate, we assign more individuals to the ensemble surrogate than its base learners. Empirical results on six dynamic optimization benchmark problems demonstrate the effectiveness of the proposed algorithm compared with four state-of-the-art offline data-driven optimization algorithms. Code is available at https://github.com/Peacefulyang/DSE_MFS.git.
Cuie Yang, Jinliang Ding, Yaochu Jin, Tianyou Chai
Evol. Comput.1
2023 Contrastive Learning Assisted-Alignment for Partial Domain Adaptation
abstract
This work addresses unsupervised partial domain adaptation (PDA), in which classes in the target domain are a subset of the source domain. The key challenges of PDA are how to leverage source samples in the shared classes to promote positive transfer and filter out the irrelevant source samples to mitigate negative transfer. Existing PDA methods based on adversarial DA do not consider the loss of class discriminative representation. To this end, this article proposes a contrastive learning-assisted alignment (CLA) approach for PDA to jointly align distributions across domains for better adaptation and to reweight source instances to reduce the contribution of outlier instances. A contrastive learning-assisted conditional alignment (CLCA) strategy is presented for distribution alignment. CLCA first exploits contrastive losses to discover the class discriminative information in both domains. It then employs a contrastive loss to match the clusters across the two domains based on adversarial domain learning. In this respect, CLCA attempts to reduce the domain discrepancy by matching the class-conditional and marginal distributions. Moreover, a new reweighting scheme is developed to improve the quality of weights estimation, which explores information from both the source and the target domains. Empirical results on several benchmark datasets demonstrate that the proposed CLA outperforms the existing state-of-the-art PDA methods.
Cuie Yang, Yiu-Ming Cheung, Jinliang Ding, Kay Chen Tan, Bing Xue 0001, Mengjie Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Evolutionary Sequential Transfer Optimization for Objective-Heterogeneous Problems
abstract
Evolutionary sequential transfer optimization is a paradigm that leverages search experience from solved source optimization tasks to accelerate the evolutionary search of a target task. Even though many algorithms have been developed, they mainly focus on objective-homogeneous problems, where the source and target tasks possess a similar number of objectives. In this work, we explore objective-heterogeneous problems, in which knowledge transfers across single-objective optimization problems (SOPs), multiobjective optimization problems (MOPs), and many-objective optimization problems (MaOPs). Objective-heterogeneous problems challenge the existing methods due to the diverse search and objective spaces between the source and the target task. To address this issue, we present a decision variable analysis-based transfer method that can conduct knowledge transfer across problems with the different numbers of objectives. We first separate decision variables of MOPs and MaOPs into convergence-related variables (CVs) and diversity-related variables (DVs), according to their roles while treating variables of SOPs as CVs. Then, we propose a convergence transfer module to transfer knowledge of CVs to speed up the convergence. It aligns both solutions and fitness ranks for preserving fitness rank consistency between the source and target tasks, whereby accelerating search speed. Besides, a diversity transfer module is presented to refine the distribution of DVs to maintain the population diversity. The experimental results on objective-heterogeneous test problems and a real-world case study have demonstrated the effectiveness of the proposed algorithm.
Xiaoming Xue 0001, Cuie Yang, Yao Hu 0001, Kai Zhang 0029, Yiu-Ming Cheung, Linqi Song, Kay Chen Tan
IEEE Trans. Evol. Comput.2
2022 Concept Drift-Tolerant Transfer Learning in Dynamic Environments
abstract
Existing transfer learning methods that focus on problems in stationary environments are not usually applicable to dynamic environments, where concept drift may occur. To the best of our knowledge, the concept drift-tolerant transfer learning (CDTL), whose major challenge is the need to adapt the target model and knowledge of source domains to the changing environments, has yet to be well explored in the literature. This article, therefore, proposes a hybrid ensemble approach to deal with the CDTL problem provided that data in the target domain are generated in a streaming chunk-by-chunk manner from nonstationary environments. At each time step, a class-wise weighted ensemble is presented to adapt the model of target domains to new environments. It assigns a weight vector for each classifier generated from the previous data chunks to allow each class of the current data leveraging historical knowledge independently. Then, a domain-wise weighted ensemble is introduced to combine the source and target models to select useful knowledge of each domain. The source models are updated with the source instances performed by the proposed adaptive weighted CORrelation ALignment (AW-CORAL). AW-CORAL iteratively minimizes domain discrepancy meanwhile decreases the effect of unrelated source instances. In this way, positive knowledge of source domains can be potentially promoted while negative knowledge is reduced. Empirical studies on synthetic and real benchmark data sets demonstrate the effectiveness of the proposed algorithm.
Cuie Yang, Yiu-Ming Cheung, Jinliang Ding, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.1
2020 Offline Data-Driven Multiobjective Optimization: Knowledge Transfer Between Surrogates and Generation of Final Solutions
abstract
In offline data-driven optimization, only historical data is available for optimization, making it impossible to validate the obtained solutions during the optimization. To address these difficulties, this paper proposes an evolutionary algorithm assisted by two surrogates, one coarse model and one fine model. The coarse surrogate (CS) aims to guide the algorithm to quickly find a promising subregion in the search space, whereas the fine one focuses on leveraging good solutions according to the knowledge transferred from the CS. Since the obtained Pareto optimal solutions have not been validated using the real fitness function, a technique for generating the final optimal solutions is suggested. All achieved solutions during the whole optimization process are grouped into a number of clusters according to a set of reference vectors. Then, the solutions in each cluster are averaged and outputted as the final solution of that cluster. The proposed algorithm is compared with its three variants and two state-of-the-art offline data-driven multiobjective algorithms on eight benchmark problems to demonstrate its effectiveness. Finally, the proposed algorithm is successfully applied to an operational indices optimization problem in beneficiation processes.
Cuie Yang, Jinliang Ding, Yaochu Jin, Tianyou Chai
IEEE Trans. Evol. Comput.1
2019 Multitasking Multiobjective Evolutionary Operational Indices Optimization of Beneficiation Processes
abstract
Operational indices optimization is crucial for the global optimization in beneficiation processes. This paper presents a multitasking multiobjective evolutionary method to solve operational indices optimization, which involves a formulated multiobjective multifactorial operational indices optimization (MO-MFO) problem and the proposed multiobjective MFO algorithm for solving the established MO-MFO problem. The MO-MFO problem includes multiple level of accurate models of operational indices optimization, which are generated on the basis of a data set collected from production. Among the formulated models, the most accurate one is considered to be the original functions of the solved problem, while the remained models are the helper tasks to accelerate the optimization of the most accurate model. For the MFO algorithm, the assistant models are alternatively in multitasking environment with the accurate model to transfer their knowledge to the accurate model during optimization in order to enhance the convergence of the accurate model. Meanwhile, the recently proposed two-stage assortative mating strategy for a multiobjective MFO algorithm is applied to transfer knowledge among multitasking tasks. The proposed multitasking framework for operational indices optimization has conducted on 10 different production conditions of beneficiation. Simulation results demonstrate its effectiveness in addressing the operational indices optimization of beneficiation problem. Note to Practitioners-Operational indices optimization is a typical approach to achieve global production optimization by efficiently coordinating all the indices to improve the production indices. In this paper, a multiobjective multitasking framework is developed to address the operational indices optimization, which includes a multitasking multiobjective operational indices optimization problem formulation and a multitasking multiobjective evolutionary optimization to solve the above-formulated optimization problem. The proposed approach can achieve a solution set for the decision-making. The simulation results on a real beneficiation process in China with 10 operational conditions show that the proposed approach is able to obtain a superior solution set, which is associated with a higher grade and yield of the product.
Cuie Yang, Jinliang Ding, Yaochu Jin, Tianyou Chai
IEEE Trans Autom. Sci. Eng.1
2019 Generalized Multitasking for Evolutionary Optimization of Expensive Problems
abstract
Conventional evolutionary algorithms (EAs) are not well suited for solving expensive optimization problems due to the fact that they often require a large number of fitness evaluations to obtain acceptable solutions. To alleviate the difficulty, this paper presents a multitasking evolutionary optimization framework for solving computationally expensive problems. In the framework, knowledge is transferred from a number of computationally cheap optimization problems to help the solution of the expensive problem on the basis of the recently proposed multifactorial EA (MFEA), leading to a faster convergence of the expensive problem. However, existing MFEAs do not work well in solving multitasking problems whose optimums do not lie in the same location or when the dimensions of the decision space are not the same. To address the above issues, the existing MFEA is generalized by proposing two strategies, one for decision variable translation and the other for decision variable shuffling, to facilitate knowledge transfer between optimization problems having different locations of the optimums and different numbers of decision variables. To assess the effectiveness of the generalized MFEA (G-MFEA), empirical studies have been conducted on eight multitasking instances and eight test problems for expensive optimization. The experimental results demonstrate that the proposed G-MFEA works more efficiently for multitasking optimization and successfully accelerates the convergence of expensive optimization problems compared to single-task optimization.
Jinliang Ding, Cuie Yang, Yaochu Jin, Tianyou Chai
IEEE Trans. Evol. Comput.2
2018 Incremental data-driven optimization of complex systems in nonstationary environments
Cuie Yang, Jinliang Ding, Yaochu Jin, Tianyou Chai
Sci. China Inf. Sci.1
2016 Reference point based prediction for evolutionary dynamic multiobjective optimization
abstract
Using evolutionary algorithms (EAs) to handle dynamic multiobjective optimization problems (DMOPs) is a challenging topic. In this paper, a prediction strategy based on reference points is proposed to improve the performance of EAs in solving DMOPs. The reference point based strategy is to partition the population into several subpopulations according to the reference points. When a change is detected, a sequence of the subpopulation centers in the previous environments belonging to the same reference point are used to estimate the center of in new environment. Based on the predicted centers, the EA will generate an initial population for the new environment using a combination of a Uniform distribution for enhancing population diversity and a Gaussian distributions for accelerating convergence. Experiments on ten test instances have been carried out to evaluate the performance of proposed strategy and the results show that reference point based prediction strategy exhibits superior performance in dealing with DMOPs with nonlinear correlations between decision variables and severe environmental changes.
Yaochu Jin, Cuie Yang, Jinliang Ding, Tianyou Chai
CEC2