Bach Hoai Nguyen

dblp:151/4329 · DBLP profile ↗
← Back
23ranked-venue papers
17as first author
12since 2021 · last 2025
0000-0002-6930-6863ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 17 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Differential Evolutionary for Label Ordering in Multi-label Classification
Bach Hoai Nguyen, Binh P. Nguyen, Vinh Truong Hoang
ADMA (4)1
2025 A Consistent Lebesgue Measure for Multi-label Learning
Kaan Demir, Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
PRICAI2
2025 Evolutionary Sparsity Regularisation-Based Feature Selection for Binary Classification
abstract
In classification, feature selection is an essential preprocessing step that selects a small subset of features to improve classification performance. Existing feature selection approaches can be divided into three main approaches: wrapper approaches, filter approaches, and embedded approaches. In comparison with the two other approaches, embedded approaches usually have better trade-off between classification performance and computation time. One of the most well-known embedded approaches is sparsity regularisation-based feature selection which generates sparse solutions for feature selection. Despite its good performance, sparsity regularisation-based feature selection outputs only a feature ranking which requires the number of selected features to be predefined. More importantly, the ranking mechanism introduces a risk of ignoring feature interactions which leads to the fact that many top-ranked but redundant features are selected. This work addresses the above problems by proposing a new representation that considers the interactions between features and can automatically determine an appropriate number of selected features. The proposed representation is used in a differential evolutionary (DE) algorithm to optimise the feature subset. In addition, a novel initialisation mechanism is proposed to let DE consider various numbers of selected features at the beginning. The proposed algorithm is examined on both synthetic and real-world datasets. The results on the synthetic dataset show that the proposed algorithm can select complementary features while existing sparsity regularisation-based feature selection algorithms are at risk of selecting redundant features. The results on real-world datasets show that the proposed algorithm achieves better classification performance than well-known wrapper, filter, and embedded approaches. The algorithm is also as efficient as filter feature selection approaches.
Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
Evol. Comput.1
2025 Evolutionary Instance Selection With Multiple Partial Adaptive Classifiers for Domain Adaptation
abstract
Domain adaptation reuses the knowledge learned from an existing (source) domain to classify unlabelled data from another related (target) domain. However, the two domains have different data distributions. Common approaches to bridge the two distributions are selecting/reweighting instances, building domain-invariant feature subspaces, or directly building adaptive classifiers. Recent domain adaptation work has shown that combining the above first two approaches before applying the third approach achieves better performance than performing each approach individually. However, most existing instance selection approaches are based on a ranking mechanism, ignore interdependences between instances, and require a pre-defined number of selected instances. Furthermore, adaptive classifiers are sensitive to their parameters which are challenging to optimise due to the lack of target labelled instances. This paper introduces a novel evolutionary instance selection approach for domain adaptation. We propose a compacted representation and an efficient fitness function for Particle Swarm Optimisation to automatically determine the number of selected instances while considering the interdependencies among instances. This paper also proposes to use multiple partial classifiers to build a more reliable and robust adaptive classifier. The results show that evolutionary instance selection selects better instances than the ranking approach. In cooperation with multiple partial classifiers, the proposed algorithm achieves better performance than nine state-of-the-art and well-known domain adaptation approaches.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2024 Improving Inference of Biochemical Composition in Marine Biomass via Genetic Algorithm-Based Feature Selection on Raman Spectroscopic Data
abstract
Assessing biochemical compositions of biomass from the fishing industry is challenging due to the complexity and seasonal variability of biological samples. Raman spectroscopy is often applied to measure complex biological samples, thereby enabling rapid quality control by associating the spectra with biochemical reference data using methods such as partial least squares regression. However, a small number of samples, noisy or misleading signals, and collinearity, often seen in real-world spectroscopic data, can negatively impact the fitting quality and inference capability of partial least squares regression. Feature selection is widely used to select a small and informative subset of the original features that can improve modeling performance, however, this is not always easy to achieve especially due to the aforementioned issues inherent to spectroscopic data. We address these issues by proposing a Genetic Algorithm-based feature selection approach for spectroscopic data acquired from New Zealand hoki and mackerel species. First, we apply a mathematical correction to the Raman signal most suited for each target composition, thereby reducing the effect of noise, irrelevant optical artifacts, and misleading signals. Next, we carefully curate a cross-validated feature selection process to circumvent the low number of samples using a new representation and fitness function to reduce regression error and balance model complexity. Our findings indicate that the proposed method can improve the fitting quality and inference capability of partial least squares regression over using the full set spectroscopic data. Lastly, we analyse the density of selected features to highlight the most salient signals.
Kaan Demir, Bach Hoai Nguyen, Jeremy S. Rooney, Bing Xue 0001, Mengjie Zhang 0001, Kirill Lagutin, Andrew MacKenzie, Keith C. Gordon, Daniel Killeen
CEC2
2024 Evolutionary Label Selection for Multi-label Classification
abstract
Multi-label classification is an emerging research area that aims to assign multiple class labels to each object or instance. In the era of big data, numerous multi-label problems, such as gene function prediction, involve a large number of labels, which most multi-label algorithms cannot cope with. One approach to reduce the number of labels is to construct a small set of new labels. However, creating new labels poses the challenge of preserving label meanings. This work proposes to decrease the number of labels by selecting a small subset of original labels, referred to as label selection, which preserves the original label meanings. Since label selection is a discrete optimisation problem with a large search space, we employ evolutionary computation (EC), specifically differential evolution, to achieve label selection. The contribution of this work lies in establishing a foundation for EC-based label selection, encompassing representation, fitness function, and the testing process. Experimental results demon-strate that the proposed evolutionary label selection algorithm significantly reduces the number of labels while maintaining or even improving performance compared to using all labels. Furthermore, the proposed algorithm exhibits high stability in terms of classification performance and the selected labels.
Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
CEC1
2024 A Survey on Evolutionary Multiobjective Feature Selection in Classification: Approaches, Applications, and Challenges
abstract
Maximizing the classification accuracy and minimizing the number of selected features are two primary objectives in feature selection, which is inherently a multiobjective task. Multiobjective feature selection enables us to gain various insights from complex data in addition to dimensionality reduction and improved accuracy, which has attracted increasing attention from researchers and practitioners. Over the past two decades, significant advancements in multiobjective feature selection in classification have been achieved in both the methodologies and applications, but have not been well summarized and discussed. To fill this gap, this paper presents a broad survey on existing research on multiobjective feature selection in classification, focusing on up-to-date approaches, applications, current challenges, and future directions. To be specific, we categorize multiobjective feature selection in classification on the basis of different criteria, and provide detailed descriptions of representative methods in each category. Additionally, we summarize a list of successful real-world applications of multiobjective feature selection from different domains, to exemplify their significant practical value and demonstrate their abilities in providing a set of trade-off feature subsets to meet different requirements of decision makers. We also discuss key challenges and shed lights on emerging directions for future developments of multiobjective feature selection.
Ruwang Jiao, Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.2
2024 A Constrained Competitive Swarm Optimizer With an SVM-Based Surrogate Model for Feature Selection
abstract
Feature selection (FS) is an important data preprocessing technique that selects a small subset of relevant features to improve learning performance. However, it is also challenging due to its large search space. Recently, a competitive swarm optimizer (CSO) has shown promising results in FS because of its potential global search ability. The main idea of CSO is to select two solutions randomly and then let the loser (worse fitness) learn from the winner (better fitness). Although such a search mechanism provides a high population diversity, it is at risk of generating unqualified solutions since the winner’s quality is not guaranteed. In this work, we propose a constrained evolutionary mechanism for CSO, which verifies the quality of all the particles and lets the infeasible (unqualified) solutions learn from the feasible (qualified) ones. We also propose a novel local search and a size-change operator that guide the population to search for smaller feature subsets with similar or better classification performance. A surrogate model, based on support vector machines, is proposed to assist both local search and the size-change operator to explore a massive number of potential feature subsets without requiring excessive computational resource. Results on 24 real-world datasets show that the proposed algorithm can select smaller feature subsets with higher classification performance than state-of-the-art evolutionary computation (EC) and non-EC benchmark algorithms.
Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2023 Co-operative Co-evolutionary Many-objective Embedded Multi-label Feature Selection with Decomposition-based PSO
abstract
Multi-label classification is an emerging machine-learning problem involving the prediction of a set of class labels based on the instance's features. In real-world problems, there are hundreds or thousands of features, many of which are irrelevant or redundant, resulting in a large search space for feature selection, i.e. to find small and discriminate feature subsets that improve the classification performance. Feature selection for multi-label classification is a many-objective optimisation problem with more than three main conflicting objectives: one of which is to reduce the number of selected features. There are many metrics for measuring multi-label classification performance, each of which can conflict with one another depending on the task. Hence, multi-label feature selection is a many-objective optimisation problem when three or more classification metrics and the number of selected features are optimised. In this paper, we propose to combine multi-label feature selection with evolutionary many-objective optimisation to address the above challenges and handle the trade-offs between multiple classification metrics and the number of selected features, using a decomposition-based algorithm. The results demonstrate that our proposed method is capable of finding discriminative and small feature subsets that can significantly improve the classification performances in comparison with other many-objective feature selection approaches.
Kaan Demir, Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
GECCO2
2021 Multi-objective Multi-label Feature Selection with an Aggregated Performance Metric and Dominance-based Initialisation
abstract
Feature selection (FS) is an essential pre-processing step that selects small and informative feature subsets. However, most FS methods focus on single-label classification, where each instance has one class label. Multi-label classification (MLC) is increasingly more common as instances possess a set of class labels. It is more challenging to perform FS for MLC since it needs to consider the interactions between labels. Furthermore, since the target output is a set of labels, there are various metrics for evaluating MLC performance, and each metric assesses the quality of a feature subset on various aspects. However, certain MLC metrics are shown to possess conflicting behaviour. To evolve high-quality feature subsets, it is essential to consider the trade-offs between metrics, making FS for MLC even more challenging. Evolutionary multi-objective optimisation (EMO) is a great technique to optimise multiple (potentially) conflicting metrics, representing different objectives. However, given a large number of objectives, EMO algorithms usually do not evolve a diverse set of good feature subsets, especially when there is a limited computational resource. This paper proposes a novel approach to reducing the number of objectives, which is expected to maintain or improve the evolved feature subsets over the original objectives. We also propose a new decomposition mechanism based on multiple reference points and a novel initialization mechanism to enhance the quality of the evolved feature subsets. The experimental results show that the proposed algorithm can evolve diverse feature subsets with better trade-off between multiple performance metrics than recent FS algorithms for MLC.
Kaan Demir, Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
CEC2
2021 A New Binary Particle Swarm Optimization Approach: Momentum and Dynamic Balance Between Exploration and Exploitation
abstract
Particle swarm optimization (PSO) is a heuristic optimization algorithm generally applied to continuous domains. Binary PSO is a form of PSO applied to binary domains but uses the concepts of velocity and momentum from continuous PSO, which leads to its limited performance. In our previous work, we reformulated momentum as a stickiness property and velocity as a flipping probability to develop sticky binary PSO. The initial design provides a good base, but many key factors need to be investigated. In this article, we propose a new algorithm called dynamic sticky binary PSO by developing a dynamic parameter control strategy based on an investigation of exploration and exploitation in the binary search spaces. The proposed algorithm is compared with four state-of-the-art dynamic binary algorithms on two types of binary problems: 1) knapsack and 2) feature selection. The experimental results on the knapsack datasets show that the new velocity and momentum assist sticky binary PSO in evolving better solutions than the benchmark algorithms. On feature selection, the dynamic strategy takes the advantages of these two newly defined movement concepts to help the proposed algorithm to produce smaller feature subsets with higher classification performance. This is the first time in the binary PSO, the four important concepts, that is, velocity, momentum, exploration, and exploitation, are investigated systematically to capture the properties of the binary search spaces to evolve better solutions for binary problems.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae, Mengjie Zhang 0001
IEEE Trans. Cybern.1
2021 A Hybrid Evolutionary Computation Approach to Inducing Transfer Classifiers for Domain Adaptation
abstract
Domain adaptation utilizes learned knowledge from an existing domain (source domain) to improve the classification performance of another related, but not identical, domain (target domain). Most existing domain adaptation methods first perform domain alignment, then apply standard classification algorithms. Transfer classifier induction is an emerging domain adaptation approach that incorporates the domain alignment into the process of building an adaptive classifier instead of using a standard classifier. Although transfer classifier induction approaches have achieved promising performance, they are mainly gradient-based approaches which can be trapped at local optima. In this article, we propose a transfer classifier induction algorithm based on evolutionary computation to address the above limitation. Specifically, a novel representation of the transfer classifier is proposed which has much lower dimensionality than the standard representation in existing transfer classifier induction approaches. We also propose a hybrid process to optimize two essential objectives in domain adaptation: 1) the manifold consistency and 2) the domain difference. Particularly, the manifold consistency is used in the main fitness function of the evolutionary search to preserve the intrinsic manifold structure of the data. The domain difference is reduced via a gradient-based local search applied to the top individuals generated by the evolutionary search. The experimental results show that the proposed algorithm can achieve better performance than seven state-of-the-art traditional domain adaptation algorithms and four state-of-the-art deep domain adaptation algorithms.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae, Mengjie Zhang 0001
IEEE Trans. Cybern.1
2020 A Decomposition based Multi-objective Evolutionary Algorithm with ReliefF based Local Search and Solution Repair Mechanism for Feature Selection
abstract
Feature selection has two main objectives which are to maximise the classification accuracy and to minimise the number of selected features. Unfortunately, the two objectives are usually in conflict, which makes feature selection a multi-objective problem. MOEA/D (multi-objective optimisation evolutionary algorithm based on decomposition) has shown to be effective in solving multi-objective feature selection, which evolves more diverse fronts than other multi-objective algorithms such as SPEA2 or NSGAII. However, sometimes the feature subsets around the middle of the evolved fronts do not have high classification performance. The goal of this work is to propose a local search for MOEA/D with an expectation of maintaining the front diversity while improving the classification performance of the feature subsets in the evolved fronts. The local search is based on three operators: insert, remove, and swap. The insert/remove operators either add/remove a single feature from the current feature subset, while the swap operator exchanges a selected feature with an unselected feature. The selection of added/removed/swapped features is based on Relief, a well-known measure which considers feature interactions. The experimental results show that the proposed local search can maintain or improve the fronts evolved by MOEA/D-DYN, a state-of-the-art MOEA/D algorithm for feature selection.
Kaan Demir, Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001
CEC2
2020 Multiple Reference Points-Based Decomposition for Multiobjective Feature Selection in Classification: Static and Dynamic Mechanisms
abstract
Feature selection is an important task in machine learning that has two main objectives: 1) reducing dimensionality and 2) improving learning performance. Feature selection can be considered a multiobjective problem. However, it has its problematic characteristics, such as a highly discontinuous Pareto front, imbalance preferences, and partially conflicting objectives. These characteristics are not easy for existing evolutionary multiobjective optimization (EMO) algorithms. We propose a new decomposition approach with two mechanisms (static and dynamic) based on multiple reference points under the multiobjective evolutionary algorithm based on decomposition (MOEA/D) framework to address the above-mentioned difficulties of feature selection. The static mechanism alleviates the dependence of the decomposition on the Pareto front shape and the effect of the discontinuity. The dynamic one is able to detect regions in which the objectives are mostly conflicting, and allocates more computational resources to the detected regions. In comparison with other EMO algorithms on 12 different classification datasets, the proposed decomposition approach finds more diverse feature subsets with better performance in terms of hypervolume and inverted generational distance. The dynamic mechanism successfully identifies conflicting regions and further improves the approximation quality for the Pareto fronts.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae, Hisao Ishibuchi, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2019 Population-based ensemble classifier induction for domain adaptation
abstract
In classification, the task of domain adaptation is to learn a classifier to classify target data using unlabeled data from the target domain and labeled data from a related, but not identical, source domain. Transfer classifier induction is a common domain adaptation approach that learns an adaptive classifier directly rather than first adapting the source data. However, most existing transfer classifier induction algorithms are gradient-based, so they can easily get stuck at local optima. Moreover, they usually generate only a single classifier which might fit the source data too well, which results in poor target accuracy. In this paper, we propose a population-based algorithm that can address the above two limitations. The proposed algorithm can re-initialize a population member to a promising region when the member is trapped at local optima. The population-based mechanism allows the proposed algorithm to output a set of classifiers which is more reliable than a single classifier. The experimental results show that the proposed algorithm achieves significantly better target accuracy than four state-of-the-art and well-known domain adaptation algorithms on three real-world domain adaptation problems.
Bach Hoai Nguyen, Bing Xue 0001, Mengjie Zhang 0001, Peter Andreae
GECCO1
2018 A particle swarm optimization based feature selection approach to transfer learning in classification
abstract
Transfer learning aims to use acquired knowledge from existing (source) domains to improve learning performance on a different but similar (target) domains. Feature-based transfer learning builds a common feature space, which can minimize differences between source and target domains. However, most existing feature-based approaches usually build a common feature space with certain assumptions about the differences between domains. The number of common features needs to be predefined. In this work, we propose a new feature-based transfer learning method using particle swarm optimization (PSO), where a new fitness function is developed to guide PSO to automatically select a number of original features and shift source and target domains to be closer. Classification performance is used in the proposed fitness function to maintain the discriminative ability of selected features in both domains. The use of classification accuracy leads to a minimum number of model assumptions. The proposed algorithm is compared with four state-of-the-art feature-based transfer learning approaches on three well-known real-world problems. The results show that the proposed algorithm is able to extract less than half of the original features with better performance than using all features and outperforms the four benchmark semi-supervised and unsupervised algorithms. This is the first time Evolutionary Computation, especially PSO, is utilized to achieve feature selection for transfer learning.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae
GECCO1
2017 Particle Swarm Optimisation with genetic operators for feature selection
abstract
Feature selection is an important task in machine learning, which aims to reduce the dataset dimensionality while at least maintaining the classification performance. Particle Swarm Optimisation (PSO) has been widely applied to feature selection because of its effectiveness and efficiency. However, since feature selection is a challenging task with a complex search space, PSO easily gets stuck at local optima. This paper aims to improve the PSO's searching ability by applying genetic operators such as crossover and mutation to assist the swarm to explore the search space better. The proposed genetic operators are specifically designed for feature selection, which not only improve the quality of current feature subsets but also make the search smoother. The proposed algorithm, called CMPSO, is tested and compared with three recent PSO based feature selection algorithms. Experimental results on eight datasets show that CMPSO can adapt with different numbers of features to evolve small feature subsets, which achieve similar or better classification performance than using all features and the three PSO based algorithms. The analysis on evolutionary processes shows that genetic operators assist CMPSO to evolve better solutions than the original PSO.
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae, Mengjie Zhang 0001
CEC1
2017 Surrogate-Model Based Particle Swarm Optimisation with Local Search for Feature Selection in Classification
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae
EvoApplications (1)1
2016 Mutual Information Estimation for Filter Based Feature Selection Using Particle Swarm Optimization
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae
EvoApplications (1)1
2016 Proceedings in Adaptation, Learning and Optimization
Bach Hoai Nguyen, Bing Xue 0001, Peter Andreae
IES1
2016 New mechanism for archive maintenance in PSO-based multi-objective feature selection
Bach Hoai Nguyen, Bing Xue 0001, Ivy Liu, Peter Andreae, Mengjie Zhang 0001
Soft Comput.1
2015 Gaussian Transformation Based Representation in Particle Swarm Optimisation for Feature Selection
Bach Hoai Nguyen, Bing Xue 0001, Ivy Liu, Peter Andreae, Mengjie Zhang 0001
EvoApplications1
2014 Filter based backward elimination in wrapper based PSO for feature selection in classification
abstract
The advances in data collection increase the dimensionality of the data (i.e. the total number of features) in many fields, which arises a challenge to many existing feature selection approaches. This paper develops a new feature selection approach based on particle swarm optimisation (PSO) and a local search that mimics the typical backward elimination feature selection method. The proposed algorithm uses a wrapper based fitness function, i.e. the classification error rate. The local search is performed only on the global best and uses a filter based measure, which aims to take the advantages of both filter and wrapper approaches. The proposed approach is tested and compared with three recent PSO based feature selection algorithms and two typical traditional feature selection methods. Experiments on eight benchmark datasets show that the proposed algorithm can be successfully used to select a significantly smaller number of features and simultaneously improve the classification performance over using all features. The proposed approach outperforms the three PSO based algorithms and the two traditional methods.
Bach Hoai Nguyen, Bing Xue 0001, Ivy Liu, Mengjie Zhang 0001
IEEE Congress on Evolutionary Computation1