Yanan Sun 0001

dblp:44/8711-1 · DBLP profile ↗
← Back
90ranked-venue papers
17as first author
69since 2021 · last 2026
0000-0001-6374-1429ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 79 · 17 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models
abstract
Spiking Large Language Models (LLMs) have emerged as an energy-efficient alternative to conventional LLMs through their event-driven computation. To effectively obtain spiking LLMs, researchers develop different ANN-to-SNN conversion methods by leveraging pre-trained ANN parameters while inheriting the energy efficiency of SNN. However, existing conversion methods struggle with extreme activation outliers and incompatible nonlinear operations of ANN-based LLMs. To address this, we propose a loss-less ANN-SNN conversion for fully spike-driven LLMs, termed LAS. Specifically, LAS introduces two novel neurons to convert the activation outlier and nonlinear operation of ANN-based LLMs. Moreover, LAS tailors the spike-equivalent Transformer components for spiking LLMs, which can ensure full spiking conversion without any loss of performance. Experimental results on six language models and two vision-language models demonstrate that LAS achieves loss-less conversion. Notably, on OPT-66B, LAS even improves the accuracy of 2% on the WSC task. In addition, the parameter and ablation studies further verify the effectiveness of LAS.
Xiaotian Song, Yanan Sun 0001
AAAI3
2026 Adapt Before Continual Learning
abstract
Continual Learning (CL) seeks to enable neural networks to incrementally acquire new knowledge (plasticity) while retaining existing knowledge (stability). Although pre-trained models (PTMs) have provided a strong foundation for CL, existing approaches face a fundamental challenge in balancing these two competing objectives. Current methods typically address stability by freezing the PTM backbone, which severely limits the model's plasticity, particularly when incoming data distribution diverges largely from the pre-training data. Alternatively, sequentially fine-tuning the entire PTM can adapt to new knowledge but often leads to catastrophic forgetting, highlighting the critical stability-plasticity trade-off in PTM-based CL. To address this limitation, we propose Adapting PTMs before the core CL process (ACL), a novel framework that introduces a plug-and-play adaptation phase prior to learning each new task. During this phase, ACL refines the PTM backbone by aligning embeddings with their original class prototypes while distancing them from irrelevant classes. This mechanism theoretically and empirically demonstrates desirable balance between stability and plasticity, significantly improving CL performance across benchmarks and integrated methods.
Aojun Lu, Tao Feng 0014, Hangjie Yuan, Chunhui Ding, Yanan Sun 0001
AAAI5
2026 When Fitness Is Cheap: Pareto-Based Evolutionary Optimisation for Lightweight Neural Architectures
abstract
Evolutionary algorithms are well suited to neural architecture search and other combinatorial design problems, but their scalability is often limited by the high cost of fitness evaluation. This paper studies evolutionary multi-objective optimisation in a regime where fitness evaluations are effectively free, enabled by a training-free proxy for neural network expressivity. We propose SWAP-Lite, a Pareto-guided evolutionary algorithm that maintains an explicit archive of non-dominated solutions over representational capacity and deployment cost, yielding an anytime optimiser that exposes budget-feasible solutions throughout the search. Using a MobileNet-style architecture space as a case study, we instantiate the fitness function with a sample-wise activation pattern proxy and perform large-scale evolutionary searches with up to 105 architecture evaluations. Experiments on CIFAR-10 and ImageNet show that SWAP-Lite discovers compact architectures that are competitive with state-of-the-art training-based and zero-shot baselines, while reducing search cost by one to four orders of magnitude. Analysis of the evolutionary dynamics demonstrates that explicit bi-objective optimisation produces higher-quality constrained Pareto fronts and superior anytime hypervolume compared with random search, greedy local search, and single-objective evolutionary baselines.
Jingyue Cong, Yameng Peng, Yanan Sun 0001, Andy Song
GECCO4
2026 A gradient-based lightweight network automated design method for facial expression recognition
Shuchao Deng, Xiaotian Song, Jiyuan Liu 0012, Yanan Sun 0001
Expert Syst. Appl.5
2026 Activation features empowered: An efficient gradient-free proxy for zero-shot neural architecture search
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001, Gary G. Yen
Pattern Recognit.3
2026 DAP: Domain Adaptive Performance Predictor for Efficient Neural Architecture Search
abstract
Neural architecture search (NAS) aims to automatically design high-performance architectures of deep neural networks, which have shown great potential in various fields. However, the search process of NAS is computationally expensive since plenty of deep neural networks are trained to get the performance on GPUs. Performance predictors can directly estimate the performance of architectures without GPU-based training, thus can overcome this barrier. However, the construction of performance predictors requires labeling plenty of architectures sampled from the corresponding NAS search space, which is still prohibitively costly. In this paper, we propose a Domain Adaptive performance Predictor (DAP), which can construct a performance predictor based on the labeled architectures provided by existing benchmarks and then enable it to other search spaces via domain adaptive techniques. To achieve this, we first propose a domain-agnostic feature extraction method to refine the domain-invariant features of neural architectures. Then, we propose a novel embedding method to learn the shared representations of architecture operations. Experimental results demonstrate that DAP outperforms eight baselines upon six popular search spaces. Notably, we only require the search cost of 0:0002 GPU Days to find the architecture with 77:10% top-1 accuracy on ImageNet and 97:86% on CIFAR-10. In addition, we show the theoretical upper bound of the generalization error in the target search space, further illustrating the generalizability of DAP. The source code is available athttps://anonymous.4open.science/r/DAP-2F1F/.
Xiaotian Song, Andy Song, Yanan Sun 0001
IEEE Trans. Computers5
2026 Runtime analysis of evolutionary neural architecture search for binary classification
Zeqiong Lv, Chao Bian 0002, Chao Qian 0001, Yanan Sun 0001
Theor. Comput. Sci.4
2026 A Multiagent Transformer-Based Algorithm for Multitask Dynamic Scheduling With Constrained Machines
abstract
Modern manufacturers often require handling multiple tasks simultaneously under dynamic environments by sharing constrained machines. Existing multitask scheduling algorithms typically focus on transferring knowledge among multiple tasks. However, these algorithms overlook the need for collaborative multitasking efforts required to share constrained machines. To overcome this limitation, we propose a multiagent transformer (MAT)-based algorithm to solve multitask dynamic scheduling with constrained machines. Specifically, we first formulate the multitask scheduling problem as a sequential multiagent decision-making process, enabling agents to make collaborative decisions by accessing the actions of others. Furthermore, a joint policy network is developed to support the agents in adaptively selecting the appropriate heuristic for each task. It improves the decision-making quality by enabling agents to leverage common and task-specific knowledge. In addition, a comprehensive reward function is designed to guide the learning of a joint policy network for collaborative decision-making across tasks. This ensures that agents holistically consider the objectives of all tasks during the learning process. With these designs, the proposed algorithm can effectively address multiple tasks through collaborative machine sharing. The proposed algorithm is evaluated against 14 state-of-the-art competitors on 270 instances with varying scales. The results confirm that the proposed algorithm outperforms all competitors on each instance. In addition, the ablation study demonstrates the effectiveness of distinct reward mechanisms, revealing that the joint policy network makes more informed decisions by leveraging both individual and common knowledge.
Sri Srinivasa Raju Modampuri, Yanan Sun 0001
IEEE Trans. Cybern.4
2026 Evolutionary Multiobjective Spiking Neural Architecture Search for Image Classification
abstract
Spiking neural networks (SNNs) have the merit of energy efficiency, and have been widely used for various real-world applications. Similar to other types of neural networks, the performance of SNN is also significantly decided by its architecture. In this paper, we propose an Evolutionary Multi-objective Spiking Neural Architecture Search (EMO-SNAS) method that completely enables the automatic design of SNN architectures with both high performance and low power consumption. To achieve this, we first design a variable-length encoding strategy for SNNs, addressing the issue that traditional encoding strategies need to manually set the depth in advance. Furthermore, we propose an exploitation operator focusing on the local search for the variable-length encoding, as well as an exploration operator focusing on the global search based on the temporal expansion. Based on NSGA-II, EMO-SNAS can greatly balance the performance and power consumption during the architecture design. Experiments on three widely used image classification datasets show that EMO-SNAS can achieve the best among the state-of-the-art methods. Specifically, EMO-SNAS gains 0.45%, 0.26%, and 7.52% in terms of classification accuracy, yet significantly contributes to 33%, 29%, and 12% fewer spike numbers on CIFAR10, CIFAR100, and TinyImageNet datasets. Ablation studies show that the temporal expansion can improve the performance of EMO-SNAS. Moreover, the measurement of power consumption and theoretical convergence of EMO-SNAS are also discussed to justify its component design. In addition, with EMO-SNAS, the impact of initial channels for SNNs is also systemically investigated, based on which a conclusion against existing consensus is achieved. The source code is available at https://github.com/songxt3/EMO-SNAS.
Xiaotian Song, Zeqiong Lv, Deng Xiong, Jiancheng Lv 0001, Jiyuan Liu 0012, Yanan Sun 0001
IEEE Trans. Evol. Comput.7
2026 Accurate and Robust Neural Architecture Search via a Flexible Supernet
abstract
Neural architecture search (NAS) has been widely adopted to design high-accuracy architectures, which are often vulnerable against adversarial attacks. To address this problem, existing robust NAS methods mainly focus on optimizing both natural accuracy and adversarial robustness in a fixed supernet, which is designed for natural accuracy. As a result, the derived architectures have the same construction scheme as the supernet, suffering from limited adversarial robustness and flexibility. In this article, we present the ARNAS++ method to search for accurate and robust neural architectures via a supernet with flexible parameter budgets and width. Specifically, we propose a parameter budget controlling loss to make architectures contain less parameters in the rear cells, based on which the adversarial robustness can be guaranteed. Moreover, we also propose a learnable filter number reduction ratio to control the filter numbers in the supernet, which can find more robust architectures beyond the fixed supernet, and make the supernet more flexible at the same time. We conduct experiments on six widely used benchmark datasets against the state of the art. The experimental results demonstrate that the proposed ARNAS++ method outperforms the competitors in terms of both natural accuracy and adversarial robustness under various popular adversarial attacks. In addition, the ablation studies show the effectiveness of the designed components and their positive contributions to the overall performance. The source code is available at: https://github.com/fyqsama/ARNASpp.
Yuwei Ou, Yanan Sun 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Revisiting Long-Tailed Learning: Insights from an Architectural Perspective
abstract
Long-Tailed (LT) recognition has been widely studied to tackle the challenge of imbalanced data distributions in real-world applications. However, the design of neural architectures for LT settings has received limited attention, despite evidence showing that architecture choices can substantially affect performance. This paper aims to bridge the gap between LT challenges and neural network design by providing an in-depth analysis of how various architectures influence LT performance. Specifically, we systematically examine the effects of key network components on LT handling, such as topology, convolutions, and activation functions. Based on these observations, we propose two convolutional operations optimized for improved performance. Recognizing that operation interactions are also crucial to network effectiveness, we apply Neural Architecture Search (NAS) to facilitate efficient exploration. We propose LT-DARTS, a NAS method with a novel search space and search strategy specifically designed for LT data. Experimental results demonstrate that our approach consistently outperforms existing architectures across multiple LT datasets, achieving parameter-efficient, state-of-the-art results when integrated with current LT methods.
Yuhan Pan, Yanan Sun 0001, Wei Gong 0001
CIKM2
2025 Evolving Comprehensive Proxies for Zero-Shot Neural Architecture Search
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001
GECCO3
2025 CARL: Causality-Guided Architecture Representation Learning for an Interpretable Performance Predictor
abstract
Performance predictors have emerged as a promising method to accelerate the evaluation stage of neural architecture search (NAS). These predictors estimate the performance of unseen architectures by learning from the correlation between a small set of trained architectures and their performance. However, most existing predictors ignore the inherent distribution shift between limited training samples and diverse test samples. Hence, they tend to learn spurious correlations as shortcuts to predictions, leading to poor generalization. To address this, we propose a Causality-guided Architecture Representation Learning (CARL) method aiming to separate critical (causal) and redundant (non-causal) features of architectures for generalizable architecture performance prediction. Specifically, we employ a substructure extractor to split the input architecture into critical and redundant substructures in the latent space. Then, we generate multiple interventional samples by pairing critical representations with diverse redundant representations to prioritize critical features. Extensive experiments on five NAS search spaces demonstrate the state-of-the-art accuracy and superior interpretability of CARL. For instance, CARL achieves 97.67% top-1 accuracy on CIFAR-10 using DARTS.
Han Ji 0003, Yanan Sun 0001
ICCV4
2025 Loss Functions for Predictor-Based Neural Architecture Search
abstract
Evaluation is a critical but costly procedure in neural architecture search (NAS). Performance predictors have been widely adopted to reduce evaluation costs by directly estimating architecture performance. The effectiveness of predictors is heavily influenced by the choice of loss functions. While traditional predictors employ regression loss functions to evaluate the absolute accuracy of architectures, recent approaches have explored various ranking-based loss functions, such as pairwise and listwise ranking losses, to focus on the ranking of architecture performance. Despite their success in NAS, the effectiveness and characteristics of these loss functions have not been thoroughly investigated. In this paper, we conduct the first comprehensive study on loss functions in performance predictors, categorizing them into three main types: regression, ranking, and weighted loss functions. Specifically, we assess eight loss functions using a range of NAS-relevant metrics on 13 tasks across five search spaces. Our results reveal that specific categories of loss functions can be effectively combined to enhance predictor-based NAS. Furthermore, our findings could provide practical guidance for selecting appropriate loss functions for various tasks. We hope this work provides meaningful insights to guide the development of loss functions for predictor-based methods in the NAS community.
Han Ji 0003, Yanan Sun 0001
ICCV4
2025 Zero-cost Proxy for Adversarial Robustness Evaluation
abstract
Deep neural networks (DNNs) easily cause security issues due to the lack of adversarial robustness. An emerging research topic for this problem is to design adversarially robust architectures via neural architecture search (NAS), i.e., robust NAS. However, robust NAS needs to train numerous DNNs for robustness estimation, making the search process prohibitively expensive. In this paper, we propose a zero-cost proxy to evaluate the adversarial robustness without training. Specifically, the proposed zero-cost proxy formulates the upper bound of adversarial loss, which can directly reflect the adversarial robustness. The formulation involves only the initialized weights of DNNs, thus the training process is no longer needed. Moreover, we theoretically justify the validity of the proposed proxy based on the theory of neural tangent kernel and input loss landscape. Experimental results show that the proposed zero-cost proxy can bring more than $20\times$ speedup compared with the state-of-the-art robust NAS methods, while the searched architecture has superior robustness and transferability under white-box and black-box attacks. Furthermore, compared with the state-of-the-art zero-cost proxies, the calculation of the proposed method has the strongest correlation with adversarial robustness. Our source code is available at https://github.com/fyqsama/Robust_ZCP.
Yuwei Ou, Yanan Sun 0001
ICLR4
2025 Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective
abstract
The quest for Continual Learning (CL) seeks to empower neural networks with the ability to learn and adapt incrementally. Central to this pursuit is addressing the stability-plasticity dilemma, which involves striking a balance between two conflicting objectives: preserving previously learned knowledge and acquiring new knowledge. While numerous CL methods aim to achieve this trade-off, they often overlook the impact of network architecture on stability and plasticity, restricting the trade-off to the parameter level. In this paper, we delve into the conflict between stability and plasticity at the architectural level. We reveal that under an equal parameter constraint, deeper networks exhibit better plasticity, while wider networks are characterized by superior stability. To address this architectural-level dilemma, we introduce a novel framework denoted Dual-Arch, which serves as a plug-in component for CL. This framework leverages the complementary strengths of two distinct and independent networks: one dedicated to plasticity and the other to stability. Each network is designed with a specialized and lightweight architecture, tailored to its respective objective. Extensive experiments demonstrate that Dual-Arch enhances the performance of existing CL methods while being up to 87% more compact in terms of parameters.
Aojun Lu, Hangjie Yuan, Tao Feng 0014, Yanan Sun 0001
ICML4
2025 Runtime Analysis of Evolutionary NAS for Multiclass Classification
abstract
Evolutionary neural architecture search (ENAS) is a key part of evolutionary machine learning, which commonly utilizes evolutionary algorithms (EAs) to automatically design high-performing deep neural architectures. During past years, various ENAS methods have been proposed with exceptional performance. However, the theory research of ENAS is still in the infant. In this work, we step for the runtime analysis, which is an essential theory aspect of EAs, of ENAS upon multiclass classification problems. Specifically, we first propose a benchmark to lay the groundwork for the analysis. Furthermore, we design a two-level search space, making it suitable for multiclass classification problems and consistent with the common settings of ENAS. Based on both designs, we consider (1+1)-ENAS algorithms with one-bit and bit-wise mutations, and analyze their upper and lower bounds on the expected runtime. We prove that the algorithm using both mutations can find the optimum with the expected runtime upper bound of $O(rM\ln{rM})$ and lower bound of $\Omega(rM\ln{M})$. This suggests that a simple one-bit mutation may be greatly considered, given that most state-of-the-art ENAS methods are laboriously designed with the bit-wise mutation. Empirical studies also support our theoretical proof.
Zeqiong Lv, Chao Qian 0001, Yanan Sun 0001
ICML5
2025 Prior Knowledge Guided Neural Architecture Generation
abstract
Automated architecture design methods, especially neural architecture search, have attracted increasing attention. However, these methods naturally need to evaluate numerous candidate architectures during the search process, thus computationally extensive and time-consuming. In this paper, we propose a prior knowledge guided neural architecture generation method to generate high-performance architectures without any search and evaluation process. Specifically, in order to identify valuable prior knowledge for architecture generation, we first quantify the contribution of each component within an architecture to its overall performance. Subsequently, a diffusion model guided by prior knowledge is presented, which can easily generate high-performance architectures for different computation tasks. Extensive experiments on new search spaces demonstrate that our method achieves superior accuracy over state-of-the-art methods. For example, we only need $0.004$ GPU Days to generate architecture with $76.1\%$ top-1 accuracy on ImageNet and $97.56\%$ on CIFAR-10. Furthermore, we can find competitive architecture for more unseen search spaces, such as TransNAS-Bench-101 and NATS-Bench, which demonstrates the broad applicability of the proposed method.
Han Ji 0003, Yanan Sun 0001
ICML3
2025 Neural Architecture Generation via Contrastive Representation Learning
abstract
Neural architecture search aims to automatically design architectures, attracting increasing attention recently. However, these methods need to individually evaluate numerous candidate architectures, thus causing a huge budget of time and computational resources. To address this issue, we propose a novel neural architecture generation method to generate optimal architectures without iterative evaluation and additional training. Specifically, in order to obtain a flexible and comprehensive representation of architectures, we propose a contrastive learning method to capture both static attributions and latent performance properties. Subsequently, a self-guided diffusion model is designed to learn these representations and generate optimal architectures for diverse tasks. Compared with existing methods, our method not only saves the computational time required for evaluating architectures but also gains better performance in different cases. For example, we only need 0.02 GPU Days to generate architecture with 77.1% top-1 accuracy on ImagNet and 97.65% on CIFAR-10. Furthermore, we can find competitive architecture from convolutional and transformer search spaces, i.e., DARTS, NAS-Bench-201, and AutoFormer, which demonstrates the broad applicability of the proposed method.
Minxiao Zhong, Yanan Sun 0001
IJCNN4
2025 A Lightweight Architecture for Predicting Neutronic Parameters via Neural Architecture Search
abstract
The online monitoring system is crucial for predicting safety-related neutronic parameters in nuclear power plants. Traditionally, the prediction of neutronic parameters relies on physics-based software, which is computationally inefficient. In recent years, neural network-based surrogate models are employed as an alternative quick predictor. Unfortunately, existing works are mostly based on widely reported handcrafted networks, which are architecturally complex and lack specialization for neutronic parameters. This results in high computational complexity and limited performance, making deployment in computational resources-constrained nuclear power plants impractical. To address these issues, we proposed a network architecture search algorithm to automatically design a convolutional neural network (CNN) for the prediction. Specifically, the pruning operation is utilized to refine architecture for less training time. Furthermore, we incorporate the Hybrid-attention Block into the search space to increase the accuracy of the prediction. We investigate the searched CNN regarding both scaler and matrix neutronic parameters. The experiments show that the automatically designed CNN outperforms other peer competitors. For example, in just 0.3 GPU days, we generated a CNN with 0.39M parameters, achieving a mean square error of 4e−4. This is less than half the parameters of Inception-ResNet and offers higher precision. This demonstrates that our proposed algorithm enables the surrogate model to achieve high prediction accuracy and short training time in limited computational resources.
Minxiao Zhong, Yanan Sun 0001
IJCNN4
2025 Vulnerable Data-Aware Adversarial Training
abstract
Fast adversarial training (FAT) has been considered as one of the most effective alternatives to the computationally-intensive adversarial training. Generally, FAT methods pay equal attention to each sample of the target task. However, the distance between each sample and the decision boundary is different, learning samples which are far from the decision boundary (i.e., less important to adversarial robustness) brings additional training cost and leads to sub-optimal results. To tackle this issue, we present vulnerable data-aware adversarial training (VDAT) in this study. Specifically, we first propose a margin-based vulnerability calculation method to measure the vulnerability of data samples. Moreover, we propose a vulnerability-aware data filtering method to reduce the training data for adversarial training thus improve the training efficiency. The experiments are conducted in terms of adversarial training and robust neural architecture search on CIFAR-10, CIFAR-100, and ImageNet-1K. The results demonstrate that VDAT is up to 76% more efficient than state-of-the-art FAT methods, while achieving improvements regarding the natural accuracy and adversarial accuracy in both scenarios. Furthermore, the visualizations and ablation studies show the effectiveness of both core components designed in VDAT.
Yanan Sun 0001
NeurIPS3
2025 Enhancing Replay-Based Continual Learning via Predictive Uncertainty Controller
Chunhui Ding, Aojun Lu, Yanan Sun 0001
PRICAI4
2025 Classification of sewer pipe defects based on an automatically designed convolutional neural network
Yu Wang 0322, Yanan Sun 0001
Expert Syst. Appl.3
2025 Automated design of neural networks with multi-scale convolutions via multi-path weight sampling
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001, Gary G. Yen
Pattern Recognit.3
2025 A Bidirectional Differential Evolution-Based Unknown Cyberattack Detection System
abstract
The evolving unknown cyberattacks, compounded by the widespread emerging technologies (say 5G, Internet of Things, etc.), have rapidly expanded the cyber threat landscape. However, most existing intrusion detection systems (IDSs) are effective in detecting only known cyberattacks, because only known cyberattack samples are usually available for IDS training. Identifying unknown cyberattacks, therefore, remains a big challenging issue. To meet this gap, in this paper, motivated by artificial immunity (AIm) and differential evolution (DE), we propose a bidirectional differential evolution based unknown cyberattack detection system, coined BDE-IDS. Specifically, we first design a bidirectional differential evolution algorithm for known nonself antigens (abnormal data), where bidirectional evolutionary directions are considered for increasing or decreasing the differences between known nonself antigens and self antigens (normal data), to create new antigens possibly used for generating cyberattack detectors. Second, a novel tolerance training mechanism is developed to eliminate invalid newly-evolved antigens falling into the coverage of either known self or nonself antigens. Third, the remaining antigens are employed to generate detectors for unknown cyberattacks. Extensive experiments demonstrate that the proposed BDE-IDS achieves outperformance in detecting unknown cyberattacks (as well as known cyberattacks) compared to state-of-the-art studies, including those AIm-based, signature-based, and anomaly-based IDSs.
Hanyuan Huang, Tao Li 0016, Beibei Li 0002, Wenhao Wang 0001, Yanan Sun 0001
IEEE Trans. Evol. Comput.5
2025 Evolutionary Trainer-Based Deep Q-Network for Dynamic Flexible Job-Shop Scheduling
abstract
Dynamic flexible job shop scheduling (DFJSS) aims to achieve the optimal efficiency for production planning in the face of dynamic events. In practice, deep Q-network (DQN) algorithms have been intensively studied for solving various DFJSS problems. However, these algorithms often cause moving targets for the given job-shop state. This will inevitably lead to unstable training and severe deterioration of the performance. In this paper, we propose a training algorithm based on genetic algorithm to efficiently and effectively address this critical issue. Specifically, a state feature extraction method is first developed, which can effectively represent different job shop scenarios. Furthermore, a genetic encoding strategy is designed, which can reduce the encoding length to enhance search ability. In addition, an evaluation strategy is proposed to calculate a fixed target for each job-shop state, which can avoid the parameter update of target networks. With the designs, the DQNs could be stably trained, thus their performance is greatly improved. Extensive experiments demonstrate that the proposed algorithm outperforms the state-of-the-art peer competitors in terms of both effectiveness and generalizability to multiple scheduling scenarios with different scales. In addition, the ablation study also reveals that the proposed algorithm can outperform the DQN algorithms with different updating frequencies of target networks.
Fangfang Zhang 0003, Yanan Sun 0001, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.3
2025 REP: An Interpretable Robustness Enhanced Plugin for Differentiable Neural Architecture Search
abstract
Neural architecture search (NAS) is widely used to automate the design of high-accuracy deep architectures, which are often vulnerable to adversarial attacks in practice due to the lack of adversarial robustness. Existing methods focus on the direct utilization of regularized optimization process to address this critical issue, which causes the lack of interpretability for the end users to learn how the robust architecture is constructed. In this paper, we introduce a robust enhanced plugin (REP) method for differentiable NAS to search for robust neural architectures. Different from existing peer methods, REP focuses on the robust search primitives in the search space of NAS methods, and naturally has the merit of contributing to understanding how the robust architectures are progressively constructed. Specifically, we first propose an effective sampling strategy to sample robust search primitives in the search space. In addition, we also propose a probabilistic enhancement method to guarantee natural accuracy and adversarial robustness simultaneously during the search process. We conduct experiments on both convolutional neural networks and graph neural networks with widely used benchmarks against state of the arts. The results reveal that REP can achieve superiority in terms of both the adversarial robustness to popular adversarial attacks and the natural accuracy of original data. REP is flexible and can be easily used by any existing differentiable NAS methods to enhance their robustness without much additional effort.
Yanan Sun 0001, Gary G. Yen, Kay Chen Tan
IEEE Trans. Knowl. Data Eng.2
2025 LRNAS: Differentiable Searching for Adversarially Robust Lightweight Neural Architecture
abstract
The adversarial robustness is critical to deep neural networks (DNNs) in deployment. However, the improvement of adversarial robustness often requires compromising with the network size. Existing approaches to addressing this problem mainly focus on the combination of model compression and adversarial training. However, their performance heavily relies on neural architectures, which are typically manual designs with extensive expertise. In this article, we propose a lightweight and robust neural architecture search (LRNAS) method to automatically search for adversarially robust lightweight neural architectures. Specifically, we propose a novel search strategy to quantify contributions of the components in the search space, based on which the beneficial components can be determined. In addition, we further propose an architecture selection method based on a greedy strategy, which can keep the model size while deriving sufficient beneficial components. Owing to these designs in LRNAS, the lightness, the natural accuracy, and the adversarial robustness can be collectively guaranteed to the searched architectures. We conduct extensive experiments on various benchmark datasets against the state of the arts. The experimental results demonstrate that the proposed LRNAS method is superior at finding lightweight neural architectures that are both accurate and adversarially robust under popular adversarial attacks. Moreover, ablation studies are also performed, which reveals the validity of the individual components designed in LRNAS and the component effects in positively deciding the overall performance.
Zeqiong Lv, Hongyang Chen 0001, Shangce Gao, Yanan Sun 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 E-3SFC: Communication-Efficient Federated Learning With Double-Way Features Synthesizing
abstract
The exponential growth in model sizes has significantly increased the communication burden in federated learning (FL). Existing methods to alleviate this burden by transmitting compressed gradients often face high compression errors, which slow down the model's convergence. To simultaneously achieve high compression effectiveness and lower compression errors, we study the gradient compression problem from a novel perspective. Specifically, we propose a systematical algorithm termed extended single-step synthetic features compressing (E-3SFC), which consists of three subcomponents, i.e., the single-step synthetic features compressor (3SFC), a double-way compression (DWC) algorithm, and a communication budget scheduler (BS). First, we regard the process of gradient computation of a model as decompressing gradients from corresponding inputs, while the inverse process is considered as compressing the gradients. Based on this, we introduce a novel gradient compression method termed 3SFC, which utilizes the model itself as a decompressor, leveraging training priors such as model weights and objective functions. The 3SFC compresses raw gradients into tiny synthetic features in a single-step simulation, incorporating error feedback (EF) to minimize overall compression errors. To further reduce communication overhead, 3SFC is extended to E-3SFC, allowing DWC and dynamic communication budget scheduling. Our theoretical analysis under both strongly convex and nonconvex conditions demonstrates that 3SFC achieves linear and sublinear convergence rates with aggregation noise. Extensive experiments across six datasets and six models reveal that 3SFC outperforms the state-of-the-art methods by up to 13.4% while reducing communication costs by 111.6 times. These findings suggest that 3SFC can significantly enhance communication efficiency in FL without compromising model performance.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Online Performance Predictor for Evolutionary Neural Architecture Search
abstract
NAS can automatically design well-performing architectures of deep neural networks (DNNs), which have been widely investigated for real-world applications. However, neural architecture search (NAS) is computationally expensive because vast DNNs are trained to obtain the performance during the search process. Performance predictors can directly obtain the performance of DNNs without training, thus having great potential to overcome this barrier. However, existing performance predictors are typically offline, and are trained with rare architectures labeled. As a result, their predictive ability is limited due to ignoring the promising DNN architectures generated in the NAS process. In this article, we propose a Gaussian process-based online performance predictor (GPOPP) tailored to evolutionary NAS. To achieve this, GPOPP selects the proper architectures to update itself during the search process upon EI as an acquisition function. Further, as architectures cannot be directly used as training data, GPOPP includes a binary encoding schema that can convert the architectures into a suitable format. Although the original intention of designing GPOPP is to provide an alternative way to build promising performance predictors that may compromise performance, the experiments show that GPOPP can outperform most state-of-the-arts. For example, GPOPP-assisted NAS gains 76.1% in accuracy on ImageNet with only 1.1 graphic processing unit (GPU) Days. In addition, the ablation studies also demonstrate the effectiveness of the components designed in GPOPP. The source code is available at https://github.com/songxt3/GPOPP.
Xiaotian Song, Han Ji 0003, Jiyuan Liu 0012, Shangce Gao, Yanan Sun 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2024 Towards Accurate and Robust Architectures via Neural Architecture Search
abstract
To defend deep neural networks from adversarial attacks, adversarial training has been drawing increasing attention for its effectiveness. However, the accuracy and robustness resulting from the adversarial training are limited by the architecture, because adversarial training improves accuracy and robustness by adjusting the weight connection affiliated to the architecture. In this work, we propose ARNAS to search for accurate and robust architectures for adversarial training. First we design an accurate and robust search space, in which the placement of the cells and the proportional relationship of the filter numbers are care-fully determined. With the design, the architectures can obtain both accuracy and robustness by deploying accurate and robust structures to their sensitive positions, re-spectively. Then we propose a differentiable multi-objective search strategy, performing gradient descent towards directions that are beneficial for both natural loss and adversar-ial loss, thus the accuracy and robustness can be guaran-teed at the same time. We conduct comprehensive experiments in terms of white-box attacks, black-box attacks, and transferability. Experimental results show that the searched architecture has the strongest robustness with the compet-itive accuracy, and breaks the traditional idea that NAS-based architectures cannot transfer well to complex tasks in robustness scenarios. By analyzing outstanding architectures searched, we also conclude that accurate and robust neural architectures tend to deploy different structures near the input and output, which has great practical significance on both hand-crafting and automatically designing of accurate and robust architectures.
Yuwei Ou, Yanan Sun 0001
CVPR3
2024 Generalizable Symbolic Optimizer Learning
Xiaotian Song, Yanan Sun 0001, Andy Song
ECCV (81)3
2024 Runtime Analysis of Population-based Evolutionary Neural Architecture Search for a Binary Classification Problem
abstract
Evolutionary neural architecture search (ENAS) employs evolutionary techniques, e.g., evolutionary algorithm (EA), to design high-performing neural network architectures, and has achieved great success. However, compared to the application, its theoretical analysis is still in its infancy and only touches the ENAS without populations. In this work, we consider the (μ+λ)-ENAS algorithm (based on a general population-based EA with mutation only, i.e., (μ+λ)-EA) to find an optimal neural network architecture capable of solving a binary classification problem Uniform (with problem size n), and obtain the following mathematical runtime results: 1) by applying a local mutation, it can find the optimum in an expected runtime of O(μ + nλ/(1 - e-λ/μ)) and Ω(μ + nλ/(1 - e−-λ)); 2) by applying a global mutation, it can find the optimum in an expected runtime of O(μ + λcλn/(1 - e−-λ/μ)), and Ω(μ + λn ln ln re/ln n) for some constant c > 1. The derived results reveal that the (μ+λ)-ENAS algorithm is always not asymptotically faster than (1+1)-ENAS on the Uniform problem when λ ϵ ω(ln n/(ln ln n)). The concrete theoretical analysis and proof show that increasing the population size has the potential to increase the runtime and thus should be carefully considered in the ENAS algorithm setup.
Zeqiong Lv, Chao Bian 0002, Chao Qian 0001, Yanan Sun 0001
GECCO4
2024 Physics-Informed Neural Networks with Generalized Residual-Based Adaptive Sampling
Xiaotian Song, Shuchao Deng, Yanan Sun 0001
ICIC (2)4
2024 Precisely Predicting Neutronics Parameters of Nuclear Reactor
Minxiao Zhong, Yanan Sun 0001
ICIC (2)4
2024 CAP: A Context-Aware Neural Predictor for NAS
Han Ji 0003, Yanan Sun 0001
IJCAI3
2024 Revisiting Neural Networks for Continual Learning: An Architectural Perspective
Aojun Lu, Tao Feng 0014, Hangjie Yuan, Xiaotian Song, Yanan Sun 0001
IJCAI5
2024 One-step Spiking Transformer with a Linear Complexity
Xiaotian Song, Andy Song, Yanan Sun 0001
IJCAI4
2024 Robust Neural Architecture Search under Long-Tailed Distribution
abstract
Robust neural architecture search (NAS) has emerged as a promising approach to automatically design robust neural architectures against adversarial attacks. However, existing robust NAS methods suffer from a performance degradation on the real-world data with long-tailed distribution, due to the limited number of samples from the tail classes. To push robust NAS methods towards more realistic scenarios, we present a novel robust NAS method called ALTNAS in this paper. Specifically, we propose a simple but effective data augmentation method to augment the data in tail classes and make the dataset more balanced. The proposed augmentation method can generate both natural data and adversarial examples in the search process. Moreover, we propose a neural architecture search method to search for architectures with the help of augmented data. Because the data in tail classes is enriched by both natural data and adversarial examples, the derived architecture can achieve promising performance in the adversarial long-tailed recognition task. We conduct extensive experiments on CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT benchmark datasets against the state-of-the-arts. The experimental results show that ALTNAS is superior at designing well-performing neural architectures for adversarial long-tailed recognition. In addition, the analysis and ablation studies are also performed, demonstrating the validity and effectiveness of the designed components in ALTNAS.
Yanan Sun 0001
IJCNN3
2024 HZS-NAS: Neural Architecture Search with Hybrid Zero-Shot Proxy for Facial Expression Recognition
abstract
Facial expression recognition (FER) involves analyzing and interpreting human sentiment from facial data, which has been an active topic in human-computer interaction. In recent years, convolutional neural networks (CNNs) have been intensively studied for FER tasks. Unfortunately, high-accuracy FER models typically contain numerous parameters, making them unsuitable for deployment on resource-constrained devices. Designing lightweight models that strike a balance between model complexity and accuracy is challenging. Neural architecture search (NAS) is a promising approach to automatically achieve the balance. However, previous NAS-based works required heavy training in supernet or intensive network evaluations, making the search process expensive. In this paper, we propose a NAS algorithm with a hybrid zero-shot (ZS) proxy for FER tasks, which can efficiently search for lightweight FER networks with high accuracy. To the best of our knowledge, the paper is the first to use zero-shot neural architecture search on FER tasks. Specifically, a novel performance evaluation strategy with the efficient hybrid ZS proxy is proposed, which can evaluate the accuracy of a single FER network in seconds. In addition, a hierarchical search space for FER tasks is designed to encourage the diversity of FER networks, which significantly affects recognition accuracy. Experimental results show that the proposed method consistently achieves excellent results in terms of accuracy and the number of parameters across public FER datasets (FER2013 and RAF-DB). Moreover, the proposed method takes much less time than other state-of-the-art NAS-based FER algorithms.
Jiyuan Liu 0012, Yanan Sun 0001
IJCNN4
2024 Multi-objective optimization and decision making for integrated energy system using STA and fuzzy TOPSIS
Xiaojun Zhou 0001, Wan Tan, Yanan Sun 0001, Tingwen Huang, Chunhua Yang 0001
Expert Syst. Appl.3
2024 A multiscale neural architecture search framework for multimodal fusion
Jindi Lv, Yanan Sun 0001, Wentao Feng, Jiancheng Lv 0001
Inf. Sci.2
2024 Architecture Augmentation for Performance Predictor via Graph Isomorphism
abstract
Neural architecture search (NAS) can automatically design architectures for deep neural networks (DNNs) and has become one of the hottest research topics in the current machine learning community. However, NAS is often computationally expensive because a large number of DNNs require to be trained for obtaining performance during the search process. Performance predictors can greatly alleviate the prohibitive cost of NAS by directly predicting the performance of DNNs. However, building satisfactory performance predictors highly depends on enough trained DNN architectures, which are difficult to obtain due to the high computational cost. To solve this critical issue, we propose an effective DNN architecture augmentation method named graph isomorphism-based architecture augmentation method (GIAug) in this article. Specifically, we first propose a mechanism based on graph isomorphism, which has the merit of efficiently generating a factorial of n (i.e., n ) diverse annotated architectures upon a single architecture having n nodes. In addition, we also design a generic method to encode the architectures into the form suitable to most prediction models. As a result, GIAug can be flexibly utilized by various existing performance predictors-based NAS algorithms. We perform extensive experiments on CIFAR-10 and ImageNet benchmark datasets on small-, medium- and large-scale search space. The experiments show that GIAug can significantly enhance the performance of the state-of-the-art peer predictors. In addition, GIAug can save three magnitude order of computation cost at most on ImageNet yet with similar performance when compared with state-of-the-art NAS algorithms.
Xiangning Xie, Yanan Sun 0001, Yuqiao Liu 0002, Mengjie Zhang 0001, Kay Chen Tan
IEEE Trans. Cybern.2
2024 Benchmarking Analysis of Evolutionary Neural Architecture Search
abstract
Evolutionary computation-based neural architecture search (ENAS) is a popular technique for automating the architecture design of deep neural networks. For any evolutionary computation-based algorithm, the runtime and convergence are the most important aspects concerned by theoretical analysis. However, because of the lacking of benchmarked fitness functions specialized for ENAS, the corresponding theoretical work is rarely available. To address this issue, we propose three different benchmark functions in this paper based on NAS-Bench-101. Specifically, we first propose a correlation-based feature extraction method, to capture the accuracy relationship between neural architectures and their fitness values. Furthermore, we propose a function toolkit, which allows combining different architecture features to specific benchmark functions. In addition, three benchmark functions are derived upon the toolkit by considering the features of neural net topologies, the features of neural operations, and their combinations. Based on these designs, the search space partition and transition probability calculation could be easily established, which in turn greatly promote the runtime and convergence analysis. We perform the experiments of ranking correlation, and the experimental results demonstrate the correctness of the proposed benchmark functions. To the best of our knowledge, this is the first work focusing on ENAS benchmark functions.
Zeqiong Lv, Chao Qian 0001, Yanan Sun 0001
IEEE Trans. Evol. Comput.3
2024 Guest Editorial Evolutionary Neural Architecture Search
abstract
Deep neural networks (DNNs) have shown remarkable performance in solving a wide variety of real-world problems, ranging from image recognition to natural language processing and self-driving vehicles. In principle, the achievements of DNNs are mainly contributed by their deep architectures, which can learn meaningful representations at different levels. This can greatly enhance the performance of the subsequent machine-learning algorithms. However, manually designing an optimal deep architecture for a particular problem requires a rich knowledge of both the investigated problem domain and the DNNs, which is not necessarily held by every end user interested in this area.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Evol. Comput.1
2024 GPU-Based Genetic Programming for Faster Feature Extraction in Binary Image Classification
abstract
Genetic programming (GP) has been applied to various binary image classification tasks and achieved promising results. However, existing approaches are difficult to be applied to large binary classification tasks due to the huge computational cost in fitness evaluations. To address this issue, we introduce a highly efficient method that enables fitness evaluations to be entirely conducted on graphics processing units (GPUs). Specifically, a prefix notation is used as program representation on the GPU device side, and column-major storage is used for the training dataset on the device side to achieve coalesced global memory access on the GPU. The evaluation of multiple GP programs in each generation can be simultaneous, which increases the parallelism of the algorithm. In addition, a parallel reduction is performed to maximize the use of the powerful parallel computing capability of GPU devices. Furthermore, the hoist mutation is also added to the proposed approach to help eliminate stack overflow on the device side. We compare training time and classification accuracy on various datasets with several GP and non-GP approaches. Experimental results indicate that the proposed approach significantly speeds up the existing GP-based binary image classification approaches without degradation in classification accuracy. We also analyze the influence of the batch size on the training time and investigate the classification accuracy in different settings of the max program depth and the number of generations. The code is available at https://github.com/RayZhhh/CupaGP for reference.
Yanan Sun 0001, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.2
2024 Neural Networks Learn Specified Information for Imbalanced Data Classification
abstract
Imbalanced data problem is a classic topic in artificial intelligence. Neural network approaches to solve this problem mostly rely on resampling or reweighting strategies. However, these methods severely suffer from the learning bias in most cases when the empirical representation of known samples is insufficient. One-class learning can provide an ideal classification property to alleviate this critical issue. However, extending one-class learning to imbalanced data presents problems of hypersphere collapse, ambiguous interclass relations, and compact representations. In this paper, a new one-class learning paradigm is proposed for binary imbalanced data classification. Specifically, a neural network is employed to map known samples to a specified attribute space to solve the problems of hypersphere collapse and ambiguous interclass relations. Then, to alleviate the compact representation problem, a dynamic information potential energy is developed to disperse the mapped majority samples to fill the specified region as much as possible. The proposed method is validated on 34 imbalanced datasets with imbalanced ratios ranging from 16.90 to 100.14. The test results show that the proposed method achieves the best performance on more than half of the test datasets.
Zhan ao Huang, Yongsheng Sang, Yanan Sun 0001, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.3
2024 Neural Network With a Preference Sampling Paradigm for Imbalanced Data Classification
abstract
Most data in real life are characterized by imbalance problems. One of the classic models for dealing with imbalanced data is neural networks. However, the data imbalance problem often causes the neural network to display negative class preference behavior. Using an undersampling strategy to reconstruct a balanced dataset is one of the methods to alleviate the data imbalance problem. However, most existing undersampling methods focus more on the data or aim to preserve the overall structural characteristics of the negative class through potential energy estimation, while the problems of gradient inundation and insufficient empirical representation of positive samples have not been well considered. Therefore, a new paradigm for solving the data imbalance problem is proposed. Specifically, to solve the problem of gradient inundation, an informative undersampling strategy is derived from the performance degradation and used to restore the ability of neural networks to work under imbalanced data. In addition, to alleviate the problem of insufficient empirical representation of positive samples, a boundary expansion strategy with linear interpolation and the prediction consistency constraint is considered. We tested the proposed paradigm on 34 imbalanced datasets with imbalance ratios ranging from 16.90 to 100.14. The test results show that our paradigm obtained the best area under the receiver operating characteristic curve (AUC) on 26 datasets.
Zhan ao Huang, Yongsheng Sang, Yanan Sun 0001, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Split-Level Evolutionary Neural Architecture Search With Elite Weight Inheritance
abstract
Neural architecture search (NAS) has recently gained extensive interest in the deep learning community because of its great potential in automating the construction process of deep models. Among a variety of NAS approaches, evolutionary computation (EC) plays a pivotal role with its merit of gradient-free search ability. However, a massive number of the current EC-based NAS approaches evolve neural architectures in an absolutely discrete manner, which makes it tough to flexibly handle the number of filters for each layer, since they often reduce it to a limit set rather than searching for all possible values. Moreover, EC-based NAS methods are often criticized for their inefficiency in performance evaluation, which usually requires laborious full training for hundreds of candidate architectures generated. To address the inflexible search issue on the number of filters, this work proposes a split-level particle swarm optimization (PSO) approach. Each dimension of the particle is subdivided into an integer part and a fractional part, encoding the configurations of the corresponding layer, and the number of filters within a large range, respectively. In addition, the evaluation time is greatly saved by a novel elite weight inheritance method based on an online updating weight pool, and a customized fitness function considering multiple objectives is developed to well control the complexity of the searched candidate architectures. The proposed method, termed split-level evolutionary NAS (SLE-NAS), is computationally efficient, and outperforms many state-of-the-art peer competitors at much lower complexity across three popular image classification benchmark datasets.
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Neural Networks Learn. Syst.3
2023 Communication-efficient Federated Learning with Single-Step Synthetic Features Compressor for Faster Convergence
abstract
Reducing communication overhead in federated learning (FL) is challenging but crucial for large-scale distributed privacy-preserving machine learning. While methods utilizing sparsification or other techniques can largely reduce the communication overhead, the convergence rate is also greatly compromised. In this paper, we propose a novel method named Single-Step Synthetic Features Compressor (3SFC) to achieve communication-efficient FL by directly constructing a tiny synthetic dataset containing synthetic features based on raw gradients. Therefore, 3SFC can achieve an extremely low compression rate when the constructed synthetic dataset contains only one data sample. Additionally, the compressing phase of 3SFC utilizes a similarity-based objective function so that it can be optimized with just one step, considerably improving its performance and robustness. To minimize the compressing error, error feedback (EF) is also incorporated into 3SFC. Experiments on multiple datasets and models suggest that 3SFC has significantly better convergence rates compared to competing methods with lower compression rates (i.e., up to 0.02%). Furthermore, ablation studies and visualizations show that 3SFC can carry more information than competing methods for every communication round, further validating its effectiveness.
Yuhao Zhou 0004, Mingjia Shi, Yanan Sun 0001, Jiancheng Lv 0001
ICCV4
2023 Neural network with absent minority class samples and boundary shifting for imbalanced data classification
Zhan ao Huang, Yongsheng Sang, Yanan Sun 0001, Jiancheng Lv 0001
Neural Comput. Appl.3
2023 Particle Swarm Optimization for Compact Neural Architecture Search for Image Classification
abstract
Convolutional neural networks (CNNs) are a superb computing paradigm in deep learning, and their architectures are considered to be the key to performance breakthroughs in various tasks. Recently, neural architecture search (NAS) methods have been proposed to automate the process of network architecture design, many of which have discovered novel CNN architectures that are superior to human-designed ones. However, most of the current NAS methods suffer from either prohibitively high computational complexity of resulting architectures which have inadvertently affected the deep model deployment, or limitations which impede the flexibility of architecture design. To address these deficiencies, this work proposes an evolutionary computation (EC)-based method for compact and flexible NAS. A valuable search space with parameter-efficient mobile-inverted bottleneck convolution blocks as the primitive components is proposed to ensure the initial quality of the compact architectures. In addition, a two-level variable-length particle swarm optimization (PSO) approach is devised to evolve both the microarchitecture and macroarchitecture of CNNs. Furthermore, this study proposes an effective scheme by integrating multiple computationally reduced methods to greatly speed up the evaluation process. Experimental results on the CIFAR-10, CIFAR-100, and ImageNet datasets show the superiority of the proposed method against the state-of-the-art algorithms in terms of the classification performance, search cost, and resulting architecture complexity.
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Evol. Comput.3
2023 Automatic Design of Convolutional Neural Network Architectures Under Resource Constraints
abstract
With the rise of various smart electronics and mobile/edge devices, many existing high-accuracy convolutional neural network (CNN) models are difficult to be applied in practice due to the limited resources, such as memory capacity, power consumption, and spectral efficiency. In order to meet these constraints, researchers have carefully designed some lightweight networks. Meanwhile, to reduce the reliance on manual design on expert experience, some researchers also work to improve neural architecture search (NAS) algorithms to automatically design small networks, exploiting the multiobjective approaches that consider both accuracy and other important goals during optimization. However, simply searching for smaller network models is not consistent with the current research belief of "the deeper the better" and may affect the effectiveness of the model and thus waste the limited resources available. In this article, we propose an automatic method for designing CNNs architectures under constraint handling, which can search for optimal network models meeting the preset constraint. Specifically, an adaptive penalty algorithm is used for fitness evaluation, and a selective repair operation is developed for infeasible individuals to search for feasible CNN architectures. As a case study, we set the complexity (the number of parameters) as a resource constraint and perform multiple experiments on CIFAR-10 and CIFAR-100, to demonstrate the effectiveness of the proposed method. In addition, the proposed algorithm is compared with a state-of-the-art algorithm, NSGA-Net, and several manual-designed models. The experimental results show that the proposed algorithm can successfully solve the problem of the uncertain size of the optimal CNN model under the random search strategy, and the automatically designed CNN model can satisfy the predefined resource constraint while achieving better accuracy.
Yanan Sun 0001, Gary G. Yen, Mengjie Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 A Survey on Evolutionary Neural Architecture Search
abstract
Deep neural networks (DNNs) have achieved great success in many applications. The architectures of DNNs play a crucial role in their performance, which is usually manually designed with rich expertise. However, such a design process is labor-intensive because of the trial-and-error process and also not easy to realize due to the rare expertise in practice. Neural architecture search (NAS) is a type of technology that can design the architectures automatically. Among different methods to realize NAS, the evolutionary computation (EC) methods have recently gained much attention and success. Unfortunately, there has not yet been a comprehensive summary of the EC-based NAS algorithms. This article reviews over 200 articles of most recent EC-based NAS methods in light of the core components, to systematically discuss their design principles and justifications on the design. Furthermore, current challenges and issues are also discussed to identify future research in this emerging field.
Yuqiao Liu 0002, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.2
2022 Adaptively Joint Pixel-wise Semantic Correlation in Surface Normal Estimation
abstract
Surface normal estimation is a fundamental but challenging pixel-wise task in the field of computer vision. With the advance of convolutional neural networks, the data-driven methods have demonstrated its superiority in end-to-end predicting surface normal from a single image, rather than manually extracting geometric clues. Yet, most existing methods usually result in poorly predicted details such as the smoothness and the shape, due to the lack of pixel-wise constraints. In view of this, we propose a coarse-to-fine surface normal estimation approach, called Seman-sn, to estimate the surface normal from a single RGB image. The proposed method can adaptively joint the semantic information by neural architecture search, and then further improve the surface normal estimation in details by exploring the pixel-wise semantic constraints. Specifically, the initial predictions are first made by automatically selecting the extraction layers. Then, the pixel-wise constraint between surface normal estimation and semantic segmentation is modeled to further refine the surface normal features in the final prediction stage. The experiments have been carried on a widely used data set, and the results show the great superiority compared with the recent methods. In addition, we have proved the effectiveness of our proposed method through reasonable ablation experiments, including the quantitative results and the visualization performance.
Yanan Sun 0001
IJCNN2
2022 Exploring the Effectiveness of Appearance Descriptor in DeepSORT
abstract
Tracking-by-detection approaches have demonstrated their strength in addressing Multiple Object Tracking (MOT) problems. DeepSORT, one of the classical tracking-by-detection MOT methods, relies on a deep appearance descriptor to extract global appearance features of identities. Although the appearance descriptor acts as a key component of such tracking-by-detection methods, which is responsible for modeling appearance information, the relationship between it and tracking performance remains unclear, especially whether further improvements to it will be reflected in the tracking performance. To explore that, extensive experiments are conducted on the appearance descriptor by applying various traditional optimization methods. Furthermore, we propose an Evolutionary Neural Architecture Search (ENAS) strategy for the appearance descriptor named Genetic-SORT to assist exploration. The experimental results demonstrate that tracking performance fails to follow the improvements applied on the appearance descriptor and even shows a negative correlation, which is contrary to our intuition.
Zilin Xiao, Yanan Sun 0001
IJCNN2
2022 Bridge the Gap Between Architecture Spaces via A Cross-Domain Predictor
abstract
Neural Architecture Search (NAS) can automatically design promising neural architectures without artificial experience. Though it achieves great success, prohibitively high search cost is required to find a high-performance architecture, which blocks its practical implementation. Neural predictor can directly evaluate the performance of neural networks based on their architectures and thereby save much budget. However, existing neural predictors require substantial annotated architectures trained from scratch, which still consume many computational resources. To solve this issue, we propose a Cross-Domain Predictor (CDP), which is trained based on the existing NAS benchmark datasets (e.g., NAS-Bench-101), but can be used to find high-performance architectures in large-scale search spaces. Particularly, we propose a progressive subspace adaptation strategy to address the domain discrepancy between the source architecture space and the target space. Considering the large difference between two architecture spaces, an assistant space is developed to smooth the transfer process. Compared with existing NAS methods, the proposed CDP is much more efficient. For example, CDP only requires the search cost of 0.1 GPU Days to find architectures with 76.9% top-1 accuracy on ImageNet and 97.51% on CIFAR-10.
Yuqiao Liu 0002, Yehui Tang 0001, Zeqiong Lv, Yunhe Wang 0001, Yanan Sun 0001
NeurIPS5
2022 Speeding up Genetic Programming Based Symbolic Regression Using GPUs
Andrew Lensen, Yanan Sun 0001
PRICAI (1)3
2022 A neural network learning algorithm for highly imbalanced data classification
Zhan ao Huang, Yongsheng Sang, Yanan Sun 0001, Jiancheng Lv 0001
Inf. Sci.3
2022 BenchENAS: A Benchmarking Platform for Evolutionary Neural Architecture Search
abstract
Neural architecture search (NAS), which automatically designs the architectures of deep neural networks, has achieved breakthrough success over many applications in the past few years. Among different classes of NAS methods, evolutionary computation-based NAS (ENAS) methods have recently gained much attention. Unfortunately, the development of ENAS is hindered by unfair comparison between different ENAS algorithms due to different training conditions and high computational cost caused by expensive performance evaluation. This article develops a platform named BenchENAS, in short for benchmarking evolutionary NAS, to address these issues. BenchENAS makes it easy to achieve fair comparisons between different algorithms by keeping them under the same settings. To accelerate the performance evaluation in a common lab environment, BenchENAS designs a novel and generic efficient evaluation method for the population characteristics of evolutionary computation. This method has greatly improved the efficiency of the evaluation. Furthermore, BenchENAS is easy to install and highly configurable and modular, which brings benefits in good usability and easy extensibility. This article conducts efficient comparison experiments on eight ENAS algorithms with high GPU utilization on this platform. The experiments validate that the fair comparison issue does exist in the current ENAS algorithms, and BenchENAS can alleviate this issue. A Website has been built to promote BenchENAS athttps://benchenas.com, where interested researchers can obtain the source code and document of BenchENAS for free.
Xiangning Xie, Yuqiao Liu 0002, Yanan Sun 0001, Gary G. Yen, Bing Xue 0001, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.3
2021 A Flexible Variable-length Particle Swarm Optimization Approach to Convolutional Neural Network Architecture Design
abstract
The great success of convolutional neural networks (CNNs) in image classification benefits from the powerful feature learning abilities of their architectures. However, arbitrarily constructing these architectures is very costly in terms of manpower and computation resources. Recently, particle swarm optimization (PSO), as a promising evolutionary computation (EC) method, has been used to automatically search for the CNN architectures and achieved encouraging results on image classification, but these existing methods are not well-designed and/or do not have good search capabilities. In this paper, a flexible variable-length PSO algorithm is proposed for automated CNN architecture design for image classification tasks. Particularly, an improved encoding scheme is proposed to truly break the fixed-length representation constraint of PSO and encode the parameters in a more meaningful way. In addition, a novel velocity and position updating approach is developed to update variable-length particles. Experiments on four benchmark datasets are carried out to confirm the superiority of the proposed algorithm. It is shown that the proposed method leads to better results comparing to the state-of-the-art algorithms in terms of classification performance, parameter size and computational complexity.
Junhao Huang 0002, Bing Xue 0001, Yanan Sun 0001, Mengjie Zhang 0001
CEC3
2021 A Survey of Advances in Evolutionary Neural Architecture Search
abstract
Deep neural networks (DNNs) have been frequently and widely applied for intelligent systems such as object detection, natural language understanding and speech recognition. Given a specific problem, we always aim to construct the most suitable DNN to solve it, which requires choosing the most appropriate model architecture and seeking the best model parameters values. However, most existing works focus on model parameters learning under the assumption that the model architecture can be manually specified as per prior knowledge and/or trial-and-error experimentation. To overcome this problem, evolutionary algorithms (EAs) have been widely used to design model architectures automatically. Further, EAs have been used for neural network optimization for more than 30 years. Therefore, in this paper, we review the evolutionary neural architecture search (ENAS) from the view of the advanced techniques. We hope this work can provide a comprehensive understanding of EAs' roles for the readers and focus themselves on ENAS.
A. K. Qin 0001, Yanan Sun 0001, Kay Chen Tan
CEC3
2021 Evolving Deep Parallel Neural Networks for Multi-Task Learning
Yanan Sun 0001
ICA3PP (2)2
2021 Homogeneous Architecture Augmentation for Neural Predictor
abstract
Neural Architecture Search (NAS) can automatically design well-performed architectures of Deep Neural Networks (DNNs) for the tasks at hand. However, one bottleneck of NAS is the prohibitively computational cost largely due to the expensive performance evaluation. The neural predictors can directly estimate the performance without any training of the DNNs to be evaluated, thus have drawn increasing attention from researchers. Despite their popularity, they also suffer a severe limitation: the shortage of annotated DNN architectures for effectively training the neural predictors. In this paper, we proposed Homogeneous Architecture Augmentation for Neural Predictor (HAAP) of DNN architectures to address the issue aforementioned. Specifically, a homogeneous architecture augmentation algorithm is proposed in HAAP to generate sufficient training data taking the use of homogeneous representation. Furthermore, the one-hot encoding strategy is introduced into HAAP to make the representation of DNN architectures more effective. The experiments have been conducted on both NAS-Benchmark-101 and NAS-Bench-201 dataset. The experimental results demonstrate that the proposed HAAP algorithm outperforms the state of the arts compared, yet with much less training data. In addition, the ablation studies on both benchmark datasets have also shown the universality of the homogeneous architecture augmentation. Our code has been made available at https://github.com/lyq998/HAAP.
Yuqiao Liu 0002, Yehui Tang 0001, Yanan Sun 0001
ICCV3
2021 Evolving Deep Neural Networks for Collaborative Filtering
Yuhan Fang, Yuqiao Liu 0002, Yanan Sun 0001
ICONIP (6)3
2021 Heart-Darts: Classification of Heartbeats Using Differentiable Architecture Search
abstract
Arrhythmia is a cardiovascular disease that manifests irregular heartbeats. In arrhythmia detection, the electrocardiogram (ECG) signal is an important diagnostic technique. However, manually evaluating ECG signals is a complicated and time-consuming task. With the application of convolutional neural networks (CNNs), the evaluation process has been accelerated and the performance is improved. It is noteworthy that the performance of CNNs heavily depends on their architecture design, which is a complex process grounded on expert experience and trial-and-error. In this paper, we propose a novel approach, Heart-Darts, to efficiently classify the ECG signals by automatically designing the CNN model with the differentiable architecture search (i.e., Darts, a cell-based neural architecture search method). Specifically, we initially search a cell architecture by Darts and then customize a novel CNN model for ECG classification based on the obtained cells. To investigate the efficiency of the proposed method, we evaluate the constructed model on the MIT-BIH arrhythmia database. Additionally, the extensibility of the proposed CNN model is validated on two other new databases. Extensive experimental results demonstrate that the proposed method outperforms several state-of-the-art CNN models in ECG classification in terms of both performance and generalization capability.
Jindi Lv, Yanan Sun 0001, Jiancheng Lv 0001
IJCNN3
2021 Evolving Deep Convolutional Variational Autoencoders for Image Classification
abstract
Variational autoencoders (VAEs) have demonstrated their superiority in unsupervised learning for image processing in recent years. The performance of the VAEs highly depends on their architectures, which are often handcrafted by the human expertise in deep neural networks (DNNs). However, such expertise is not necessarily available to each of the end users interested. In this article, we propose a novel method to automatically design optimal architectures of VAEs for image classification, called evolving deep convolutional VAE (EvoVAE), based on a genetic algorithm (GA). In the proposed EvoVAE algorithm, the traditional VAEs are first generalized to a more generic and asymmetrical one with four different blocks, and then a variable-length gene encoding mechanism of the GA is presented to search for the optimal network depth. Furthermore, an effective genetic operator is designed to adapt to the proposed variable-length gene encoding strategy. To verify the performance of the proposed algorithm, nine variants of AEs and VAEs are chosen as the peer competitors to perform the comparisons on MNIST, street view house numbers, and CIFAR-10 benchmark datasets. The experiments reveal the superiority of the proposed EvoVAE algorithm, which wins 21 times out of the 24 comparisons and outperforms the best competitors by 1.39%, 14.21%, and 13.03% on the three benchmark datasets, respectively.
Yanan Sun 0001, Mengjie Zhang 0001, Dezhong Peng
IEEE Trans. Evol. Comput.2
2021 A Novel Training Protocol for Performance Predictors of Evolutionary Neural Architecture Search Algorithms
abstract
Evolutionary neural architecture search (ENAS) can automatically design the architectures of deep neural networks (DNNs) using evolutionary computation algorithms. However, most ENAS algorithms require an intensive computational resource, which is not necessarily available to the users interested. Performance predictors are a type of regression models which can assist to accomplish the search, while without exerting much computational resource. Despite various performance predictors have been designed, they employ the same training protocol to build the regression models: 1) sampling a set of DNNs with performance as the training dataset; 2) training the model with the mean square error criterion; and 3) predicting the performance of DNNs newly generated during the ENAS. In this article, we point out that the three steps constituting the training protocol are not well thought-out through intuitive and illustrative examples. Furthermore, we propose a new training protocol to address these issues, consisting of designing a pairwise ranking indicator to construct the training target, proposing to use the logistic regression to fit the training samples, and developing a differential method to build the training instances. To verify the effectiveness of the proposed training protocol, four widely used regression models in the field of machine learning have been chosen to perform the comparisons on two benchmark datasets. The experimental results of all the comparisons demonstrate that the proposed training protocol can significantly improve the performance prediction accuracy against the traditional training protocols.
Yanan Sun 0001, Yuhan Fang, Gary G. Yen, Yuqiao Liu 0002
IEEE Trans. Evol. Comput.1
2021 A Distributed Framework for EA-Based NAS
abstract
Evolutionary Algorithms (EA) are widely applied in Neural Architecture Search (NAS) and have achieved appealing results. Different EA-based NAS algorithms may utilize different encoding schemes for network representation, while they have the same workflow. Specifically, the first step is the initialization of the population with different encoding schemes, and the second step is the evaluation of the individuals by the fitness function. Then, the EA-based NAS algorithm executes evolution operations, e.g., selection, mutation, and crossover, to eliminate weak individuals and generate more competitive ones. Lastly, evolution continues until the max generation and the best neural architectures will be chosen. Because each individual needs complete training and validation on the target dataset, the EA-based NAS always consumes significant computation and time inevitably, which results in the bottleneck of this approach. To ameliorate this issue, this article proposes a distributed framework to boost the computation of the EA-based NAS algorithm. This framework is a server/worker model where the server distributes individuals requested by the computing nodes and collects the validated individuals and hosts the evolution operations. Meanwhile, the most time-consuming phase (i.e., individual evaluation) of the EA-based NAS is allocated to the computing nodes, which send requests asynchronously to the server and evaluate the fitness values of the individuals. Additionally, a new packet structure of the message delivered in the cluster is designed to encapsulate various network representations and support different EA-based NAS algorithms. We design an EA-based NAS algorithm as a case to investigate the efficiency of the proposed framework. Extensive experiments are performed on an illustrative cluster with different scales, and the results reveal that the framework can achieve a nearly linear reduction of the search time with the increase of the computational nodes. Furthermore, the length of the exchanged messages among the cluster is tiny, which benefits the framework expansion.
Yanan Sun 0001, Jixin Zhang, Jiancheng Lv 0001
IEEE Trans. Parallel Distributed Syst.2
2020 Evolving Deep Convolutional Neural Networks for Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) are susceptible to various noise factors leading to the loss of information, and the noise restricts the subsequent HSIs object detection and classification tasks. In recent years, learning-based methods have demonstrated their superior strengths in denoising the HSIs. Unfortunately, most of the methods are manually designed based on the extensive expertise that is not necessarily available to the users interested. In this paper, we propose a novel algorithm to automatically build an optimal Convolutional Neural Network (CNN) to effectively denoise HSIs. Particularly, the proposed algorithm focuses on the architectures and the initialization of the connection weights of the CNN. The experiments of the proposed algorithm have been well-designed and compared against the state-of-the-art peer competitors, and the experimental results demonstrate the competitive performance of the proposed algorithm in terms of the different evaluation metrics, visual assessments, and the computational complexity.
Yuqiao Liu 0002, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001
IJCNN2
2020 PSO-PS: Parameter Synchronization with Particle Swarm Optimization for Distributed Training of Deep Neural Networks
abstract
Parameter updating is an important stage in parallelism-based distributed deep learning. Synchronous methods are widely used in distributed training the Deep Neural Networks (DNNs). To reduce the communication and synchronization overhead of synchronous methods, decreasing the synchronization frequency (e.g., every n mini-batches) is a straightforward approach. However, it often suffers from poor convergence. In this paper, we propose a new algorithm of integrating Particle Swarm Optimization (PSO) into the distributed training process of DNNs to automatically compute new parameters. In the proposed algorithm, a computing work is encoded by a particle, the weights of DNNs and the training loss are modeled by the particle attributes. At each synchronization stage, the weights are updated by PSO from the sub weights gathered from all workers, instead of averaging the weights or the gradients. To verify the performance of the proposed algorithm, the experiments are performed on two commonly used image classification benchmarks: MNIST and CIFAR10, and compared with the peer competitors at multiple different synchronization configurations. The experimental results demonstrate the competitiveness of the proposed algorithm.
Yanan Sun 0001, Jiancheng Lv 0001
IJCNN3
2020 Automatically Designing CNN Architectures Using the Genetic Algorithm for Image Classification
abstract
Convolutional neural networks (CNNs) have gained remarkable success on many image classification tasks in recent years. However, the performance of CNNs highly relies upon their architectures. For the most state-of-the-art CNNs, their architectures are often manually designed with expertise in both CNNs and the investigated problems. Therefore, it is difficult for users, who have no extended expertise in CNNs, to design optimal CNN architectures for their own image classification problems of interest. In this article, we propose an automatic CNN architecture design method by using genetic algorithms, to effectively address the image classification tasks. The most merit of the proposed algorithm remains in its "automatic" characteristic that users do not need domain knowledge of CNNs when using the proposed algorithm, while they can still obtain a promising CNN architecture for the given images. The proposed algorithm is validated on widely used benchmark image classification datasets, compared to the state-of-the-art peer competitors covering eight manually designed CNNs, seven automatic + manually tuning, and five automatic CNN architecture design algorithms. The experimental results indicate the proposed algorithm outperforms the existing automatic CNN architecture design algorithms in terms of classification accuracy, parameter numbers, and consumed computational resources. The proposed algorithm also shows the very comparable classification accuracy to the best one from manually designed and automatic + manually tuning CNNs, while consuming fewer computational resources.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen, Jiancheng Lv 0001
IEEE Trans. Cybern.1
2020 Surrogate-Assisted Evolutionary Deep Learning Using an End-to-End Random Forest-Based Performance Predictor
abstract
Convolutional neural networks (CNNs) have shown remarkable performance in various real-world applications. Unfortunately, the promising performance of CNNs can be achieved only when their architectures are optimally constructed. The architectures of state-of-the-art CNNs are typically handcrafted with extensive expertise in both CNNs and the investigated data, which consequently hampers the widespread adoption of CNNs for less experienced users. Evolutionary deep learning (EDL) is able to automatically design the best CNN architectures without much expertise. However, the existing EDL algorithms generally evaluate the fitness of a new architecture by training from scratch, resulting in the prohibitive computational cost even operated on high-performance computers. In this paper, an end-to-end offline performance predictor based on the random forest is proposed to accelerate the fitness evaluation in EDL. The proposed performance predictor shows the promising performance in term of the classification accuracy and the consumed computational resources when compared with 18 state-of-the-art peer competitors by integrating into an existing EDL algorithm as a case study. The proposed performance predictor is also compared with the other two representatives of existing performance predictors. The experimental results show the proposed performance predictor not only significantly speeds up the fitness evaluations but also achieves the best prediction among the peer performance predictors.
Yanan Sun 0001, Handing Wang, Bing Xue 0001, Yaochu Jin, Gary G. Yen, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2020 Evolving Deep Convolutional Neural Networks for Image Classification
abstract
Evolutionary paradigms have been successfully applied to neural network designs for two decades. Unfortunately, these methods cannot scale well to the modern deep neural networks due to the complicated architectures and large quantities of connection weights. In this paper, we propose a new method using genetic algorithms for evolving the architectures and connection weight initialization values of a deep convolutional neural network to address image classification problems. In the proposed algorithm, an efficient variable-length gene encoding strategy is designed to represent the different building blocks and the potentially optimal depth in convolutional neural networks. In addition, a new representation scheme is developed for effectively initializing connection weights of deep convolutional neural networks, which is expected to avoid networks getting stuck into local minimum that is typically a major issue in the backward gradient-based optimization. Furthermore, a novel fitness evaluation method is proposed to speed up the heuristic search with substantially less computational resource. The proposed algorithm is examined and compared with 22 existing algorithms on nine widely used image classification tasks, including the state-of-the-art methods. The experimental results demonstrate the remarkable superiority of the proposed algorithm over the state-of-the-art designs in terms of classification error rate and the number of parameters (weights).
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Evol. Comput.1
2020 Completely Automated CNN Architecture Design Based on Blocks
abstract
The performance of convolutional neural networks (CNNs) highly relies on their architectures. In order to design a CNN with promising performance, extensive expertise in both CNNs and the investigated problem domain is required, which is not necessarily available to every interested user. To address this problem, we propose to automatically evolve CNN architectures by using a genetic algorithm (GA) based on ResNet and DenseNet blocks. The proposed algorithm is completely automatic in designing CNN architectures. In particular, neither preprocessing before it starts nor postprocessing in terms of CNNs is needed. Furthermore, the proposed algorithm does not require users with domain knowledge on CNNs, the investigated problem, or even GAs. The proposed algorithm is evaluated on the CIFAR10 and CIFAR100 benchmark data sets against 18 state-of-the-art peer competitors. Experimental results show that the proposed algorithm outperforms the state-of-the-art CNNs hand-crafted and the CNNs designed by automatic peer competitors in terms of the classification performance and achieves a competitive classification accuracy against semiautomatic peer competitors. In addition, the proposed algorithm consumes much less computational resource than most peer competitors in finding the best CNN architectures.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Neural Networks Learn. Syst.1
2019 A Graph-Based Encoding for Evolutionary Convolutional Neural Network Architecture Design
abstract
Convolutional neural networks (CNNs) have demonstrated highly effective performance in image classification across a range of data sets. The best performance can only be obtained with CNNs when the appropriate architecture is chosen, which depends on both the volume and nature of the training data available. Many of the state-of-the-art architectures in the literature have been hand-crafted by human researchers, but this requires expertise in CNNs, domain knowledge, or trial-and-error experimentation, often using expensive resources. Recent work based on evolutionary deep learning has offered an alternative, in which evolutionary computation (EC) is applied to automatic architecture search. A key component in evolutionary deep learning is the chosen encoding strategy; however, previous approaches to CNN encoding in EC typically have restrictions in the architectures that can be represented. Here, we propose an encoding strategy based on a directed acyclic graph representation, and introduce an algorithm for random generation of CNN architectures using this encoding. In contrast to previous work, our proposed encoding method is more general, enabling representation of CNNs of arbitrary connectional structure and unbounded depth. We demonstrate its effectiveness using a random search, in which 200 randomly generated CNN architectures are evaluated. To improve the computational efficiency, the 200 CNNs are trained using only 10% of the CIFAR-10 training data; the three bestperforming CNNs are then re-trained on the full training set. The results show that the proposed representation and initialisation method can achieve promising accuracy compared to manually designed architectures, despite the simplicity of the random search approach and the reduced data set. We intend that future work can improve on these results by applying evolutionary search using this encoding.
William Irwin-Harris, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001
CEC2
2019 Evolving deep neural networks by multi-objective particle swarm optimization for image classification
abstract
In recent years, convolutional neural networks (CNNs) have become deeper in order to achieve better classification accuracy in image classification. However, it is difficult to deploy the state-of-the-art deep CNNs for industrial use due to the difficulty of manually fine-tuning the hyperparameters and the trade-off between classification accuracy and computational cost. This paper proposes a novel multi-objective optimization method for evolving state-of-the-art deep CNNs in real-life applications, which automatically evolves the non-dominant solutions at the Pareto front. Three major contributions are made: Firstly, a new encoding strategy is designed to encode one of the best state-of-the-art CNNs; With the classification accuracy and the number of floating point operations as the two objectives, a multi-objective particle swarm optimization method is developed to evolve the non-dominant solutions; Last but not least, a new infrastructure is designed to boost the experiments by concurrently running the experiments on multiple GPUs across multiple machines, and a Python library is developed and released to manage the infrastructure. The experimental results demonstrate that the non-dominant solutions found by the proposed method form a clear Pareto front, and the proposed infrastructure is able to almost linearly reduce the running time.
Bin Wang 0044, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001
GECCO2
2019 A Hybrid GA-PSO Method for Evolving Architecture and Short Connections of Deep Convolutional Neural Networks
Bin Wang 0044, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001
PRICAI (3)2
2019 A New Two-Stage Evolutionary Algorithm for Many-Objective Optimization
abstract
Convergence and diversity are interdependently handled during the evolutionary process by most existing many-objective evolutionary algorithms (MaOEAs). In such a design, the degraded performance of one would deteriorate the other, and only solutions with both are able to improve the performance of MaOEAs. Unfortunately, it is not easy to constantly maintain a population of solutions with both convergence and diversity. In this paper, an MaOEA based on two independent stages is proposed for effectively solving many-objective optimization problems (MaOPs), where the convergence and diversity are addressed in two independent and sequential stages. To achieve this, we first propose a nondominated dynamic weight aggregation method by using a genetic algorithm, which is capable of finding the Pareto-optimal solutions for MaOPs with concave, convex, linear and even mixed Pareto front shapes, and then these solutions are employed to learn the Pareto-optimal subspace for the convergence. Afterward, the diversity is addressed by solving a set of single-objective optimization problems with reference lines within the learned Pareto-optimal subspace. To evaluate the performance of the proposed algorithm, a series of experiments are conducted against six state-of-the-art MaOEAs on benchmark test problems. The results show the significantly improved performance of the proposed algorithm over the peer competitors. In addition, the proposed algorithm can focus directly on a chosen part of the objective space if the preference area is known beforehand. Furthermore, the proposed algorithm can also be used to effectively find the nadir points.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Evol. Comput.1
2019 Evolving Unsupervised Deep Neural Networks for Learning Meaningful Representations
abstract
Deep learning (DL) aims at learning the meaningful representations. A meaningful representation gives rise to significant performance improvement of associated machine learning (ML) tasks by replacing the raw data as the input. However, optimal architecture design and model parameter estimation in DL algorithms are widely considered to be intractable. Evolutionary algorithms are much preferable for complex and nonconvex problems due to its inherent characteristics of gradient-free and insensitivity to the local optimal. In this paper, we propose a computationally economical algorithm for evolving unsupervised deep neural networks to efficiently learn meaningful representations, which is very suitable in the current big data era where sufficient labeled data for training is often expensive to acquire. In the proposed algorithm, finding an appropriate architecture and the initialized parameter values for an ML task at hand is modeled by one computational efficient gene encoding approach, which is employed to effectively model the task with a large number of parameters. In addition, a local search strategy is incorporated to facilitate the exploitation search for further improving the performance. Furthermore, a small proportion labeled data is utilized during evolution search to guarantee the learned representations to be meaningful. The performance of the proposed algorithm has been thoroughly investigated over classification tasks. Specifically, error classification rate on MNIST with 1.15% is reached by the proposed algorithm consistently, which is considered a very promising result against state-of-the-art unsupervised DL algorithms.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.1
2019 IGD Indicator-Based Evolutionary Algorithm for Many-Objective Optimization Problems
abstract
Inverted generational distance (IGD) has been widely considered as a reliable performance indicator to concurrently quantify the convergence and diversity of multiobjective and many-objective evolutionary algorithms. In this paper, an IGD indicator-based evolutionary algorithm for solving many-objective optimization problems (MaOPs) has been proposed. Specifically, the IGD indicator is employed in each generation to select the solutions with favorable convergence and diversity. In addition, a computationally efficient dominance comparison method is designed to assign the rank values of solutions along with three newly proposed proximity distance assignments. Based on these two designs, the solutions are selected from a global view by linear assignment mechanism to concern the convergence and diversity simultaneously. In order to facilitate the accuracy of the sampled reference points for the calculation of IGD indicator, we also propose an efficient decomposition-based nadir point estimation method for constructing the Utopian Pareto front (PF) which is regarded as the best approximate PF for real-world MaOPs at the early stage of the evolution. To evaluate the performance, a series of experiments is performed on the proposed algorithm against a group of selected state-of-the-art many-objective optimization algorithms over optimization problems with 8-, 15-, and 20-objective. Experimental results measured by the chosen performance metrics indicate that the proposed algorithm is very competitive in addressing MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.1
2019 A Particle Swarm Optimization-Based Flexible Convolutional Autoencoder for Image Classification
abstract
Convolutional autoencoders (CAEs) have shown their remarkable performance in stacking to deep convolutional neural networks (CNNs) for classifying image data during the past several years. However, they are unable to construct the state-of-the-art CNNs due to their intrinsic architectures. In this regard, we propose a flexible CAE (FCAE) by eliminating the constraints on the numbers of convolutional layers and pooling layers from the traditional CAE. We also design an architecture discovery method by exploiting particle swarm optimization, which is capable of automatically searching for the optimal architectures of the proposed FCAE with much less computational resource and without any manual intervention. We test the proposed approach on four extensively used image classification data sets. Experimental results show that our proposed approach in this paper significantly outperforms the peer competitors including the state-of-the-art algorithms.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
IEEE Trans. Neural Networks Learn. Syst.1
2018 An Experimental Study on Hyper-parameter Optimization for Stacked Auto-Encoders
abstract
Deep learning algorithms have shown their superiority especially in addressing challenging machine learning tasks. The best performance of deep learning algorithms can be reached only when their hyper-parameters have been successfully optimized. However, the hyper-parameter optimization problem is non-convex and non-differentiable, and traditional optimization algorithms are incapable of addressing them well. Evolutionary algorithms are a class of meta-heuristic search algorithms, preferred for optimizing real-world problems due largely to their no mathematical requirements on the problems to be optimized. Although most researchers from the community of deep learning are aware of the effectiveness of evolutionary algorithms in optimizing the hyper-parameters of deep learning algorithms, they still believe that the grid search method is more effective when the number of hyper-parameters is small. To clarify this, we design a hyper-parameter optimization method by using particle swarm optimization that is a widely used evolutionary algorithm, to perform 192 experimental comparisons for stacked auto-encoders that are a class of deep learning algorithms with a relative small number of hyper-parameters, investigate and compare the classification accuracy and computational complexity with those of the grid search method on eight widely used image classification benchmark datasets. The experimental results show that the proposed algorithm can achieve the comparative classification accuracy but saving 10×-100× computational complexity compared with the grid search method.
Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001, Gary G. Yen
CEC1
2018 Evolving Deep Convolutional Neural Networks by Variable-Length Particle Swarm Optimization for Image Classification
abstract
Convolutional neural networks (CNNs) are one of the most effective deep learning methods to solve image classification problems, but the best architecture of a CNN to solve a specific problem can be extremely complicated and hard to design. This paper focuses on utilising Particle Swarm Optimisation (PSO) to automatically search for the optimal architecture of CNNs without any manual work involved. In order to achieve the goal, three improvements are made based on traditional PSO. First, a novel encoding strategy inspired by computer networks which empowers particle vectors to easily encode CNN layers is proposed; Second, in order to allow the proposed method to learn variable-length CNN architectures, a Disabled layer is designed to hide some dimensions of the particle vector to achieve variable-length particles; Third, since the learning process on large data is slow, partial datasets are randomly picked for the evaluation to dramatically speed it up. The proposed algorithm is examined and compared with 12 existing algorithms including the state-of-art methods on three widely used image classification benchmark datasets. The experimental results show that the proposed algorithm is a strong competitor to the state-of-art algorithms in terms of classification error. This is the first work using PSO for automatically evolving the architectures of CNNs.
Bin Wang 0044, Yanan Sun 0001, Bing Xue 0001, Mengjie Zhang 0001
CEC2
2018 Improved Regularity Model-Based EDA for Many-Objective Optimization
abstract
The performance of multiobjective evolutionary algorithms deteriorates appreciably in solving many-objective optimization problems (MaOPs) which encompass more than three objectives. One of the known rationales is the loss of selection pressure which leads to the selected parents not generating promising offspring toward Pareto-optimal front (PF) with diversity. Estimation of distribution algorithms sample new solutions with a probabilistic model built from the statistics extracting over the existing solutions so as to mitigate the adverse impact of genetic operators. In this paper, an improved regularity-based estimation of distribution algorithm is proposed to effectively tackle unconstrained MaOPs. In the proposed algorithm, diversity repairing mechanism is utilized to mend the areas, where need nondominated solutions with a closer proximity to the PF. Then favorable solutions are generated by the model built from the regularity of the solutions surrounding a group of representatives. These two steps collectively enhance the selection pressure which gives rise to the superior convergence of the proposed algorithm. In addition, dimension reduction technique is employed in the decision space to speed up the estimation search of the proposed algorithm. Finally, by assigning the Pareto-optimal solutions to the uniformly distributed reference vectors, a set of solutions with excellent diversity and convergence is obtained. To measure the performance, NSGA-III, GrEA, MOEA/D, HypE, MBN-EDA, and RM-MEDA are selected to perform comparison experiments over DTLZ and DTLZ-test suites with 3-, 5-, 8-, 10-, and 15-objective. Experimental results quantified by the selected performance metrics reveal that the proposed algorithm shows considerable competitiveness in addressing unconstrained MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.1
2017 Global view-based selection mechanism for many-objective evolutionary algorithms
abstract
In traditional many-objective evolutionary algorithms (MaOEAs), solutions survived to the next generation are individually selected which leads to the favorable quality upon the population composed of these selected solutions not necessarily to be gained. However, MaOEAs which are widely used in solving many-objective optimization problems (MaOPs) are considerably preferred due to their population-based nature. In this paper, a global view-based selection mechanism has been proposed to concern the quality of the selected solutions from a global prospective, which is capable of simultaneously facilitating the performance of the entire selected solutions. Indeed, the proposed selection mechanism is equivalent to solve a linear assignment problem whose cost matrix is constructed by the entries concurrently measuring the convergence and diversity of each solution. In addition, this design principle is also utilized for the mating selection to guarantee the convergence and diversity of the selected parents to generate promising offspring. As a case study, the proposed global view-based selection mechanism is integrated into NSGA-III (i.e., GS-NSGA-III). To validate the proposed selection mechanism, extensive experiments are performed by GS-NSGA-III against four state-of-the-art MaOEAs over 8-, 10-, and 15-objective DTLZ1-DTLZ7 test problems. The results measured by the selected performance metric reveal that GS-NSGA-III shows considerable competitiveness in addressing MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
CEC1
2017 Reference line-based Estimation of Distribution Algorithm for many-objective optimization
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
Knowl. Based Syst.1
2017 Explicit guiding auto-encoders for learning meaningful representation
Yanan Sun 0001, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001
Neural Comput. Appl.1
2016 Manifold dimension reduction based clustering for multi-objective evolutionary algorithm
abstract
Real world optimization problems always possess multiple objectives which are conflict in nature. Multi-objective evolutionary algorithms (MOEAs), which provide a group of solutions in region of Pareto front, increasingly draw researchers attention for their excellent performance. In this regard, solutions with a wide diversity would be more favored as they give decision makers more choices to evaluate upon their problems. Based on the insight of investigating the evolution, the Pareto front often lies in a manifold space, not Euclidian space. However, most MOEAs utilize Euclidian distance as a sole mechanism to keep a wide range of diversity for solutions, which is not suitable somewhat from this aspect. To this end, manifold dimension reduction algorithm which has the ability to map solutions in the same front of objective space into Euclidian space is adapted in further. And then, general clustering algorithm are utilized. At the end, we use this technology to replace the crowding distance technology in NSGA-II to choose individuals when there is not enough slots in mating selection process. Based on a range of experiments over benchmark problems against state-of-the-art, it is fully expected benefit of performance improvement will be more significant when applied in many objectives optimization problems. This will be pursuit in our future study.
Yanan Sun 0001, Gary G. Yen, Hua Mao 0001, Zhang Yi 0001
CEC1
2016 Learning a good representation with unsymmetrical auto-encoder
Yanan Sun 0001, Hua Mao 0001, Quan Guo, Zhang Yi 0001
Neural Comput. Appl.1