Yaqing Hou

dblp:180/1448 · DBLP profile ↗
← Back
78ranked-venue papers
12as first author
65since 2021 · last 2026
0000-0002-9929-2650ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 9 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BernO: A Breath-Driven Odor Display for Spatial Olfactory Interaction in VR
abstract
We present a breath-driven odor display device that enables spatial odor perception in virtual reality, relying on users’ natural inhalation rather than pumps or fans. The device supports rapid concentration adjustment through two models: a continuous gradient (monotonic concentration change with position) and a plume (intermittent, fluctuating patterns resembling natural dispersal). To explore its potential, we conducted a proof-of-concept evaluation across four tasks: concentration discrimination, direction and distance localization, and integrated position searching. Results show that the device can dynamically modulate odor concentration for spatial olfactory perception. Our findings further reveal complementary strengths and limitations of the two models—gradients support stable, precise cues, whereas plumes better emulate natural variability. This work introduces a simple yet effective method for simulating spatial odor experiences in VR, offering a lightweight, energy-efficient pathway that expands the design space for olfactory interaction research in virtual environments.
Yu Zhang 0001, Chih-Hung Lee, Jingtong Cai, Yaqing Hou, Qianyao Xu, Qi Lu 0001
CHI4
2026 Regularity model-driven large-scale multi-objective evolutionary algorithm based on dual-information offspring reproduction strategy
Ziliang Du, Gonglin Yuan, Zhenzhou Tang, Ferrante Neri, Yaqing Hou
Expert Syst. Appl.6
2026 Enhancing neural combinatorial optimization by progressive training paradigm
Yaoxin Wu, Yaqing Hou, Hong-Wei Ge
Neurocomputing3
2026 Evolutionary Content Generation via Multimodal LLM-Based Fitness Evaluation
abstract
Evolutionary algorithms (EAs) have gained prominence as a powerful optimization tool inspired by biological evolution, excelling in various complex domains. In the context of Generative Artificial Intelligence (GAI), EAs have shown promise in generating diverse, high-quality solutions. However, traditional EAs heavily rely on human-designed fitness functions, which may often lack flexibility and comprehensiveness for different GAI scenarios. Recently, the emergence of Large Language Models (LLMs) has opened new avenues for enhancing the evolutionary process. This paper proposes a formal framework named ECG-LFit, which utilizes an LLM (e.g., GPT-4-Turbo) for multimodal fitness evaluations. We validate the effectiveness of our framework using EAs (CMA-ES, MAP-Elites, and CMA-ME) in our case study on Super Mario game level generation. The results show that our framework improves the quality and playability of the generated levels. Additionally, user studies indicate that participants prefer the levels generated by the ECG-LFit framework, particularly regarding aesthetics, challenge, and playability. Furthermore, to enhance evaluation efficiency, we design a distilled model to simulate the scoring process of the LLM, enabling rapid and effective content evaluation in resource-constrained environments.
Yaqing Hou, Zhaoping Yu, David M. Bossens, Qiang Zhang 0008, Yew-Soon Ong
IEEE Trans. Evol. Comput.1
2025 A Biogeography-based Dual-Strategy Particle Swarm Algorithm for Numerical and Engineering Design Optimization
abstract
Recently, researchers developed numerous nature-inspired meta-heuristics algorithms, such as particle swarm optimization (PSO), to solve numerical and engineering optimization problems. However, the PSO still suffers from complex optimization problems, such as slow convergence speed and getting stuck in local optima. This paper proposes a Biogeography-based Dual-strategy PSO algorithm (DBPSO). Firstly, a Dual-Strategy Search (DSS) is proposed, which utilizes an improved random learning strategy for global search and a modified quasi-Newton search strategy for local search. The proportion of resources allocated to the modified quasi-Newton search strategy is associated with the inertia weight value of the improved random learning, achieving a balance between exploration and exploitation. Secondly, an improved migration operator of biogeography-based optimization is embedded in DSS. Finally, an iteration-based hybridization strategy that allows one-way information transmission is proposed to fuse the two algorithms effectively. Experiments are conducted on CEC2013 and CEC2017 test suites, and the results show that DBPSO ranks first by comparing it with multiple PSO variants and other meta-heuristics variants. In addition, experiments are conducted on two real-world engineering optimization problems to demonstrate the applicability of DBPSO in solving practical optimization problems, such as the car side impact design problem.
Hong-Wei Ge, Yaqing Hou, Mengyue Wang
CEC3
2025 Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesis
abstract
Human motion synthesis aims to generate plausible human motion sequences, which has raised widespread attention in computer animation. Recent score-based generative models (SGMs) have demonstrated impressive results on this task. However, their training process involves complex curvature trajectories, leading to unstable training process. In this paper, we propose a Deterministic-to-Stochastic Diverse Latent Feature Mapping (DSDFM) method for human motion synthesis. DSDFM consists of two stages. The first human motion reconstruction stage aims to learn the latent space distribution of human motions. The second diverse motion generation stage aims to build connections between the Gaussian distribution and the latent space distribution of human motions, thereby enhancing the diversity and accuracy of the generated human motions. This stage is achieved by the designed deterministic feature mapping procedure with DerODE and stochastic diverse output generation procedure with DivSDE. DSDFM is easy to train compared to previous SGMs-based methods and can enhance diversity without introducing additional training parameters. Through qualitative and quantitative experiments, DSDFM achieves state-of-the-art results surpassing the latest methods, validating its superiority in human motion synthesis.
Hua Yu 0006, Weiming Liu 0005, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008
CVPR4
2025 Efficient Decision Sequence Modeling via Feature-LevelMasking
abstract
Decision models based on sequence modeling have become prevalent in the field of offline reinforcement learning. However, existing approaches such as Decision Transformer and Trajectory Transformer suffer from poor data efficiency. One important reason is that they fail to extract useful information from potentially high-dimensional and noisy states. To resolve this issue, we propose a data-efficient decision sequence modeling method called Data Efficient Decision Sequence Model (DEDS), which dynamically identifies and filters out task-irrelevant state features for more compact state representations. Specifically, DEDS employs a state mask model guided by the bisimulation metric to ensure that only the most task-relevant state features are preserved for decision-making. Extensive experiments on various environments demonstrate that DEDS outperforms existing offline RL methods and achieves a significantly high data efficiency, especially in tasks with high-dimensional and complex state spaces.
Xinzhi Zhang 0009, Yaqing Hou, Xuechuan Liu, Zengyang Wang, Yi Cai 0001, Bo An 0001, Mengchen Zhao
DAI3
2025 EFormer: An Effective Edge-based Transformer for Vehicle Routing Problems
abstract
Recent neural heuristics for the Vehicle Routing Problem (VRP) primarily rely on node coordinates as input, which may be less effective in practical scenarios where real cost metrics—such as edge-based distances—are more relevant. To address this limitation, we introduce EFormer, an Edge-based Transformer model that uses edge as the sole input for VRPs. Our approach employs a precoder module with a mixed-score attention mechanism to convert edge information into temporary node embeddings. We also present a parallel encoding strategy characterized by a graph encoder and a node encoder, each responsible for processing graph and node embeddings in distinct feature spaces, respectively. This design yields a more comprehensive representation of the global relationships among edges. In the decoding phase, parallel context embedding and multi-query integration are used to compute separate attention mechanisms over the two encoded embeddings, facilitating efficient path construction. We train EFormer using reinforcement learning in an autoregressive manner. Extensive experiments on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) reveal that EFormer outperforms established baselines on synthetic datasets, including large-scale and diverse distributions. Moreover, EFormer demonstrates strong generalization on real-world instances from TSPLib and CVRPLib. These findings confirm the effectiveness of EFormer’s core design in solving VRPs.
Dian Meng, Zhiguang Cao, Yaoxin Wu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008
IJCAI4
2025 Preference-based Deep Reinforcement Learning for Historical Route Estimation
abstract
Recent Deep Reinforcement Learning (DRL) techniques have advanced solutions to Vehicle Routing Problems (VRPs). However, many of these methods focus exclusively on optimizing distance-oriented objectives (i.e., minimizing route length), often overlooking the implicit drivers' preferences for routes. These preferences, which are crucial in practice, are challenging to model using traditional DRL approaches. To address this gap, we propose a preference-based DRL method characterized by its reward design and optimization objective, which is specialized to learn historical route preferences. Our experiments demonstrate that the method aligns generated solutions more closely with human preferences. Moreover, it exhibits strong generalization performance across a variety of instances, offering a robust solution for different VRP scenarios.
Boshen Pan, Yaoxin Wu, Zhiguang Cao, Yaqing Hou, Guangyu Zou, Qiang Zhang 0008
IJCAI4
2025 Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person Retrieval
abstract
Text-to-image person retrieval aims to identify target person images using natural language descriptions. Current state-of-the-art methods predominantly rely on single-round retrieval frameworks, where retrieval accuracy heavily depends on the quality of the initial textual descriptions. However, users sometimes struggle to provide detailed and distinctive descriptions in a single attempt, resulting in generic initial queries that lack discriminative details. This fundamental limitation of the single-round retrieval framework frequently leads to the misinterpretation of user intent and suboptimal retrieval performance. To address this limitation, we propose Dialogue-driven Interactive Dynamic Learning (DIDL) for text-to-image person retrieval. Specifically, we first introduce Collaborative Query Refinement (CQR), which progressively refines retrieval conditions through multi-round dialogues. Then, we design Dynamic Context Resampling (DCR) based on a bi-granular mask strategy that enhances the model's adaptation to dialogue-style contexts and effectively balances its attention between initial descriptions and supplementary information. Based on these components, we further propose cross-modal Probabilistic Context Matching Modeling (ProCMM) that establishes effective associations between static visual features and dynamic contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across all three benchmark datasets.
Hong-Wei Ge, Yuxuan Liu 0015, Yaqing Hou
ACM Multimedia4
2025 UniteFormer: Unifying Node and Edge Modalities in Transformers for Vehicle Routing Problems
abstract
Neural solvers for the Vehicle Routing Problem (VRP) have typically relied on either node or edge inputs, limiting their flexibility and generalization in real-world scenarios. We propose UniteFormer, a unified neural solver that supports node-only, edge-only, and hybrid input types through a single model trained via joint edge-node modalities. UniteFormer introduces: (1) a mixed encoder that integrates graph convolutional networks and attention mechanisms to collaboratively process node and edge features, capturing cross-modal interactions between them; and (2) a parallel decoder enhanced with query mapping and a feed-forward layer for improved representation. The model is trained with REINFORCE by randomly sampling input types across batches. Experiments on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) demonstrate that UniteFormer achieves state-of-the-art performance and generalizes effectively to TSPLib and CVRPLib instances. These results underscore UniteFormer’s ability to handle diverse input modalities and its strong potential to improve performance across various VRP tasks.
Dian Meng, Zhiguang Cao, Jie Gao 0010, Yaoxin Wu, Yaqing Hou
NeurIPS5
2025 MTRec: Learning to Align with User Preferences via Mental Reward Models
abstract
Recommendation models are predominantly trained using implicit user feedback, since explicit feedback is often costly to obtain. However, implicit feedback, such as clicks, does not always reflect users' real preferences. For example, a user might click on a news article because of its attractive headline, but end up feeling uncomfortable after reading the content. In the absence of explicit feedback, such erroneous implicit signals may severely mislead recommender systems. In this paper, we propose MTRec, a novel sequential recommendation framework designed to align with real user preferences by uncovering their internal satisfaction on recommended items. Specifically, we introduce a mental reward model to quantify user satisfaction and propose a distributional inverse reinforcement learning approach to learn it. The learned mental reward model is then used to guide recommendation models to better align with users’ real preferences. Our experiments show that MTRec brings significant improvements to a variety of recommendation models. We also deploy MTRec on an industrial short video platform and observe a 7\% increase in average user viewing time.
Mengchen Zhao, Yaqing Hou, Xiangyang Li 0004, Pengjie Gu, Zhenhua Dong, Ruiming Tang, Yi Cai 0001
NeurIPS3
2025 GAM: A Generative Autoencoder for Diverse Human Motion Prediction
Jiapeng Bai, Hua Yu 0006, Yaqing Hou, Qiang Zhang 0008
PRICAI3
2025 Niche-based Memetic algorithm with adaptive parameters for optimizing order delivery strategies in O2O platforms
Guangyu Zou, Heng Qi, Jiafu Tang, Yaqing Hou
Appl. Intell.5
2025 An evolutionary multitasking algorithm for multi-objective feature selection using dual-perspective reduction
Mengyue Wang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou
Eng. Appl. Artif. Intell.5
2025 MARIC: an efficient multi-agent real-time intention-based communication model for team cooperation
Hong-Wei Ge, Zhangang Hao, Yaqing Hou
Neural Comput. Appl.4
2025 DNA sequence design model for multi-scene fusion
Yanfen Zheng, Yaqing Hou, Qiang Zhang 0008, Xiaopeng Wei
Neural Comput. Appl.4
2025 A Spatio-Temporal Continuous Network for Stochastic 3D Human Motion Prediction
abstract
Stochastic Human Motion Prediction (HMP) has received increasing attention due to its wide applications. Despite the rapid progress in generative fields, existing methods often face challenges in learning continuous temporal dynamics and predicting stochastic motion sequences. They tend to overlook the flexibility inherent in complex human motions and are prone to mode collapse. To alleviate these issues, we propose a novel method called STCN, for stochastic and continuous human motion prediction, which consists of two stages. Specifically, in the first stage, we propose a spatio-temporal continuous network to generate smoother human motion sequences. In addition, the anchor set is innovatively introduced into the stochastic HMP task to prevent mode collapse, which refers to the potential human motion patterns. In the second stage, STCN endeavors to acquire the Gaussian mixture distribution (GMM) of observed motion sequences with the aid of the anchor set. It also focuses on the probability associated with each anchor, and employs the strategy of sampling multiple sequences from each anchor to alleviate intra-class differences in human motions. Experimental results on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy.
Hua Yu 0006, Yaqing Hou, Xu Gui 0001, Shanshan Feng 0001, Qiang Zhang 0008
IEEE Trans. Circuits Syst. Video Technol.2
2025 Cooperative Multiagent Learning and Exploration With Min-Max Intrinsic Motivation
abstract
In the field of multiagent reinforcement learning (MARL), the ability to effectively explore unknown environments and collect information and experiences that are most beneficial for policy learning represents a critical research area. However, existing work often encounters difficulties in addressing the uncertainties caused by state changes and the inconsistencies between agents' local observations and global information, which presents significant challenges to coordinated exploration among multiple agents. To address this issue, this article proposes a novel MARL exploration method with Min-Max intrinsic motivation (E2M) that promotes the learning of joint policies of agents by introducing surprise minimization and social influence maximization. Since the agent is subject to unstable state changes in the environment, we introduce surprise minimization by computing state entropy to encourage the agents to cope with more stable and familiar situations. This method enables surprise estimation based on the low-dimensional representation of states obtained from random encoders. Furthermore, to prevent surprise minimization from leading to conservative policies, we introduce mutual information between agents' behaviors as social influence. By maximizing social influence, the agents are encouraged to interact to facilitate the emergence of cooperative behavior. The performance of our proposed E2M is testified across a range of popular StarCraft II and Multiagent MuJoCo tasks. Comprehensive results demonstrate its effectiveness in enhancing the cooperative capability of the multiple agents.
Yaqing Hou, Haiyin Piao, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Qiang Zhang 0008
IEEE Trans. Cybern.1
2025 Learning to Evolve With Guiding Solutions Generated by Generative Adversarial Network
abstract
Many search strategies have been designed to generate a promising offspring population for efficiently solving large-scale multiobjective optimization problems (LSMOPs). The effectiveness of existing search strategies relies on the quality of good parent solutions. However, especially in early generations, the current population does not always include high-quality solutions. This article proposes a generative adversarial network (GAN)-guided search (G2S) strategy for learning to evolve with guiding solutions. Its main idea is to employ GAN for mapping a set of guiding points in the objective space with good convergence and diversity back to the decision space to guide evolution. Specifically, the current population is used as real data, and the guiding points consisting of nondominated solutions and reference vectors are used as virtual data. The trained GAN generates guiding solutions in the decision space to guide the population to evolve efficiently. A large-scale multiobjective evolutionary framework using G2S is also proposed which can be embedded into multiobjective evolutionary algorithms (MOEAs) to improve their ability to handle LSMOPs. Experimental studies on several benchmark problems with the highest-5000-D decision space show that the proposed G2S is competitive compared with the state-of-the-art algorithms and has impressive efficiency as the component to improve the performance of MOEAs for solving LSMOPs.
Hong-Wei Ge, Yaqing Hou, Hisao Ishibuchi
IEEE Trans. Evol. Comput.3
2025 DG-SMOTE: A Distance-Angle-Based Genetic Synthetic Minority Over-Sampling Technique for Unbalanced Data Learning
abstract
Many real-world applications often generate unbalanced data. Learning from such data may lead to biased classifiers that perform poorly on the class of interest. Oversampling methods have been shown to be effective in rebalancing unbalanced data to help classifiers avoid performance bias. However, many existing oversampling methods rely on a predesigned linear model structure and the neighborhood information of an original instance. This may lead to the generation of noisy instances when the original data has noise. In this study, we develop a novel oversampling method in which genetic programming is introduced to automatically select good-quality instances and evolve a model structure that combines the selected instances to create a new instance. In the proposed oversampling method, an individual is used to represent a generated instance, which is evaluated by the fitness function designed based on the Euclidean distance and the cosine theorem. In the experiments, we examine the effectiveness of the proposed oversampling method in assisting different types of classifiers to solve the issue of class imbalance, and compare it with popular sampling methods in unbalanced classification. The results have been analyzed comprehensively, indicating that the new method successfully addressed the class imbalance issue by generating a group of good-quality instances for the minority class and outperformed the compared sampling methods in almost all cases.
Wenbin Pei, Yuyang Cui, Bing Xue 0001, Mengjie Zhang 0001, Jiqing Zhang, Yaqing Hou, Guangyu Zou, Qiang Zhang 0008
IEEE Trans. Evol. Comput.6
2025 DivDiff: A Conditional Diffusion Model for Diverse Human Motion Prediction
abstract
Diverse human motion prediction (HMP) aims to predict multiple plausible future motions given an observed human motion sequence. It is a challenging task due to the diversity of potential human motions while ensuring an accurate description of future human motions. Current solutions are either low-diversity or limited in expressiveness. Recent denoising diffusion probabilistic models (DDPM) demonstrate promising performance in various generative tasks. However, introducing DDPM directly into diverse HMP incurs some issues. While DDPM can enhance the diversity of potential human motion patterns, the predicted human motions gradually become implausible over time due to significant noise disturbances in the forward process of DDPM. This phenomenon leads to the predicted human motions being unrealistic, seriously impacting the quality of predicted motions and restricting their practical applicability in real-world scenarios. To alleviate this, we propose a novel conditional diffusion-based generative model, called DivDiff, to predict more diverse and realistic human motions. Specifically, the DivDiff employs DDPM as our backbone and incorporates Discrete Cosine Transform (DCT) and Transformer mechanisms to encode the observed human motion sequence as a condition to instruct the reverse process of DDPM. More importantly, we design a diversified reinforcement sampling function (DRSF) to enforce human skeletal constraints on the predicted human motions. DRSF utilizes the acquired information from human skeletal as prior knowledge, thereby reducing significant disturbances introduced during the forward process. Extensive results received in the experiments on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy.
Hua Yu 0006, Yaqing Hou, Wenbin Pei, Yew-Soon Ong, Qiang Zhang 0008
IEEE Trans. Multim.2
2025 A Group-Based Many-Task Collaborative Optimization Framework for Evolutionary Robots Design
abstract
In evolutionary robotics (ER), the evolution of a robot’s morphology (i.e., physical structure) or controller (i.e., control algorithm or instruction sequence) often entails tackling an extensive number of tasks. The use of evolutionary multitasking (EMT) in ER, which optimizes multiple tasks simultaneously by reusing potentially useful knowledge across diverse tasks, could improve the performance of problem-solving to each task. However, existing EMT methods do not fully use intertask correlations, limiting knowledge sharing. In view of this, this study introduces a novel framework, termed adaptive group-based collaborative optimization, tailored for handling optimization problems involving a large number of tasks within the ER domain simultaneously. The proposed framework divides tasks into groups according to their similarity and then proceeds through two principal stages, namely, intergroup knowledge separation and intragroup knowledge reunion. During intergroup knowledge separation stage, an adaptive method for selecting crossover operators enables source tasks to share useful knowledge to the target task across groups. During intragroup knowledge reunion stage, an adaptive knowledge combination strategy facilitates the target task in assimilating knowledge from multiple sources intragroup. We validated the efficacy of the proposed framework in both planar manipulators and hexapod robot experiments. The results indicate that our method outperforms existing state-of-the-art algorithms (i.e., MME, MMKT) on several metrics (e.g., mean fitness and quality diversity metrics). The proposed method can effectively improve the effectiveness and diversity of solutions in solving ER problems with a large number of tasks (e.g., 5 000 or 10 000), and has broad potential in practical ER applications.
Yaqing Hou, Zhaoping Yu, Wenbin Pei, Yaoxin Wu, Hong-Wei Ge, Bing Xue 0001, Mengjie Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Similar Locality Based Transfer Evolutionary Optimization for Minimalistic Attacks
abstract
Deep neural networks are powerful and popular learning models; however, recent studies have shown that deep neural network-based policies are susceptible to deception by adversarial attacks. A minimalistic attack is a specialized form of adversarial attack that aims to accomplish successful attacks at the lowest possible cost. Recently, transfer optimization algorithms have been applied to deceive previously trained policies by acquiring knowledge from previously solved tasks. Experiments indicate that the transfer optimization algorithms perform well compared to traditional optimization algorithms. However, current transfer algorithms for addressing minimalistic attacks not only select a single source task for knowledge transfer but also tend to overly rely on identified appropriate source tasks. To address this issue, this paper introduces a similar locality based transfer evolutionary optimization algorithm. It can adaptively select multiple source tasks and extract valuable knowledge from these source tasks. Moreover, by leveraging the concept of similar locality, the algorithm alleviates its excessive dependence on familiar tasks, thereby providing fresh knowledge for the optimization of the target task. On this basis, the algorithm can mine more valuable knowledge from the large source task space to achieve a successful attack in a shorter period. The algorithm is tested on three Atari games-BeamRider, Qbert, and Seaquest-demonstrating its ability and potential to outperform other transfer optimization algorithms currently available in solving this problem.
Wenqiang Ma, Yaqing Hou, Hua Yu 0006, Xiangrong Tong, Zexuan Zhu 0001, Qiang Zhang 0008
CEC2
2024 Enhancing Imbalanced Classification with Support Vector Machines via Evolutionary Oversampling Algorithms
abstract
Support Vector Machines (SVMs), as well-known algorithms, have been successfully applied to classification problems. However, when dealing with imbalanced data, the classification performance of SVMs could be significantly compromised. One approach to tackle the class imbalance is oversampling the minority class, exemplified by methods like SMOTE and its variants. These methods generate new samples by interpolation between existing ones and determine the weights based on the ratio of samples from different classes, leading to inaccurate weight assignment, limited generation scope, and indiscriminate sample generation. To address these limitations, we propose novel evolutionary oversampling algorithms based on Support Vector Machine (SVM) and Evolutionary Algorithms (EAs) called SEOA. SEOA leverages the inherent capability of SVM to identify the samples that have a critical influence on the decision boundary and assign them appropriate weights, thereby eliminating the reliance on human experience. Furthermore, SEOA utilizes a novel approach for sample generation and emphasizes the significance of margin for classification, introducing a mechanism that employs margin as the metric to evaluate the quality of generated samples. To assess the performance of SEOA, we conducted a comprehensive comparison against various oversampling methods across 19 real-world datasets. The results underscore SEOA's superiority, showcasing its distinct strengths in addressing the challenges posed by imbalanced classification.
Yongchao Chen, Yaqing Hou, Xiangrong Tong, Qiang Zhang 0008
CEC4
2024 SAEIR: Sequentially Accumulated Entropy Intrinsic Reward for Cooperative Multi-Agent Reinforcement Learning with Sparse Reward
Hong-Wei Ge, Yaqing Hou
IJCAI3
2024 Self-Attention Guided Advice Distillation in Multi-Agent Deep Reinforcement Learning
abstract
Advising is an effective method to enhance agent learning performance in multi-agent deep reinforcement learning. Existing advising methods typically rely on a teacher-student framework where a teacher agent provides student agents with action or Q-value advice. However, they share a common limitation: the advice from a teacher agent can only assist a student in making a one-time decision in the current state and cannot be internalized into the student agent’s knowledge to intrinsically change the student agent’s decision model. Consequently, the advice acts more like a one-time instruction from the teacher rather than a learning aid. If the student agent encounters the same problem again, it may still be unable to make a sound decision and need to request advice. This not only fails to rapidly enhance the agent’s decision-making ability fundamentally but also leads to a considerable waste of communication costs. Hence, we propose a multi-agent advice distillation framework through attention that allows the student agent to request advice from the experienced teacher and distill that advice into their own decision model via the self-attention mechanism. As a result, advice is fully utilized, allowing for a rapid and intrinsic improvement in the agent’s decision-making capabilities. Our empirical evaluations demonstrate that, compared to existing advising methods, our method significantly improves learning performance while reducing the communication cost.
Sihan Zhou, Yaqing Hou, Liran Zhou, Hong-Wei Ge, Liang Feng 0001
IJCNN3
2024 Safety-guided Deep Reinforcement Learning for Path Planning of Autonomous Mobile Robots
abstract
The past decade has witnessed the blooming of autonomous mobile robot (AMR) applications with an increasing trend from closed to open working environments with unexpected obstacles. In this context, safety becomes a critical concern in the path planning of AMRs. However, it is a challenging task to guarantee the safety while maintaining the efficiency of path planning algorithm due to the uncertainty of open environments. Therefore, we propose in this paper to categorize the working environment of AMRs into three safety levels, i.e. low-risk, medium-risk and high-risk areas, according to the real-time distance between the AMR and the nearest obstacle. In particular, we incorporate safety levels into the famous reinforcement learning algorithm, deep deterministic policy gradient (DDPG), and develop a safety-guided DDPG algorithm for the path planning of AMRs. In low-risk areas, we adopt the conventional DDPG path planning algorithm directly to guarantee the efficiency (since the safety is generally not an issue in this case). In the medium-risk areas, we design a velocity threshold adjustment method, an OU noise with bias and propose a new reward function, aiming at encouraging AMRs to perform flexible actions as early as possible in order to avoid collisions. In the high-risk areas, we re-design the reward function based on the potential collision risk and shield misleading rewards that may cause local optimum problem, so as to ensure the safety of AMRs in dangerous situations. Simulation results confirm the satisfactory performance of the proposed scheme.
Zhuoru Yu, Yaqing Hou, Qiang Zhang 0008, Qian Liu 0001
IJCNN2
2024 Towards Efficient and Diverse Generative Model for Unconditional Human Motion Synthesis
abstract
Recent generative methods have revolutionized the way of human motion synthesis, such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Denoising Diffusion Probabilistic Models (DMs). These methods have gained significant attention in human motion fields. However, there are still challenges in unconditionally generating highly diverse human motions from a given distribution. To enhance the diversity of synthesized human motions, previous methods usually employ deep neural networks (DNNs) to train a transport map that transforms Gaussian noise distribution into real human motion distribution. According to Figalli's regularity theory, the optimal transport map computed by DNNs frequently exhibits discontinuities. This is due to the inherent limitation of DNNs in representing only continuous maps. Consequently, the generated human motions tend to heavily concentrate on densely populated regions of the data distribution, resulting in mode collapse or mode mixture. To address the issues, we propose an efficient method called MOOT for unconditional human motion synthesis. First, we utilize a reconstruction network based on GRU and transformer to map human motions to latent space. Next, we employ convex optimization to match the noise distribution with the latent space distribution of human motions through the Optimal Transport (OT) map. Then, we combine the extended OT map with the generator of reconstruction network to generate new human motions. Thereby overcoming the issues of mode collapse and mode mixture. MOOT generates a latent code distribution that is well-behaved and highly structured, providing a strong motion prior for various applications in the field of human motion. Through qualitative and quantitative experiments, MOOT achieves state-of-the-art results surpassing the latest methods, validating its superiority in unconditional human motion generation.
Hua Yu 0006, Weiming Liu 0005, Jiapeng Bai, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008
ACM Multimedia5
2024 Stochastic online decisioning hyper-heuristic for high dimensional optimization
Hong-Wei Ge, Mingde Zhao 0002, Yaqing Hou
Appl. Intell.4
2024 Progressive reconstruction-decoupled face super-resolution framework with controllable knowledge guidance
Guozhi Tang, Hong-Wei Ge, Enxuan Gu, Yaqing Hou, Mingde Zhao 0002
Knowl. Based Syst.4
2024 Temporal prediction model with context-aware data augmentation for robust visual reinforcement learning
Xinkai Yue, Hong-Wei Ge, Yaqing Hou
Neural Comput. Appl.4
2024 A Multiagent Cooperative Learning System With Evolution of Social Roles
abstract
Recent developments in reinforcement learning (RL) have been able to derive optimal policies for sophisticated and capable agents, and shown to achieve human-level performance on a number of challenging tasks. Unfortunately, when it comes to multiagent systems (MASs), complexities, such as nonstationarity and partial observability bring new challenges to the field. Building a flexible and efficient multiagent RL (MARL) algorithm capable of handling complex tasks has to date remained an open challenge. This article presents a multiagent learning system with the evolution of social roles (eSRMA). The main interest is placed on solving the key issues in the definition and evolution of suitable roles, and optimizing the policies accompanied by social roles in MAS efficiently. Specifically, eSRMA incorporates and cultivates role division awareness of agents to improve the ability to deal with complex cooperative tasks. Each agent is assigned a role module, which can dynamically generate roles based on the individuals’ local observations. A novel MARL algorithm is designed as the principal driving force that governs the role-policy learning process by a role-attention credit assignment mechanism. Moreover, a role evolution process is developed to help agents dynamically choose appropriate roles in decision making. Comprehensive experiments on the StarCraft II micromanagement benchmarkhave demonstrated that eSRMA exhibits superiority in achieving higher learning capability and efficiency for multiple agents compared to the state-of-the-art MARL methods.
Yaqing Hou, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Hong-Wei Ge, Qiang Zhang 0008
IEEE Trans. Evol. Comput.1
2024 A Virtual-Sensor Construction Network Based on Physical Imaging for Image Super-Resolution
abstract
Image imaging in the real world is based on physical imaging mechanisms. Existing super-resolution methods mainly focus on designing complex network structures to extract and fuse image features more effectively, but ignore the guiding role of physical imaging mechanisms for model design, and cannot mine features from a physical perspective. Inspired by the mechanism of physical imaging, we propose a novel network architecture called Virtual-Sensor Construction network (VSCNet) to simulate the sensor array inside the camera. Specifically, VSCNet first generates different splitting directions to distribute photons to construct virtual sensors, and then performs a multi-stage adaptive fine-tuning operation to fine-tune the number of photons on the virtual sensors to increase the photosensitive area and eliminate photon cross-talk, and finally converts the obtained photon distributions into RGB images. These operations can naturally be regarded as the virtual expansion of the camera's sensor array in the feature space, which makes our VSCNet bridge the physical space and feature space, and uses their complementarity to mine more effective features to improve performance. Extensive experiments on various datasets show that the proposed VSCNet achieves state-of-the-art performance with fewer parameters. Moreover, we perform experiments to validate the connection between the proposed VSCNet and the physical imaging mechanism. The implementation code is available at https://github.com/GZ-T/VSCNet.
Guozhi Tang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou, Mingde Zhao 0002
IEEE Trans. Image Process.4
2024 Discriminative Identity-Feature Exploring and Differential Aware Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to learn discriminative representations for person retrieval from unlabeled data. Currently, state-of-the-art techniques accomplish this task by using instance contrastive learning, which contrasts the similarities of the instances in different views. However, existing contrastive methods only focus on the positive effects of inter-instance relationships, while neglecting the negative effects of intra-instance redundancy information. This redundancy information can generate invalid or spurious intra-class relationships during the instance contrasting process, which enlarges the intra-class gaps and increases the noisy pseudo-labels. To address this issue, we propose a discriminative identity-feature exploring and differential aware learning (DiDAL) framework to learn more discriminative intra-identity representations. Specifically, the DiDAL extracts intra-instance salient features by synthetic complementary attention, and further explores the discriminative identity features by modeling the relationship among these salient features based on graph neural networks. This strategy aims to reduce the intra-instance redundancy information. Moreover, DiDAL explores hard instances by leveraging the extracted intra-instance salient features, and matches an anchor with multiple hard positive instances to enhance the robustness of the model to noisy pseudo-labels. Extensive experiment results on two widely used person re-identification datasets and a vehicle re-identification dataset demonstrate the superiority of the proposed method compared with existing state-of-the-art methods.
Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002
IEEE Trans. Multim.4
2024 Clothes-Changing Person Re-Identification via Universal Framework With Association and Forgetting Learning
abstract
Clothes-changing person re-identification (Re-ID) aims at learning identity-relevant feature representations among clothing-changed persons. Currently, the state-of-the-art methods accomplish this task by using additional assistance (e.g., silhouettes, sketches, clothes labels, etc.) to explore identity-relevant information. However, humans do not require redundant assistance information to retrieve clothing-changed persons. It is commonly known that humans can recall targets they have seen before with a simple reminder. Inspired by human perception, we propose an association and forgetting learning (AFL) framework for clothes-changing person re-identification. Specifically, on the one hand, during the association learning process, the AFL framework constructs association factors for each identity to simulate the reminders found in human perception. Then, the original instances and the explored hardest positive instances are cross-correlated by the association factors to learn identity-relevant features. On the other hand, the model is forced to forget the identity-irrelevant features by the proposed forgetting learning module, which improves the intra-class compactness. Finally, we further propose a clustering relationship exploration (CRE) module to optimize the cluster distribution of clothes-changing instances, which enables AFL to also be effectively applied in unsupervised settings, improving the universal applicability of the model. Extensive experiment results obtained on clothes-changing person Re-ID datasets under supervised and unsupervised settings demonstrate the superiority of the proposed method over the existing state-of-the-art methods.
Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002
IEEE Trans. Multim.4
2023 A Knowledge Transfer-Based Genetic Algorithm for Multi-Target Robotic Arm Control
abstract
The ability to swiftly and precisely reach any user-specified target location is necessary for a robotic arm that can be used in real-world scenarios. To date, many evolutionary optimization algorithms have been used to design controllers for robotic arms. However, when designing a robotic arm to reach multiple targets, most existing methods need to evolve the control strategy from scratch for each target, rather than trying to reuse existing experience. Therefore, computational resources are repeatedly and meaninglessly consumed. To this end, this paper proposes a genetic algorithm based on knowledge transfer (GAKT) dedicated to reusing existing knowledge to optimize a new robotic arm control task. Specifically, the knowledge transfer process can be summarized into the following two steps. First, through sequential transfer, GAKT initializes the population with the help of a knowledge base constructed by a quality diversity algorithm. Second, underperforming individuals are encouraged to acquire knowledge from excellent individuals in the same generation during the optimization process. We tested the effectiveness of GAKT and investigated its average performance by selecting multiple target points in different dimensions. The results show that GAKT can find the most advantageous arrival strategy (that is, make the end of the manipulator the closest to the target) on most of the selected targets. Moreover, we conducted ablation experiments and demonstrated the effectiveness of the knowledge transfer processes.
Zhaoping Yu, Wenbin Pei, Yaqing Hou, Zexuan Zhu 0001, Xianneng Li
CEC5
2023 Surrogate-Assisted Morphology Optimization by Genetic Algorithms
abstract
Deep reinforcement learning has attracted wide interest because of its extraordinary capabilities in multiple fields. However, morphology optimization by using evolutionary computation techniques has not been intensively investigated. In this paper, we explore the use of genetic algorithms (GA) to automatically design the morphology of an agent. Evaluating the performance of an agent is very time-consuming because it needs to be trained from scratch. Moreover, it is computationally infeasible to train separate controllers for all possible different morphologies of agents to identify the optimal ones and is difficult to obtain the accurate cumulative reward of an agent to estimate the performance of the morphologies. To address these issues, we use a morphology comparator as a surrogate model to estimate the probability of one morphology being better than the other, instead of directly predicting the performance of each morphology. A set of surrogate models based on a radial basis function network are developed before evolution to make full use of the data to guide the search. Experimental results indicate that the proposed method is able to efficiently find out optimal morphologies to achieve better performance than the default morphology.
Jinlin Jiang, Yongchao Chen, Wenbin Pei, Junxiang Zhang, Yaqing Hou, Hong-Wei Ge, Liang Feng 0001
CEC5
2023 Co-speech Gesture Synthesis by Reinforcement Learning with Contrastive Pretrained Rewards
abstract
There is a growing demand of automatically synthesizing co-speech gestures for virtual characters. However, it remains a challenge due to the complex relationship between input speeches and target gestures. Most existing works focus on predicting the next gesture that fits the data best, however, such methods are myopic and lack the ability to plan for future gestures. In this paper, we propose a novel reinforcement learning (RL) framework called RACER to generate sequences of gestures that maximize the overall satisfactory. RACER employs a vector quantized variational autoencoder to learn compact representations of gestures and a GPT-based policy architecture to generate coherent sequence of gestures autoregressively. In particular, we propose a contrastive pre-training approach to calculate the rewards, which integrates contextual information into action evaluation and successfully captures the complex relationships between multi-modal speech-gesture data. Experimental results show that our method significantly outperforms existing baselines in terms of both objective metrics and subjective human judgements. Demos can be found at https://github.com/RLracer/RACER.git.
Mengchen Zhao, Yaqing Hou, Minglei Li 0001, Huang Xu 0003, Songcen Xu, Jianye Hao
CVPR3
2023 BRGR: Multi-agent cooperative reinforcement learning with bidirectional real-time gain representation
Hong-Wei Ge, Liang Sun 0003, Yaqing Hou
Appl. Intell.5
2023 Camera-aware progressive learning for unsupervised person re-identification
Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou
Neural Comput. Appl.4
2023 SCAU-net: 3D self-calibrated attention U-Net for brain tumor segmentation
Ning Sheng, Yutong Han, Yaqing Hou, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008
Neural Comput. Appl.4
2023 Attention-guided spatial-temporal graph relation network for video-based person re-identification
Hong-Wei Ge, Wenbin Pei, Yuxuan Liu 0015, Yaqing Hou, Liang Sun 0003
Neural Comput. Appl.5
2023 Multi-agent air combat with two-stage graph-attention communication
Zhixiao Sun, Huahua Wu, Yandong Shi, Xiangchao Yu, Wenbin Pei, Zhen Yang 0011, Haiyin Piao, Yaqing Hou
Neural Comput. Appl.9
2023 Complementary Attention-Driven Contrastive Learning With Hard-Sample Exploring for Unsupervised Domain Adaptive Person Re-ID
abstract
Unsupervised domain adaptive (UDA) methods for person re-identification (Re-ID) aim to transfer the knowledge of the labeled source domain to the unlabeled target domain without further annotations, which is challenging due to the drift of label distribution and the missing of target domain labels. Improving the clustering accuracy of pseudo-labels can help the model fit the target domain. However, the errors of pseudo-label noise will be accumulated during training, which is harmful to the model performance. Moreover, the hard samples can lead to a large gap between intra-class features and a small gap between inter-class features. To address these problems, this paper proposes a complementary attention-driven contrastive learning with hard-sample exploring (CACHE) algorithm. In CACHE, on one hand, the complementary attention module is used to improve the discriminability of the features. The obtained discriminative features can reduce noisy pseudo-labels and improve the clustering accuracy of pseudo labels; On the other hand, we explore the hard samples based on the instance relationship and cluster relationship for contrastive learning. This way can make the cluster more compact. Extensive experiments on three large-scale person re-identification benchmarks demonstrate the effectiveness of the proposed method, which significantly outperforms state-of-the-art methods in terms of mAP and CMC.
Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Toward Realistic 3D Human Motion Prediction With a Spatio-Temporal Cross- Transformer Approach
abstract
Human motion prediction intends to predict how humans move given a historical sequence of 3D human motions. Recent transformer-based methods have attracted increasing attentions and demonstrated their promising performance in 3D human motion prediction. However, existing methods generally decompose the input of human motion information into spatial and temporal branches in a separate way and seldom consider their inherent coherence between the two branches, hence often failing to register the dynamic spatio-temporal information during the training process. Motivated by these issues, we propose a spatio-temporal cross-transformer network (STCT) for 3D human motion predictions. Specifically, we investigate various types of interaction methods (i.e., Concatenation Interaction, Msg token interaction, and Cross-transformer) to capture the coherence of the spatial and temporal branches. According to the obtained results, the proposed cross-transformer interaction method shows its superiority over other methods. Meanwhile, considering that most existing works treat the human body as a set of 3D human joint positions, the predicted human joints are proportionally less appropriate to the realistic human body due to unreasonable bone length and non-plausible poses as time progresses. We further resort to the bone constraints of human mesh to produce more realistic human motions. By fitting a parametric body model (i.e., SMPL-X model) to the predicted human joints, a reconstruction loss function is proposed to remedy the unreasonable bone length and pose errors. Comprehensive experiments on AMASS and Human3.6M datasets have demonstrated that our method achieves superior performance over compared methods.
Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Wenbin Pei, Hong-Wei Ge, Xin Yang 0011, Qiang Zhang 0008, Mengjie Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 A Preliminary Study of Multi-task MAP-Elites with Knowledge Transfer for Robotic Arm Design
abstract
The structure design of robotic arms is of great importance on completing industrial tasks successfully. This is a typical multi-task optimization problem when considering different constraints as different tasks. However, mainstream methods for multi-task optimization such as evolutionary multitasking and Multi-task MAP-Elites algorithms tend to encounter problems such as high computational cost and slow convergence when solving large-scale robotic arm tasks. To this end, this paper proposes a new framework based on the MAP-Elites algorithms for solving large-scale robot arm design tasks, called Multi-task MAP-Elites with Knowledge Transfer (MMKT). Specifically, this paper designs the group-based knowledge transfer process for large-scale task optimization in which all tasks are classified into different groups according to their similarity to generate multiple knowledge transfer areas; and knowledge transfer strategies are designed to enhance the quality of solutions with low fitness value. We test the effectiveness of the MMKT framework in planar robotic arm experiments (2000, 5000, and 10,000 tasks; 10, 15-dimensional search space). The experimental results prove that the MMKT outperforms the MME, CMA-ES, and classical ES algorithms.
Hua Yu 0006, Han Linghu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008
CEC6
2022 Towards Efficient 3D Human Motion Prediction using Deformable Transformer-based Adversarial Network
abstract
Human motion prediction is a crucial step for achieving human-robot interactions. While recent transformer-based methods have shown great potentials in 3D human motion prediction, they still suffer from mode collapse to non-plausible poses and quadratically computational complexity with respect to the increasing length of input sequences. In this paper, we propose a novel spatio-temporal deformable transformer-based adversarial network (STDTA) for 3D human motion prediction. First, we design a spatio-temporal deformable transformer module to capture the correlations between human joints while reducing the computational costs. Second, we introduce the adversarial training mechanism and design fidelity and continuity discriminators to maintain smoothness and stability for the long-term prediction. Finally, extensive experiments on Human 3.6M and AMASS benchmarks demonstrate that the proposed STDTA achieves state-of-the-art performance.
Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Cai Kang, Qiang Zhang 0008
ICRA3
2022 Multi-Omics Data Integration Patient Classification Method Based on Deep Dense Residual Shrinkage Network
abstract
Represented by high-throughput sequencing technology, rapid advances in omics technologies have allowed scholars to understand human diseases in a more profound and comprehensive manner. However, omics data are difficult to analyze accurately on account of their characters: massive redundancy, full of noise and high dimension. Meanwhile, it is crucial to improve the performance of omics analysis from multiple modalities using data from different omics. To this end, we propose a novel multi-omics integration model for patient classification. Firstly, anti-noisy deep dense residual shrinkage neural networks (DDRSNN) are trained and utilized to make preliminary predictions on various omics data of patients. Then confidence scores are calculated for each sample through the confidence distance to evaluate the reliability of the preliminary prediction result. Finally, referring to uncertain evidence fusion theory, the preliminary prediction results from different omics are integrated based on confidence scores and the final patient classification is obtained. The algorithm combines multiple omics, treating each omics as evidence of a separate modality, and improves the accuracy and reliability of patient classification by integrating evidence from multiple modalities.
Yaqing Hou, Qiang Zhang 0008
ICTAI3
2022 Budgeted Sequence Submodular Maximization
abstract
The problem of selecting a sequence of items that maximizes a given submodular function appears in many real-world applications. Existing study on the problem only considers uniform costs over items, but non-uniform costs on items are more general. Taking this cue, we study the problem of budgeted sequence submodular maximization (BSSM), which introduces non-uniform costs of items into the sequence selection. This problem can be found in a number of applications such as movie recommendation, course sequence design and so on. Non-uniform costs on items significantly increase the solution complexity and we prove that BSSM is NP-hard. To solve the problem, we first propose a greedy algorithm GBM with an error bound. We also design an anytime algorithm POBM based on Pareto optimization to improve the quality of solutions. Moreover, we prove that POBM can obtain approximate solutions in expected polynomial running time, and converges faster than a state-of-the-art algorithm POSEQSEL for sequence submodular maximization with cardinality constraints. We further introduce optimizations to speed up POBM. Experimental results on both synthetic and real-world datasets demonstrate the performance of our new algorithms.
Xuefeng Chen 0001, Liang Feng 0001, Xin Cao 0001, Yifeng Zeng, Yaqing Hou
IJCAI5
2022 Deep Relationship Graph Reinforcement Learning for Multi-Aircraft Air Combat
abstract
Air combat Artificial Intelligence (AI) has attracted increasing attentions from aeronautics engineers and artificial intelligence researchers. However, it is often of great difficulties for the existing methods to solve the collaboration problems in multi-aircraft air combat due to their high complexity incurred by combination explosion. In view of this, we propose a Deep Relationship Graph Reinforcement Learning (DRGRL) algorithm for multi-aircraft collaboration. Specifically, DRGRL significantly simplifies the complex situation space via abstracting the original problem into a symbolic form. Besides, a novel Air Combat Relationship Graph (ACRG) is introduced to represent the learned collaboration pattern, which concentrates on the most important combat relationships for tactic decision making. Consequently, experiments are conducted in an air combat simulation environment named WUKONG. The comprehensive experimental results demonstrate that DRGRL could evidently learn some valuable collaboration patterns and achieve better combat performance than state-of-the-art air combat AI methods.
Haiyin Piao, Yaqing Hou, Zhixiao Sun, Shengqi Yang, Xuanqi Peng, Songyuan Fan
IJCNN3
2022 ICMiF: Interactive cascade microformers for cross-domain person re-identification
Jiajian Huang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou
Inf. Sci.4
2022 Adaptive kernel selection network with attention constraint for surgical instrument classification
abstract
Abstract Computer vision (CV) technologies are assisting the health care industry in many respects, i.e., disease diagnosis. However, as a pivotal procedure before and after surgery, the inventory work of surgical instruments has not been researched with the CV-powered technologies. To reduce the risk and hazard of surgical tools’ loss, we propose a study of systematic surgical instrument classification and introduce a novel attention-based deep neural network called SKA-ResNet which is mainly composed of: (a) A feature extractor with selective kernel attention module to automatically adjust the receptive fields of neurons and enhance the learnt expression and (b) A multi-scale regularizer with KL-divergence as the constraint to exploit the relationships between feature maps. Our method is easily trained end-to-end in only one stage with few additional calculation burdens. Moreover, to facilitate our study, we create a new surgical instrument dataset called SID19 (with 19 kinds of surgical tools consisting of 3800 images) for the first time. Experimental results show the superiority of SKA-ResNet for the classification of surgical tools on SID19 when compared with state-of-the-art models. The classification accuracy of our method reaches up to 97.703%, which is well supportive for the inventory and recognition study of surgical tools. Also, our method can achieve state-of-the-art performance on four challenging fine-grained visual classification datasets.
Yaqing Hou, Qian Liu 0001, Hong-Wei Ge, Jun Meng, Qiang Zhang 0008, Xiaopeng Wei
Neural Comput. Appl.1
2022 Multi-space evolutionary search with dynamic resource allocation strategy for large-scale optimization
Qingxia Shang, Junwei Dong, Yaqing Hou, Yu Wang 0108, Min Li 0056, Liang Feng 0001
Neural Comput. Appl.4
2022 Perception-oriented Single Image Super-Resolution Network with Receptive Field Block
Yaqing Hou, Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Neural Comput. Appl.2
2022 Multi-Agent Transfer Reinforcement Learning With Multi-View Encoder for Adaptive Traffic Signal Control
abstract
Multi-agent reinforcement learning (MARL) based methods for adaptive traffic signal control (ATSC) have shown promising potentials to solve the heavy traffic problems. The existing MARL methods adopt centralized or distributed strategies. The former only models the environment as an agent and suffers from the exponential growth of action and state space. The latter extends the independent reinforcement learning methods, such as DQN, to multiple interactions directly or propagates information, such as state and policy, without taking their qualities into account. In this paper, we propose a multi-agent transfer reinforcement learning method to enhance the performance of MARL for ATSC, which is termed as multi-agent transfer soft actor-critic with the multi-view encoder (MT-SAC). The MT-SAC combines centralized and distributed strategies. In MT-SAC, we propose a multi-view state encoder and a transfer learning paradigm with guidance. The encoder processes input states from multiple perspectives and uses an attention mechanism to weigh the neighborhood information. While the paradigm enables the agents to handle different conditions for improving generalization abilities by transfer learning. Experimental studies on different scale road networks show that the MT-SAC outperforms the state-of-the-art algorithms and makes the traffic signal controllers more collaborative and robust.
Hong-Wei Ge, Dongwan Gao, Liang Sun 0003, Yaqing Hou, Chao Yu 0004, Yuxin Wang 0001, Guozhen Tan
IEEE Trans. Intell. Transp. Syst.4
2021 Multi-task Actor-Critic with Knowledge Transfer via a Shared Critic
abstract
Multi-task actor-critic is a learning paradigm proposed in the literature to improve the learning efficiency of multiple actor-critics by sharing the learned policies across tasks while the reinforcement learning progresses online. However, existing multi-task actor-critic algorithms can only handle reinforcement learning tasks within the same problem domain, they may fail in cases where tasks possessing diverse state-action spaces. Taking this cue, in this paper, we embark a study on multi-task actor-critic with knowledge transfer via a share critic to enable the multi-task learning of actor-critic in heterogeneous state-action environments. Further, for efficient learning of the proposed multi-task actor-critic, a new formula for calculating the gradient of the actor network is also presented. To evaluate the performance of our approach, comprehensive empirical studies on continuous robotic tasks with different numbers of links. The experimental results confirmed the effectiveness of the proposed multi-task actor-critic algorithm.
Gengzhi Zhang, Liang Feng 0001, Yaqing Hou
ACML3
2021 EMT-ReMO: Evolutionary Multitasking for High-Dimensional Multi-Objective Optimization via Random Embedding
abstract
Since multi-objective optimization (MOO) involves multiple conflicting objectives, the high dimensionality of the solution space has a much more severe impact on multi-objective problems than single-objective optimization. Taking the advantage of random embedding, some related works have been proposed to scale derivative-free MOO methods to high-dimensional functions. However, with the premise of "low effective dimensionality", a single randomly embedded subspace cannot guarantee the effectiveness of obtained solutions. Taking this cue, we propose an evolutionary multitasking paradigm for multi-objective optimization via random embedding (EMT-ReMO) to enhance the efficiency and effectiveness of current embedding-based methods in solving high-dimensional optimization problems with low effective dimensions. In EMT-ReMO, the target problem is firstly embedded into multiple low-dimensional subspaces by using different random embeddings, aiming to build up a multi-task environment for identifying the underlying effective subspace. Then the implicit multi-objective evolutionary multitasking is performed with seamless knowledge transfer to enhance the optimization process. Experimental results obtained on six high-dimensional MOO functions with or without low effective dimensions have confirmed the effectiveness as well as the efficiency of the proposed EMT-ReMO.
Yinglan Feng, Liang Feng 0001, Yaqing Hou, Kay Chen Tan, Sam Kwong
CEC3
2021 A Study on Realtime Task Selection Based on Credit Information Updating in Evolutionary Multitasking
Yumeng Cao, Yaqing Hou, Liang Feng 0001, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei
EMO2
2021 Multi-Scale Attention Constraint Network for Fine-Grained Visual Classification
abstract
Capturing subtle yet discriminative features constitutes a great challenge in fine-grained visual classification due to the large intra-class and small inter-class variances. Main-stream works for this problem localize at attention mechanism and feature relationship learning. However, existing methods treat the features in isolation while neglecting the effect of attention-enhanced features on relationships between different network layers. In this paper, we propose a novel attention-based method by Multi-Scale Attention Constraint network composed of two important components: (1) a feature extractor with lightweight group-wise enhanced attention blocks that guides the generation of high representation features; and (2) a multi-scale regularizer that explores the relationships between different features. Extensive experiments show that our approach achieves state-of-the-art performance on standard benchmark datasets. Moreover, we introduce a new dataset, consisting of comprehensive surgical instrument categories based on three common surgeries, to support the classification and inventory work of surgical instruments.
Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei
ICME1
2021 Weakly Supervised Gleason Grading of Prostate Cancer Slides using Graph Neural Network
Yaqing Hou, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008
ICPRAM2
2021 Brain Tumor Segmentation based on Knowledge Distillation and Adversarial Training
abstract
3D MRI brain tumor segmentation is a reliable method for disease diagnosis and treatment plans in the future. Early on, the segmentation of brain tumors is mostly done manually. However, manual segmentation of 3D MRI brain tumor requires professional anatomical knowledge and may be inaccurate. In this paper, we propose a 3D MRI brain tumor segmentation architecture based on the encoder-decoder structure. Specially, we introduce knowledge distillation and adversarial training methods, which compresses and improves the accuracy and robustness of the model. Furthermore, we obtain soft targets by designing multiple teacher network training and then apply them to the student network. Finally, we evaluate our method on a challenging BraTS dataset. As a result, the performance of our proposed model is superior to state-of-the-art methods.
Yaqing Hou, Tianbo Li, Qiang Zhang 0008, Hua Yu 0006, Hong-Wei Ge
IJCNN1
2021 A Contextual Attention Network for Multimodal Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) is a challenging task due to the complexity of emotions and dynamics in dialogues. Current studies for emotion recognition mostly focus on the modeling of a single utterance in dialogue, which neglects self and inter-speaker influence. This paper presents a contextual attention neural network based on the multimodal framework that leverages the conversational information from both target and the other speaker for utterance-level emotion detection. Specifically, we utilize recurrent neural networks based on contextual attention for modeling the transaction and dependence between speakers. Further, the feature fusion is proposed to unite the important modal information extracted from multiple modalities, including audio, text and video, hence providing more useful and comprehensive knowledge for emotion recognition. The proposed approach shows its superiority in extracting contexts for self and inter-speaker influence and synthesizing them as global features that are beneficial to detect individual emotion state. Experiment result on the IEMOCAP corpus reports an accuracy of 64.6%, demonstrating the superiority of the proposed method in emotion recognition comparing to the state-of-the-arts.
Tana Wang, Yaqing Hou, Qiang Zhang 0008
IJCNN2
2021 Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognition
abstract
Abstract In the study of human action recognition, two-stream networks have made excellent progress recently. However, there remain challenges in distinguishing similar human actions in videos. This paper proposes a novel local-aware spatio-temporal attention network with multi-stage feature fusion based on compact bilinear pooling for human action recognition. To elaborate, taking two-stream networks as our essential backbones, the spatial network first employs multiple spatial transformer networks in a parallel manner to locate the discriminative regions related to human actions. Then, we perform feature fusion between the local and global features to enhance the human action representation. Furthermore, the output of the spatial network and the temporal information are fused at a particular layer to learn the pixel-wise correspondences. After that, we bring together three outputs to generate the global descriptors of human actions. To verify the efficacy of the proposed approach, comparison experiments are conducted with the traditional hand-engineered IDT algorithms, the classical machine learning methods (i.e., SVM) and the state-of-the-art deep learning methods (i.e., spatio-temporal multiplier networks). According to the results, our approach is reported to obtain the best performance among existing works, with the accuracy of 95.3% and 72.9% on UCF101 and HMDB51, respectively. The experimental results thus demonstrate the superiority and significance of the proposed architecture in solving the task of human action recognition.
Yaqing Hou, Hua Yu 0006, Pengfei Wang 0013, Hong-Wei Ge, Jianxin Zhang 0001, Qiang Zhang 0008
Neural Comput. Appl.1
2021 Evolutionary Multiagent Transfer Learning With Model-Based Opponent Behavior Prediction
abstract
This article embarks a study on multiagent transfer learning (TL) for addressing the specific challenges that arise in complex multiagent systems where agents have different or even competing objectives. Specifically, beyond the essential backbone of a state-of-the-art evolutionary TL framework (eTL), this article presents the novel TL framework with prediction (eTL-P) as an upgrade over existing eTL to endow agents with abilities to interact with their opponents effectively by building candidate models and accordingly predicting their behavioral strategies. To reduce the complexity of candidate models, eTL-P constructs a monotone submodular function, which facilitates to select Top-${K}$models from all available candidate models based on their representativeness in terms of behavioral coverage as well as reward diversity. eTL-P also integrates social selection mechanisms for agents to identify their better-performing partners, thus improving their learning performance and reducing the complexity of behavior prediction by reusing useful knowledge with respect to their partners’ mind universes. Experiments based on a partner-opponent minefield navigation task (PO-MNT) have shown that eTL-P exhibits the superiority in achieving higher learning capability and efficiency of multiple agents when compared to the state-of-the-art multiagent TL approaches.
Yaqing Hou, Yew-Soon Ong, Jing Tang 0001, Yifeng Zeng
IEEE Trans. Syst. Man Cybern. Syst.1
2020 Large-Scale optimization via Evolutionary Multitasking assisted Random Embedding
abstract
Evolutionary algorithms (EAs) often lose their superiority and effectiveness when applied to large-scale optimization problems. In the literature, many research studies have been proposed to improve the search performance of EAs, such as cooperative co-evolution, embedding, and new search operator design. Among those, memetic multi-agent optimization (MeMAO) is a recently proposed paradigm for high-dimensional problems by using random embeddings. It demonstrated high efficacy with the assumption of “effective dimension However, as prior knowledge is always unknown for a given problem, this method may fail on the large-scale problems that do not have low effective dimensions. Taking this cue, we propose an evolutionary multitasking (EMT) assisted random embedding method (EMT-RE) for solving large-scale optimization problems. Instead of conducting a search on the randomly embedded space directly, we treat the embedded task as the auxiliary task for the given problem. By performing EMT with both the given problem and the randomly embedded task, not only the useful solutions found along the search can be transferred across tasks toward efficient problem solving, but the effectiveness of the search on problems Without a low effective dimensionality is also guaranteed. To evaluate the performance of newly proposed EMT-RE, comprehensive empirical studies are carried out on 8 synthetic continuous optimization functions with up to 2,000 dimensions.
Yinglan Feng, Liang Feng 0001, Yaqing Hou, Kay Chen Tan
CEC3
2020 Memetic Multi-agent optimization with Problem Reformulation by Coordinate Rotation
abstract
Memetic multi-agent system (MeMAS) is recently proposed as an enhanced version that integrates meme concept into multi-agent system (MAS) wherein all meme-inspired agents have an improvement in learning performance via meme evolution independently or social interaction. In the process of solving the black box optimization problem, the potential advantages of MeMAS have not been utilized well, which makes it a fertile area for further exploration. This paper presents a memetic multi-agent optimization paradigm through coordinate rotation (MeMAO-R) to combine MeMAS with evolutionary algorithms (EAs) to improve optimization efficiency. Based on MeMAS, the particular interest of MeMAO-R is placed on assisting original complex optimization task with new tasks generated by coordinate rotation. Further, MeMAO-R constructs the social interaction mechanism which facilitates to improve their convergence speed for solving the target optimization problem by utilizing meaningful information transferred across multiple agents with differing views of the target problem. Besides, MeMAO-R employs one or more classical EAs as the fundamental population based evolutionary solvers for multiple agents to optimize multiple tasks in a multi-agent scenario. Lastly, to testify the efficacy of the proposed MeMAO-R, comprehensive empirical studies on basic optimization problems are provided.
Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Abhishek Gupta 0001, Xianneng Li
CEC2
2020 D2D-Enabled Reliable Data Collection for Mobile Crowd Sensing
abstract
With increasing more powerful sensing capacities of mobile devices, the Mobile Crowd Sensing (MCS) system requires to collect larger sensing data from participants. Nevertheless, collecting such large volume of data will cost a lot for participants, base stations and MCS server. Even worse, some sensing data cannot satisfy the MCS sensing requirement due to the low quality and are filtered by the MCS server in clouds. Inspired by the D2D technique, where mobile devices can communicate directly with the help of the nearby base station, in 5G networks, we propose the Reliable Data Collection (RDC) algorithm to validate the generated sensing data at device sides in this paper. To be specific, the whole progress is formulated as a Probability problem of Discovering Reliable sensing data (PDR) at client sides, and Expectation Maximization (EM) is leveraged to devise the algorithm. Finally, the extensive simulations and real-world use case are conducted to evaluate the performance of RDC algorithm, and the result shows that RDC outperforms the other two benchmarks in estimating accuracy and saving data collection cost.
Pengfei Wang 0013, Chi Lin 0001, Leyou Yang, Yaqing Hou, Qiang Zhang 0008
ICPADS5
2020 Two-stage Automatic Image Annotation Based on Latent Semantic Scene Classification
abstract
The rapid growth of multimedia content makes existing automatic image annotation techniques difficult to satisfy the demands of real-world applications. In this paper, we propose a two-stage automatic image annotation algorithm (TAIA) based on latent semantic scene classification. In the offline training phase, the hidden connectivity of labels is firstly excavated by a directed-weighed graph based on label co-occurrence relation matrix, and then the latent scene categories are detected among the labels by using nonnegative matrix factorization. Further, we propose a multi-view extreme learning machine (MELM) to learn the probability that the multiple visual feature maps to the semantic scenes. In the online annotation phase, the image to be annotated is fed to the scene classifier MELM to identify its relevant scenes. Then k-nearest neighbor based annotator is conducted on the relevant scenes to predict labels for the unannotated images. The TAIA is formulated in such a framework so that the relationship between labels and semantic scenes is fully considered, and the hard classification problem is solved. The experimental results on multiple datasets have demonstrated that the proposed framework TAIA is both effective and efficient.
Hong-Wei Ge, Kai Zhang 0050, Yaqing Hou, Chao Yu 0004, Mingde Zhao 0002, Zhen Wang 0004, Liang Sun 0003
IJCNN3
2020 A Preliminary Study of Fusion ARTs with Adaptively Information Intensity Attenuation Controlling
abstract
Fusion ART is an enhanced version of Adaptive Resonance Theory (ART) which is derived from a biologically-plausible theory of human cognitive information processing. Due to its well-established ability of learning associative mappings across multimodal pattern channels in an online and incremental manner, fusion ART has been widely applied in many real world learning problems. In this paper, we take a Fusion Architecture for Learning, Cognition, and Navigation (FALCON) as the specification and essential backbone of fusion ART and introduce an intensity attenuation controller δ for adaptively adjusting the intensity of information captured from the environment, by taking inspiration from Broadbent-Treisman Filter-Attenuation's perceptual model of environmental attention. Particularly, we propose both an adaptive δ detection algorithm as well as a δ-based pruning algorithm to enhance the learning performance of FALCON while reduce the redundant memory storage incurred by the "detrimental δ". To verify the effectiveness and efficiency of our proposed method, comprehensive experimental studies are carried out on a classical minefield navigation task.
Wenxuan Zhu, Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Liang Feng 0001, Xinghua Qu
IJCNN2
2020 A Preliminary Study of Improving Evolutionary Multi-Objective Optimization via Knowledge Transfer from Single-Objective Problems
abstract
In the last decades, evolutionary algorithms (EAs) have demonstrated strong search capabilities in solving multi-objective optimization problems (MOPs). To improve the search performance of EAs, as problems seldom exist in isolation, transferring knowledge from related problems have attracted considerable attentions in recent years. In this paper, we present a preliminary study to enhance existing evolutionary algorithms (MOEAs) by transferring knowledge from the process of solving the single objectives involved in a given MOP of interest. As the single objectives are the objectives of the MOP, they naturally share great similarity with the given MOP, which thus could yield useful traits for enhancing the problem-solving of the MOP. To the best of our knowledge, this work severs as the first attempt to improve evolutionary multi-objective optimization via transferring knowledge from single objective problems. To evaluate the performance of the proposed method, empirical studies using a popular MOEA, i.e., NSGAII, on commonly used multi-objective benchmarks are conducted. The obtained results confirmed the efficacy of the proposed method in terms of both convergence speed and solution quality.
Lingyu Huang, Liang Feng 0001, Handing Wang, Yaqing Hou, Kai Liu 0001, Chao Chen 0004
SMC4
2019 Memetic Multi-agent Optimization in High Dimensions using Random Embeddings
abstract
In this paper, we propose a memetic multi-agent optimization (MeMAO) paradigm to enhance the search efficacy of classical EAs (i.e., Differential Evolution (DE)) in solving the complex optimization problems. The essential backbone of MeMAO is a recently proposed memetic multi-agent learning system wherein agents acquire increasing learning capabilities by interacting with the environment mainly in a reinforcement learning manner. Differing from MeMAS, the particular interest of MeMAO is placed on addressing the specific challenges when applying classical EAs to optimize the high dimensional optimization problems with a "low effective dimensionality". To achieve this, the target optimization problem is firstly re-formulated into multiple low dimensional tasks via random embedding methods. Further, MeMAO employs DE as the fundamental population based evolutionary solver for multiple agents to optimize multiple low dimensional tasks in a multi-agent scenario. Importantly, MeMAO constructs the social interaction mechanisms among multiple agents, hence improves their convergence speed for solving the target optimization problem by sharing the beneficial information across multiple agents. Lastly, to testify the efficacy of the proposed MeMAO, comprehensive empirical studies on 8 synthetic optimization problems with a dimensionality of 2,000 are provided.
Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xinghua Qu, L. Feng, Abhishek Gupta 0001
CEC1
2019 Memetic Evolution Strategy for Reinforcement Learning
abstract
Neuroevolution (i.e., training neural network with Evolution Computation) has successfully unfolded a range of challenging reinforcement learning (RL) tasks. However, existing neuroevolution methods suffer from high sample complexity, as the black-box evaluations (i.e., accumulated rewards of complete Markov Decision Processes (MDPs)) discard bunches of temporal frames (i.e., time-step data instances in MDP). Actually, these temporal frames hold the Markov property of the problem, that benefits the training of neural network as well by temporal difference (TD) learning. In this paper, we propose a memetic reinforcement learning (MRL) framework that optimizes the RL agent by leveraging both black-box evaluations and temporal frames. To this end, an evolution strategy (ES) is associated with Q learning, where ES provides diversified frames to globally train the agent, and Q learning locally exploits the Markov property within frames to refresh the agent. Therefore, MRL conveys a novel memetic framework that allows evaluation free local search by Q learning. Experiments on classical control problem verify the efficiency of the proposed MRL, that achieves significantly faster convergence than canonical ES.
Xinghua Qu, Yew-Soon Ong, Yaqing Hou, Xiaobo Shen 0001
CEC3
2019 A Preliminary Study of Adaptive Task Selection in Explicit Evolutionary Many-Tasking
abstract
Recently, evolutionary multi-tasking (EMT) has been proposed as a new evolutionary search paradigm that op-timizes multiple problems simultaneously. Due to the knowledge transfer across optimization tasks occurs along the evolutionary search process, EMT has been demonstrated to outperform the traditional single-task evolutionary search algorithms on many complex optimization problems, such as multimodal continuous optimization problems, NP-hard combinatorial optimization problems, and constrained optimization problems. Today, EMT has attracted lots of attentions, and many EMT algorithms have been proposed in the literature. The explicit EMT algorithm (EEMTA) is a recent proposed new EMT algorithm. In contrast to most of existing EMT algorithms, which employ a single population using unified space and common search operators for solving multiple problems, the EEMTA uses multiple populations which possess problem-specific solution representations and search mechanisms for different problems in evolutionary multi-tasking, which thus could lead to enhanced optimization performance. However, the original EEMTA was proposed for solving only two tasks. As knowledge transfer from inappropriate tasks may lead to negative effect on the evolutionary optimization process, additional designs of identifying task pairs for knowledge transfer is necessary in EEMTA for evolutionary multi-tasking with tasks more than two. To the best of our knowledge, there is no research effort has been conducted on this issue. Keeping this in mind, in this paper, we present a preliminary study on the task selection in EEMTA for many-task optimization. As task similarity may lose to capture the usefulness between tasks in evolutionary search, instead of using similarity measures for task selection, here we propose a credit assignment approach for selecting proper task to conduct knowledge transfer in explicit evolutionary many-tasking. The proposed approach is based on the feedbacks from the transferred solutions across tasks, which is adaptively updated along the evolutionary search. To confirm the efficacy of the proposed method, empirical studies on the many-task optimization problem, which consists of 7 commonly used optimization benchmarks, have been presented and discussed.
Qingxia Shang, Liang Feng 0001, Yaqing Hou, J. Zhong, Abhishek Gupta 0001, Kay Chen Tan, H.-L. Liu
CEC4
2018 A Preliminary Study of Adaptive Indicator Based Evolutionary Algorithm for Dynamic Multiobjective Optimization via Autoencoding
abstract
Dynamic multi-objective optimization problem (D-MOP) is widely existed in many real-world applications. Over the years, DMOP has attracted many research attentions in the literature. The adaptive indicator-based evolutionary algorithm (IBEA2) is a recently proposed multi-objective evolutionary algorithm (MOEA). It has demonstrated strong search capability on commonly used multi-objective benchmarks over state-of-the-art MOEAs. However, as the adaptation of parameter$k$is based on the selected solutions with maximum hypervolume, this mechanism will be inappropriate if the problem changes over time. The reason is that the solutions with high hypervolume at one particular time instance may not be with high hypervolume at another if the problem changed. Keeping this in mind, inspired by the recent autoencoding evolutionary search, which is able to transfer the past search experiences to improve the evolutionary search on unseen problems, in this paper, we propose to extend the IBEA2 by adapting k with transferred high hypervolume solutions obtained before the dynamic change occurs, for solving DMOP. To evaluate the proposed method, empirical comparisons on the commonly used Farina-Deb-Amato (FDA) DMOP benchmarks, against both the IBEA2 and one recently proposed dynamic MOEA, are presented.
Wei Zhou 0001, Liang Feng 0001, Siwei Jiang, Shu Zhang 0003, Yaqing Hou, Yew-Soon Ong, Zexuan Zhu 0001, Kai Liu 0001
CEC5
2017 An Evolutionary Transfer Reinforcement Learning Framework for Multiagent Systems
abstract
In this paper, we present an evolutionary transfer reinforcement learning framework (eTL) for developing intelligent agents capable of adapting to the dynamic environment of multiagent systems (MASs). Specifically, we take inspiration from Darwin's theory of natural selection and Universal Darwinism as the principal driving forces that govern the evolutionary knowledge transfer process. The essential backbone of our proposed eTL comprises several meme-inspired evolutionary mechanisms, namely meme representation, meme expression, meme assimilation, meme internal evolution, and meme external evolution. Our proposed approach constructs social selection mechanisms that are modeled after the principles of human learning to identify appropriate interacting partners. eTL also models the intrinsic parallelism of natural evolution and errors that are introduced due to the physiological limits of the agents' ability to perceive differences, so as to generate “growth” and “variation” of knowledge that agents have of the world, thus exhibiting higher adaptivity capabilities on solving complex problems. To verify the efficacy of the proposed paradigm, comprehensive investigations of the proposed eTL against existing state-of-the-art TL methods in MAS, are conducted on the “minefield navigation tasks” platform and the “Unreal Tournament 2004” first person shooter computer game, in which homogeneous and heterogeneous learning machines are considered.
Yaqing Hou, Yew-Soon Ong, Liang Feng 0001, Jacek M. Zurada
IEEE Trans. Evol. Comput.1
2016 A conceptual modeling of flocking-regulated multi-agent reinforcement learning
abstract
In this paper, we present a multi-agent reinforcement learning (MARL) framework that leverages the emergent behaviors from swarm intelligence (SI). The essential backbone of our framework is an flocking-regulated cooperative learning paradigm in which the cooperation among learning agents is realized via the self-organizing principles derived from natural interaction of flocking boids. In the proposed MARL, each reinforcement learner learns and evolves in the dynamic environment, and is steered by flocking behavior rules such as cohesion, separation, alignment, fear, etcs. The use of the flocking rules provides distributed sensing and communication content for the cooperation of multiple learning agents in the context of pursuit game. The effectiveness of the MARL framework is studied by its application of the multi-agent pursuit game.
Yaqing Hou, Yew-Soon Ong
IJCNN2
2016 Creating human-like non-player game characters using a Memetic Multi-Agent System
abstract
Memetic Multi-Agent System (MeMAS) has recently emerged as a combination of memetic automaton and multi-agent system (MAS), wherein all meme-inspired agents acquire increasing learning capacity and intelligence through meme evolution. This paper further presents a study of MeMAS in developing human-like non-player characters in complex first-person shooter (FPS) games. In particular, we consider a well-known commercial FPS game, known as Unreal Tournament 2004 (UT2004), as our game of interest and discuss the details of non-player characters based on a manifestation of “Temporal Difference - Fusion Architecture for Learning and Cognition” (TDFALCON) neural network. In addition, we present a brief cross-domain study of MeMAS wherein the useful knowledge in the form of memes learned from different yet related simple domains previously solved are used to enhance learning performance of the non-player characters in UT2004. Benchmark experiments are studied to investigate the efficacy of MeMAS in UT2004 from various aspects, including learning efficiency, generalization capability, and computational cost. The empirically results indicate that the MeMAS could clearly improve the learning effectiveness and efficiency of designed non-player characters in UT2004.
Yaqing Hou, Liang Feng 0001, Yew-Soon Ong
IJCNN1