Xin Xu 0001

dblp:66/3874-1 · DBLP profile ↗
← Back
138ranked-venue papers
19as first author
76since 2021 · last 2026
0000-0003-3238-745XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 83 · 16 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 3 since 2021Systems, architecture and hardware · 9 · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VALA: Virtual anchor-guided label assignment for tiny object detection
Jiyun Zhang, Qiang Fang 0001, Shuohao Shi, Yujun Zeng, Xin Xu 0001
Neurocomputing5
2026 Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation
abstract
Learning neural implicit fields of 3D shapes is a rapidly emerging field that enables shape representation at arbitrary resolutions. Due to the flexibility, neural implicit fields have succeeded in many research areas, including shape reconstruction, novel view image synthesis, and more recently, object pose estimation. Neural implicit fields enable learning dense correspondences between the camera space and the object's canonical space - including unobserved regions in camera space - significantly boosting object pose estimation performance in challenging scenarios like highly occluded objects and novel shapes. Despite progress, predicting canonical coordinates for unobserved camera-space regions remains challenging due to the lack of direct observational signals. This necessitates heavy reliance on the model's generalization ability, resulting in high uncertainty. Consequently, densely sampling points across the entire camera space may yield inaccurate estimations that hinder the learning process and compromise performance. To alleviate this problem, we propose a method combining an SO(3)-equivariant convolutional implicit network and a positive-incentive point sampling (PIPS) strategy. The SO(3)-equivariant convolutional implicit network estimates point-level attributes with SO(3)-equivariance at arbitrary query locations, demonstrating superior performance compared to most existing baselines. The PIPS strategy dynamically determines sampling locations based on the input, thereby boosting the network's accuracy and training efficiency. The PIPS strategy is implemented with a PIPS estimation network which generates sparse sample points with distinctive features capable of determining all object pose DoFs with high certainty. To collect the training data of the PIPS estimation network, we propose to automatically generate the pseudo ground-truth with a teacher model. Our method outperforms the state-of-the-art on three pose estimation datasets. It achieves 0.63 in the $5^{\circ }2$5∘2 cm metric on NOCS-REAL275, 0.62 in the $5^{\circ }5$5∘5 cm metric on ShapeNet-C, and 77.3 in the AR metric on LineMOD-O. Notably, it demonstrates significant improvements in challenging scenarios, such as objects captured with unseen pose, high occlusion, novel geometry, and severe noise.
Boyan Wan, Xin Xu 0001, Kai Xu 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Learning Task-Preferred Inference Routes for Gradient De-Conflict in Multi-Output DNNs
abstract
Multi-output deep neural networks (MONs) contain multiple output branches of various tasks, and these tasks typically share partial network filters, resulting in entangled inference routes between different tasks within the networks. Due to the divergent optimization objectives, the task gradients during training usually interfere with each other along the shared routes, which decreases the overall model performance. To address this issue, we propose a novel gradient de-conflict algorithm named DR-MGF (Dynamic Routes and Meta-weighted Gradient Fusion). Different from existing de-conflict methods, DR-MGF achieves gradient de-conflict in MONs by learning task-preferred inference routes. The proposed method is motivated by our experimental findings that the shared filters are not equally important for different tasks. By designing learnable task-specific importance variables, DR-MGF evaluates the importance of filters for different tasks. Through making the dominance of tasks over filters proportional to the task-specific importance of filters, DR-MGF can effectively reduce inter-task interference. These task-specific importance variables ultimately determine task-preferred inference routes at the end of training iterations. Extensive experimental results on CIFAR, ImageNet, and NYUv2 demonstrate that DR-MGF outperforms existing de-conflict methods. Furthermore, DR-MGF can be extended to general MONs without modifying the overall network structures.
Xiaochang Hu, Xin Xu 0001, Jian Li 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 AdaPL: Adaptive Pseudo Labeling for deep active learning in image classification
Qiang Fang 0001, Xin Xu 0001
Pattern Recognit. Lett.2
2026 STAGE: Spatio-Temporal Aggregation via Graph Embedding for Multi-Agent Reinforcement Learning in Industrial Optimization
Chiqiang Liu, Dazi Li, Xin Xu 0001
IEEE Trans Autom. Sci. Eng.3
2026 CHTracker: Confidence-Guided Hierarchical Association Paradigm for Multi-Object Tracking
abstract
Multi-object tracking (MOT) has garnered considerable attention due to its relevance in practical applications such as automated devices in smart cities. However, under complex conditions, existing trackers often fail to accurately capture or characterize target motion patterns, exhibiting limitations in flexibility and interpretability. To address these challenges, this paper introduces CHTracker, a confidence-guided hierarchical association paradigm for MOT. By integrating spatial features with varying confidence levels, CHTracker enhances the granularity of motion pattern modeling in edge-case scenarios where conventional trackers are prone to association ambiguity. Our paradigm adaptively utilizes distinct tracking cues and assignment metrics tailored to hierarchical target structures, thereby enabling collaborative tracking. Additionally, CHTracker incorporates the diagonal length of the target bounding box as a state variable during position prediction, which significantly improves the robustness against diverse motion noise. Extensive experimental results on multiple benchmarks, including Dance-Track, MOT17, MOT20, and Singapore Maritime Dataset (SMD), demonstrate that CHTracker achieves the state-of-the-art performance in accuracy, robustness, and generalization. Furthermore, our association paradigm is extended to a visible-infrared fusion version for evaluation on the multimodal CAMEL dataset, underscoring its practical potential to fulfill heterogeneous modality requirements in real-world scenarios. Our code will be available at https://github.com/ZyanChenyang/CHTracker.
Chenyang Yan, Yueying Wang, Yuhao Qing, Weidong Zhang 0007, Xin Xu 0001
IEEE Trans Autom. Sci. Eng.5
2026 Token Calibration for Transformer-Based Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain by learning domain-invariant representations. Motivated by the recent success of Vision Transformers (ViTs), several UDA approaches have adopted ViT architectures to exploit fine-grained patch-level representations, which are unified as Transformer-based $D$ omain $A$ daptation (TransDA) independent of CNN-based. However, we have a key observation in TransDA: due to inherent domain shifts, patches (tokens) from different semantic categories across domains may exhibit abnormally high similarities, which can mislead the self-attention mechanism and degrade adaptation performance. To solve that, we propose a novel $P$ atch- $A$ daptation Transformer (PATrans), which first identifies similarity-anomalous patches and then adaptively suppresses their negative impact to domain alignment, i.e. token calibration. Specifically, we introduce a $P$ atch- $A$ daptation $A$ ttention (PAA) mechanism to replace the standard self-attention mechanism, which consists of a weight-shared triple-branch mixed attention mechanism and a patch-level domain discriminator. The mixed attention integrates self-attention and cross-attention to enhance intra-domain feature modeling and inter-domain similarity estimation. Meanwhile, the patch-level domain discriminator quantifies the anomaly probability of each patch, enabling dynamic reweighting to mitigate the impact of unreliable patch correspondences. Furthermore, we introduce a contrastive attention regularization strategy, which leverages category-level information in a contrastive learning framework to promote class-consistent attention distributions. Extensive experiments on four benchmark datasets demonstrate that PATrans attains significant improvements over existing state-of-the-art UDA methods (e.g., 89.2% on the VisDA-2017). Code is available at: https://github.com/YSY145/PATrans.
Xiaowei Fu, Shiyu Ye, Chenxu Zhang 0001, Fuxiang Huang, Xin Xu 0001, Lei Zhang 0038
IEEE Trans. Image Process.5
2026 Energy-Aware Collaborative AAV Target Tracking via Reinforcement Learning-Based Predictive Control With Asynchronous Policy Iteration
abstract
Autonomous aerial vehicle (AAV) target tracking technology is an essential component for enabling diverse low-altitude activities. Due to the constraints on energy and computing resources of AAVs, current approaches face challenges in balancing prolonged flight duration with precise tracking while avoiding high computational complexity. Therefore, this paper proposes an energy-aware formation control algorithm for multiple AAVs to cooperatively track a target while retaining a desired formation pattern. Firstly, to achieve a balanced outcome in terms of tracking performance and control effort, an actor-critic based learning predictive rule is explored to develop a near-optimal control protocol that stabilizes error dynamics and minimizes value functions for discrete-time AAV systems. By decomposing the infinite-horizon target tracking problem into a sequence of finite-horizon sub-problems, the reinforcement learning (RL)-based predictive control algorithm can achieve fast convergence in approximating the solution of Hamilton-Jacobi-Bellman (HJB) equation. Furthermore, by employing a delicately designed asynchronous policy iteration mechanism with adjustable learning intervals in RL, the cumbersome learning process can be effectively mitigated, thereby attaining both high learning efficiency and a reduced computational burden simultaneously. The involved errors are proven to be convergent and simulation results validate the optimality of our method.
Xiangwang Hou, Xin Xu 0001, Jingjing Wang 0001, Chunxiao Jiang, Dusit Niyato
IEEE Trans. Mob. Comput.4
2025 Skill Expansion and Composition in Parameter Space
abstract
Humans excel at reusing prior knowledge to address new challenges and developing skills while solving problems. This paradigm becomes increasingly popular in the development of autonomous agents, as it develops systems that can self-evolve in response to new challenges like human beings. However, previous methods suffer from limited training efficiency when expanding new skills and fail to fully leverage prior knowledge to facilitate new task learning. We propose Parametric Skill Expansion and Composition (PSEC), a new framework designed to iteratively evolve the agents' capabilities and efficiently address new challenges by maintaining a manageable skill library. This library can progressively integrate skill primitives as plug-and-play Low-Rank Adaptation (LoRA) modules in parameter-efficient finetuning, facilitating efficient and flexible skill expansion. This structure also enables the direct skill compositions in parameter space by merging LoRA modules that encode different skills, leveraging shared information across skills to effectively program new skills. Based on this, we propose a context-aware modular to dynamically activate different skills to collaboratively handle new tasks. Empowering diverse applications including multi-objective composition, dynamics shift, and continual policy shift, the results on D4RL, DSRL benchmarks, and the DeepMind Control Suite show that PSEC exhibits superior capacity to leverage prior knowledge to efficiently tackle new challenges, as well as expand its skill libraries to evolve the capabilities. Project website: https://ltlhuuu.github.io/PSEC/.
Tenglong Liu, Yinan Zheng, Yixing Lan, Xin Xu 0001, Xianyuan Zhan
ICLR6
2025 Can LLMs Solve Longer Math Word Problems Better?
abstract
Math Word Problems (MWPs) play a vital role in assessing the capabilities of Large Language Models (LLMs), yet current research primarily focuses on questions with concise contexts. The impact of longer contexts on mathematical reasoning remains under-explored. This study pioneers the investigation of Context Length Generalizability (CoLeG), which refers to the ability of LLMs to solve MWPs with extended narratives. We introduce Extended Grade-School Math (E-GSM), a collection of MWPs featuring lengthy narratives, and propose two novel metrics to evaluate the efficacy and resilience of LLMs in tackling these problems. Our analysis of existing zero-shot prompting techniques with proprietary LLMs along with open-source LLMs reveals a general deficiency in CoLeG. To alleviate these issues, we propose tailored approaches for different categories of LLMs. For proprietary LLMs, we introduce a new instructional prompt designed to mitigate the impact of long contexts. For open-source LLMs, we develop a novel auxiliary task for fine-tuning to enhance CoLeG. Our comprehensive results demonstrate the effectiveness of our proposed methods, showing improved performance on E-GSM. Additionally, we conduct an in-depth analysis to differentiate the effects of semantic understanding and reasoning efficacy, showing that our methods improves the latter. We also establish the generalizability of our methods across several other MWP benchmarks. Our findings highlight the limitations of current LLMs and offer practical solutions correspondingly, paving the way for further exploration of model generalizability and training methodologies.
Xin Xu 0001, Tong Xiao 0004, Zitong Chao, Zhenya Huang, Can Yang 0002, Yang Wang 0020
ICLR1
2025 UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in solving complex reasoning tasks, particularly in mathematics. However, the domain of physics reasoning presents unique challenges that have received significantly less attention. Existing benchmarks often fall short in evaluating LLMs’ abilities on the breadth and depth of undergraduate-level physics, underscoring the need for a comprehensive evaluation. To fill this gap, we introduce UGPhysics, a large-scale and diverse benchmark specifically designed to evaluate **U**nder**G**raduate-level **Physics** (**UGPhysics**) reasoning with LLMs. UGPhysics includes 5,520 undergraduate-level physics problems in both English and Chinese across 13 subjects with seven different answer types and four distinct physics reasoning skills, all rigorously screened for data leakage. Additionally, we develop a Model-Assistant Rule-based Judgment (**MARJ**) pipeline specifically tailored for assessing physics problems, ensuring accurate evaluation. Our evaluation of 31 leading LLMs shows that the highest overall accuracy, 49.8% (achieved by OpenAI-o1-mini), emphasizes the need for models with stronger physics reasoning skills, beyond math abilities. We hope UGPhysics, along with MARJ, will drive future advancements in AI for physics reasoning. Codes and data are available at \href{https://github.com/YangLabHKUST/UGPhysics}{https://github.com/YangLabHKUST/UGPhysics}.
Xin Xu 0001, Qiyun Xu, Tong Xiao 0004, Tianhao Chen, Shizhe Diao, Can Yang 0002, Yang Wang 0020
ICML1
2025 Versatile Distributed Maneuvering With Generalized Formations Using Guiding Vector Fields
abstract
This paper presents a unified approach to realize versatile distributed maneuvering with generalized formations. Specifically, we decompose the robots' maneuvers into two independent components, i.e., interception and enclosing, which are parameterized by two independent virtual coordinates. Treating these two virtual coordinates as dimensions of an abstract manifold, we derive the corresponding singularity-free guiding vector field (GVF), which, along with a distributed coordination mechanism based on the consensus theory, guides robots to achieve various motions (i.e., versatile maneuvering), including (a) formation tracking, (b) target enclosing, and (c) circumnavigation. Additional motion parameters can generate more complex cooperative robot motions. Based on GVFs, we design a controller for a nonholonomic robot model. Besides the theoretical results, extensive simulations and experiments are performed to validate the effectiveness of the approach.
Sha Luo, Pengming Zhu, Weijia Yao, Héctor García de Marina, Xin Xu 0001
ICRA7
2025 Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
abstract
In offline reinforcement learning, value overestimation caused by out-of-distribution (OOD) actions significantly limits policy performance. Recently, diffusion models have been leveraged for their strong distribution-matching capabilities, enforcing conservatism through behavior policy constraints. However, existing methods often apply indiscriminate regularization to redundant actions in low-quality datasets, resulting in excessive conservatism and an imbalance between the expressiveness and efficiency of diffusion modeling. To address these issues, we propose DIffusion policies with Value-conditional Optimization (DIVO), a novel approach that leverages diffusion models to generate high-quality, broadly covered in-distribution state-action samples while facilitating efficient policy improvement. Specifically, DIVO introduces a binary-weighted mechanism that utilizes the advantage values of actions in the offline dataset to guide diffusion model training. This enables a more precise alignment with the dataset’s distribution while selectively expanding the boundaries of high-advantage actions. During policy improvement, DIVO dynamically filters high-return-potential actions from the diffusion model, effectively guiding the learned policy toward better performance. This approach achieves a critical balance between conservatism and explorability in offline RL. We evaluate DIVO on the D4RL benchmark and compare it against state-of-the-art baselines. Empirical results demonstrate that DIVO achieves superior performance, delivering significant improvements in average returns across locomotion tasks and outperforming existing methods in the challenging AntMaze domain, where sparse rewards pose a major difficulty.
Yunchang Ma, Tenglong Liu, Yixing Lan, Changxin Zhang, Xin Xu 0001
IROS7
2025 Learning Predictive Control with Online Modeling for Agile Maneuvering of Autonomous Vehicles
abstract
The agile maneuvering control of autonomous vehicles (AVs) requires the tracking of reference trajectories characterized by high acceleration, sharp curvature, considerable disturbances, and significant time-varying, all while ensuring stability and accuracy. The inherent uncertainty and time-varying nature of both the vehicle model and its environment pose significant challenges to achieving high-performance tracking during agile maneuvers. Developing a control algorithm that enables solving the optimal policy for nonlinear systems with uncertainties is critical. In this paper, we propose a learning-based predictive control approach, namely, an adaptive model predictive control (AMPC) with Actor-Critic Learning (ACL) for generating closed-loop MPC policies for agile maneuvering of AVs. The proposed approach leverages neural networks to model the dynamics uncertainties online. The control policy and model are updated simultaneously to realize performance op-timization under time-varying uncertainties. Simulation results demonstrate that our proposed algorithm outperforms other leading ACL methods, as well as MPC and Linear Quadratic Regulator (LQR). Furthermore, field test experiment results validate its effectiveness on the HongQi-EHS3 electric vehicle, showing superior control performance compared to MPC both on paved roads and curved off-roads with excellent stability performance.
Zengyi Zhang, Tenglong Liu, Yixing Lan, Xin Xu 0001
IROS6
2025 GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
abstract
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002
NeurIPS2
2025 Adaptive generative adversarial maximum entropy inverse reinforcement learning
Li Song 0003, Dazi Li, Xin Xu 0001
Inf. Sci.3
2025 Denser Teacher: Rethinking Dense Pseudo-Label for Semi-Supervised Oriented Object Detection
abstract
Oriented object detection, which aims to detect multi-oriented objects, is a fundamental task for visual analysis in complex scenarios, such as aerial images. However, powerful detection performance relies on abundant and accurate annotations. Therefore, semi-supervised oriented object detection, which utilizes unlabeled data to improve performance, is a promising method to address this problem. In this work, we explore Dense Pseudo-Label (DPL), which directly selects pseudo labels from the original output of the teacher model without any complicated post-processing steps, and expose the shortcomings of existing methods. Through analysis, we identify that the imbalance between obtaining potential positive samples and removing the interference of inaccurate pseudo labels hinders the effectiveness of DPL. To further improve DPL efficiency, we propose Denser Teacher, a new semi-supervised oriented object detection method. In this method, we design a simple yet effective adaptive mechanism called global dynamic k estimation to guide the selection of DPLs in densely-distributed scenes. Additionally, to improve scale adaptation, we introduce dense multi-scale learning for DPL, where DPLs from different scales are utilized to bridge the scale gap. We conduct extensive experiments on several benchmarks to demonstrate the effectiveness of our proposed method in leveraging unlabeled data for performance improvement. Our code will be available athttps://github.com/Haru-zt/DenserTeacher.
Qiang Fang 0001, Xin Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multiscale Gaussian Attention Mechanism for Tiny-Object Detection in Remote Sensing Images
abstract
Tiny object detection is increasingly crucial in the fields such as remote sensing, traffic monitoring, and robotics. Inspired by human visual perception, attention mechanism has become a widely used method for enhancing object detection performance. While existing attention mechanisms have significantly advanced general object detection performance, they often fall short in adapting to the characteristics in tiny object datasets, including huge object size variations and concentrated distributions. In detailed, most current attention mechanisms rely on convolutional or linear layers with fixed receptive fields to compute attention vectors. Some methods attempt to enlarge the receptive fields by using multiscale structures, but they often simply sum feature maps, leading to information interference and increased computational costs. To address these issues, we propose a novel Multiscale Gaussian Attention Mechanism (MGAM). This mechanism integrates multiscale receptive fields with dynamic feature weighting and a Gaussian attention module, replacing traditional convolutional layers to reduce training and inference overhead. In additional, our mechanism can be easily embedded into various detectors without any hyperparameters. Extensive experiments on six object detection datasets demonstrate the effectiveness and robustness of our method. Code is available at: https://github.com/cszzshi/MGAM.
Shuohao Shi, Qiang Fang 0001, Xin Xu 0001, Dezun Dong
IEEE Trans. Geosci. Remote. Sens.3
2025 Receding-Horizon Reinforcement Learning for Time-Delayed Human-Machine Shared Control of Intelligent Vehicles
abstract
Human–machine shared control has recently been regarded as a promising paradigm to improve safety and performance in complex driving scenarios. One crucial task in shared control is dynamically optimizing the driving weights between the driver and the intelligent vehicle to adapt to dynamic driving scenarios. However, designing an optimal human–machine shared controller with guaranteed performance and stability is challenging due to nonnegligible time delays caused by communication protocols and uncertainties in driver behavior. This article proposes a novel receding-horizon reinforcement learning approach for time-delayed human–machine shared control of intelligent vehicles. First, we build a multikernel-based data-driven model of vehicle dynamics and driving behavior, considering time delays and uncertainties of drivers' actions. Second, a model-based receding horizon actor–critic learning algorithm is presented to learn an explicit policy for time-delayed human–machine shared control online. Unlike classic reinforcement learning, policy learning of the proposed approach is performed according to a receding-horizon strategy to enhance learning efficiency and adaptability. In theory, the closed-loop stability under time delays is analyzed. Hardware-in-the-loop experiments on the time-delayed human–machine shared control of intelligent vehicles have been conducted in variable curvature road scenarios. The results demonstrate that our approach has significant improvements in driving performance and driver workload compared with pure manual driving and previous shared control methods.
Xinxin Yao, Xin Xu 0001
IEEE Trans. Hum. Mach. Syst.4
2025 How to Enhance the Interpretability of Learning-Based Motion Planning for Intelligent Vehicles - A Survey
abstract
With the advancement of deep learning, the learning-based motion planning (MP) approach exhibits immense potential in intelligent vehicles (IVs). Because the principle and framework of the learning-based MP method differ from the traditional MP methods, exploring effective strategies to enhance interpretability plays an important role. This survey fills the gaps in the IV field’s learning-based motion planning and interpretability enhancement. Our study aims to explore two fundamental inquiries. Firstly, how can we design learning-based MP to achieve high performance? Secondly, how can we enhance the interpretability of learning-based MP? To this end, this paper provides an extensive overview of more than 200 papers employed in learning-based MP techniques within the last 10 years. By summarizing these techniques, a taxonomy for integrating learning-based MP techniques into an IV architecture is presented as three modes: learning-based key-module generator, learning-based trajectory generator, and learning-based policy generator. Interpretability enhancement has different considerations for different modes. Additionally, we compile a summary of resources utilized in learning-based MP. Finally, we discuss critical challenges and make suggestions.
Tao Wu 0001, Huijing Zhao, Xin Xu 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Multi-Kernel Enhanced Receding-Horizon Reinforcement Learning for Steering Control of Intelligent Vehicles
abstract
Achieving optimal control in the path-tracking of intelligent vehicles is crucial for enhancing driving performance, yet it remains challenging due to model uncertainties and nonlinear dynamics. Reinforcement learning (RL), as a class of approximated optimal control methods, has gained attention for tackling complex problems with uncertain or nonlinear dynamics. However, effective feature representation and online learning efficiency are two major issues that persist in RL methods for adaptive optimal control. To address these challenges, this paper proposes a Receding-horizon Multi-kernel Reinforcement Learning (RM-RL) algorithm, which integrates an efficient online learning mechanism with improved feature representations. RM-RL operates within a receding-horizon control framework that facilitates online policy optimization and deployment. Meanwhile, employing a multi-kernel features learning approach for the actor-critic structure and the dynamics model further improves learning efficiency and generalization. Besides, the convergence and closed-loop stability are analyzed in depth. Real-world experiments conducted on the Hongqi E-HS3 vehicle and simulations on the CarSim platform demonstrate the effectiveness and the superiority of RM-RL over advanced comparative methods.
Changxin Zhang, Xin Xu 0001, Ruizhuo Wu
IEEE Trans. Intell. Transp. Syst.4
2025 A Unified and Quality-Guaranteed Approach for Dubins Vehicle Path Planning With Obstacle Avoidance and Curvature Constraint
abstract
Robotic technologies and applications have recently witnessed remarkable advancements. A major challenge is the shortest-path planning problem of a curvature-bounded vehicle from the known starting configuration to visit a target point and finally return to the starting configuration in an obstacle environment. Spurred by this significant issue in robotic surveillance and patrolling applications, this paper proposed the Two-trip Obstacle-environment Relaxed Dubins Problem (TORDP). In TORDP, the vehicle’s target-visiting heading is a critical variable. Analytical approaches have existed for simpler scenarios than TORDP. However, these approaches are unavailable when solving the complex TORDP simultaneously with bounded curvature, variable target heading and unified ability to tackle with- or without- obstacle cases. Hence, we develop the mixed-integer piecewise-linear program (MIPWLP) approach, making the otherwise intractable complex scenario unifiedly solved with guaranteed good quality. Extensive experiments demonstrate that the proposed approach demonstrates effective performance. Furthermore, the objective approximation error in some cases was analyzed to achieve a length near the optimal length within$h^{2}/(2\sqrt {2})$tolerance wherehis the approximation piece length. The proposed MIPWLP approach could also offer a generalizable optimization framework for broader robotic path-planning applications in constrained environments.
Xing Zhou 0004, Lin Li 0075, Hao Gao 0014, Kangxing Yao, Xin Xu 0001
IEEE Trans. Intell. Transp. Syst.6
2025 Enhancing Graph Reconstruction: Uniting Dual-Level Graph Structure With Graph Reinforcement Learning
abstract
A combinatorial optimization problem is typically regarded as a 1-D sorting problem in most existing research. The representation ignores some information about the problem because of dimension compression. When applying reinforcement learning (RL) to this problem, convolutional neural networks (CNNs) used in conventional RL cannot directly extract the connection information between two elements in the feature matrix. A typical class of combinatorial optimization problems, the job shop scheduling problem (JSSP), is used in this article as an example. Considering the limitations in previous research, this article reexamines the task from the perspective of graph reconstruction and proposes a graph RL (GRL) method that combines a double deep Q-network (DDQN) and graph attention network (GAT) to achieve breakthroughs beyond the constraints of CNN performance. Moreover, a dual-level graph representation structure is constructed to comprehensively learn the features of scheduling information and overcome the difficulty of learning dynamic graphs. Experiments show that the quality of the obtained solution and generalization performance are both improved compared with models based on original deep RL (DRL) algorithms.
Dazi Li, Yanyang Bao, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Improved Dual Correlation Reduction Network With Affinity Recovery
abstract
Deep graph clustering, which aims to reveal the underlying graph structure and divide the nodes into different clusters without human annotations, is a fundamental yet challenging task. However, we observe that the existing methods suffer from the representation collapse problem and tend to encode samples with different classes into the same latent embedding. Consequently, the discriminative capability of nodes is limited, resulting in suboptimal clustering performance. To address this problem, we propose a novel deep graph clustering algorithm termed improved dual correlation reduction network (IDCRN) through improving the discriminative capability of samples. Specifically, by approximating the cross-view feature correlation matrix to an identity matrix, we reduce the redundancy between different dimensions of features, thus improving the discriminative capability of the latent space explicitly. Meanwhile, the cross-view sample correlation matrix is forced to approximate the designed clustering-refined adjacency matrix to guide the learned latent representation to recover the affinity matrix even across views, thus enhancing the discriminative capability of features implicitly. Moreover, we avoid the collapsed representation caused by the oversmoothing issue in graph convolutional networks (GCNs) through an introduced propagation regularization term, enabling IDCRN to capture the long-range information with the shallow network structure. Extensive experimental results on six benchmarks have demonstrated the effectiveness and efficiency of IDCRN compared with the existing state-of-the-art deep graph clustering algorithms. The code of IDCRN is released at IDCRN. Besides, we share a collection of deep graph clustering, including papers, codes, and datasets at ADGC.
Yue Liu 0008, Sihang Zhou 0001, Xihong Yang, Xinwang Liu 0002, Wenxuan Tu, Liang Li 0041, Xin Xu 0001, Fuchun Sun 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 Game-Theoretic Constrained Policy Optimization for Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) aims to optimize the task performance with safety guarantees. One common modeling scheme to study safe RL problems is the constrained Markov decision process (CMDP). However, current safe RL methods within the CMDP framework face challenges in tradeoffs among various objectives and gradient conflicts of policy updating. To cope with these challenges, this article presents a novel safe RL approach called game-theoretic constrained policy optimization (GCPO). The proposed approach formulates the CMDP problem as a general-sum Markov game with multiple players, where a task player seeks to maximize the reward objective, while constraint players aim to minimize constraint objectives until they are fulfilled. By doing so, GCPO adopts the learning mode with multiple subpolicies, each aligned with a distinct objective, that collectively constitute the overall behavior of the agent. The learning convergence of the GCPO can be ensured with the contraction mapping to the Nash equilibrium. Furthermore, a novel dominant timescale update rule is presented for multiplayer policy learning to guarantee constraint satisfaction. The learning convergence and constraint satisfaction of GCPO are theoretically analyzed. Consequently, GCPO eliminates the necessity of tuning tradeoff parameters and mitigates gradient conflicts during multiobjective policy updating. Experimental results show that GCPO outperforms state-of-the-art safe RL algorithms in a quadrotor trajectory tracking task and various high-dimensional robot locomotion benchmarks. Moreover, GCPO exhibits robustness to diverse scales of task rewards and constraint costs without the need for intricate tradeoffs.
Changxin Zhang, Yixing Lan, Hao Gao 0014, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC
abstract
Distributed model predictive control (DMPC) is promising in achieving optimal cooperative control in multirobot systems (MRS). However, real-time DMPC implementation relies on numerical optimization tools to periodically calculate local control sequences online. This process is computationally demanding and lacks scalability for large-scale, nonlinear MRS. This article proposes a novel distributed learning-based predictive control framework for scalable multirobot control. Unlike conventional DMPC methods that calculate open-loop control sequences, our approach centers around a computationally fast and efficient distributed policy learning algorithm that generates explicit closed-loop DMPC policies for MRS without using numerical solvers. The policy learning is executed incrementally and forward in time in each prediction interval through an online distributed actor–critic implementation. The control policies are successively updated in a receding-horizon manner, enabling fast and efficient policy learning with the closed-loop stability guarantee. The learned control policies could be deployed online to MRS with varying robot scales, enhancing scalability and transferability for large-scale MRS. Furthermore, we extend our methodology to address the multirobot safe learning challenge through a force field-inspired policy learning approach. We validate our approach's effectiveness, scalability, and efficiency through extensive experiments on cooperative tasks of large-scale wheeled robots and multirotor drones. Our results demonstrate the rapid learning and deployment of DMPC policies for MRS with scales up to 10 000 units.
Wei Pan 0004, Cong Li 0015, Xin Xu 0001, Xiangke Wang, Dewen Hu
IEEE Trans. Robotics4
2024 Density-Guided Dense Pseudo Label Selection for Semi-Supervised Oriented Object Detection
abstract
Recently, dense pseudo-label, which directly selects pseudo labels from the original output of the teacher model without any complicated post-processing steps, has received considerable attention in semi-supervised object detection (SSOD). However, for the multi-oriented and dense objects that are common in aerial scenes, existing dense pseudolabel selection methods are inefficient because they ignore the significant density difference. Therefore, we propose Density-Guided Dense Pseudo Label Selection (DDPLS) for semi-supervised oriented object detection. In DDPLS, we design a simple but effective adaptive mechanism to guide the selection of dense pseudo labels. Specifically, we propose the Pseudo Density Score (PDS) to estimate the density of potential objects and use this score to select reliable dense pseudo labels. On the DOTA-v1.5 benchmark, the proposed method outperforms previous methods especially when labeled data are scarce. For example, it achieves 49.78 mAP given only 5% of annotated data, which surpasses previous state-of-the-art method given 10% of annotated data by 1.15 mAP. Our codes is available at https://github.com/Haru-zt/DDPLS.
Qiang Fang 0001, Shuohao Shi, Xin Xu 0001
ICIP4
2024 Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
abstract
In offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy regularization. However, these methods often suffer from the issue of unnecessary conservativeness, hampering policy improvement. This occurs due to the indiscriminate use of all actions from the behavior policy that generates the offline dataset as constraints. The problem becomes particularly noticeable when the quality of the dataset is suboptimal. Thus, we propose Adaptive Advantage-guided Policy Regularization (A2PR), obtaining high-advantage actions from an augmented behavior policy combined with VAE to guide the learned policy. A2PR can select high-advantage actions that differ from those present in the dataset, while still effectively maintaining conservatism from OOD actions. This is achieved by harnessing the VAE capacity to generate samples matching the distribution of the data points. We theoretically prove that the improvement of the behavior policy is guaranteed. Besides, it effectively mitigates value overestimation with a bounded performance gap. Empirically, we conduct a series of experiments on the D4RL benchmark, where A2PR demonstrates state-of-the-art performance. Furthermore, experimental results on additional suboptimal mixed datasets reveal that A2PR exhibits superior performance. Code is available at https://github.com/ltlhuuu/A2PR.
Tenglong Liu, Yang Li 0116, Yixing Lan, Hao Gao 0014, Wei Pan 0004, Xin Xu 0001
ICML6
2024 HDKI: A Hierarchical Deep Koopman Framework for Spatio-Temporal Prediction with Image Observations
Haibin Xie, Junheng Liu, Wei Jiang 0006, Xin Xu 0001
ICONIP (7)6
2024 Similarity Distance-Based Label Assignment for Tiny Object Detection
abstract
Tiny object detection is becoming one of the most challenging tasks in computer vision because of the limited object size and lack of information. The label assignment strategy is a key factor affecting the accuracy of object detection. Although there are some effective label assignment strategies for tiny objects, most of them focus on reducing the sensitivity to the bounding boxes to increase the number of positive samples and have some fixed hyperparameters need to set. However, more positive samples may not necessarily lead to better detection results, in fact, excessive positive samples may lead to more false positives. In this paper, we introduce a simple but effective strategy named the Similarity Distance (SimD) to evaluate the similarity between bounding boxes. This proposed strategy not only considers both location and shape similarity but also learns hyperparameters adaptively, ensuring that it can adapt to different datasets and various object sizes in a dataset. Our approach can be simply applied in common anchor-based detectors in place of the IoU for label assignment and Non Maximum Suppression (NMS). Extensive experiments on four mainstream tiny object detection datasets demonstrate superior performance of our method, especially, 1.8 AP points and 4.1 AP points of very tiny higher than the state-of-the-art competitors on AI-TOD. Code is available at: https://github.com/cszzshi/simd.
Shuohao Shi, Qiang Fang 0001, Xin Xu 0001
IROS3
2024 M3-GMN: A Multi-environment, Multi-LiDAR, Multi-task dataset for Grid Map based Navigation
abstract
In this paper, we propose a multi-environment, multi-LiDAR, multi-task dataset to promote the grid map-based navigation capability for autonomous vehicles. The dataset comprises structured and unstructured environmental data captured by different types of LiDAR and contains various challenging scenarios, including moving objects, negative obstacles, steep slopes, cliffs, overhangs, etc. Further, we have devised an innovative method for generating ground truth, facilitating the creation of dense, accurate, and stable grid maps with a minimal requirement for human annotation efforts. A new baseline method and two existing approaches are evaluated on this dataset. Results indicate that existing approaches perform much worse than the proposed baseline. The dataset will be made publicly available at https://github.com/guanglei96/M3-GMN.
Guanglei Xie, Hao Fu 0001, Hanzhang Xue, Bokai Liu, Xin Xu 0001, Xiaohui Li 0007, Zhenping Sun
IROS5
2024 Deep Learning for Visual Speech Analysis: A Survey
abstract
Visual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper will present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions.
Changchong Sheng, Gangyao Kuang, Liang Bai 0003, Chenping Hou, Yulan Guo, Xin Xu 0001, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Efficient Reinforcement Learning With the Novel N-Step Method and V-Network
abstract
The application of reinforcement learning (RL) in artificial intelligence has become increasingly widespread. However, its drawbacks are also apparent, as it requires a large number of samples for support, making the enhancement of sample efficiency a research focus. To address this issue, we propose a novel N-step method. This method extends the horizon of the agent, enabling it to acquire more long-term effective information, thus resolving the issue of data inefficiency in RL. Additionally, this N-step method can reduce the estimation variance of Q-function, which is one of the factors contributing to estimation errors in Q-function estimation. Apart from high variance, estimation bias in Q-function estimation is another factor leading to estimation errors. To mitigate the estimation bias of Q-function, we design a regularization method based on the V-function, which has been underexplored. The combination of these two methods perfectly addresses the problems of low sample efficiency and inaccurate Q-function estimation in RL. Finally, extensive experiments conducted in discrete and continuous action spaces demonstrate that the proposed novel N-step method, when combined with classical deep Q-network, deep deterministic policy gradient, and TD3 algorithms, is effective, consistently outperforming the classical algorithms.
Miaomiao Zhang 0001, Shuo Zhang 0023, Zhiyi Shi, Xiangyang Deng, Qi Wu 0003, Xin Xu 0001
IEEE Trans. Cybern.7
2024 Robust Depth Estimation Based on Parallax Attention for Aerial Scene Perception
abstract
Given the precalibrated image pairs, stereo matching aims to infer the scene depth information in real-time, which has important research value in the fields of high-precision 3-D reconstruction of the Earth’s surface, automatic driving and unmanned aerial vehicle (UAV) navigation. The cost volume-based stereo matching method adopts a coarse-to-fine manner to construct cascaded cost volume, and applies 3-D convolution to capture the correspondence of feature matching to infer the disparity map, which achieves comparable performance. However, the existing method has difficulty dealing with jitter regions with disparity change, and direct disparity regression easily leads to overfitting of cost volume regularization. To alleviate the above two problems, this work proposes an end-to-end disparity estimation network based on Transformer. Its specific improvements are as follows. 1) The cross-view feature interaction module based on Transformer is introduced to realize the feature interaction of global context information. 2) A parallax attention mechanism is designed to impose global geometric constraints on the epipolar line to improve the reliability of feature matching. 3) Focal loss is applied for the training of the disparity classification model to emphasize one-hot supervision in ambiguous regions. Comprehensive experiments on public datasets Sceneflow, KITTI2015, ETH3D, and aerial WHU datasets validate that the proposed work can effectively enhance the performance of disparity estimation.
Kevin W. Tong, Miaomiao Zhang 0001, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003
IEEE Trans. Ind. Informatics4
2024 AMARL: An Attention-Based Multiagent Reinforcement Learning Approach to the Min-Max Multiple Traveling Salesmen Problem
abstract
In recent years, the multiple traveling salesmen problem (MTSP or multiple TSP) has received increasing research interest and one of its main applications is coordinated multirobot mission planning, such as cooperative search and rescue tasks. However, it is still challenging to solve MTSP with improved inference efficiency as well as solution quality in varying situations, e.g., different city positions, different numbers of cities, or agents. In this article, we propose an attention-based multiagent reinforcement learning (AMARL) approach, which is based on the gated transformer feature representations for min-max multiple TSPs. The state feature extraction network in our proposed approach adopts the gated transformer architecture with reordering layer normalization (LN) and a new gate mechanism. It aggregates fixed-dimensional attention-based state features irrespective of the number of agents and cities. The action space of our proposed approach is designed to decouple the interaction of agents' simultaneous decision-making. At each time step, only one agent is assigned to a non-zero action so that the action selection strategy can be transferred across tasks with different numbers of agents and cities. Extensive experiments on min-max multiple TSPs were conducted to illustrate the effectiveness and advantages of the proposed approach. Compared with six representative algorithms, our proposed approach achieves state-of-the-art performance in solution quality and inference efficiency. In particular, the proposed approach is suitable for tasks with different numbers of agents or cities without extra learning, and experimental results demonstrate that the proposed approach realizes powerful transfer capability across tasks.
Hao Gao 0014, Xing Zhou 0004, Xin Xu 0001, Yixing Lan, Yongqian Xiao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Sample Efficient Deep Reinforcement Learning With Online State Abstraction and Causal Transformer Model Prediction
abstract
Deep reinforcement learning (RL) typically requires a tremendous number of training samples, which are not practical in many applications. State abstraction and world models are two promising approaches for improving sample efficiency in deep RL. However, both state abstraction and world models may degrade the learning performance. In this article, we propose an abstracted model-based policy learning (AMPL) algorithm, which improves the sample efficiency of deep RL. In AMPL, a novel state abstraction method via multistep bisimulation is first developed to learn task-related latent state spaces. Hence, the original Markov decision processes (MDPs) are compressed into abstracted MDPs. Then, a causal transformer model predictor (CTMP) is designed to approximate the abstracted MDPs and generate long-horizon simulated trajectories with a smaller multistep prediction error. Policies are efficiently learned through these trajectories within the abstracted MDPs via a modified multistep soft actor-critic algorithm with a λ -target. Moreover, theoretical analysis shows that the AMPL algorithm can improve sample efficiency during the training process. On Atari games and the DeepMind Control (DMControl) suite, AMPL surpasses current state-of-the-art deep RL algorithms in terms of sample efficiency. Furthermore, DMControl tasks with moving noises are conducted, and the results demonstrate that AMPL is robust to task-irrelevant observational distractors and significantly outperforms the existing approaches.
Yixing Lan, Xin Xu 0001, Qiang Fang 0001, Jianye Hao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Deep Reinforcement Learning: A Survey
abstract
Deep reinforcement learning (DRL) integrates the feature representation ability of deep learning with the decision-making ability of reinforcement learning so that it can achieve powerful end-to-end learning control capabilities. In the past decade, DRL has made substantial advances in many tasks that require perceiving high-dimensional input and making optimal or near-optimal decisions. However, there are still many challenging problems in the theory and applications of DRL, especially in learning control tasks with limited samples, sparse rewards, and multiple agents. Researchers have proposed various solutions and new theories to solve these problems and promote the development of DRL. In addition, deep learning has stimulated the further development of many subfields of reinforcement learning, such as hierarchical reinforcement learning (HRL), multiagent reinforcement learning, and imitation learning. This article gives a comprehensive overview of the fundamental theories, key algorithms, and primary research domains of DRL. In addition to value-based and policy-based DRL algorithms, the advances in maximum entropy-based DRL are summarized. The future research topics of DRL are also analyzed and discussed.
Xu Wang 0043, Xingxing Liang, Dawei Zhao 0003, Jincai Huang 0001, Xin Xu 0001, Bin Dai 0001, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.6
2024 Collision-Avoiding Flocking With Multiple Fixed-Wing UAVs in Obstacle-Cluttered Environments: A Task-Specific Curriculum- Based MADRL Approach
abstract
Multiple unmanned aerial vehicles (UAVs) are able to efficiently accomplish a variety of tasks in complex scenarios. However, developing a collision-avoiding flocking policy for multiple fixed-wing UAVs is still challenging, especially in obstacle-cluttered environments. In this article, we propose a novel curriculum-based multiagent deep reinforcement learning (MADRL) approach called task-specific curriculum-based MADRL (TSCAL) to learn the decentralized flocking with obstacle avoidance policy for multiple fixed-wing UAVs. The core idea is to decompose the collision-avoiding flocking task into multiple subtasks and progressively increase the number of subtasks to be solved in a staged manner. Meanwhile, TSCAL iteratively alternates between the procedures of online learning and offline transfer. For online learning, we propose a hierarchical recurrent attention multiagent actor-critic (HRAMA) algorithm to learn the policies for the corresponding subtask(s) in each learning stage. For offline transfer, we develop two transfer mechanisms, i.e., model reload and buffer reuse, to transfer knowledge between two neighboring stages. A series of numerical simulations demonstrate the significant advantages of TSCAL in terms of policy optimality, sample efficiency, and learning stability. Finally, the high-fidelity hardware-in-the-loop (HITL) simulation is conducted to verify the adaptability of TSCAL. A video about the numerical and HITL simulations is available at https://youtu.be/R9yLJNYRIqY.
Chang Wang 0005, Xiaojia Xiang, Huat Kin Low, Xiangke Wang, Xin Xu 0001, Lincheng Shen
IEEE Trans. Neural Networks Learn. Syst.6
2024 Efficient Incremental Offline Reinforcement Learning With Sparse Broad Critic Approximation
abstract
Offline reinforcement learning (ORL) has been getting increasing attention in robot learning, benefiting from its ability to avoid hazardous exploration and learn policies directly from precollected samples. Approximate policy iteration (API) is one of the most commonly investigated ORL approaches in robotics, due to its linear representation of policies, which makes it fairly transparent in both theoretical and engineering analysis. One open problem of API is how to design efficient and effective basis functions. The broad learning system (BLS) has been extensively studied in supervised and unsupervised learning in various applications. However, few investigations have been conducted on ORL. In this article, a novel incremental ORL approach with sparse broad critic approximation (BORL) is proposed with the advantages of BLS, which approximates the critic function in a linear manner with randomly projected sparse and compact features and dynamically expands its broad structure. The BORL is the first extension of API with BLS in the field of robotics and ORL. The approximation ability and convergence performance of BORL are also analyzed. Comprehensive simulation studies are then conducted on two benchmarks, and the results demonstrate that the proposed BORL can obtain comparable or better performance than conventional API methods without laborious hyperparameter fine-tuning work. To further demonstrate the effectiveness of BORL in practical robotic applications, a variable force tracking problem in robotic ultrasound scanning (RUSS) is investigated, and a learning-based adaptive impedance control (LAIC) algorithm is proposed based on BORL. The experimental results demonstrate the advantages of LAIC compared with conventional force tracking methods.
Baoliang Zhao, Xin Xu 0001, Ziwen Wang 0002, Pak-Kin Wong 0001, Ying Hu 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 SQIX: QMIX Algorithm Activated by General Softmax Operator for Cooperative Multiagent Reinforcement Learning
abstract
Multiagent cooperative systems can be used to conceptualize many real-world problems. Reinforcement learning is a particularly effective tool. The issue of bias in$Q$-function value estimation in single-agent reinforcement learning has garnered a lot of interest and substantial study. Indeed, this challenge endures in multiagent reinforcement learning, primarily owing to the inclusion of maximization operations. The crux of the matter lies in the inability to seamlessly extrapolate single-agent reinforcement learning algorithms to their multiagent counterparts. In this article, we introduce a more encompassing and straightforward principle: the notion of appropriate value correction. We suggest replacing the maximization operation with a monotonically nondecreasing function to obtain more accurate value estimates. We theoretically demonstrate that this operation effectively reduces the potential overestimation bias in the QMIX algorithm. Ultimately, our methodology, dubbed the SMIX algorithm—a fusion of the QMIX algorithm empowered by the Softmax operator, attains state-of-the-art outcomes across diverse multiagent cooperative tasks. This success extends to challenging domains such as StarCraft II, marking it as one of the most formidable games to date.
Miaomiao Zhang 0001, Kevin W. Tong, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003
IEEE Trans. Syst. Man Cybern. Syst.4
2023 DDK: A Deep Koopman Approach for Longitudinal and Lateral Control of Autonomous Ground Vehicles
abstract
Autonomous driving has attracted lots of attention in recent years. For some tasks, e.g., trajectory prediction, motion planning, and trajectory tracking, an accurate vehicle model can reduce the difficulty of these tasks and improve task completion performance. Prior works focused on parameter estimation of physical models or modeling nonlinear dynamics using neural networks. Still, these methods rely on internal parameters of vehicles or are not friendly for control due to the strong nonlinearity of models. This paper proposes a data-driven method to approximate vehicle dynamics based on the Koopman operator. The resulting model is an interpretable linear time-invariant model, facilitating controller design and solving related optimization problems. In the proposed approach, the state transition matrix is constructed based on the learned Koopman eigenvalues, while the input matrix is trained as a tensor. Based on the resulting model, a linear model predictive controller is designed to implement coupled longitudinal and lateral trajectory tracking. Simulations and experiments, including vehicle dynamics modeling and coupled longitudinal and lateral trajectory tracking, are performed in a high-fidelity CarSim environment and a real vehicle platform. An oil-driven D-Class SUV is selected in the simulation, while a real electric SUV is utilized in the experiment. Simulation and experiment results illustrate that the model of the nonlinear vehicle dynamics can be identified effectively via the proposed method, and high-quality trajectory tracking performance can be obtained with the resulting model.
Yongqian Xiao, Xin Xu 0001
ICRA3
2023 Kernel-based multiagent reinforcement learning for near-optimal formation control of mobile robots
Xin Xu 0001, Quan Xiong, Qingwen Ma, Yaoqian Peng
Appl. Intell.2
2023 Learning to Detect 3D Symmetry From Single-View RGB-D Images With Weak Supervision
abstract
3D symmetry detection is a fundamental problem in computer vision and graphics. Most prior works detect symmetry when the object model is fully known, few studies symmetry detection on objects with partial observation, such as single RGB-D images. Recent work addresses the problem of detecting symmetries from incomplete data with a deep neural network by leveraging the dense and accurate symmetry annotations. However, due to the tedious labeling process, full symmetry annotations are not always practically available. In this work, we present a 3D symmetry detection approach to detect symmetry from single-view RGB-D images without using symmetry supervision. The key idea is to train the network in a weakly-supervised learning manner to complete the shape based on the predicted symmetry such that the completed shape be similar to existing plausible shapes. To achieve this, we first propose a discriminative variational autoencoder to learn the shape prior in order to determine whether a 3D shape is plausible or not. Based on the learned shape prior, a symmetry detection network is present to predict symmetries that produce shapes with high shape plausibility when completed based on those symmetries. Moreover, to facilitate end-to-end network training and multiple symmetry detection, we introduce a new symmetry parametrization for the learning-based symmetry estimation of both reflectional and rotational symmetry. The proposed approach, coupled symmetry detection with shape completion, essentially learns the symmetry-aware shape prior, facilitating more accurate and robust symmetry detection. Experiments demonstrate that the proposed method is capable of detecting reflectional and rotational symmetries accurately, and shows good generality in challenging scenarios, such as objects with heavy occlusion and scanning noise. Moreover, it achieves state-of-the-art performance, improving the F1-score over the existing supervised learning method by 2%-11% on the ShapeNet and ScanNet datasets.
Xin Xu 0001, Junhua Xi, Xiaochang Hu, Dewen Hu, Kai Xu 0004
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Anti-Disturbance Path-Following Control for Snake Robots With Spiral Motion
abstract
Three-dimensional spiral gait enables a snake robot to climb over obstacles, cross caves, and adapt to complex environments. This article reports an antidisturbance path-following control method for a snake robot with a spiral gait. This method reduces the deviation of the robot's position in following the ideal path by estimating the time-varying parameters, the external disturbances, and the viscous friction coefficients. The estimations are used to compensate for the control inputs of the system, which can improve the adaptability of the robot to the environment. Then, the attitude and position errors can rapidly converge to the origin. An appropriate Lyapunov function is adopted to explore the stability of following errors. Experimental results show that the proposed method can accelerate the convergence rate of errors, reduce the fluctuation peak, and improve the following stability of snake robots.
Dongfang Li 0001, Kevin W. Tong, Ping Li 0044, Rob Law 0001, Xin Xu 0001, Limin Zhu 0001, Qi Wu 0003
IEEE Trans. Ind. Informatics6
2023 Robust Neural Dynamics Method for Redundant Robot Manipulator Control With Physical Constraints
abstract
Redundant robot manipulators play a significant role in modern industry. In this article, we propose a solution scheme to the trajectory tracking problem of the redundant robot manipulator with physical constraints through the Zhang neural dynamics method. Such problem is integrated into a time-varying system consisting of time-varying nonlinear equation (TVNE) and time-varying linear inequality (TVLI) and solved online by the varying-parameter Zhang neural dynamics (VPZND) model. It is ensured that the redundant robot manipulator can still perform the tracking task perfectly under the coexistence of time-varying bounded noise and physical constraints. Theoretical analysis proves that this VPZND model also has an explicit fixed convergence time. Numerical experiments confirm the feasibility of our VPZND model for TVLI. The trajectory tracking problem of the redundant robot manipulator with six or three degrees of freedom under the dual influence of physical constraints and noise is perfectly solved by the VPZND model, which is enough to verify its practical value.
Miaomiao Zhang 0001, Kevin W. Tong, Ping Li 0044, Yuhong Hou, Xin Xu 0001, Limin Zhu 0001, Qi Wu 0003
IEEE Trans. Ind. Informatics5
2023 Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge Distillation
abstract
Class-Incremental Learning (CIL) aims at incrementally learning novel classes without forgetting old ones. This capability becomes more challenging when novel tasks contain one or a few labeled training samples, which leads to a more practical learning scenario,i.e., Few-Shot Class- Incremental Learning (FSCIL). The dilemma on FSCIL lies in serious overfitting and exacerbated catastrophic forgetting caused by the limited training data from novel classes. In this paper, excited by the easy accessibility of unlabeled data, we conduct a pioneering work and focus on a Semi-Supervised Few-Shot Class-Incremental Learning (Semi-FSCIL) problem, which requires the model incrementally to learn new classes from extremely limited labeled samples and a large number of unlabeled samples. To address this problem, a simple but efficient framework is first constructed based on the knowledge distillation technique to alleviate catastrophic forgetting. To efficiently mitigate the overfitting problem on novel categories with unlabeled data, uncertainty-guided semi-supervised learning is incorporated into this framework to select unlabeled samples into incremental learning sessions considering the model uncertainty. This process provides extra reliable supervision for the distillation process and contributes to better formulating the class means. Our extensive experiments on CIFAR100, miniImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction.
Yawen Cui, Wanxia Deng, Xin Xu 0001, Zhen Liu 0004, Zhong Liu 0002, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Multim.3
2023 Weakly-Supervised 3D Human Pose Estimation With Cross-View U-Shaped Graph Convolutional Network
abstract
Although monocular 3D human pose estimation methods have made significant progress, it is far from being solved due to the inherent depth ambiguity. Instead, exploiting multi-view information is a practical way to achieve absolute 3D human pose estimation. In this paper, we propose a simple yet effective pipeline for weakly-supervised cross-view 3D human pose estimation. By only using two camera views, our method can achieve state-of-the-art performance in a weakly-supervised manner, requiring no 3D ground truth but only 2D annotations. Specifically, our method contains two steps: triangulation and refinement. First, given the 2D keypoints that can be obtained through any classic 2D detection methods, triangulation is performed across two views to lift the 2D keypoints into coarse 3D poses. Then, a novel cross-view U-shaped graph convolutional network (CV-UGCN), which can explore the spatial configurations and cross-view correlations, is designed to refine the coarse 3D poses. In particular, the refinement progress is achieved through weakly-supervised learning, in which geometric and structure-aware consistency checks are performed. We evaluate our method on the standard benchmark dataset, Human3.6M. The Mean Per Joint Position Error on the benchmark dataset is 27.4 mm, which outperforms existing state-of-the-art methods remarkably (27.4 mm vs 30.2 mm).
Guoliang Hua, Hong Liu 0008, Wenhao Li 0002, Runwei Ding, Xin Xu 0001
IEEE Trans. Multim.6
2023 Multiple Kernel Clustering With Compressed Subspace Alignment
abstract
Multiple kernel clustering (MKC) has recently achieved remarkable progress in fusing multisource information to boost the clustering performance. However, the$\mathcal {O}({n}^{2})$memory consumption and$\mathcal {O}({n}^{3})$computational complexity prohibit these methods from being applied into median- or large-scale applications, where$n$denotes the number of samples. To address these issues, we carefully redesign the formulation of subspace segmentation-based MKC, which reduces the memory and computational complexity to$\mathcal {O}({n})$and$\mathcal {O}({n}^{2})$, respectively. The proposed algorithm adopts a novel sampling strategy to enhance the performance and accelerate the speed of MKC. Specifically, we first mathematically model the sampling process and then learn it simultaneously during the procedure of information fusion. By this way, the generated anchor point set can better serve data reconstruction across different views, leading to improved discriminative capability of the reconstruction matrix and boosted clustering performance. Although the integrated sampling process makes the proposed algorithm less efficient than the linear complexity algorithms, the elaborate formulation makes our algorithm straightforward for parallelization. Through the acceleration of GPU and multicore techniques, our algorithm achieves superior performance against the compared state-of-the-art methods on six datasets with comparable time cost to the linear complexity algorithms.
Sihang Zhou 0001, Qiyuan Ou, Xinwang Liu 0002, Siqi Wang 0001, Luyan Liu, Siwei Wang 0001, En Zhu, Jianping Yin, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.9
2023 Parameter Estimation and Anti-Sideslip Line-of-Sight Method-Based Adaptive Path-Following Controller for a Multijoint Snake Robot
abstract
This work reports an adaptive path-following controller for a multijoint snake robot (MSR) to improve the adaptability of the robot to the environment. The new strategy estimates the time-varying parameters of the system and the external interference to adjust the motion state of the robot in real time. Estimations are used to compensate for the joint torque of an MSR, thus reducing the fluctuation peak of path-following errors. In addition, this work designs an anti-sideslip line-of-sight (LOS) guidance strategy to avoid the deviation of the direction angle. The method can improve the tracking accuracy of an MSR, and the position errors enable the system to achieve uniformly ultimate boundedness (UUB). The angle errors converge to the origin to achieve stability. Experimental results demonstrate that the novel method can accurately estimate the time-dependent parameters, sideslip, and interference, raise the convergent speed of errors, and reduce the fluctuation peak.
Dongfang Li 0001, Binxin Zhang, Ping Li 0044, Qi Wu 0003, Rob Law 0001, Xin Xu 0001, Aiguo Song, Limin Zhu 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2023 Receding Horizon Actor-Critic Learning Control for Nonlinear Time-Delay Systems With Unknown Dynamics
abstract
With the development of modern mechatronics and networked systems, the controller design of time-delay systems has received notable attention. Time delays can greatly influence the stability and performance of the systems, especially for optimal control design. In this article, we propose a receding horizon actor–critic learning control approach for near-optimal control of nonlinear time-delay systems (RACL-TD) with unknown dynamics. In the proposed approach, a data-driven predictor for nonlinear time-delay systems is first learned based on the Koopman theory using precollected samples. Then, a receding horizon actor–critic architecture is designed to learn a near-optimal control policy. In RACL-TD, the terminal cost is determined by using the Lyapunov–Krasovskii approach so that the influences of the delayed states and control inputs can be well addressed. Furthermore, a relaxed terminal condition is present to reduce the computational cost. The convergence and optimality of RACL-TD in each prediction interval as well as the closed-loop property of the system are discussed and analyzed. Simulation results on a two-stage time-delayed chemical reactor illustrate that RACL-TD can achieve better control performance than nonlinear model predictive control (MPC) and infinite-horizon adaptive dynamic programming. Moreover, RACL-TD can have less computational cost than nonlinear MPC.
Xin Xu 0001, Quan Xiong
IEEE Trans. Syst. Man Cybern. Syst.3
2023 A Dual-Level Model Predictive Control Scheme for Multitimescale Dynamical Systems
abstract
So far, many control algorithms have been developed for singularly perturbed systems. However, in many industrial processes, enforcing closed-loop fast-slow dynamics for peculiarly nonseparable ones is a prior request and a crucial issue to be resolved. Aiming at the above problem, this article presents two dual-level model predictive control (MPC) algorithms for multitimescale dynamical systems with unknown bounded disturbances and input constraints. The proposed algorithms, each one composed of two regulators working in slow and fast time scales, are designed to generate closed-loop separable dynamics at high and low levels. As a prominent feature, the proposed algorithms are not only suitable for singularly perturbed systems but also capable of imposing separable closed-loop performance for dynamics that are nonseparable and strongly coupled. The recursive feasibility and convergence properties are proven under suitable assumptions. The simulation results on controlling a boiler turbine (BT) system, including the comparisons with other classic controllers, are demonstrated, which show the effectiveness of the proposed algorithms.
Wei Jiang 0006, Shuyou Yu 0001, Xin Xu 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Meta-GF: Training Dynamic-Depth Neural Networks Harmoniously
Jian Li 0003, Xin Xu 0001
ECCV (11)3
2022 Barrier Function-based Safe Reinforcement Learning for Formation Control of Mobile Robots
abstract
Distributed model predictive control (DMPC) concerns how to online control multiple robotic systems with constraints effectively. However, the nonlinearity, nonconvexity, and strong interconnections of dynamic system models and constraints can make the real-time and real-world DMPC implementations nontrivial. Reinforcement learning (RL) algorithms are promising for control policy design. However, how to ensure safety in terms of state constraints in RL remains a significant issue. This paper proposes a barrier function-based safe reinforcement learning algorithm for DMPC of nonlinear multi-robot systems under state constraints. The proposed approach is composed of several local learning-based MPC regulators. Each regulator, associated with a local system, learns and deploys the local control policy using a safe reinforcement learning algorithm in a distributed manner, i.e., with state information only among the neighbor agents. As a prominent feature of the proposed algorithm, we present a novel barrier-based policy structure to ensure safety, which has a clear mechanistic interpretation. Both simulated and real-world experiments on the formation control of mobile robots with collision avoidance show the effectiveness of the proposed safe reinforcement learning algorithm for DMPC.
Yaoqian Peng, Wei Pan 0004, Xin Xu 0001, Haibin Xie
ICRA4
2022 Learning practically feasible policies for online 3D bin packing
Hang Zhao 0018, Chenyang Zhu 0002, Xin Xu 0001, Hui Huang 0004, Kai Xu 0004
Sci. China Inf. Sci.3
2022 Extended clustering algorithm based on cluster shape boundary
abstract
Based on the shape characteristics of the sample distribution in the clustering problem, this paper proposes an extended clustering algorithm based on cluster shape boundary (ECBSB). The algorithm automatically determines the number of clusters and classification discrimination boundaries by finding the boundary closures of the clusters from a global perspective of the sample distribution. Since ECBSB is insensitive to local features of the sample distribution, it can accurately identify clusters on complex shape and uneven density distribution. ECBSB first detects the shape boundary points of the cluster in the sample set with edge noise points eliminated, and then generates boundary closures around the cluster based on the boundary points. Finally, the cluster labels of the boundary are propagated to the entire sample set by a nearest neighbor search. The proposed method is evaluated on multiple benchmark datasets. Exhaustive experimental results show that the proposed method achieves highly accurate and robust clustering results, and is superior to the classical clustering baselines on most of the test data.
Haibin Xie, Xin Xu 0001
Intell. Data Anal.4
2022 Facial Kinship Verification: A Comprehensive Review and Outlook
abstract
The goal of Facial Kinship Verification (FKV) is to automatically determine whether two individuals have a kin relationship or not from their given facial images or videos. It is an emerging and challenging problem that has attracted increasing attention due to its practical applications. Over the past decade, significant progress has been achieved in this new field. Handcrafted features and deep learning techniques have been widely studied in FKV. The goal of this paper is to conduct a comprehensive review of the problem of FKV. We cover different aspects of the research, including problem definition, challenges, applications, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. In retrospect of what has been achieved so far, we identify gaps in current research and discuss potential future research directions.
Xiaoting Wu, Xiaoyi Feng, Xiaochun Cao, Xin Xu 0001, Dewen Hu, Miguel Bordallo López, Li Liu 0002
Int. J. Comput. Vis.4
2022 AdaBoost maximum entropy deep inverse reinforcement learning with truncated gradient
Li Song 0003, Dazi Li, Xiao Wang 0002, Xin Xu 0001
Inf. Sci.4
2022 Transfer reinforcement learning via meta-knowledge extraction using auto-pruned decision trees
Yixing Lan, Xin Xu 0001, Qiang Fang 0001, Yujun Zeng, Xinwang Liu 0002, Xianjian Zhang
Knowl. Based Syst.2
2022 Sparse online maximum entropy inverse reinforcement learning via proximal optimization and truncated gradient
Li Song 0003, Dazi Li, Xin Xu 0001
Knowl. Based Syst.3
2022 Cross-Modal Cross-Domain Dual Alignment Network for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared cross-modal person re-identification (Re-ID) has drawn increasing attention due to its application value in practice. Most of the current works rely on a supervised training manner. However, in real-world applications, manual collection of pair-wise RGB-Infrared (IR) person data is labor-intensive and time-consuming. Moreover, when a trained model is directly used in another domain, there is usually a significant performance drop. To overcome the above problems, we make the first attempt to transfer the learned model to a new RGB-IR domain which is unlabeled. The practical problem covers two kinds of challenges, i.e., cross-modal (RGB-Infrared) and cross-domain (different dataset) person Re-ID. Previous works have often considered only one of them either cross-modal or cross-domain. In this work, we propose a dual alignment network (DAN) to solve the RGB-Infrared cross-modal cross-domain person Re-ID problem. This network consists of three parts: Domain Adversarial Alignment component (DAA), Pseudo Label Generation module for target domain (PLG), and Cross-Modal Alignment component (CMA). These three modules complement and promote the model to learn domain-invariant and modality-invariant person representations. Further, we propose a protocol of cross-modal cross-domain person Re-ID by synthesizing target domains by adding random noise, adjusting the lighting intensity, and changing the background color, respectively. Experiments on real and synthetic datasets under the same cross-modalities across domains demonstrate the effectiveness of our method.
Xiaowei Fu, Fuxiang Huang, Huimin Ma 0001, Xin Xu 0001, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.5
2022 Inferring Cognitive State of Pilot's Brain Under Different Maneuvers During Flight
abstract
This work designs an adversarial Bayesian deep network to solve the cognitive detection of pilot fatigue. Batch normalization and data enhancement are adopted in the posterior inference of the proposed model parameters to effectively improve the generalization of neural networks. The generator is used to enhance the brain power map generated from three cognitive indicators and improve the accuracy of fatigue state recognition. This work also adds adversarial noise in the vicinity of each brain electrode to form an adversarial image, which further reveals the correlation between the cognitive state of brain and the location of brain regions. Compared with other deep models and parameter optimization methods, our model achieves better detection accuracy.
Qi Wu 0003, Zhengtao Cao, Zhao-Hui Sun, Dongfang Li 0001, Rob Law 0001, Xin Xu 0001, Limin Zhu 0001, Mengsun Yu
IEEE Trans. Intell. Transp. Syst.6
2022 Multi-View Spectral Clustering With High-Order Optimal Neighborhood Laplacian Matrix
abstract
Multi-view spectral clustering can effectively reveal the intrinsic cluster structure among data by performing clustering on the learned optimal embedding across views. Though demonstrating promising performance in various applications, most of existing methods usually linearly combine a group of pre-specified first-order Laplacian matrices to construct the optimal Laplacian matrix, which may result in limited representation capability and insufficient information exploitation. Also, storing and implementing complex operations on the{$n\times n}$Laplacian matrices incurs intensive storage and computation complexity. To address these issues, this paper first proposes a multi-view spectral clustering algorithm that learns a high-order optimal neighborhood Laplacian matrix, and then extends it to the late fusion version for accurate and efficient multi-view clustering. Specifically, our proposed algorithm generates the optimal Laplacian matrix by searching the neighborhood of the linear combination of both the first-order and high-order base Laplacian matrices simultaneously. By this way, the representative capacity of the learned optimal Laplacian matrix is enhanced, which is helpful to better utilize the hidden high-order connection information among data, leading to improved clustering performance. We design an efficient algorithm with proved convergence to solve the resultant optimization problem. Extensive experimental results on nine datasets demonstrate the superiority of the proposed algorithm
Weixuan Liang, Sihang Zhou 0001, Jian Xiong 0002, Xinwang Liu 0002, Siwei Wang 0001, En Zhu, Zhiping Cai, Xin Xu 0001
IEEE Trans. Knowl. Data Eng.8
2022 Coordinated Path-Following Control of Fixed-Wing Unmanned Aerial Vehicles
abstract
This article investigates the problem of coordinated path following for fixed-wing unmanned aerial vehicles (UAVs) with speed constraints in the two-dimensional plane. The objective is to steer a fleet of UAVs along the path(s) while achieving the desired sequenced inter-UAV arc distance. In contrast to the previous coordinated path-following studies, we are able through our proposed hybrid control law to deal with the forward speed and the angular speed constraints of fixed-wing UAVs. More specifically, the hybrid control law makes all the UAVs work at two different levels: 1) those UAVs whose path-following errors are within an invariant set (i.e., the designed coordination set) work at the coordination level and 2) the other UAVs work at the single-agent level. At the coordination level, we prove that even with speed constraints, the proposed control law can make sure the path-following errors reduce to zero, while the inter-UAV arc distances converge to the desired value. At the single-agent level, analysis for the path-following error entering the coordination set is provided. We develop a hardware-in-the-loop simulation testbed of the multi-UAV system by using actual autopilots and the X-Plane simulator. The effectiveness of the proposed approach is corroborated with both numerical simulation and the testbed.
Hao Chen 0044, Yirui Cong, Xiangke Wang, Xin Xu 0001, Lincheng Shen
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Online Sparse Temporal Difference Learning Based on Nested Optimization and Regularized Dual Averaging
abstract
In policy evaluation of reinforcement learning tasks, the temporal difference (TD) learning with value function approximation has been widely studied. However, feature representation has a decisive influence on both accuracy of value function approximation and convergence rate. Therefore, it is important to develop the feature selection theory and methods that can efficiently prevent overfitting and improve estimation accuracy in TD learning algorithms. In this article, we propose an online sparse TD learning algorithm for policy evaluation by using$\ell _{1}$-regualrization for feature selection. The per-step-time runtime computational complexity of the proposed algorithm is linear with respect to feature dimension. The loss function is defined as a nested optimization with$\ell _{1}$-regularization penalty, and the solver minimizes two suboptimization problems by running stochastic gradient descent and regularized dual averaging method, alternately. The convergence results for the fixed points are also established. The experiments on benchmarks with high-dimensional features show the abilities of learning and generalization of the proposed algorithms.
Tianheng Song, Dazi Li, Xin Xu 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Robust Learning-Based Predictive Control for Discrete-Time Nonlinear Systems With Unknown Dynamics and State Constraints
abstract
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and frequent access to onboard computational resources. In this article, we propose an efficient robust MPC solution based on receding horizon reinforcement learning, called r-LPC, for unknown nonlinear systems with state constraints and disturbances. The proposed r-LPC utilizes a Koopman operator-based prediction model obtained offline from precollected input–output datasets. Unlike classic tube-based MPC, in each prediction time interval of r-LPC, we use an actor–critic structure to learn a near-optimal feedback control policy rather than a control sequence. The resulting closed-loop control policy can be learned offline and deployed online or learned online in an asynchronous way. In the latter case, online learning can be activated whenever necessary; for instance, the safety constraint is violated with the deployed policy. The closed-loop recursive feasibility, robustness, and asymptotic stability are proven under function approximation errors of the actor–critic networks. Simulation and experimental results on two nonlinear systems with unknown dynamics and disturbances have demonstrated that our approach has better or comparable performance when compared with tube-based MPC and linear quadratic regulator, and outperforms a recently developed actor–critic learning approach.
Xin Xu 0001, Shuyou Yu 0001, Hong Chen 0003
IEEE Trans. Syst. Man Cybern. Syst.3
2021 StablePose: Learning 6D Object Poses From Geometrically Stable Patches
abstract
We introduce the concept of geometric stability to the problem of 6D object pose estimation and propose to learn pose inference based on geometrically stable patches extracted from observed 3D point clouds. According to the theory of geometric stability analysis, a minimal set of three planar/cylindrical patches are geometrically stable and determine the full 6DoFs of the object pose. We train a deep neural network to regress 6D object pose based on geometrically stable patch groups via learning both intra-patch geometric features and inter-patch contextual features. A subnetwork is jointly trained to predict per-patch poses. This auxiliary task is a relaxation of the group pose prediction: A single patch cannot determine the full 6DoFs but is able to improve pose accuracy in its corresponding DoFs. Working with patch groups makes our method generalize well for random occlusion and unseen instances. The method is easily amenable to resolve symmetry ambiguities. Our method achieves the state-of-the-art results on public benchmarks compared not only to depth-only but also to RGBD methods. It also performs well in category-level pose estimation.
Junwen Huang 0001, Xin Xu 0001, Kai Xu 0004
CVPR3
2021 Accelerating Deep Reinforcement Learning via Hierarchical State Encoding with ELMs
Qiang Fang 0001, Xin Xu 0001, Yujun Zeng
ICIC (2)3
2021 Deep Q-learning with Explainable and Transferable Domain Rules
Junkai Ren, Qiang Fang 0001, Xin Xu 0001
ICIC (2)5
2021 Event-triggered shared lateral control for safe-maneuver of intelligent vehicles
Xin Xu 0001, Xing Zhou 0004, Zhengzheng Dong
Sci. China Inf. Sci.3
2021 Multi-target tracking for unmanned aerial vehicle swarms using deep reinforcement learning
Wenhong Zhou, Xin Xu 0001, Lincheng Shen
Neurocomputing4
2021 Robust semi-supervised classification based on data augmented online ELMs with deep features
Xiaochang Hu, Yujun Zeng, Xin Xu 0001, Sihang Zhou 0001, Li Liu 0002
Knowl. Based Syst.3
2021 Label Disentangled Analysis for unsupervised visual domain adaptation
Ni Xiao, Lei Zhang 0038, Xin Xu 0001, Tan Guo, Huimin Ma 0001
Knowl. Based Syst.3
2021 Dual-branch combination network (DCN): Towards accurate diagnosis and lesion segmentation of COVID-19 using CT images
Kai Gao 0011, Jianpo Su, Zhongbiao Jiang, Zhichao Feng, Hui Shen 0004, Pengfei Rong, Xin Xu 0001, Yuexiang Yang, Wei Wang 0434, Dewen Hu
Medical Image Anal.8
2021 Actor-Critic Learning Control With Regularization and Feature Selection in Policy Gradient Estimation
abstract
Actor-critic (AC) learning control architecture has been regarded as an important framework for reinforcement learning (RL) with continuous states and actions. In order to improve learning efficiency and convergence property, previous works have been mainly devoted to solve regularization and feature learning problem in the policy evaluation. In this article, we propose a novel AC learning control method with regularization and feature selection for policy gradient estimation in the actor network. The main contribution is that ℓ1-regularization is used on the actor network to achieve the function of feature selection. In each iteration, policy parameters are updated by the regularized dual-averaging (RDA) technique, which solves a minimization problem that involves two terms: one is the running average of the past policy gradients and the other is the ℓ1-regularization term of policy parameters. Our algorithm can efficiently calculate the solution of the minimization problem, and we call the new adaptation of policy gradient RDApolicy gradient (RDA-PG). The proposed RDA-PG can learn stochastic and deterministic near-optimal policies. The convergence of the proposed algorithm is established based on the theory of two-timescale stochastic approximation. The simulation and experimental results show that RDA-PG performs feature selection successfully in the actor and learns sparse representations of the actor both in stochastic and deterministic cases. RDA-PG performs better than existing AC algorithms on standard RL benchmark problems with irrelevant features or redundant features.
Luntong Li, Dazi Li, Tianheng Song, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Multi-Kernel Online Reinforcement Learning for Path Tracking Control of Intelligent Vehicles
abstract
Path tracking control of intelligent vehicles has to deal with the difficulties of model uncertainties and nonlinearities. As a class of adaptive optimal control methods, reinforcement learning (RL) has received increasing attention in solving difficult control problems. However, feature representation and online learning ability are two major problems to be solved for learning control of uncertain dynamic systems. In this article, we propose a multi-kernel online RL approach for path tracking control of intelligent vehicles. In the proposed approach, a multiple kernel feature learning framework is designed for online learning control based on dual heuristic programming (DHP) and the new online learning control algorithm is called multi-kernel DHP (MKDHP). In MKDHP, instead of the expert knowledge for selecting and fine-tuning of a suitable kernel function, only a set of basic kernel functions is required to be predefined and the multi-kernel features can be learned for value function approximation in the critic. The simulation studies on path tracking control for intelligent vehicles have been conducted under$S$-curve and urban road conditions. The results demonstrated that compared with other typical path tracking controllers for intelligent vehicles, such as the linear quadratic regulator (LQR), the pure pursuit controller and the ribbon-based controller, the proposed multi-kernel learning controller can achieve better performance in terms of tracking precision and smoothness.
Zhenhua Huang 0004, Xin Xu 0001, Shiliang Sun, Dazi Li
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Efficient Batch-Mode Reinforcement Learning Using Extreme Learning Machines
abstract
As a class of batch-mode reinforcement learning (RL) methods for Markov decision problems with large or continuous state spaces, approximate policy iteration (API) has received increasing attention in the past decades. One open problem in the design of API algorithms is how to construct the basis functions or features for value function approximation (VFA). In this paper, we propose a novel batch-mode RL approach with randomly projected features for VFA. The proposed approach can be viewed as an extension of extreme learning machines (ELMs) to RL problems so it can be called ELM-API. The ELMs have been popularly studied in supervised learning problems, but there is not much work on the extension of ELMs to learning control problems. The proposed approach has advantages over the previous API algorithms in that the features for VFA can be quickly generated without complex parameter selection and the performance will be adaptive to different sample sets in batch-mode RL. In particular, the ELM-API approach can realize fast and efficient feature reconstruction when training sample sets are relatively small. Comprehensive simulation studies on two benchmark learning control problems were carried out to test the performance of API algorithms with different feature construction methods. It is shown that the ELM-API algorithm can obtain comparable or better performance than the previous API approaches. To further show the effectiveness of ELM-API in real-world applications, the simulation results on a more challenging high-dimensional lane-changing decision problem in dynamic traffic environment are also reported, which show the capability of the ELM-API algorithm in learning satisfactory lane-changing policies with high data efficiency.
Lei Zuo 0002, Xin Xu 0001, Junkai Ren, Qiang Fang 0001, Xinwang Liu 0002
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Multi-View Deep Attention Network for Reinforcement Learning (Student Abstract)
abstract
The representation approximated by a single deep network is usually limited for reinforcement learning agents. We propose a novel multi-view deep attention network (MvDAN), which introduces multi-view representation learning into the reinforcement learning task for the first time. The proposed model approximates a set of strategies from multiple representations and combines these strategies based on attention mechanisms to provide a comprehensive strategy for a single-agent. Experimental results on eight Atari video games show that the MvDAN has effective competitive performance than single-view reinforcement learning methods.
Yueyue Hu, Shiliang Sun, Xin Xu 0001, Jing Zhao 0015
AAAI3
2020 Deep reinforcement learning for pedestrian collision avoidance and human-machine cooperative driving
Xin Xu 0001, Bang Cheng, Junkai Ren
Inf. Sci.3
2020 Drosophila-inspired 3D moving object detection based on point clouds
Dawei Zhao 0003, Tao Wu 0001, Hao Fu 0001, Liang Xiao 0007, Xin Xu 0001, Bin Dai 0001
Inf. Sci.7
2020 Distributed Multiagent Coordinated Learning for Autonomous Driving in Highways Based on Dynamic Coordination Graphs
abstract
Autonomous driving is one of the most important AI applications and has attracted extensive interest in recent years. A large number of studies have successfully applied reinforcement learning techniques in various aspects of autonomous driving, ranging from low-level control of driving maneuvers to higher level of strategic decision-making. However, comparatively less progress has been made in investigating how co-existing autonomous vehicles would interact with each other in a common environment and how reinforcement learning can be helpful in such situations by applying multiagent reinforcement learning techniques in the high-level strategic decision-making of the following or overtaking for a group of autonomous vehicles in highway scenarios. Learning to achieve coordination among vehicles in such situations is challenging due to the unique feature of vehicular mobility, which renders it infeasible to directly apply the existing coordinated learning approaches. To solve this problem, we propose using dynamic coordination graph to model the continuously changing topology during vehicles' interactions and come up with two basic learning approaches to coordinate the driving maneuvers for a group of vehicles. Several extension mechanisms are then presented to make these approaches workable in a more complex and realistic setting with any number of vehicles. The experimental evaluation has verified the benefits of the proposed coordinated learning approaches, compared with other approaches that learn without coordination or rely on some traditional mobility models based on some expert driving rules.
Chao Yu 0004, Xin Wang 0077, Xin Xu 0001, Minjie Zhang 0001, Hong-Wei Ge, Jiankang Ren, Liang Sun 0003, Bingcai Chen, Guozhen Tan
IEEE Trans. Intell. Transp. Syst.3
2020 Adaptive Self-Paced Deep Clustering with Data Augmentation
abstract
Deep clustering gains superior performance than conventional clustering by jointly performing feature learning and cluster assignment. Although numerous deep clustering algorithms have emerged in various applications, most of them fail to learn robust cluster-oriented features which in turn hurts the final clustering performance. To solve this problem, we propose a two-stage deep clustering algorithm by incorporating data augmentation and self-paced learning. Specifically, in the first stage, we learn robust features by training an autoencoder with examples that are augmented by random shifting and rotating the given clean examples. Then, in the second stage, we encourage the learned features to be cluster-oriented by alternatively finetuning the encoder with the augmented examples and updating the cluster assignments of the clean examples. During finetuning the encoder, the target of each augmented example in the loss function is the center of the cluster to which the clean example is assigned. The targets may be computed incorrectly, and the examples with incorrect targets could mislead the encoder network. To stabilize the network training, we select most confident examples in each iteration by utilizing the adaptive self-paced learning. Extensive experiments validate that our algorithm outperforms the state of the arts on four image datasets.
Xifeng Guo 0001, Xinwang Liu 0002, En Zhu, Xinzhong Zhu, Miaomiao Li 0001, Xin Xu 0001, Jianping Yin
IEEE Trans. Knowl. Data Eng.6
2020 SymmetryNet: learning to predict reflectional and rotational symmetries of 3D shapes from single-view RGB-D images
abstract
We study the problem of symmetry detection of 3D shapes from single-view RGB-D images, where severely missing data renders geometric detection approach infeasible. We propose an end-to-end deep neural network which is able to predict both reflectional and rotational symmetries of 3D objects present in the input RGB-D image. Directly training a deep model for symmetry prediction, however, can quickly run into the issue of overfitting. We adopt a multi-task learning approach. Aside from symmetry axis prediction, our network is also trained to predict symmetry correspondences. In particular, given the 3D points present in the RGB-D image, our network outputs for each 3D point its symmetric counterpart corresponding to a specific predicted symmetry. In addition, our network is able to detect for a given shape multiple symmetries of different types. We also contribute a benchmark of 3D symmetry detection based on single-view RGB-D images. Extensive evaluation on the benchmark demonstrates the strong generalization ability of our method, in terms of high accuracy of both symmetry axis prediction and counterpart estimation. In particular, our method is robust in handling unseen object instances with large variation in shape, multi-symmetry composition, as well as novel object categories.
Junwen Huang 0001, Xin Xu 0001, Szymon Rusinkiewicz, Kai Xu 0004
ACM Trans. Graph.4
2020 A Reinforcement Learning Approach to Autonomous Decision Making of Intelligent Vehicles on Highways
abstract
Autonomous decision making is a critical and difficult task for intelligent vehicles in dynamic transportation environments. In this paper, a reinforcement learning approach with value function approximation and feature learning is proposed for autonomous decision making of intelligent vehicles on highways. In the proposed approach, the sequential decision making problem for lane changing and overtaking is modeled as a Markov decision process with multiple goals, including safety, speediness, smoothness, etc. In order to learn optimized policies for autonomous decision-making, a multiobjective approximate policy iteration (MO-API) algorithm is presented. The features for value function approximation are learned in a data-driven way, where sparse kernel-based features or manifold-based features can be constructed based on data samples. Compared with previous RL algorithms such as multiobjective Q-learning, the MO-API approach uses data-driven feature representation for value and policy approximation so that better learning efficiency can be achieved. A highway simulation environment using a 14 degree-of-freedom vehicle dynamics model was established to generate training data and test the performance of different decision-making methods for intelligent vehicles on highways. The results illustrate the advantages of the proposed MO-API method under different traffic conditions. Furthermore, we also tested the learned decision policy on a real autonomous vehicle to implement overtaking decision and control under normal traffic on highways. The experimental results also demonstrate the effectiveness of the proposed method.
Xin Xu 0001, Lei Zuo 0002, Lilin Qian, Junkai Ren, Zhenping Sun
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Augmenting cascaded correlation filters with spatial-temporal saliency for visual tracking
Dawei Zhao 0003, Liang Xiao 0007, Hao Fu 0001, Tao Wu 0001, Xin Xu 0001, Bin Dai 0001
Inf. Sci.5
2019 Large-scale gesture recognition with a fusion of RGB-D data based on optical flow and the C3D model
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Zhenxin Ma, Jianfeng Song
Pattern Recognit. Lett.5
2019 Parameterized Batch Reinforcement Learning for Longitudinal Control of Autonomous Land Vehicles
abstract
This paper presents a parameterized batch reinforcement learning algorithm for near-optimal longitudinal control of autonomous land vehicles (ALVs). The proposed approach uses an actor-critic architecture, where parameterized feature vectors based on kernels are learned from collected samples for approximating the value functions and policies. One difference between the parameterized batch actor-critic (PBAC) algorithm and previous actor-critic learning approaches is that the critic and actor in PBAC share the same linear features, which has been theoretically proved to be a beneficial property for the convergence of actor-critic learning approaches. In order to obtain better learning efficiency, least-squares-based batch updating rules are designed for the critic and actor, respectively. Based on the PBAC learning algorithm, a data-driven longitudinal control method is presented for ALVs to obtain near-optimal control policies which adaptively tune the fuel/brake control signals to track different speeds. A multiobjective reward function is designed so that both tracking precision and driving smoothness are considered. Extensive experiments were conducted on a real ALV platform while driving on flat, slippery, sloping, and bumpy roads. The experimental results illustrate the superiority of the PBAC-based self-learning controller over conventional longitudinal control methods such as proportional-integral (PI) control and learning-based PI control.
Zhenhua Huang 0004, Xin Xu 0001, Haibo He, Zhenping Sun
IEEE Trans. Syst. Man Cybern. Syst.2
2018 Toward Autonomous Driving in Highway and Urban Environment: HQ3 and IVFC 2017
abstract
The 2017 Intelligent Vehicle Future Challenge of China (IVFC) was held in Changshu between 24th November and 26th November, 2017. As the ninth series of this event, last year's competition has introduced many new features and has attracted 21 teams to join this competition. The HQ3 autonomous vehicle, jointly developed by National University of Defense Technology, Jilin University and Central South University, took part in this competition. This paper mainly describes the key modules of HQ3, including GPS-free localization, environment perception and behavior planning. All of these modules together enable HQ3 to perform well during the competition.
Lilin Qian, Hao Fu 0001, Xiaohui Li 0007, Bang Cheng, Tingbo Hu, Zengping Sun, Tao Wu 0001, Bin Dai 0001, Xin Xu 0001
Intelligent Vehicles Symposium9
2018 Large-Scale Gesture Recognition With a Fusion of RGB-D Data Based on Saliency Theory and C3D Model
abstract
Gesture recognition has raised wide attention in computer vision owing to its many applications. However, the task of video-based large-scale gesture recognition yet faces many challenges, since many gesture-irrelevant factors like the background may disturb the recognition accuracy. To better recognize gestures with large-scale videos, we propose a method based on RGB-D data in this paper, where the “RGB-D” means RGB and depth data captured simultaneously by specific devices like Kinect. To learn gesture details better, we first use an adaptive frame unification strategy to unify the frame number of inputs, and then the RGB and depth data are sent to the C3D model to extract spatiotemporal features, respectively. In order to alleviate the interference of gesture-irrelevant factors, the saliency theory is also employed to generate auxiliary data. Next the features of these data are combined to boost the performance, which can also avoid unreasonable synthetic data, since the dimension of C3D features is uniform. Finally the performances of several classifiers are tested and the best one of SVM classifier is selected to output the ultimate accuracy. Our approach achieves 52.04% and 59.43% accuracy on the validation and testing subset of the Chalearn LAP IsoGD, respectively, both of which outperform our results in the chalearn LAP Large-scale Gesture Recognition Challenge as reported in ICPR 2016.
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song
IEEE Trans. Circuits Syst. Video Technol.5
2018 Actor-Critic Learning Control Based on ℓ2-Regularized Temporal-Difference Prediction With Gradient Correction
abstract
Actor-critic based on the policy gradient (PG-based AC) methods have been widely studied to solve learning control problems. In order to increase the data efficiency of learning prediction in the critic of PG-based AC, studies on how to use recursive least-squares temporal difference (RLS-TD) algorithms for policy evaluation have been conducted in recent years. In such contexts, the critic RLS-TD evaluates an unknown mixed policy generated by a series of different actors, but not one fixed policy generated by the current actor. Therefore, this AC framework with RLS-TD critic cannot be proved to converge to the optimal fixed point of learning problem. To address the above problem, this paper proposes a new AC framework named critic-iteration PG (CIPG), which learns the state-value function of current policy in an on-policy way and performs gradient ascent in the direction of improving discounted total reward. During each iteration, CIPG keeps the policy parameters fixed and evaluates the resulting fixed policy by -regularized RLS-TD critic. Our convergence analysis extends previous convergence analysis of PG with function approximation to the case of RLS-TD critic. The simulation results demonstrate that the -regularization term in the critic of CIPG is undamped during the learning process, and CIPG has better learning efficiency and faster convergence rate than conventional AC learning control methods.
Luntong Li, Dazi Li, Tianheng Song, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Learning-Based Predictive Control for Discrete-Time Nonlinear Systems With Stochastic Disturbances
abstract
In this paper, a learning-based predictive control (LPC) scheme is proposed for adaptive optimal control of discrete-time nonlinear systems under stochastic disturbances. The proposed LPC scheme is different from conventional model predictive control (MPC), which uses open-loop optimization or simplified closed-loop optimal control techniques in each horizon. In LPC, the control task in each horizon is formulated as a closed-loop nonlinear optimal control problem and a finite-horizon iterative reinforcement learning (RL) algorithm is developed to obtain the closed-loop optimal/suboptimal solutions. Therefore, in LPC, RL and adaptive dynamic programming (ADP) are used as a new class of closed-loop learning-based optimization techniques for nonlinear predictive control with stochastic disturbances. Moreover, LPC also decomposes the infinite-horizon optimal control problem in previous RL and ADP methods into a series of finite horizon problems, so that the computational costs are reduced and the learning efficiency can be improved. Convergence of the finite-horizon iterative RL algorithm in each prediction horizon and the Lyapunov stability of the closed-loop control system are proved. Moreover, by using successive policy updates between adjoint time horizons, LPC also has lower computational costs than conventional MPC which has independent optimization procedures between two different prediction horizons. Simulation results illustrate that compared with conventional nonlinear MPC as well as ADP, the proposed LPC scheme can obtain a better performance both in terms of policy optimality and computational efficiency.
Xin Xu 0001, Hong Chen 0003, Chuanqiang Lian, Dazi Li
IEEE Trans. Neural Networks Learn. Syst.1
2017 Traffic Sign Recognition Using Kernel Extreme Learning Machines With Deep Perceptual Features
abstract
Traffic sign recognition plays an important role in autonomous vehicles as well as advanced driver assistance systems. Although various methods have been developed, it is still difficult for the state-of-the-art algorithms to obtain high recognition precision with low computational costs. In this paper, based on the investigation on the influence that color spaces have on the representation learning of convolutional neural network, a novel traffic sign recognition approach called DP-KELM is proposed by using a kernel-based extreme learning machine (KELM) classifier with deep perceptual features. Unlike the previous approaches, the representation learning process in DP-KELM is implemented in the perceptual Lab color space. Based on the learned deep perceptual feature, a kernel-based ELM classifier is trained with high computational efficiency and generalization performance. Through the experiments on the German traffic sign recognition benchmark, the proposed method is demonstrated to have higher precision than most of the state-of-the-art approaches. In particular, when compared with the hinge loss stochastic gradient descent method which has the highest precision, the proposed method can achieve a comparable recognition rate with significantly fewer computational costs.
Yujun Zeng, Xin Xu 0001, Dayong Shen, Yuqiang Fang, Zhipeng Xiao
IEEE Trans. Intell. Transp. Syst.2
2017 Manifold-Based Reinforcement Learning via Locally Linear Reconstruction
abstract
Feature representation is critical not only for pattern recognition tasks but also for reinforcement learning (RL) methods to solve learning control problems under uncertainties. In this paper, a manifold-based RL approach using the principle of locally linear reconstruction (LLR) is proposed for Markov decision processes with large or continuous state spaces. In the proposed approach, an LLR-based feature learning scheme is developed for value function approximation in RL, where a set of smooth feature vectors is generated by preserving the local approximation properties of neighboring points in the original state space. By using the proposed feature learning scheme, an LLR-based approximate policy iteration (API) algorithm is designed for learning control problems with large or continuous state spaces. The relationship between the value approximation error of a new data point and the estimated values of its nearest neighbors is analyzed. In order to compare different feature representation and learning approaches for RL, a comprehensive simulation and experimental study was conducted on three benchmark learning control problems. It is illustrated that under a wide range of parameter settings, the LLR-based API algorithm can obtain better learning control performance than the previous API methods with different feature representation schemes.
Xin Xu 0001, Zhenhua Huang 0004, Lei Zuo 0002, Haibo He
IEEE Trans. Neural Networks Learn. Syst.1
2016 Large-scale gesture recognition with a fusion of RGB-D data based on the C3D model
abstract
The gesture recognition has raised attention in computer vision owing to its many applications. However, video-based large-scale gesture recognition still faces many challenges, since many factors like background may disturb the accuracy. To achieve gesture recognition with large-scale videos, we propose a method based on RGB-D data. To learn gesture details better, the inputs are expanded into 32-frame videos first, and then the RGB and depth videos are sent to the C3D model to extract spatiotemporal features respectively. Next these features are combined to boost the performance, which can also avoid unreasonable synthetic data due to the uniform dimension of C3D features. Our approach achieves 49.2% accuracy on the validation subset of the Chalearn LAP IsoGD Database just with a linear SVM classifier. It also outperforms the baseline and other methods in the challenge and wins the first place at 56.9% on testing set.
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song
ICPR5
2016 Near-Optimal Tracking Control of Mobile Robots Via Receding-Horizon Dual Heuristic Programming
abstract
Trajectory tracking control of wheeled mobile robots (WMRs) has been an important research topic in control theory and robotics. Although various tracking control methods with stability have been developed for WMRs, it is still difficult to design optimal or near-optimal tracking controller under uncertainties and disturbances. In this paper, a near-optimal tracking control method is presented for WMRs based on receding-horizon dual heuristic programming (RHDHP). In the proposed method, a backstepping kinematic controller is designed to generate desired velocity profiles and the receding horizon strategy is used to decompose the infinite-horizon optimal control problem into a series of finite-horizon optimal control problems. In each horizon, a closed-loop tracking control policy is successively updated using a class of approximate dynamic programming algorithms called finite-horizon dual heuristic programming (DHP). The convergence property of the proposed method is analyzed and it is shown that the tracking control system based on RHDHP is asymptotically stable by using the Lyapunov approach. Simulation results on three tracking control problems demonstrate that the proposed method has improved control performance when compared with conventional model predictive control (MPC) and DHP. It is also illustrated that the proposed method has lower computational burden than conventional MPC, which is very beneficial for real-time tracking control.
Chuanqiang Lian, Xin Xu 0001, Hong Chen 0003, Haibo He
IEEE Trans. Cybern.2
2016 Fuzzy-Based Goal Representation Adaptive Dynamic Programming
abstract
In this paper, a novel nonlinear learning controller called fuzzy-based goal representation adaptive dynamic programming (Fuzzy-GrADP) is proposed. In the proposed GrADP method, a goal representation network is introduced to generate an adaptive internal reinforcement signal to the critic network to help the controller provide a general mapping between the input and output actions. Moreover, in the proposed architecture, the action network in the GrADP is improved by using the fuzzy hyperbolic model, which combines the merits of the fuzzy model and the neural network model. Based on the back-propagation technique, the parameters in the membership functions and the fuzzy rules are all undergo training and online adapting. The proposed controller is tested on two numerical benchmarks, and the simulation results show that the proposed controller outperforms the original adaptive dynamic fuzzy controller and the pure neural network-based GrADP controller. In addition, the proposed controller is further applied on a large multimachine power system for static var compensator damping control, where simulation results demonstrate the effectiveness of the proposed approach on real applications. Furthermore, in order to demonstrate the theoretical guarantee of the proposed method, Lyapunov stability analysis to support the proposed Fuzzy-GrADP approach has also been carried out.
Yufei Tang, Haibo He, Zhen Ni, Xiangnan Zhong, Dongbin Zhao, Xin Xu 0001
IEEE Trans. Fuzzy Syst.6
2015 A hierarchical path planning approach based on A⁎ and least-squares policy iteration for mobile robots
Lei Zuo 0002, Xin Xu 0001, Hao Fu 0001
Neurocomputing3
2015 GrDHP: A General Utility Function Representation for Dual Heuristic Dynamic Programming
abstract
A general utility function representation is proposed to provide the required derivable and adjustable utility function for the dual heuristic dynamic programming (DHP) design. Goal representation DHP (GrDHP) is presented with a goal network being on top of the traditional DHP design. This goal network provides a general mapping between the system states and the derivatives of the utility function. With this proposed architecture, we can obtain the required derivatives of the utility function directly from the goal network. In addition, instead of a fixed predefined utility function in literature, we conduct an online learning process for the goal network so that the derivatives of the utility function can be adaptively tuned over time. We provide the control performance of both the proposed GrDHP and the traditional DHP approaches under the same environment and parameter settings. The statistical simulation results and the snapshot of the system variables are presented to demonstrate the improved learning and controlling performance. We also apply both approaches to a power system example to further demonstrate the control capabilities of the GrDHP approach.
Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001, Danil V. Prokhorov
IEEE Trans. Neural Networks Learn. Syst.4
2015 Multiobjective Reinforcement Learning: A Comprehensive Overview
abstract
Reinforcement learning (RL) is a powerful paradigm for sequential decision-making under uncertainties, and most RL algorithms aim to maximize some numerical value which represents only one long-term objective. However, multiple long-term objectives are exhibited in many real-world decision and control systems, so recently there has been growing interest in solving multiobjective reinforcement learning (MORL) problems where there are multiple conflicting objectives. The aim of this paper is to present a comprehensive overview of MORL. The basic architecture, research topics, and naïve solutions of MORL are introduced at first. Then, several representative MORL approaches and some important directions of recent research are comprehensively reviewed. The relationships between MORL and other related research are also discussed, which include multiobjective optimization, hierarchical RL, and multiagent RL. Moreover, research challenges and open problems of MORL techniques are suggested.
Chunming Liu, Xin Xu 0001, Dewen Hu
IEEE Trans. Syst. Man Cybern. Syst.2
2014 Self-learning PD algorithms based on approximate dynamic programming for robot motion planning
abstract
Motion planning is a key technology of the navigation and control for mobile robots. However, when considering the complexity of exterior environment and mobile robot's kinematics and dynamics, the motion planning results obtained by some traditional methods are often hard to optimize. In this paper, we propose two self-learning PD algorithms to solve motion planning for mobile robots. We firstly utilize a virtual Proportional Derivative (PD) control strategy to transform the motion planning problem into an optimization problem of the virtual control policy. Afterwards, two approximate dynamic programming algorithms, which are the Least Squares Policy Iteration (LSPI) algorithm and the Dual Heuristic Programming (DHP) algorithm, are incorporated into the virtual control strategy to tune the PD parameters automatically, namely the LSPI-PD algorithm and the DHP-PD algorithm. Simulations have been performed to validate the effectiveness of the two algorithms, where the LSPI-PD algorithm is suitable for solving problems with discrete action spaces while the DHP-PD algorithm has an advantage in solving problems with continuous action spaces.
Huiyuan Yang, Xin Xu 0001, Chuanqiang Lian
IJCNN3
2014 Event-triggered reinforcement learning approach for unknown nonlinear continuous-time system
abstract
This paper provides an adaptive event-triggered method using adaptive dynamic programming (ADP) for the nonlinear continuous-time system. Comparing to the traditional method with fixed sampling period, the event-triggered method samples the state only when an event is triggered and therefore the computational cost is reduced. We demonstrate the theoretical analysis on the stability of the event-triggered method, and integrate it with the ADP approach. The system dynamics are assumed unknown. The corresponding ADP algorithm is given and the neural network techniques are applied to implement this method. The simulation results verify the theoretical analysis and justify the efficiency of the proposed event-triggered technique using the ADP approach.
Xiangnan Zhong, Zhen Ni, Haibo He, Xin Xu 0001, Dongbin Zhao
IJCNN4
2014 Reinforcement learning with automatic basis construction based on isometric feature mapping
Zhenhua Huang 0004, Xin Xu 0001, Lei Zuo 0002
Inf. Sci.2
2014 Reinforcement learning algorithms with function approximation: Recent advances and applications
Xin Xu 0001, Lei Zuo 0002, Zhenhua Huang 0004
Inf. Sci.1
2014 A Clustering-Based Graph Laplacian Framework for Value Function Approximation in Reinforcement Learning
abstract
In order to deal with the sequential decision problems with large or continuous state spaces, feature representation and function approximation have been a major research topic in reinforcement learning (RL). In this paper, a clustering-based graph Laplacian framework is presented for feature representation and value function approximation (VFA) in RL. By making use of clustering-based techniques, that is, K-means clustering or fuzzy C-means clustering, a graph Laplacian is constructed by subsampling in Markov decision processes (MDPs) with continuous state spaces. The basis functions for VFA can be automatically generated from spectral analysis of the graph Laplacian. The clustering-based graph Laplacian is integrated with a class of approximation policy iteration algorithms called representation policy iteration (RPI) for RL in MDPs with continuous state spaces. Simulation and experimental results show that, compared with previous RPI methods, the proposed approach needs fewer sample points to compute an efficient set of basis functions and the learning control performance can be improved for a variety of parameter settings.
Xin Xu 0001, Zhenhua Huang 0004, Daniel Graves, Witold Pedrycz
IEEE Trans. Cybern.1
2013 A combined hierarchical reinforcement learning based approach for multi-robot cooperative target searching in complex unknown environments
abstract
Effective cooperation of multi-robots in unknown environments is essential in many robotic applications, such as environment exploration and target searching. In this paper, a combined hierarchical reinforcement learning approach, together with a designed cooperation strategy, is proposed for the real-time cooperation of multi-robots in completely unknown environments. Unlike other algorithms that need an explicit environment model or select parameters by trial and error, the proposed cooperation method obtains all the required parameters automatically through learning. By integrating segmental options with the traditional MAXQ algorithm, the cooperation hierarchy is built. In new tasks, the designed cooperation method can control the multi-robot system to complete the task effectively. The simulation results demonstrate that the proposed scheme is able to effectively and efficiently lead a team of robots to cooperatively accomplish target searching tasks in completely unknown environments.
Simon X. Yang, Xin Xu 0001
ADPRL3
2013 Real-time tracking on adaptive critic design with uniformly ultimately bounded condition
abstract
In this paper, we proposed a new nonlinear tracking controller based on heuristic dynamic programming (HDP) with the tracking filter. Specifically, we integrate a goal network into the regular HDP design and provide the critic network with detailed internal reward signal to help the value function approximation. The architecture is explicitly explained with the tracking filter, goal network, critic network and action network, respectively. We provide the stability analysis of our proposed controller with Lyapunov approach. It is shown that the filtered tracking errors and the weights estimation errors in neural networks are all uniformly ultimately bounded (UUB) under certain conditions. Finally, we compare our proposed approach with regular HDP approach in virtual reality (VR)/Simulink environment to justify the improved control performance.
Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001
ADPRL5
2013 A novel approach for constructing basis functions in approximate dynamic programming for feedback control
abstract
This paper presents a novel approach for constructing basis functions in approximate dynamic programming (ADP) through the locally linear embedding (LLE) process. It considers the experience (sample) data as a high-dimensional space and the basis functions to be solved as a low-dimensional space. Through mapping the high-dimensional data into a single global coordinate system of lower dimensionality, the solved basis functions in low-dimensional space have the property that nearby experience data in the high dimensional space remain nearby and similarly co-located with respect to one in the low dimensional space. Thus, the obtained basis functions can precisely approximate the real value/action-value function. The simulation results show that the basis functions obtained by LLE can represent the final policy with a higher precision.
Jian Wang 0011, Zhenhua Huang 0004, Xin Xu 0001
ADPRL3
2013 A hierarchical reinforcement learning approach for optimal path tracking of wheeled mobile robots
Lei Zuo 0002, Xin Xu 0001, Chunming Liu, Zhenhua Huang 0004
Neural Comput. Appl.2
2013 Visual Saliency Based on Scale-Space Analysis in the Frequency Domain
abstract
We address the issue of visual saliency from three perspectives. First, we consider saliency detection as a frequency domain analysis problem. Second, we achieve this by employing the concept of nonsaliency. Third, we simultaneously consider the detection of salient regions of different size. The paper proposes a new bottom-up paradigm for detecting visual saliency, characterized by a scale-space analysis of the amplitude spectrum of natural images. We show that the convolution of the image amplitude spectrum with a low-pass Gaussian kernel of an appropriate scale is equivalent to an image saliency detector. The saliency map is obtained by reconstructing the 2D signal using the original phase and the amplitude spectrum, filtered at a scale selected by minimizing saliency map entropy. A Hypercomplex Fourier Transform performs the analysis in the frequency domain. Using available databases, we demonstrate experimentally that the proposed model can predict human fixation data. We also introduce a new image database and use it to show that the saliency detector can highlight both small and large salient regions, as well as inhibit repeated distractors in cluttered images. In addition, we show that it is able to predict salient regions on which people focus their attention.
Jian Li 0003, Martin D. Levine, Xiangjing An, Xin Xu 0001, Hangen He
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Goal Representation Heuristic Dynamic Programming on Maze Navigation
abstract
Goal representation heuristic dynamic programming (GrHDP) is proposed in this paper to demonstrate online learning in the Markov decision process. In addition to the (external) reinforcement signal in literature, we develop an adaptively internal goal/reward representation for the agent with the proposed goal network. Specifically, we keep the actor-critic design in heuristic dynamic programming (HDP) and include a goal network to represent the internal goal signal, to further help the value function approximation. We evaluate our proposed GrHDP algorithm on two 2-D maze navigation problems, and later on one 3-D maze navigation problem. Compared to the traditional HDP approach, the learning performance of the agent is improved with our proposed GrHDP approach. In addition, we also include the learning performance with two other reinforcement learning algorithms, namely Sarsa(λ) and Q-learning, on the same benchmarks for comparison. Furthermore, in order to demonstrate the theoretical guarantee of our proposed method, we provide the characteristics analysis toward the convergence of weights in neural networks in our GrHDP approach.
Zhen Ni, Haibo He, Jinyu Wen, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2013 Online Learning Control Using Adaptive Critic Designs With Sparse Kernel Machines
abstract
In the past decade, adaptive critic designs (ACDs), including heuristic dynamic programming (HDP), dual heuristic programming (DHP), and their action-dependent ones, have been widely studied to realize online learning control of dynamical systems. However, because neural networks with manually designed features are commonly used to deal with continuous state and action spaces, the generalization capability and learning efficiency of previous ACDs still need to be improved. In this paper, a novel framework of ACDs with sparse kernel machines is presented by integrating kernel methods into the critic of ACDs. To improve the generalization capability as well as the computational efficiency of kernel machines, a sparsification method based on the approximately linear dependence analysis is used. Using the sparse kernel machines, two kernel-based ACD algorithms, that is, kernel HDP (KHDP) and kernel DHP (KDHP), are proposed and their performance is analyzed both theoretically and empirically. Because of the representation learning and generalization capability of sparse kernel machines, KHDP and KDHP can obtain much better performance than previous HDP and DHP with manually designed neural networks. Simulation and experimental results of two nonlinear control problems, that is, a continuous-action inverted pendulum problem and a ball and plate control problem, demonstrate the effectiveness of the proposed kernel ACD methods.
Xin Xu 0001, Zhongsheng Hou, Chuanqiang Lian, Haibo He
IEEE Trans. Neural Networks Learn. Syst.1
2012 A Novel Feature Sparsification Method for Kernel-Based Approximate Policy Iteration
Zhenhua Huang 0004, Chunming Liu, Xin Xu 0001, Chuanqiang Lian
ISNN (1)3
2012 A Rapid Sparsification Method for Kernel Machines in Approximate Policy Iteration
Chunming Liu, Zhenhua Huang 0004, Xin Xu 0001, Lei Zuo 0002
ISNN (1)3
2012 Stereo matching using weighted dynamic programming on a single-direction four-connected tree
Tingbo Hu, Baojun Qi, Tao Wu 0001, Xin Xu 0001, Hangen He
Comput. Vis. Image Underst.4
2012 A sparse Gaussian process regression model for tourism demand forecasting in Hong Kong
Qi Wu 0003, Rob Law 0001, Xin Xu 0001
Expert Syst. Appl.3
2011 Adaptive sample collection using active learning for kernel-based approximate policy iteration
abstract
Approximate policy iteration (API) has been shown to be a class of reinforcement learning methods with stability and sample efficiency. However, sample collection is still an open problem which is critical to the performance of API methods. In this paper, a novel adaptive sample collection strategy using active learning-based exploration is proposed to enhance the performance of kernel-based API. In this strategy, an online kernel-based least squares policy iteration (KLSPI) method is adopted to construct nonlinear features and approximate the Q-function simultaneously. Therefore, more representative samples can be obtained for value function approximation. Simulation results on typical learning control problems illustrate that by using the proposed strategy, the performance of KLSPI can be improved remarkably.
Chunming Liu, Xin Xu 0001, Haiyun Hu, Bin Dai 0001
ADPRL2
2011 Adaptive Dual Heuristic Programming Based on Delta-Bar-Delta Learning Rule
Xin Xu 0001, Chuanqiang Lian
ISNN (3)2
2011 Adaptive Kernel-Width Selection for Kernel-Based Least-Squares Policy Iteration Algorithm
Xin Xu 0001, Lei Zuo 0002, Zhaobin Li, Jian Wang 0011
ISNN (2)2
2011 A novel multi-agent reinforcement learning approach for job scheduling in Grid computing
Xin Xu 0001, Chunming Liu
Future Gener. Comput. Syst.2
2011 Continuous-action reinforcement learning with fast policy search and adaptive basis function selection
Xin Xu 0001, Chunming Liu, Dewen Hu
Soft Comput.1
2011 A Complementary Modularized Ramp Metering Approach Based on Iterative Learning Control and ALINEA
abstract
Ramp metering is an effective tool for traffic management on freeway networks. In this paper, we apply iterative learning control (ILC) to address ramp metering in a macroscopic-level freeway environment. By formulating the original ramp metering problem as an output regulating and disturbance rejection problem, ILC has been applied to control the traffic response. The learning mechanism is further combined with Asservissement Linéaire d'Entrée Autoroutière (ALINEA) in a complementary manner to achieve the desired control performance. The ILC-based ramp metering strategy and the modified modularized ramp metering approach based on ILC and ALINEA in the presence of input constraints are also analyzed to highlight the advantages and the robustness of the proposed methods. Extensive simulations are given to verify the effectiveness of the proposed approaches.
Zhongsheng Hou, Xin Xu 0001, Jianxin Xu 0001, Gang Xiong 0001
IEEE Trans. Intell. Transp. Syst.2
2011 Variational Inference for Infinite Mixtures of Gaussian Processes With Applications to Traffic Flow Prediction
abstract
This paper proposes a new variational approximation for infinite mixtures of Gaussian processes. As an extension of the single Gaussian process regression model, mixtures of Gaussian processes can characterize varying covariances or multimodal data and reduce the deficiency of the computationally cubic complexity of the single Gaussian process model. The infinite mixture of Gaussian processes further integrates a Dirichlet process prior to allowing the number of mixture components to automatically be determined from data. We use variational inference and a truncated stick-breaking representation of the Dirichlet process to approximate the posterior of hidden variables involved in the model. To fix the hyperparameters of the model, the variational EM algorithm and a greedy algorithm are employed. In addition to presenting the variational infinite-mixture model, we apply it to the problem of traffic flow prediction. Experiments with comparisons to other approaches show the effectiveness of the proposed model.
Shiliang Sun, Xin Xu 0001
IEEE Trans. Intell. Transp. Syst.2
2011 Data-Driven Intelligent Transportation Systems: A Survey
abstract
For the last two decades, intelligent transportation systems (ITS) have emerged as an efficient way of improving the performance of transportation systems, enhancing travel security, and providing more choices to travelers. A significant change in ITS in recent years is that much more data are collected from a variety of sources and can be processed into various forms for different stakeholders. The availability of a large amount of data can potentially lead to a revolution in ITS development, changing an ITS from a conventional technology-driven system into a more powerful multifunctional data-driven intelligent transportation system (D2ITS) : a system that is vision, multisource, and learning algorithm driven to optimize its performance. Furthermore, D2ITS is trending to become a privacy-aware people-centric more intelligent system. In this paper, we provide a survey on the development of D2ITS, discussing the functionality of its key components and some deployment issues associated with D2ITS Future research directions for the development of D2ITS is also presented.
Junping Zhang, Fei-Yue Wang 0001, Kunfeng Wang, Wei-Hua Lin, Xin Xu 0001
IEEE Trans. Intell. Transp. Syst.5
2011 Incremental Learning From Stream Data
abstract
Recent years have witnessed an incredibly increasing interest in the topic of incremental learning. Unlike conventional machine learning situations, data flow targeted by incremental learning becomes available continuously over time. Accordingly, it is desirable to be able to abandon the traditional assumption of the availability of representative training data during the training period to develop decision boundaries. Under scenarios of continuous data flow, the challenge is how to transform the vast amount of stream raw data into information and knowledge representation, and accumulate experience over time to support future decision-making process. In this paper, we propose a general adaptive incremental learning framework named ADAIN that is capable of learning from continuous raw data, accumulating experience over time, and using such knowledge to improve future learning and prediction performance. Detailed system level architecture and design strategies are presented in this paper. Simulation results over several real-world data sets are used to validate the effectiveness of this method.
Haibo He, Sheng Chen 0005, Kang Li 0002, Xin Xu 0001
IEEE Trans. Neural Networks4
2011 Hierarchical Approximate Policy Iteration With Binary-Tree State Space Decomposition
abstract
In recent years, approximate policy iteration (API) has attracted increasing attention in reinforcement learning (RL), e.g., least-squares policy iteration (LSPI) and its kernelized version, the kernel-based LSPI algorithm. However, it remains difficult for API algorithms to obtain near-optimal policies for Markov decision processes (MDPs) with large or continuous state spaces. To address this problem, this paper presents a hierarchical API (HAPI) method with binary-tree state space decomposition for RL in a class of absorbing MDPs, which can be formulated as time-optimal learning control tasks. In the proposed method, after collecting samples adaptively in the state space of the original MDP, a learning-based decomposition strategy of sample sets was designed to implement the binary-tree state space decomposition process. Then, API algorithms were used on the sample subsets to approximate local optimal policies of sub-MDPs. The original MDP was decomposed into a binary-tree structure of absorbing sub-MDPs, constructed during the learning process, thus, local near-optimal policies were approximated by API algorithms with reduced complexity and higher precision. Furthermore, because of the improved quality of local policies, the combined global policy performed better than the near-optimal policy obtained by a single API algorithm in the original MDP. Three learning control problems, including path-tracking control of a real mobile robot, were studied to evaluate the performance of the HAPI method. With the same setting for basis function selection and sample collection, the proposed HAPI obtained better near-optimal policies than previous API methods such as LSPI and KLSPI.
Xin Xu 0001, Chunming Liu, Simon X. Yang, Dewen Hu
IEEE Trans. Neural Networks1
2010 An adaptive roadmap guided Multi-RRTs strategy for single query path planning
abstract
During the past decade, Rapidly-exploring Random Tree (RRT) and its variants are shown to be powerful sampling based single query path planning approaches for robots in high-dimensional configuration space. However, the performance of such tree-based planners that rely on uniform sampling strategy degrades significantly when narrow passages are contained in the configuration space. Given the assumption that computation resources should be allocated in proportion the geometric complexity of local region, we present a novel single-query Multi-RRTs path planning framework that employs an improved Bridge Test algorithm to identify global important roadmaps in narrow passages. Multiple trees can grown from these sampled roadmaps to explore sub-regions which are difficult to reach. The probability of selecting one particular tree for expansion and connection, which can dynamically updated by on-line learning algorithm based on the historic results of exploration, guides the tree through narrow passage rapidly. Experimental results show that the proposed approach gives substantial improvement in planning efficiency over a wide range of single-query path planning problems.
Wei Wang 0434, Xin Xu 0001, Simon X. Yang
ICRA3
2009 Reordering Sparsification of Kernel Machines in Approximate Policy Iteration
Chunming Liu, Jinze Song, Xin Xu 0001
ISNN (2)3
2009 Reinforcement Learning Control of a Real Mobile Robot Using Approximate Policy Iteration
Xin Xu 0001, Chunming Liu, Qiping Yuan
ISNN (3)2
2008 Self-learning path-tracking control of autonomous vehicles using kernel-based approximate dynamic programming
abstract
With the fast development of robotics and intelligent vehicles, there has been much research work on modeling and motion control of autonomous vehicles. However, due to model complexity, and unknown disturbances from dynamic environment, the motion control of autonomous vehicles is still a difficult problem. In this paper, a novel self-learning path-tracking control method is proposed for a car-like robotic vehicle, where kernel-based approximate dynamic programming (ADP) is used to optimize the controller performance with little prior knowledge on vehicle dynamics. The kernel-based ADP method is a recently developed reinforcement learning algorithm called kernel least-squares policy iteration (KLSPI), which uses kernel methods with automatic feature selection in policy evaluation to get better generalization performance and learning efficiency. By using KLSPI, the lateral control performance of the robotic vehicle can be optimized in a self-learning and data-driven style. Compared with previous learning control methods, the proposed method has advantages in learning efficiency and automatic feature selection. Simulation results show that the proposed method can obtain an optimized path-tracking control policy only in a few iterations, which will be very practical for real applications.
Xin Xu 0001, Bin Dai 0001, Hangen He
IJCNN1
2007 Classification of Business Travelers Using SVMs Combined with Kernel Principal Component Analysis
Xin Xu 0001, Rob Law 0001, Tao Wu 0001
ADMA1
2007 A Kernel-Based Reinforcement Learning Approach to Dynamic Behavior Modeling of Intrusion Detection
Xin Xu 0001, Yirong Luo
ISNN (1)1
2007 Kernel-Based Least Squares Policy Iteration for Reinforcement Learning
abstract
In this paper, we present a kernel-based least squares policy iteration (KLSPI) algorithm for reinforcement learning (RL) in large or continuous state spaces, which can be used to realize adaptive feedback control of uncertain dynamic systems. By using KLSPI, near-optimal control policies can be obtained without much a priori knowledge on dynamic models of control plants. In KLSPI, Mercer kernels are used in the policy evaluation of a policy iteration process, where a new kernel-based least squares temporal-difference algorithm called KLSTD-Q is proposed for efficient policy evaluation. To keep the sparsity and improve the generalization ability of KLSTD-Q solutions, a kernel sparsification procedure based on approximate linear dependency (ALD) is performed. Compared to the previous works on approximate RL methods, KLSPI makes two progresses to eliminate the main difficulties of existing results. One is the better convergence and (near) optimality guarantee by using the KLSTD-Q algorithm for policy evaluation with high precision. The other is the automatic feature selection using the ALD-based kernel sparsification. Therefore, the KLSPI algorithm provides a general RL method with generalization performance and convergence guarantee for large-scale Markov decision problems (MDPs). Experimental results on a typical RL task for a stochastic chain problem demonstrate that KLSPI can consistently achieve better learning efficiency and policy quality than the previous least squares policy iteration (LSPI) algorithm. Furthermore, the KLSPI method was also evaluated on two nonlinear feedback control problems, including a ship heading control problem and the swing up control of a double-link underactuated pendulum called acrobot. Simulation results illustrate that the proposed method can optimize controller performance using little a priori information of uncertain dynamic systems. It is also demonstrated that KLSPI can be applied to online learning control by incorporating an initial controller to ensure online performance.
Xin Xu 0001, Dewen Hu, Xicheng Lu
IEEE Trans. Neural Networks1
2006 Local Stability Analysis of Maximum Nongaussianity Estimation in Independent Component Analysis
Xin Xu 0001, Dewen Hu
ISNN (1)2
2005 An Adaptive Network Intrusion Detection Method Based on PCA and Support Vector Machines
Xin Xu 0001
ADMA1
2005 A Reinforcement Learning Approach for Host-Based Intrusion Detection Using Sequences of System Calls
Xin Xu 0001, Tao Xie 0005
ICIC (1)1
2005 Self-adaptive FastICA Based on Generalized Gaussian Model
Xin Xu 0001, Dewen Hu
ISNN (1)2
2004 Mobile Robot Path-Tracking Using an Adaptive Critic Learning PD Controller
Xin Xu 0001, Dewen Hu
ISNN (2)1
2002 Evolutionary adaptive-critic methods for reinforcement learning
abstract
In this paper, a novel hybrid learning method is proposed for reinforcement learning problems with continuous state and action spaces. The reinforcement learning problems are modeled as Markov decision processes (MDPs) and the hybrid learning method combines evolutionary algorithms with gradient-based adaptive heuristic critic (AHC) algorithms to approximate the optimal policy of MDPs. The suggested method takes the advantages of evolutionary learning and gradient-based reinforcement learning to solve reinforcement learning problems. Simulation results on the learning control of an acrobot illustrate the efficiency of the presented method.
Xin Xu 0001, Hangen He, Dewen Hu
IEEE Congress on Evolutionary Computation1
2002 Efficient Reinforcement Learning Using Recursive Least-Squares Methods
abstract
The recursive least-squares (RLS) algorithm is one of the most well-known algorithms used in adaptive filtering, system identification and adaptive control. Its popularity is mainly due to its fast convergence speed, which is considered to be optimal in practice. In this paper, RLS methods are used to solve reinforcement learning problems, where two new reinforcement learning algorithms using linear value function approximators are proposed and analyzed. The two algorithms are called RLS-TD(lambda) and Fast-AHC (Fast Adaptive Heuristic Critic), respectively. RLS-TD(lambda) can be viewed as the extension of RLS-TD(0) from lambda=0 to general lambda within interval [0,1], so it is a multi-step temporal-difference (TD) learning algorithm using RLS methods. The convergence with probability one and the limit of convergence of RLS-TD(lambda) are proved for ergodic Markov chains. Compared to the existing LS-TD(lambda) algorithm, RLS-TD(lambda) has advantages in computation and is more suitable for online learning. The effectiveness of RLS-TD(lambda) is analyzed and verified by learning prediction experiments of Markov chains with a wide range of parameter settings. The Fast-AHC algorithm is derived by applying the proposed RLS-TD(lambda) algorithm in the critic network of the adaptive heuristic critic method. Unlike conventional AHC algorithm, Fast-AHC makes use of RLS methods to improve the learning-prediction efficiency in the critic. Learning control experiments of the cart-pole balancing and the acrobot swing-up problems are conducted to compare the data efficiency of Fast-AHC with conventional AHC. From the experimental results, it is shown that the data efficiency of learning control can also be improved by using RLS methods in the learning-prediction process of the critic. The performance of Fast-AHC is also compared with that of the AHC method using LS-TD(lambda). Furthermore, it is demonstrated in the experiments that different initial values of the variance matrix in RLS-TD(lambda) are required to get better performance not only in learning prediction but also in learning control. The experimental results are analyzed based on the existing theoretical work on the transient phase of forgetting factor RLS methods.
Xin Xu 0001, Hangen He, Dewen Hu
J. Artif. Intell. Res.1