EDBT 2026 Demo / reviewers in the wild / expert
Dongbin Zhao
dblp:40/255
· DBLP profile ↗
187ranked-venue papers
18as first author
65since 2021 · last 2026
0000-0001-8218-9633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 155 · 15 first-author · 50 since 2021Human-computer interaction and ubiquitous computing · 17 · 1 first-author · 9 since 2021Systems, architecture and hardware · 16 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral InspectionabstractMinghui Jia, Qichao Zhang, Ali Luo, Linjing Li, Shuo Ye, Hailing Lu, Wen Hou, Dongbin Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Minghui Jia, A-Li Luo, Linjing Li, Shuo Ye, Hailing Lu, Wen Hou, Dongbin Zhao |
ACL (1) | 8 |
| 2026 | Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical MatchingabstractBo Lv, Jingbo Sun, Jianwei Lv, Chen Tang, Shaojie Zhang, Nayu Liu, Guoxin Yu, Zihao Li, Qichao Zhang, Dongbin Zhao, Ping Luo, Yue Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingbo Sun 0001, Jianwei Lv, Nayu Liu, Guoxin Yu, Dongbin Zhao, Ping Luo 0002, Yue Yu 0001 |
ACL (1) | 10 |
| 2026 | Hybrid Event-Triggered Tracking Control With Critic Learning for Nonlinear Networked SystemsabstractIn this article, a novel hybrid event-triggered (ET) control framework is constructed based on the adaptive critic technique, aiming to address the optimal tracking issue of discrete-time nonlinear networked control systems. First, an augmented plant is created by combining the system state with the reference trajectory, transforming the optimal tracking control design into the optimal regulation problem of the reconstructed nonlinear error system. Subsequently, to conserve communication network resources and ensure the stability of the error system, a hybrid ET mechanism is developed to determine a constant interval for event silence. This approach not only alleviates the limited network bandwidth but also eliminates the need for continuous evaluation of triggering conditions, as seen in traditional event-based methods. Regarding algorithm implementation, the model, critic, and action networks are established to execute the online adaptive critic algorithm, which allows the tracking control policy to be adjusted in real-time to reach the optimal level. Finally, an experimental plant with nonlinear characteristics is presented to illustrate the overall performance of the proposed online tracking control method with the hybrid ET mechanism. Ding Wang 0001, Lingzhi Hu, Dongbin Zhao |
IEEE Trans. Cybern. | 3 |
| 2026 | TeViR: Text-to-Video Reward With Diffusion Models for Efficient Reinforcement LearningabstractDeveloping scalable and generalizable reward engineering for reinforcement learning (RL) is crucial for creating general-purpose agents, especially in the challenging domain of robotic manipulation. While recent advances in reward engineering with vision–language models (VLMs) have shown promise, their sparse reward nature significantly limits sample efficiency. This article introduces text-to-video reward (TeViR), a novel method that leverages a pretrained text-to-video diffusion model to generate dense rewards by comparing the predicted image sequence with current observations. Experimental results across 13 simulation and real-world robotic tasks demonstrate that TeViR outperforms traditional methods leveraging sparse rewards and other state-of-the-art (SOTA) methods, achieving better sample efficiency and performance without ground truth environmental rewards. TeViR’s ability to efficiently guide agents in complex environments highlights its potential to advance RL applications in robotic manipulation. Yuhui Chen, Haoran Li 0010, Zhennan Jiang, Haowei Wen, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2026 | SoAD: Safety-Oriented Value Estimation for Enhanced Closed-Loop End-to-End Autonomous DrivingabstractEnd-to-end (E2E) autonomous driving systems, which map sensory inputs directly to vehicle planning, have garnered attention for harnessing the potential of data-driven methodologies in motion planning. However, current methods face two limitations that undermine their safety performance in closed-loop driving tasks. First, the predominant imitation learning (IL) paradigm overlooks long-term safety beyond predefined planning horizons, potentially guiding the ego vehicle into hazardous states. Second, the lack of reliable online evaluation mechanisms limits real-time responses to safety risks. To overcome these challenges, we propose SoAD, a safety-oriented E2E framework that integrates long-term safety awareness into planning. SoAD is distinguished by a reinforcement learning (RL)-based value estimation module to quantify the safety of planned trajectories, and a vector world model (VWM) to generate interaction-aware future rollouts. During training, the system benefits from value-guided fine-tuning (VFT) that optimizes the planning distribution to favor safer trajectories. In closed-loop deployment, a planning rescoring (PRS) mechanism is designed to perform reliable online evaluation by combining ego-conditional predictions from the VWM with corresponding safety value estimates. Experimental results on the Bench2Drive closed-loop benchmark demonstrate the state-of-the-art (SoTA) performance of SoAD, achieving a 66.40% improvement in driving score (DS) compared to the vectorized scene representation for efficient autonomous driving (VAD) baseline, while zero-shot evaluation on the driving in occlusion simulation (DOS) benchmark further highlights its strong generalization ability. Yinfeng Gao, Deqing Liu, Yupeng Zheng, Dawei Ding 0001, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2025 | In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement LearningabstractOffline preference-based reinforcement learning (PbRL) typically operates in two phases: first, use human preferences to learn a reward model and annotate rewards for a reward-free offline dataset; second, learn a policy by optimizing the learned reward via offline RL. However, accurately modeling step-wise rewards from trajectory-level preference feedback presents inherent challenges. The reward bias introduced, particularly the overestimation of predicted rewards, leads to optimistic trajectory stitching, which undermines the pessimism mechanism critical to the offline RL phase. To address this challenge, we propose In-Dataset Trajectory Return Regularization (DTR) for offline PbRL, which leverages conditional sequence modeling to mitigate the risk of learning inaccurate trajectory stitching under reward bias. Specifically, DTR employs Decision Transformer and TD-Learning to strike a balance between maintaining fidelity to the behavior policy with high in-dataset trajectory returns and selecting optimal actions based on high reward labels. Additionally, we introduce an ensemble normalization technique that effectively integrates multiple reward models, balancing the trade-off between reward differentiation and accuracy. Empirical evaluations on various benchmarks demonstrate the superiority of DTR over other state-of-the-art baselines. Songjun Tu, Jingbo Sun 0001, Yaocheng Zhang, Dongbin Zhao |
AAAI | 7 |
| 2025 | ARAC: Adaptive Regularized Multi-Agent Soft Actor-Critic in Graph-Structured Adversarial GamesabstractIn graph-structured multi-agent reinforcement learning (MARL) adversarial tasks such as pursuit and confrontation, agents must coordinate under highly dynamic interactions, where sparse rewards hinder efficient policy learning. We propose Adaptive Regularized Multi-Agent Soft Actor-Critic (ARAC), which integrates an attention-based graph neural network (GNN) for modeling agent dependencies with an adaptive divergence regularization mechanism. The GNN enables expressive representation of spatial relations and state features in graph environments. Divergence regularization can serve as policy guidance to alleviate the sparse reward problem, but it may lead to suboptimal convergence when the reference policy itself is imperfect. The adaptive divergence regularization mechanism enables the framework to exploit reference policies for efficient exploration in the early stages, while gradually reducing reliance on them as training progresses to avoid inheriting their limitations. Experiments in pursuit and confrontation scenarios demonstrate that ARAC achieves faster convergence, higher final success rates, and stronger scalability across varying numbers of agents compared with MARL baselines, highlighting its effectiveness in complex graph-structured environments. Ruochuan Shi, Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
DAI | 4 |
| 2025 | RLAE: Reinforcement Learning-Assisted Ensemble for LLMsabstractEnsembling large language models (LLMs) can effectively combine diverse strengths of different models, offering a promising approach to enhance performance across various tasks.However, existing methods typically rely on fixed weighting strategies that fail to adapt to the dynamic, context-dependent characteristics of LLM capabilities.In this work, we propose Reinforcement Learning-Assisted Ensemble for LLMs (RLAE), a novel framework that reformulates LLM ensemble through the lens of a Markov Decision Process (MDP).Our approach introduces a RL agent that dynamically adjusts ensemble weights by considering both input context and intermediate generation states, with the agent being trained using rewards that directly correspond to the quality of final outputs.We implement RLAE using both single-agent and multi-agent reinforcement learning algorithms (RLAE PPO and RLAE MAPPO ), demonstrating substantial improvements over conventional ensemble methods.Extensive evaluations on a diverse set of tasks show that RLAE outperforms existing approaches by up to 3.3% accuracy points, offering a more effective framework for LLM ensembling.Furthermore, our method exhibits superior generalization capabilities across different tasks without the need for retraining, while simultaneously achieving lower time latency.The source code is available at here. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin, Dongbin Zhao |
EMNLP | 7 |
| 2025 | World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World ModelabstractEnd-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an informative driving world model to enable perception annotation-free, end-to-end planning via self-supervised learning. In this paper, we present World4Drive, an end-to-end autonomous driving framework that employs vision foundation models to build latent world models for generating and evaluating multi-modal planning trajectories. Specifically, World4Drive first extracts scene features, including driving intention and world latent representations enriched with spatial-semantic priors provided by vision foundation models. It then generates multi-modal planning trajectories based on current scene features and driving intentions and predicts multiple intention-driven future states within the latent space. Finally, it introduces a world model selector module to evaluate and select the best trajectory. We achieve perception annotation-free, end-to-end planning through self-supervised alignment between actual future observations and predicted observations reconstructed from the latent space. World4Drive achieves state-of-the-art performance without manual perception annotations on both the open-loop nuScenes and closed-loop NavSim benchmarks, demonstrating an 18.1\% relative reduction in L2 error, 46.7% lower collision rate, and 3.75 faster training convergence. Codes will be accessed at https://github.com/ucaszyp/World4Drive. Yupeng Zheng, Pengxuan Yang, Zebin Xing, Yuhang Zheng 0004, Yinfeng Gao, Pengfei Li 0007, Zhongpu Xia, Peng Jia 0007, Xianpeng Lang, Dongbin Zhao |
ICCV | 12 |
| 2025 | Empowering LLM Agents with Zero-Shot Optimal Decision-Making through Q-learningabstractLarge language models (LLMs) are trained on extensive text data to gain general comprehension capability. Current LLM agents leverage this ability to make zero- or few-shot decisions without reinforcement learning (RL) but fail in making optimal decisions, as LLMs inherently perform next-token prediction rather than maximizing rewards. In contrast, agents trained via RL could make optimal decisions but require extensive environmental interaction. In this work, we develop an algorithm that combines the zero-shot capabilities of LLMs with the optimal decision-making of RL, referred to as the Model-based LLM Agent with Q-Learning (MLAQ). MLAQ employs Q-learning to derive optimal policies from transitions within memory. However, unlike RL agents that collect data from environmental interactions, MLAQ constructs an imagination space fully based on LLM to perform imaginary interactions for deriving zero-shot policies. Our proposed UCB variant generates high-quality imaginary data through interactions with the LLM-based world model, balancing exploration and exploitation while ensuring a sub-linear regret bound. Additionally, MLAQ incorporates a mixed-examination mechanism to filter out incorrect data. We evaluate MLAQ in benchmarks that present significant challenges for existing LLM agents. Results show that MLAQ achieves a optimal rate of over 90\% in tasks where other methods struggle to succeed. Additional experiments are conducted to reach the conclusion that introducing model-based RL into LLM agents shows significant potential to improve optimal decision-making ability. Our interactive website is available at http://mlaq.site. Jiajun Chai, Yuqian Fu, Dongbin Zhao, Yuanheng Zhu |
ICLR | 4 |
| 2025 | INS: Interaction-aware Synthesis to Enhance Offline Multi-agent Reinforcement LearningabstractData scarcity in offline multi-agent reinforcement learning (MARL) is a key challenge for real-world applications. Recent advances in offline single-agent reinforcement learning (RL) demonstrate the potential of data synthesis to mitigate this issue.
However, in multi-agent systems, interactions between agents introduce additional challenges. These interactions complicate the synthesis of multi-agent datasets, leading to data distortion when inter-agent interactions are neglected. Furthermore, the quality of the synthetic dataset is often constrained by the original dataset. To address these challenges, we propose **INteraction-aware Synthesis (INS)**, which synthesizes high-quality multi-agent datasets using diffusion models. Recognizing the sparsity of inter-agent interactions, INS employs a sparse attention mechanism to capture these interactions, ensuring that the synthetic dataset reflects the underlying agent dynamics. To overcome the limitation of diffusion models requiring continuous variables, INS implements a bit action module, enabling compatibility with both discrete and continuous action spaces. Additionally, we incorporate a select mechanism to prioritize transitions with higher estimated values, further enhancing the dataset quality. Experimental results across multiple datasets in MPE and SMAC environments demonstrate that INS consistently outperforms existing methods, resulting in improved downstream policy performance and superior dataset metrics. Notably, INS can synthesize high-quality data using only 10% of the original dataset, highlighting its efficiency in data-limited scenarios. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Dongbin Zhao |
ICLR | 5 |
| 2025 | Divergence-Regularized Discounted Aggregation: Equilibrium Finding in Multiplayer Partially Observable Stochastic GamesabstractThis paper presents Divergence-Regularized Discounted Aggregation (DRDA), a multi-round learning system for solving partially observable stochastic games (POSGs). DRDA is based on action values and applicable to multiplayer POSGs, which can unify normal-form games (NFGs), extensive-form games (EFGs) with perfect recall, and Markov games (MGs). In each single round, DRDA can be viewed as a discounted variant of Follow the Regularized Leader (FTRL) under a general value function for POSGs. While previous studies on discounted FTRL have demonstrated its last-iterate convergence towards quantal response equilibrium (QRE) in NFGs, this paper extends the theoretical results to POSGs under divergence regularization and generalizes the QRE concept of Nash distribution. The linear last-iterate convergence of single-round DRDA to its rest point is proved under the assumption on the hypomonotonicity of the game. When the rest point is unique, it induces the unique Nash distribution defined in the POSG, which has a bounded deviation from Nash equilibrium (NE). Under multiple learning rounds, DRDA keeps replacing the base policy for divergence regularization with the policy at the rest point in the previous round. It is further proved that the limit point of multi-round DRDA must be an exact NE (rather than a QRE). In experiments, discrete-time DRDA can converge to NE at a near-exponential rate in (multiplayer) NFGs and outperform the existing baselines for EFGs, MGs, and typical POSGs. Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
ICLR | 3 |
| 2025 | Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward RepresentationabstractOnline unsupervised reinforcement learning (URL) can discover diverse skills via reward-free pre-training and exhibits impressive downstream task adaptation abilities through further fine-tuning.
However, online URL methods face challenges in achieving zero-shot generalization, i.e., directly applying pre-trained policies to downstream tasks without additional planning or learning.
In this paper, we propose a novel Dual-Value Forward-Backward representation (DVFB) framework with a contrastive entropy intrinsic reward to achieve both zero-shot generalization and fine-tuning adaptation in online URL.
On the one hand, we demonstrate that poor exploration in forward-backward representations can lead to limited data diversity in online URL, impairing successor measures, and ultimately constraining generalization ability.
To address this issue, the DVFB framework learns successor measures through a skill value function while promoting data diversity through an exploration value function, thus enabling zero-shot generalization.
On the other hand, and somewhat surprisingly, by employing a straightforward dual-value fine-tuning scheme combined with a reward mapping technique, the pre-trained policy further enhances its performance through fine-tuning on downstream tasks, building on its zero-shot performance.
Through extensive multi-task generalization experiments, DVFB demonstrates both superior zero-shot generalization (outperforming on all 12 tasks) and fine-tuning adaptation (leading on 10 out of 12 tasks) abilities, surpassing state-of-the-art URL methods. Jingbo Sun 0001, Songjun Tu, Haoran Li 0010, Xin Liu 0039, Yaran Chen, Dongbin Zhao |
ICLR | 8 |
| 2025 | LeAffordNav: Enhancing Open-vocabulary Mobile Manipulation with LLM-guided Exploration and Affordance-aware NavigationabstractOpen-vocabulary mobile manipulation is a fundamental task for robotic assistants. However, Inefficient exploration and hand-off errors between different skills pose significant challenges to completing mobile manipulation tasks. In this paper, we propose a novel method named LeAffordNav, which is composed of LLM-guided exploration and Affordance-aware Navigation to address these challenges. LLM-guided exploration introduces LLMs to combine commonsense inference and frontier-based exploration, and achieves the balance between exploration and finding the target object. Considering the manipulability of the robot arm and the accessibility of the robot, we propose Affordance-aware Navigation which predicts the affordance of the mobile manipulation to reduce the hand-off errors between navigation and manipulation. Experiments on the HomeRobot benchmark show that LeAffordNav achieves new state-of-the-art performance, with a 20% higher success rate than the previous best. The code is available at https: //github.com/Cyuanwen/LeAffordNav. Yuanwen Chen, Haoran Li 0010, Yaran Chen, Dongbin Zhao |
ICME | 4 |
| 2025 | Constrained Exploitability Descent: An Offline Reinforcement Learning Method for Finding Mixed-Strategy Nash EquilibriumabstractThis paper proposes Constrained Exploitability Descent (CED), a model-free offline reinforcement learning (RL) algorithm for solving adversarial Markov games (MGs). CED combines the game-theoretical approach of Exploitability Descent (ED) with policy constraint methods from offline RL. While policy constraints can perturb the optimal pure-strategy solutions in single-agent scenarios, we find the side effect less detrimental in adversarial games, where the optimal policy can be a mixed-strategy Nash equilibrium. We theoretically prove that, under the uniform coverage assumption on the dataset, CED converges to a stationary point in deterministic two-player zero-sum Markov games. We further prove that the min-player policy at the stationary point follows the property of mixed-strategy Nash equilibrium in MGs. Compared to the model-based ED method that optimizes the max-player policy, our CED method no longer relies on a generalized gradient. Experiments in matrix games, a tree-form game, and an infinite-horizon soccer game verify that CED can find an equilibrium policy for the min-player as long as the offline dataset guarantees uniform coverage. Besides, CED achieves a significantly lower NashConv compared to an existing pessimism-based method and can gradually improve the behavior policy even under non-uniform data coverages. When combined with neural networks, CED also outperforms behavior cloning and offline self-play in a large-scale two-team robotic combat game. Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
ICML | 3 |
| 2025 | DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyabstractDiplomacy is a complex multiplayer game that re- quires both cooperation and competition, posing significant challenges for AI systems. Traditional methods rely on equilibrium search to generate extensive game data for training, which demands substantial computational resources. Large Lan- guage Models (LLMs) offer a promising alterna- tive, leveraging pre-trained knowledge to achieve strong performance with relatively small-scale fine-tuning. However, applying LLMs to Diplo- macy remains challenging due to the exponential growth of possible action combinations and the intricate strategic interactions among players. To address this challenge, we propose DipLLM, a fine-tuned LLM-based agent that learns equilib- rium policies for Diplomacy. DipLLM employs an autoregressive factorization framework to sim- plify the complex task of multi-unit action assign- ment into a sequence of unit-level decisions. By defining an equilibrium policy within this frame- work as the learning objective, we fine-tune the model using only 1.5% of the data required by the state-of-the-art Cicero model, surpassing its per- formance. Our results demonstrate the potential of fine-tuned LLMs for tackling complex strategic decision-making in multiplayer games. Kaixuan Xu, Jiajun Chai, Yuqian Fu, Yuanheng Zhu, Dongbin Zhao |
ICML | 6 |
| 2025 | UncAD: Towards Safe End-to-end Autonomous Driving via Online Map UncertaintyabstractEnd-to-end autonomous driving aims to produce planning trajectories from raw sensors directly. Currently, most approaches integrate perception, prediction, and planning modules into a fully differentiable network, promising great scalability. However, these methods typically rely on deterministic modeling of online maps in the perception module for guiding or constraining vehicle planning, which may incorporate erroneous perception information and further compromise planning safety. To address this issue, we delve into the importance of online map uncertainty for enhancing autonomous driving safety and propose a novel paradigm named UncAD. Specifically, UncAD first estimates the uncertainty of the online map in the perception module. It then leverages the uncertainty to guide motion prediction and planning modules to produce multi-modal trajectories. Finally, to achieve safer autonomous driving, UncAD proposes an uncertainty-collision-aware planning selection strategy according to the online map uncertainty to evaluate and select the best trajectory. In this study, we incorporate UncAD into various state-of-the-art (SOTA) end-to-end methods. Experiments on the nuScenes dataset show that integrating UncAD, with only a 1.9% increase in parameters, can reduce collision rates by up to 26% and drivable area conflict rate by up to 42%. Codes, pre-trained models, and demo videos can be accessed at https://github.com/pengxuanyang/UncAD. Pengxuan Yang, Yupeng Zheng, Kefei Zhu, Zebin Xing, Yun-Fu Liu, Zhiguo Su, Dongbin Zhao |
ICRA | 9 |
| 2025 | Consistency Policy with Categorical Critic for Autonomous Driving
Haoran Li 0010, Dongbin Zhao |
AAMAS | 4 |
| 2025 | Salience-Invariant Consistent Policy Learning for Generalization in Visual Reinforcement Learning
Jingbo Sun 0001, Songjun Tu, Dongbin Zhao |
AAMAS | 5 |
| 2025 | Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
Songjun Tu, Jingbo Sun 0001, Xiangyuan Lan, Dongbin Zhao |
AAMAS | 5 |
| 2025 | Offline Goal-Conditioned Reinforcement Learning with Elastic-Subgoal Diffused Policy Learning
Yaocheng Zhang, Yuanheng Zhu, Yuqian Fu, Songjun Tu, Dongbin Zhao |
AAMAS | 5 |
| 2025 | Advancing Object-Goal Navigation through LLM-enhanced Object Affinities TransferabstractObject-goal navigation requires mobile robots to efficiently locate targets with visual and spatial information, yet existing methods struggle with generalization in unseen environments. Heuristic approaches with naive metrics fail in complex layouts, while graph-based and learning-based methods suffer from environmental biases and limited generalization. Although Large Language Models (LLMs) as planners or agents offer a rich knowledge base, they are cost-inefficient and lack targeted historical experience. To address these challenges, we propose the LLM-enhanced Object Affinities Transfer (LOAT) framework, integrating LLM-derived semantics with learning-based approaches to leverage experiential object affinities for better generalization in unseen settings. LOAT employs a dual-module strategy: one module accesses LLMs’ vast knowledge, and the other applies learned object semantic relationships, dynamically fusing these sources based on context. Evaluations in AI2-THOR and Habitat simulators show significant improvements in navigation success and efficiency, and real-world deployment demonstrates the zero-shot ability of LOAT to enhance object-goal navigation systems. Mengying Lin, Shugao Liu, Dingxi Zhang, Yaran Chen, Zhaoran Wang 0001, Haoran Li 0010, Dongbin Zhao |
IROS | 7 |
| 2025 | Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent RepresentationsabstractHumans can efficiently extract knowledge and learn skills from the videos within only a few trials and errors. However, it poses a big challenge to replicate this learning process for autonomous agents, due to the complexity of visual input, the absence of action or reward signals, and the limitations of interaction steps. In this paper, we propose a novel, unsupervised, and sample-efficient framework to achieve imitation learning from videos (ILV), named Behavior Cloning from Videos via Latent Representations (BCV-LR). BCV-LR extracts action-related latent features from high-dimensional video inputs through self-supervised tasks, and then leverages a dynamics-based unsupervised objective to predict latent actions between consecutive frames. The pre-trained latent actions are fine-tuned and efficiently aligned to the real action space online (with collected interactions) for policy behavior cloning. The cloned policy in turn enriches the agent experience for further latent action finetuning, resulting in an iterative policy improvement that is highly sample-efficient.
We conduct extensive experiments on a set of challenging visual tasks, including both discrete control and continuous control. BCV-LR enables effective (even expert-level on some tasks) policy performance with only a few interactions, surpassing state-of-the-art ILV baselines and reinforcement learning methods (provided with environmental rewards) in terms of sample efficiency across 24/28 tasks. To the best of our knowledge, this work for the first time demonstrates that videos can support extremely sample-efficient visual policy learning, without the need to access any other expert supervision. Xin Liu 0039, Haoran Li 0010, Dongbin Zhao |
NeurIPS | 3 |
| 2025 | Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion GamesabstractEquilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately solved. When the underlying graph structure varies, even the state-of-the-art RL methods require recomputation or at least fine-tuning, which can be time-consuming and impair real-time applicability. This paper proposes an Equilibrium Policy Generalization (EPG) framework to effectively learn a generalized policy with robust cross-graph zero-shot performance. In the context of PEGs, our framework is generally applicable to both pursuer and evader sides in both no-exit and multi-exit scenarios. These two generalizability properties, to our knowledge, are the first to appear in this domain. The core idea of the EPG framework is to train an RL policy across different graph structures against the equilibrium policy for each single graph. To construct an equilibrium oracle for single-graph policies, we present a dynamic programming (DP) algorithm that provably generates pure-strategy Nash equilibrium with near-optimal time complexity. To guarantee scalability with respect to pursuer number, we further extend DP and RL by designing a grouping mechanism and a sequence model for joint policy decomposition, respectively. Experimental results show that, using equilibrium guidance and a distance feature proposed for cross-graph PEG training, the EPG framework guarantees desirable zero-shot performance in various unseen real-world graphs. Besides, when trained under an equilibrium heuristic proposed for the graphs with exits, our generalized pursuer policy can even match the performance of the fine-tuned policies from the state-of-the-art PEG methods. Runyu Lu, Peng Zhang 0127, Ruochuan Shi, Yuanheng Zhu, Dongbin Zhao, Yang Liu 0066, Dong Wang 0004, Cesare Alippi |
NeurIPS | 5 |
| 2025 | Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLabstractLarge reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities—enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity.
Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping.
AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks.
Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy–efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4\% while reducing token usage by 52\% on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs.
Project Page: https://github.com/ScienceOne-AI/AutoThink. Songjun Tu, Xiangyu Tian, Linjing Li, Xiangyuan Lan, Dongbin Zhao |
NeurIPS | 7 |
| 2025 | Learning and Planning Multi-Agent Tasks via an MoE-based World ModelabstractMulti-task multi-agent reinforcement learning (MT-MARL) aims to develop a single model capable of solving a diverse set of tasks. However, existing methods often fall short due to the substantial variation in optimal policies across tasks, making it challenging for a single policy model to generalize effectively. In contrast, we find that many tasks exhibit **bounded similarity** in their underlying dynamics—highly similar within certain groups (e.g., door-open/close) diverge significantly between unrelated tasks (e.g., door-open \& object-catch). To leverage this property, we reconsider the role of modularity in multi-task learning, and propose **M3W**, a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning. For learning, it uses a SoftMoE-based dynamics model alongside a SparseMoE-based predictor to facilitate knowledge reuse across similar tasks while avoiding gradient conflicts across dissimilar tasks. For planning, it evaluates and optimizes actions using the predicted rollouts from the world model, without relying directly on a explicit policy model, thereby overcoming the limitations of policy-centric methods. As the first MoE-based multi-task world model, M3W demonstrates superior performance, sample efficiency, and multi-task adaptability, as validated on Bi-DexHands with 14 tasks and MA-Mujoco with 24 tasks. The demos and anonymous code are available at \url{https://github.com/zhaozijie2022/m3w-marl}. Zijie Zhao 0001, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
NeurIPS | 7 |
| 2025 | FusionNav: Enhancing Zero-Shot Object-Goal Navigation via 3D Semantic Fusion and Farsight Value ReasoningabstractZero-Shot Object-Goal Navigation (ZSON) tasks require agents to efficiently find target objects in unfamiliar environments, demanding strong semantic understanding and generalization. We propose FusionNav, a novel zero-shot navigation framework that combines 3D point cloud semantics with farsight value estimation. By integrating both local semantic information and semantic cues from regions outside the agent’s current field of view, FusionNav enables reasoning about unexplored areas and improves navigation efficiency. Experiments on the HM3D and MP3D benchmarks show that FusionNav outperforms strong baselines in both success rate and path efficiency. Moreover, FusionNav can be directly deployed on real-world robots without additional training or fine-tuning, and operates efficiently with moderate computational resources. These results demonstrate the effectiveness and practicality of FusionNav for real-world zero-shot object-goal navigation. Shugao Liu, Haoran Li 0010, Dongbin Zhao |
SMC | 4 |
| 2025 | Adaptive search for broad attention based vision transformers
Nannan Li 0003, Yaran Chen, Dongbin Zhao |
Neurocomputing | 3 |
| 2025 | Multi-Task Multi-Agent Reinforcement Learning With Task-Entity Transformers and Value Decomposition TrainingabstractMulti-task multi-agent reinforcement learning aims to control multiple agents to perform well on multiple tasks. It encounters three core challenges: the varying number of agents and entities, the disparities in cooperative behaviors among different tasks, and the training imbalance caused by varying task difficulty levels. To address these issues, we propose a novel framework named Task-Entity Transformer Qmix (TETQmix), which employs pretrained language models for task encoding, utilizes proposed Task-Entity Transformer to handle observations across various tasks, and adjusts task learning weights to achieve balanced multi-task training. Task-Entity Transformer not only enables handling multi-task scenarios with varying numbers of agents and entities, but also leverages cross-attention modules to integrate observation and task embeddings, so that each agent can obtain individual values and decisions for multiple tasks. We then utilize a transformer-based mixer to monotonically combine the individual values, and train the whole network’s parameters using temporal-difference errors. To facilitate multi-task training, we define task regret as the difference between the current-stage return and the candidate best one, and adjust the learning weight of each task based on its task regret. Experiments are conducted on both simulated multi-particle environments and real-world multi-robot systems. Compared with existing baselines, our method not only is superior in multi-task learning efficiency, but also shows promising transfer ability on unseen tasks. Note to Practitioners—The flexibility of multi-agent systems makes them quite fit to multiple tasks. Compared to designing different decision models for different tasks, it is more convenient if one can use just one decision model to resolve multiple tasks. Besides, it can make the maximum utilization of trajectory data coming from similar tasks when the data are integrated for multi-task decision model training. Natural language provides a powerful tool to describe the task context and emphasize the similarities or differences among different tasks. Pretrained language models can encode the task context, based on which the decision model can adjust its output distribution for different tasks and even synthesize the decisions from existing and similar tasks to achieve promising zero-shot and few-shot transfer performance for unseen tasks. With our proposed TETQmix, practitioners are able to realize multi-task capability in multi-agent systems and increase the generalization in a variety of scenarios. Yuanheng Zhu, Shangjing Huang, Binbin Zuo, Dongbin Zhao, Changyin Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Balancing State Exploration and Skill Diversity in Unsupervised Skill DiscoveryabstractUnsupervised skill discovery seeks to acquire different useful skills without extrinsic reward via unsupervised reinforcement learning (RL), with the discovered skills efficiently adapting to multiple downstream tasks in various ways. However, recent advanced skill discovery methods struggle to well balance state exploration and skill diversity, particularly when the potential skills are rich and hard to discern. In this article, we propose contrastive dynamic skill discovery (ComSD) which generates diverse and exploratory unsupervised skills through a novel intrinsic incentive, named contrastive dynamic reward. It contains a particle-based exploration reward to make agents access far-reaching states for exploratory skill acquisition, and a novel contrastive diversity reward to promote the discriminability between different skills. Moreover, a novel dynamic weighting mechanism between the above two rewards is proposed to balance state exploration and skill diversity, which further enhances the quality of the discovered skills. Extensive experiments and analysis demonstrate that ComSD can generate diverse behaviors at different exploratory levels for multijoint robots, enabling state-of-the-art adaptation performance on challenging downstream tasks. It can also discover distinguishable and far-reaching exploration skills in the challenging tree-like 2-D maze. Xin Liu 0039, Yaran Chen, Guixing Chen, Haoran Li 0010, Dongbin Zhao |
IEEE Trans. Cybern. | 5 |
| 2025 | Last-Iterate Convergence to Approximate Nash Equilibria in Multiplayer Imperfect Information GamesabstractImperfect information and multiple players are the two common features of real-world games. However, few of the existing game-theoretic methods are applicable to multiplayer imperfect information games (IIGs) when it comes to finding Nash equilibria. Moreover, the commonly used methods that rely on average-iterate convergence are not conducive to deep reinforcement learning (DRL), which is widely applied to large-scale problems, as it is costly to preserve average policies under function approximation. To deal with these problems, we construct a continuous-time dynamic named imperfect-information exponential-decay score-based learning (IESL) by considering the concept of Nash distribution [a type of quantal response equilibrium (QRE)] in IIGs. Theoretically, we prove the last-iterate convergence of IESL to approximate Nash equilibria in multiplayer IIGs under the assumption of individual concavity. Empirically, we verify that IESL converges in six poker scenarios, with the ultimate NashConv lower than that of the comparative methods (including counterfactual regret minimization (CFR), replicator dynamics (RDs), and their variants) in multiplayer Leduc hold'em. When compared with the existing equilibrium-finding algorithms in multiplayer normal-form games (NFGs), IESL also demonstrates a more stable performance. In addition, we observe a trade-off between the difficulty of IESL's last-iterate convergence and the NashConv of the convergent policies, which aligns with our convergence analysis based on the hypomonotonicity of the game. Runyu Lu, Yuanheng Zhu, Dongbin Zhao, Yu Liu 0005, You He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Meta Learning Task Representation in Multiagent Reinforcement Learning: From Global Inference to Local InferenceabstractMultiagent meta reinforcement learning (MAMRL) enables multiagent systems (MASs) to adapt to multiple tasks. However, partial observability poses a significant challenge by hindering efficient task inference from agents' limited local experiences. To address this, we propose MG2L, a novel algorithm featuring a global-to-local (G2L) training scheme based on mutual information optimization (MIO). We first extend the centralized training and decentralized execution (CTDE) framework to MAMRL, and introduce a multilevel task encoder for joint global and local task inference. Building on this encoder, the MG2L scheme employs tailored loss functions to optimize task representations. For global inference, the MAS learns a centralized global representation by maximizing the MI between the representation and the task context. For local inference, we formulate conditional MI reduction to quantify the G2L gap. Agents then learn the local representation by minimizing this reduction. The MG2L scheme effectively harmonizes centralized training with decentralized execution, offering a versatile solution for MAMRL challenges. Additionally, we integrate a permutation-invariant attention (PIA) module into the task encoder to reduce sensitivity to behavior policy variations. Extensive experiments-including comparative analyses, ablation studies, meta-test evaluations, and visualizations-demonstrate MG2L's effectiveness. The implementation of MG2L is publicly available at https://github.com/zhaozijie2022/mg2l. Zijie Zhao 0001, Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Discretizing Continuous Action Space With Unimodal Probability Distributions for On-Policy Reinforcement LearningabstractFor on-policy reinforcement learning (RL), discretizing action space for continuous control can easily express multiple modes and is straightforward to optimize. However, without considering the inherent ordering between the discrete atomic actions, the explosion in the number of discrete actions can possess undesired properties and induce a higher variance for the policy gradient (PG) estimator. In this article, we introduce a straightforward architecture that addresses this issue by constraining the discrete policy to be unimodal using Poisson probability distributions. This unimodal architecture can better leverage the continuity in the underlying continuous action space using explicit unimodal probability distributions. We conduct extensive experiments to show that the discrete policy with the unimodal probability distribution provides significantly faster convergence and higher performance for on-policy RL algorithms in challenging control tasks, especially in highly complex tasks such as Humanoid. We provide theoretical analysis on the variance of the PG estimator, which suggests that our attentively designed unimodal discrete policy can retain a lower variance and yield a stable learning process. Yuanyang Zhu, Zhi Wang 0001, Yuanheng Zhu, Chunlin Chen 0001, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | LDR: Learning Discrete Representation to Improve Noise Robustness in Multiagent TasksabstractIn real-world applications of multiagent reinforcement learning (MARL), agents often face inaccurate environments due to unavoidable noise, presenting a challenge to their robustness. However, limited prior work focuses on addressing such noise in observations, hindering the deployment of multiagent systems. In this article, we propose a method named learning discrete representation (LDR) to improve robustness against noise in multiagent tasks. Specifically, LDR employs a quantization module with a segment mechanism to encode observations and teammate actions, generating discrete representations from learnable codebooks. These representations are subsequently processed via a combiner for decision-making. Through discretization, LDR is able to mitigate the impact of minor noise on decision-making. To enhance the learning efficiency, we incorporate a set-input block that treats the joint observations of agents as a permutation-invariant set, thereby reducing the complexity of the joint observation space. Additionally, we theoretically analyze the expressiveness of discrete representation and the boundedness of discrete distortion. We evaluate the proposed method on StarCraft II micromanagement tasks and multiagent MuJoCo with noisy observations. Empirical results demonstrate that LDR outperforms existing algorithms, improving robustness in noisy cooperative MARL tasks while maintaining superior performance in clean observations. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Cross-Domain Random Pretraining With Prototypes for Reinforcement LearningabstractUnsupervised cross-domain reinforcement learning (RL) pretraining shows great potential for challenging continuous visual control but poses a big challenge. In this article, we propose cross-domain random pretraining with prototypes (CRPTpro), a novel, efficient, and effective self-supervised cross-domain RL pretraining framework. CRPTpro decouples data sampling from encoder pretraining, proposing decoupled random collection to easily and quickly generate a qualified cross-domain pretraining dataset. Moreover, a novel prototypical self-supervised algorithm is proposed to pretrain an effective visual encoder that is generic across different domains. Without finetuning, the cross-domain encoder can be implemented for challenging downstream tasks defined in different domains, either seen or unseen. Compared with recent advanced methods, CRPTpro achieves better performance on downstream policy learning without extra training on exploration agents for data collection, greatly reducing the burden of pretraining. We conduct extensive experiments across multiple challenging continuous visual-control domains, including balance control, robot locomotion, and manipulation. CRPTpro significantly outperforms the next best Proto-RL(C) on 11/12 cross-domain downstream tasks with only 54.5% wall-clock pretraining time, exhibiting state-of-the-art pretraining performance with greatly improved pretraining efficiency. Xin Liu 0039, Yaran Chen, Haoran Li 0010, Boyu Li 0003, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Common Sense Language-Guided Exploration and Hierarchical Dense Perception for Instruction Following Embodied AgentsabstractEmbodied Instruction Following (EIF) involves the task of locating and manipulating objects according to language instructions. Existing methods face challenges in small object navigation due to ineffective exploration and imperfect perception, which ultimately affects their performance. This study focuses on small object navigation in the EIF domain. We propose Common Sense Language-guided exploration (CSL), a novel approach that leverages common-sense knowledge from seen scenes and information from language instructions to infer the location of objects. The proposed CSL significantly improves exploration efficiency. Additionally, we propose Hierarchical Dense Perception (HDP), which uses hierarchical features to perform semantic segmentation and depth estimation. The use of HDP significantly improves the agent’s perceptual capabilities. Experiments on the ALFRED benchmark demonstrate the effectiveness of CSL-HDP. The proposed CSL-HDP achieves an absolute improvement of 9.29% (18.45% relative) on unseen test scenes compared to the previous state-of-the-art, securing the top position on the leaderboard. Code will be available at https://github.com/Cyuanwen/CSL-HDP. Yuanwen Chen, Yaran Chen, Dongbin Zhao, Yunzhen Zhao, Pengfei Hu 0004 |
ICME | 4 |
| 2024 | High-quality Synthetic Data is Efficient for Model-based Offline Reinforcement LearningabstractRecent work has found that two types of dataset characteristics including the dataset’s coverage and data quality are critical for offline reinforcement learning (RL). To improve the policy, model-based offline RL tries to generate reliable synthetic data to expand the dataset’s coverage based on trained forward and backward dynamics models. However, the characteristic of synthetic data’s quality is ignoring, which raises a question of whether augmenting high-quality synthetic data is efficient for offline RL agents. Motivated by this, we propose a novel forward High-quality Imagination and backward Reliable Check (HIRC), which is an effective data augmentation method to generate high-quality and reliable synthetic data. Specifically, we construct a value-guided forward model to generate high-quality imaginary trajectories, and employ a backward model for reliable checking to obtain synthetic data that better match with pre-collected offline transitions. In other words, the proposed HIRC method can generate high-quality synthetic data on the premise of reliability, which can be combined with model-free offline RL methods. Experimental results on the D4RL benchmark demonstrate that high-quality synthetic data generated by HIRC boosts the performance of a base agent TD3_BC. Especially, HIRC with such a base agent achieves better scores against recent popular model-free and model-based offline RL methods. Kaixuan Xu, Weixin Zhao, Haoran Li 0010, Dongbin Zhao |
IJCNN | 6 |
| 2024 | Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience RegularizationabstractWith high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based RL, it is still an open question whether it can be extended to visual RL. In this paper, we investigate the impact of non-stationary distribution and the actor-critic framework on consistency policy in online RL, and find that consistency policy was unstable during the training, especially in visual RL with the high-dimensional state space. To this end, we suggest sample-based entropy regularization to stabilize the policy training, and propose a consistency policy with prioritized proximal experience regularization (CP3ER) to improve sample efficiency. CP3ER achieves new state-of-the-art (SOTA) performance in 21 tasks across DeepMind control suite and Meta-world. To our knowledge, CP3ER is the first method to apply diffusion/consistency models to visual RL and demonstrates the potential of consistency models in visual RL. Haoran Li 0010, Zhennan Jiang, Yuhui Chen, Dongbin Zhao |
NeurIPS | 4 |
| 2024 | Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementabstractA longstanding goal of artificial general intelligence is highly capable generalists that can learn from diverse experiences and generalize to unseen tasks. The language and vision communities have seen remarkable progress toward this trend by scaling up transformer-based models trained on massive datasets, while reinforcement learning (RL) agents still suffer from poor generalization capacity under such paradigms. To tackle this challenge, we propose Meta Decision Transformer (Meta-DT), which leverages the sequential modeling ability of the transformer architecture and robust task representation learning via world model disentanglement to achieve efficient generalization in offline meta-RL. We pretrain a context-aware world model to learn a compact task representation, and inject it as a contextual condition to the causal transformer to guide task-oriented sequence generation. Then, we subtly utilize history trajectories generated by the meta-policy as a self-guided prompt to exploit the architectural inductive bias. We select the trajectory segment that yields the largest prediction error on the pretrained world model to construct the prompt, aiming to encode task-specific information complementary to the world model maximally. Notably, the proposed framework eliminates the requirement of any expert demonstration or domain knowledge at test time. Experimental results on MuJoCo and Meta-World benchmarks across various dataset types show that Meta-DT exhibits superior few and zero-shot generalization capacity compared to strong baselines while being more practical with fewer prerequisites. Our code is available at https://github.com/NJU-RL/Meta-DT. Zhi Wang 0001, Yuanheng Zhu, Dongbin Zhao, Chunlin Chen 0001 |
NeurIPS | 5 |
| 2024 | Prototypical Context-Aware Dynamics for Generalization in Visual Control With Model-Based Reinforcement LearningabstractThe latent world model, which efficiently represents high-dimensional observations within a latent space, has shown promise in reinforcement learning-based policies for visual control tasks. Due to a lack of clear environmental context comprehension, its applicability in a variety of contexts with unknown dynamics is constrained. We propose a prototypical context- aware dynamics (ProtoCAD) model to address this issue. This model captures local dynamics using temporally consistent latent contexts and aids generalization in visual control tasks. By grouping prototypes over historical experiences, ProtoCAD collects useful contextual information that improves model-based reinforcement learning dynamics generalization in two ways. First, to guarantee the consistency of prototype assignments for various temporal segments of the same latent trajectory, a temporally consistent prototypes regularizer is used. Then, a context representation is devised to combine the aggregated prototype with the projection embedding of latent states. According to extensive trials, ProtoCAD outperforms competing approaches in terms of dynamics generalization for visual robotic control and autonomous driving applications. Yao Mu 0001, Dong Li 0016, Dongbin Zhao, Yuzheng Zhuang, Ping Luo 0002, Bin Wang 0034, Jianye Hao |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | NVIF: Neighboring Variational Information Flow for Cooperative Large-Scale Multiagent Reinforcement LearningabstractCommunication-based multiagent reinforcement learning (MARL) has shown promising results in promoting cooperation by enabling agents to exchange information. However, the existing methods have limitations in large-scale multiagent systems due to high information redundancy, and they tend to overlook the unstable training process caused by the online-trained communication protocol. In this work, we propose a novel method called neighboring variational information flow (NVIF), which enhances communication among neighboring agents by providing them with the maximum information set (MIS) containing more information than the existing methods. NVIF compresses the MIS into a compact latent state while adopting neighboring communication. To stabilize the overall training process, we introduce a two-stage training mechanism. We first pretrain the NVIF module using a randomly sampled offline dataset to create a task-agnostic and stable communication protocol, and then use the pretrained protocol to perform online policy training with RL algorithms. Our theoretical analysis indicates that NVIF-proximal policy optimization (PPO), which combines NVIF with PPO, has the potential to promote cooperation with agent-specific rewards. Experiment results demonstrate the superiority of our method in both heterogeneous and homogeneous settings. Additional experiment results also demonstrate the potential of our method for multitask learning. Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | BViT: Broad Attention-Based Vision TransformerabstractRecent works have demonstrated that transformer can achieve promising performance in computer vision, by exploiting the relationship among image patches with self-attention. They only consider the attention in a single feature layer, but ignore the complementarity of attention in different layers. In this article, we propose broad attention to improve the performance by incorporating the attention relationship of different layers for vision transformer (ViT), which is called BViT. The broad attention is implemented by broad connection and parameter-free attention. Broad connection of each transformer layer promotes the transmission and integration of information for BViT. Without introducing additional trainable parameters, parameter-free attention jointly focuses on the already available attention information in different layers for extracting useful information and building their relationship. Experiments on image classification tasks demonstrate that BViT delivers superior accuracy of 75.0%/81.6% top-1 accuracy on ImageNet with 5M/22M parameters. Moreover, we transfer BViT to downstream object recognition benchmarks to achieve 98.9% and 89.9% on CIFAR10 and CIFAR100, respectively, that exceed ViT with fewer parameters. For the generalization test, the broad attention in Swin Transformer, T2T-ViT and LVT also brings an improvement of more than 1%. To sum up, broad attention is promising to promote the performance of attention-based models. Code and pretrained models are available at https://github.com/DRL/BViT. Nannan Li 0003, Yaran Chen, Weifan Li, Zixiang Ding, Dongbin Zhao, Shuai Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Conditional Goal-Oriented Trajectory Prediction for Interacting VehiclesabstractPredicting future trajectories of pairwise traffic agents in highly interactive scenarios, such as cut-in, yielding, and merging, is challenging for autonomous driving. The existing works either treat such a problem as a marginal prediction task or perform single-axis factorized joint prediction, where the former strategy produces individual predictions without considering future interaction, while the latter strategy conducts conditional trajectory-oriented prediction via agentwise interaction or achieves conditional rollout-oriented prediction via timewise interaction. In this article, we propose a novel double-axis factorized joint prediction pipeline, namely, conditional goal-oriented trajectory prediction (CGTP) framework, which models future interaction both along the agent and time axes to achieve goal and trajectory interactive prediction. First, a goals-of-interest network (GoINet) is designed to extract fine-grained features of goal candidates via hierarchical vectorized representation. Furthermore, we propose a conditional goal prediction network (CGPNet) to produce multimodal goal pairs in an agentwise conditional manner, along with a newly designed goal interactive loss to better learn the joint distribution of the intermediate interpretable modes. Explicitly guided by the goal-pair predictions, we propose a goal-oriented trajectory rollout network (GTRNet) to predict scene-compliant trajectory pairs via timewise interactive rollouts. Extensive experimental results confirm that the proposed CGTP outperforms the state-of-the-art (SOTA) prediction models on the Waymo open motion dataset (WOMD), Argoverse motion forecasting dataset, and In-house cut-in dataset. Code is available at https://github.com/LiDinga/CGTP/. Yifeng Pan, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Dynamic-Horizon Model-Based Value Estimation With Latent ImaginationabstractExisting model-based value expansion (MVE) methods typically leverage a world model for value estimation with a fixed rollout horizon to assist policy learning. However, a proper horizon setting is essential to world-model-based policy learning. Meanwhile, choosing an appropriate horizon value is time-consuming, especially for visual control tasks. In this article, we investigate the idea of adaptively using the model knowledge for value expansion. We propose a novel world-model-based method called dynamic-horizon MVE (DMVE) to adjust the use of the world model with adaptive rollout horizon selection. Based on the reconstruction-based technique, the raw and reconstructed images are both used to obtain multihorizon rollouts by utilizing latent imagination. Then, a horizon reliability degree detection approach is given to select appropriate horizons and obtain more accurate value estimation by the reconstructed value expansion errors. Experimental results on the mainstream benchmark visual control tasks show that DMVE outperforms all baselines in sample efficiency and final performance. In addition, experiments on the autonomous driving lane-changing task further demonstrate the scalability of our method. The codes of DMVE are available at https://github.com/JunjieWang95/dmve. Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth EstimationabstractSelf-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capa-bilities of self-driving vehicles. However, it intrinsically relies upon the photometric consistency assumption, which hardly holds during nighttime. Although various supervised night-time image enhancement methods have been proposed, their generalization performance in challenging driving scenarios is not satisfactory. To this end, we propose the first method that jointly learns a nighttime image enhancer and a depth estimator, without using ground truth for either task. Our method tightly entangles two self-supervised tasks using a newly proposed uncertain pixel masking strategy. This strategy originates from the observation that nighttime images not only suffer from underexposed regions but also from overexposed regions. By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally. We benchmark the method on two established datasets: nuScenes and RobotCar and demonstrate state-of-the-art performance on both of them. Detailed ablations also reveal the mechanism of our proposal. Last but not least, to mitigate the problem of sparse ground truth of existing datasets, we provide a new photo-realistically enhanced nighttime dataset based upon CARLA. It brings meaningful new challenges to the community. Codes, data, and models are available at https://github.com/ucaszyp/STEPS. Yupeng Zheng, Chengliang Zhong, Pengfei Li 0007, Huan-ang Gao, Yuhang Zheng 0004, Bu Jin, Ling Wang 0001, Hao Zhao 0002, Guyue Zhou, Dongbin Zhao |
ICRA | 11 |
| 2023 | NeuronsMAE: A Novel Multi-Agent Reinforcement Learning Environment for Cooperative and Competitive Multi-Robot TasksabstractMulti-agent reinforcement learning (MARL) has achieved remarkable success in various challenging problems. Meanwhile, more and more benchmarks have emerged and provided some standards to evaluate the algorithms in different fields. On the one hand, the virtual MARL environments lack knowledge of real-world tasks and actuator abilities. On the other hand, the current task-specified multi-robot platform has poor support for the universality of multi-agent reinforcement learning algorithms and lacks support for transferring from simulation to the real environment. Bridging the gap between the virtual MARL environments and the real multi-robot platform becomes the key to promoting the practicability of MARL algorithms. This paper proposes a novel MARL environment for real multi-robot tasks named NeuronsMAE (Neurons Multi-Agent Environment). This environment supports cooperative and competitive multi-robot tasks and is configured with rich parameter interfaces to study the multi-agent policy transfer from simulation to reality. With this platform, we evaluate various popular MARL algorithms and build a new MARL benchmark for multi-robot tasks. We hope that this platform will facilitate the research and application of MARL algorithms for real robot tasks. Information about the benchmark and the open-source code are released at https://github.com/DRL-CASIA/NeuronsMAE. Guangzheng Hu, Haoran Li 0010, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 5 |
| 2023 | Dense Attention: A Densely Connected Attention Mechanism for Vision TransformerabstractRecently, Vision Transformer has demonstrated its impressive capability in image understanding. The multi-head self-attention mechanism is fundamental to its formidable performance. However, self-attention has the drawback of high computational effort, which makes the training of the model require powerful computational resources or more time. This paper designs a novel and efficient attention mechanism Dense Attention to overcome the above problem. Dense attention aims to focus on features from multiple views through a dense connection paradigm. Benefiting from the attention of comprehensive features, dense attention can i) remarkably strengthen the image representation of the model, and ii) partially replace the multi-head self-attention mechanism to allow model slimming. To verify the effectiveness of dense attention, we implement it in the prevalent Vision Transformer models, including non-pyramid architecture DeiT and pyramid architecture Swin Transformer. The experimental results on ImageNet classification show that dense attention indeed contributes to performance improvement,$+\mathbf{1.8}/\mathbf{1.3}\%$for DeiT-T/s and$+\mathbf{0.7}/\!\!+\mathbf{1.2}\%$for Swin-T/s, respectively. Dense attention also demonstrates its transferability on CIFAR10 and CIFAR100 recognition benchmarks with classification accuracy of 98.9% and 89.6% respectively. Furthermore, dense attention can weaken the performance sacrifice caused by the pruning in the number of heads. Code and pre-trained models will be available11https://github.com/koala719/Dense-ViT. Nannan Li 0003, Yaran Chen, Dongbin Zhao |
IJCNN | 3 |
| 2023 | Advantage Constrained Proximal Policy Optimization in Multi-Agent Reinforcement LearningabstractWe investigate the integration of value-based and policy gradient methods in multi-agent reinforcement learning (MARL). The Individual-Global-Max (IGM) principle plays an important role in value-based MARL, as it ensures consistency between joint and local action values. IGM is difficult to guarantee in multi-agent policy gradient methods due to stochastic exploration and conflicting gradient directions. In this paper, we propose a novel multi-agent policy gradient algorithm called Advantage Constrained Proximal Policy Optimization (ACPPO). ACPPO calculates each agent's current local state-action advantage based on their advantage network and estimates the joint state-action advantage based on multi-agent advantage decomposition lemma. According to the consistency of the estimated joint-action advantage and local advantage, the coefficient of each agent constrains the joint-action advantage. ACPPO, unlike previous policy gradient MARL algorithms, does not require an additional sampled baseline to reduce variance or a sequential scheme to improve accuracy. The proposed method is evaluated using the continuous matrix game, the Starcraft Multi-Agent Challenge, and the Multi-Agent MuJoCo task. ACPPO outperforms baselines such as MAPPO, MADDPG, and HATRPO, according to the results. Weifan Li, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 3 |
| 2023 | Enhanced Rolling Horizon Evolution Algorithm With Opponent Model Learning: Results for the Fighting Game AI CompetitionabstractThe Fighting Game AI Competition (FTGAIC) provides a challenging benchmark for two-player video game artificial intelligence. The challenge arises from the large action space, diverse styles of characters and abilities, and the real-time nature of the game. In this article, we propose a novel algorithm that combines the rolling horizon evolution algorithm (RHEA) with opponent model learning. The approach is readily applicable to any two-player video game. In contrast to conventional RHEA, an opponent model is proposed and is optimized by supervised learning with cross-entropy and reinforcement learning with policy gradient and Q-learning respectively, based on history observations from opponent. The model is learned during the live gameplay. With the learned opponent model, the extended RHEA is able to make more realistic plans based on what the opponent is likely to do. This tends to lead to better results. We compared our approach directly with the bots from the FTGAIC 2018 competition and found our method to significantly outperform all of them for all three characters. Furthermore, our proposed bot with the policy gradient based opponent model is the only one without using Monte Carlo tree search among the top five bots in the 2019 competition in which it achieved second place, while using much less domain knowledge than the winner. Zhentao Tang, Yuanheng Zhu, Dongbin Zhao, Simon M. Lucas |
IEEE Trans. Games | 3 |
| 2023 | Empirical Policy Optimization for n-Player Markov GamesabstractIn single-agent Markov decision processes, an agent can optimize its policy based on the interaction with the environment. In multiplayer Markov games (MGs), however, the interaction is nonstationary due to the behaviors of other players, so the agent has no fixed optimization objective. The challenge becomes finding equilibrium policies for all players. In this research, we treat the evolution of player policies as a dynamical process and propose a novel learning scheme for Nash equilibrium. The core is to evolve one's policy according to not just its current in-game performance, but an aggregation of its performance over history. We show that for a variety of MGs, players in our learning scheme will provably converge to a point that is an approximation to Nash equilibrium. Combined with neural networks, we develop an empirical policy optimization algorithm, which is implemented in a reinforcement-learning framework and runs in a distributed way, with each player optimizing its policy based on own observations. We use two numerical examples to validate the convergence property on small-scale MGs, and a pong example to show the potential on large games. Yuanheng Zhu, Weifan Li, Mengchen Zhao, Jianye Hao, Dongbin Zhao |
IEEE Trans. Cybern. | 5 |
| 2023 | UNMAS: Multiagent Reinforcement Learning for Unshaped Cooperative ScenariosabstractMultiagent reinforcement learning methods, such as VDN, QMIX, and QTRAN, that adopt centralized training with decentralized execution (CTDE) framework have shown promising results in cooperation and competition. However, in some multiagent scenarios, the number of agents and the size of the action set actually vary over time. We call these unshaped scenarios, and the methods mentioned above fail in performing satisfyingly. In this article, we propose a new method, called Unshaped Networks for Multiagent Systems (UNMAS), which adapts to the number and size changes in multiagent systems. We propose the self-weighting mixing network to factorize the joint action-value. Its adaption to the change in agent number is attributed to the nonlinear mapping from each-agent Q value to the joint action-value with individual weights. Besides, in order to address the change in an action set, each agent constructs an individual action-value network that is composed of two streams to evaluate the constant environment-oriented subset and the varying unit-oriented subset. We evaluate UNMAS on various StarCraft II micromanagement scenarios and compare the results with several state-of-the-art MARL algorithms. The superiority of UNMAS is demonstrated by its highest winning rates especially on the most difficult scenario 3s5z_vs_3s6z. The agents learn to perform effectively cooperative behaviors, while other MARL algorithms fail. Animated demonstrations and source code are provided in https://sites.google.com/view/unmas. Jiajun Chai, Weifan Li, Yuanheng Zhu, Dongbin Zhao, Kewu Sun, Jishiyu Ding |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Event-Triggered Communication Network With Limited-Bandwidth Constraint for Multi-Agent Reinforcement LearningabstractCommunicating agents with each other in a distributed manner and behaving as a group are essential in multi-agent reinforcement learning. However, real-world multi-agent systems suffer from restrictions on limited bandwidth communication. If the bandwidth is fully occupied, some agents are not able to send messages promptly to others, causing decision delay and impairing cooperative effects. Recent related work has started to address the problem but still fails in maximally reducing the consumption of communication resources. In this article, we propose an event-triggered communication network (ETCNet) to enhance communication efficiency in multi-agent systems by communicating only when necessary. For different task requirements, two paradigms of the ETCNet framework, event-triggered sending network (ETSNet) and event-triggered receiving network (ETRNet), are proposed for learning efficient sending and receiving protocols, respectively. Leveraging the information theory, the limited bandwidth is translated to the penalty threshold of an event-triggered strategy, which determines whether an agent at each step participates in communication or not. Then, the design of the event-triggered strategy is formulated as a constrained Markov decision problem and reinforcement learning finds the feasible and optimal communication protocol that satisfies the limited bandwidth constraint. Experiments on typical multi-agent tasks demonstrate that ETCNet outperforms other methods in reducing bandwidth occupancy and still preserves the cooperative performance of multi-agent systems at the most. Guangzheng Hu, Yuanheng Zhu, Dongbin Zhao, Mengchen Zhao, Jianye Hao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Hierarchical Deep Reinforcement Learning Framework for 6-DOF UCAV Air-to-Air CombatabstractUnmanned combat air vehicle (UCAV) combat is a challenging scenario with high-dimensional continuous state and action space and highly nonlinear dynamics. In this article, we propose a general hierarchical framework to resolve the within-vision-range (WVR) air-to-air combat problem under six dimensions of degree (6-DOF) dynamics. The core idea is to divide the whole decision-making process into two loops and use reinforcement learning (RL) to solve them separately. The outer loop uses a combat policy to decide the macro command according to the current combat situation. Then the inner loop uses a control policy to answer the macro command by calculating the actual input signals for the aircraft. We design the Markov decision-making process for the control policy and the Markov game between two aircraft. We present a two-stage training mechanism. For the control policy, we design an effective reward function to accurately track various macro behaviors. For the combat policy, we present a fictitious self-play mechanism to improve the combat performance by combating against the historical combat policies. Experiment results show that the control policy can achieve better tracking performance than conventional methods. The fictitious self-play mechanism can learn competitive combat policy, which can achieve high winning rates against conventional methods. Jiajun Chai, Wenzhang Chen, Yuanheng Zhu, Zong-xin Yao, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Stacked BNAS: Rethinking Broad Convolutional Neural Network for Neural Architecture SearchabstractDifferent from other deep scalable architecture-based neural architecture search (NAS) approaches, broad NAS (BNAS) proposes a broad scalable architecture which consists of convolution and enhancement blocks, dubbed broad convolutional neural network (BCNN), as the search space for amazing efficiency improvement. BCNN reuses the topologies of cells in the convolution block so that BNAS can employ few cells for efficient search. Moreover, multiscale feature fusion and knowledge embedding are proposed to improve the performance of BCNN with shallow topology. However, BNAS suffers some drawbacks: 1) insufficient representation diversity for feature fusion and enhancement and 2) time consumption of knowledge embedding design by human experts. This article proposes Stacked BNAS, whose search space is a developed broad scalable architecture named Stacked BCNN, with better performance than BNAS. On the one hand, Stacked BCNN treats mini BCNN as a basic block to preserve comprehensive representation and deliver powerful feature extraction ability. For multiscale feature enhancement, each mini BCNN feeds the outputs of deep and broad cells to the enhancement cell. For multiscale feature fusion, each mini BCNN feeds the outputs of deep, broad and enhancement cells to the output node. On the other hand, knowledge embedding search (KES) is proposed to learn appropriate knowledge embeddings in a differentiable way. Moreover, the basic unit of KES is an over-parameterized knowledge embedding module that consists of all possible candidate knowledge embeddings. Experimental results show that: 1) Stacked BNAS obtains better performance than BNAS-v2 on both CIFAR-10 and ImageNet; 2) the proposed KES algorithm contributes to reducing the parameters of the learned architecture with satisfactory performance; and 3) Stacked BNAS delivers a state-of-the-art efficiency of 0.02 GPU days. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | LILAC: Learning a Leader for Cooperative Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning,role-based learning promises to reach satisfactory policy learning through the decomposition of complicated tasks using roles. Different roles are responsible for different aspects of the task. However, how this group of roles can be quickly identified is not clear. To address this problem, we propose a novel framework, LearnIng a LeAder for Cooperative reinforcement learning (LILAC), which introduces a leader to integrate information to assign roles. Leaders take a broad view of the whole task and feed the integrated information into a Gaussian mixture model to sample role embedding distribution. It enables LILAC to assign appropriate roles to different agents and improves cooperative performance. In order to evaluate the cooperation of multiple agents, a mixing network, inputted by individual local utility networks, is constructed to estimate the global action value. Two loss functions, temporal difference loss and mean divergence loss, are adopted by LILAC to learn network parameters and to encourage diversity of policies for different roles. By virtue of the leader module, LILAC outperforms the StarCraft II micromanagement benchmark in our experiments, especially on challenging tasks. Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
CoG | 4 |
| 2022 | Neurons Perception Dataset for RoboMaster AI ChallengeabstractFrom virtual game to physical robot, games have witnessed the development of artificial intelligence (AI) technology, especially the data-driven technology represented by deep learning. Compared with virtual games, a physical robot game such as RoboMaster AI challenge needs to build a complete closed-loop architecture composed of perception, planning, control, and decision-making to support autonomous confrontation. Perception, as the eye of the robot, its performance in the complex environment depends on a massive dataset. Although there are many open perception datasets, these datasets are difficult to meet the needs of RoboMaster AI challenge due to the high dynamics of the task, the distinctiveness of the objects, and limited computing resources. In this paper, we release a dataset named Neurons11Neurons is a team dedicated to promoting the development of robot with deep neural network. We will release the code and dataset at https://github.com/DRL-CASIA/NeuronsDataset. perception dataset for RoboMaster AI challenge, which covers 3 tasks including monocular depth estimation, lightweight object detection, and multi-view 3D object detection, and makes up the data blank in this field. In addition, we also evaluate State-Of-The-Art (SOTA) methods on each task, hoping to provide an impartial benchmark for the development of perception algorithm. Haoran Li 0010, Zicheng Duan, Yaran Chen, Dongbin Zhao |
IJCNN | 6 |
| 2022 | ModuleNet: Knowledge-Inherited Neural Architecture SearchabstractAlthough neural the architecture search (NAS) can bring improvement to deep models, it always neglects precious knowledge of existing models. The computation and time costing property in NAS also means that we should not start from scratch to search, but make every attempt to reuse the existing knowledge. In this article, we discuss what kind of knowledge in a model can and should be used for a new architecture design. Then, we propose a new NAS algorithm, namely, ModuleNet, which can fully inherit knowledge from the existing convolutional neural networks. To make full use of the existing models, we decompose existing models into different modules, which also keep their weights, consisting of a knowledge base. Then, we sample and search for a new architecture according to the knowledge base. Unlike previous search algorithms, and benefiting from inherited knowledge, our method is able to directly search for architectures in the macrospace by the NSGA-II algorithm without tuning parameters in these modules. Experiments show that our strategy can efficiently evaluate the performance of a new architecture even without tuning weights in convolutional layers. With the help of knowledge we inherited, our search results can always achieve better performance on various datasets (CIFAR10, CIFAR100, and ImageNet) over original architectures. Yaran Chen, Ruiyuan Gao 0001, Fenggang Liu, Dongbin Zhao |
IEEE Trans. Cybern. | 4 |
| 2022 | BiFNet: Bidirectional Fusion Network for Road SegmentationabstractMultisensor fusion-based road segmentation plays an important role in the intelligent driving system since it provides a drivable area. The existing mainstream fusion method is mainly to feature fusion in the image space domain which causes the perspective compression of the road and damages the performance of the distant road. Considering the bird's eye views (BEVs) of the LiDAR remains the space structure in the horizontal plane, this article proposes a bidirectional fusion network (BiFNet) to fuse the image and BEV of the point cloud. The network consists of two modules: 1) the dense space transformation (DST) module, which solves the mutual conversion between the camera image space and BEV space and 2) the context-based feature fusion module, which fuses the different sensors information based on the scenes from corresponding features. This method has achieved competitive results on the KITTI dataset. Haoran Li 0010, Yaran Chen, Dongbin Zhao |
IEEE Trans. Cybern. | 4 |
| 2022 | TrajGen: Generating Realistic and Diverse Trajectories With Reactive and Feasible Agent Behaviors for Autonomous DrivingabstractRealistic and diverse simulation scenarios with reactive and feasible agent behaviors can be used for validation and verification of self-driving system performance without relying on expensive and time-consuming real-world testing. Existing simulators rely on heuristic-based behavior models for background vehicles, which cannot capture the complex interactive behaviors in real-world scenarios. Meanwhile, existing learning-based methods are typical insufficient, yielding behaviors of traffic participants that frequently collide or drive off the road especially for a long horizon. To address this issue, we propose TrajGen, a two-stage trajectory generation framework, which can capture more realistic and diverse behaviors directly from human demonstration. In particular, TrajGen consists of the multi-modal trajectory prediction stage and the reinforcement learning based trajectory modification stage. In the first stage, we propose a novel auxiliary RouteLoss for the trajectory prediction model to generate multi-modal diverse trajectories in the drivable area. In the second stage, reinforcement learning is used to track the predicted trajectories while avoiding collisions, which can improve the reactivity and feasibility of generated trajectories. In addition, we develop a simulator I-Sim that can provide support for data-driven agent behavior simulation and train reinforcement learning models in parallel based on naturalistic driving data. The vehicle model in I-Sim can guarantee that the generated trajectories by TrajGen satisfy vehicle kinematic constraints. Finally, we give comprehensive metrics to evaluate generated trajectories for simulation scenarios, which shows that TrajGen outperforms either trajectory prediction or inverse reinforcement learning in terms of fidelity, reactivity, feasibility, and diversity. Yinfeng Gao, Youtian Guo, Dawei Ding 0001, Dongbin Zhao |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2022 | Boost 3-D Object Detection via Point Clouds Segmentation and Fused 3-D GIoU-L₁ LossabstractThe 3-D object detection is crucial for many real-world applications, attracting many researchers’ attention. Beyond 2-D object detection, 3-D object detection usually needs to extract appearance, depth, position, and orientation information from light detection and ranging (LiDAR) and camera sensors. However, due to more degrees of freedom and vertices, existing detection methods that directly transform from 2-D to 3-D still face several challenges, such as exploding increase of anchors’ number and inefficient or hard-to-optimize objective. To this end, we present a fast segmentation method for 3-D point clouds to reduce anchors, which can largely decrease the computing cost. Moreover, taking advantage of 3-D generalized Intersection of Union (GIoU) and$L_{1}$losses, we propose a fused loss to facilitate the optimization of 3-D object detection. A series of experiments show that the proposed method has alleviated the abovementioned issues effectively. Yaran Chen, Haoran Li 0010, Ruiyuan Gao 0001, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | BNAS: Efficient Neural Architecture Search Using Broad Scalable ArchitectureabstractEfficient neural architecture search (ENAS) achieves novel efficiency for learning architecture with high-performance via parameter sharing and reinforcement learning (RL). In the phase of architecture search, ENAS employs deep scalable architecture as search space whose training process consumes most of the search cost. Moreover, time-consuming model training is proportional to the depth of deep scalable architecture. Through experiments using ENAS on CIFAR-10, we find that layer reduction of scalable architecture is an effective way to accelerate the search process of ENAS but suffers from a prohibitive performance drop in the phase of architecture estimation. In this article, we propose a broad neural architecture search (BNAS) where we elaborately design broad scalable architecture dubbed broad convolutional neural network (BCNN) to solve the above issue. On the one hand, the proposed broad scalable architecture has fast training speed due to its shallow topology. Moreover, we also adopt RL and parameter sharing used in ENAS as the optimization strategy of BNAS. Hence, the proposed approach can achieve higher search efficiency. On the other hand, the broad scalable architecture extracts multi-scale features and enhancement representations, and feeds them into global average pooling (GAP) layer to yield more reasonable and comprehensive representations. Therefore, the performance of broad scalable architecture can be promised. In particular, we also develop two variants for BNAS that modify the topology of BCNN. In order to verify the effectiveness of BNAS, several experiments are performed and experimental results show that 1) BNAS delivers 0.19 days which is 2.37× less expensive than ENAS who ranks the best in RL-based NAS approaches; 2) compared with small-size (0.5 million parameters) and medium-size (1.1 million parameters) models, the architecture learned by BNAS obtains state-of-the-art performance (3.58% and 3.24% test error) on CIFAR-10; and 3) the learned architecture achieves 25.3% top-1 error on ImageNet just using 3.9 million parameters. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao, Zhiquan Sun, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Online Minimax Q Network Learning for Two-Player Zero-Sum Markov GamesabstractThe Nash equilibrium is an important concept in game theory. It describes the least exploitability of one player from any opponents. We combine game theory, dynamic programming, and recent deep reinforcement learning (DRL) techniques to online learn the Nash equilibrium policy for two-player zero-sum Markov games (TZMGs). The problem is first formulated as a Bellman minimax equation, and generalized policy iteration (GPI) provides a double-loop iterative way to find the equilibrium. Then, neural networks are introduced to approximate Q functions for large-scale problems. An online minimax Q network learning algorithm is proposed to train the network with observations. Experience replay, dueling network, and double Q-learning are applied to improve the learning process. The contributions are twofold: 1) DRL techniques are combined with GPI to find the TZMG Nash equilibrium for the first time and 2) the convergence of the online learning algorithm with a lookup table and experience replay is proven, whose proof is not only useful for TZMGs but also instructive for single-agent Markov decision problems. Experiments on different examples validate the effectiveness of the proposed algorithm on TZMG problems. Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | BNAS-v2: Memory-Efficient and Performance-Collapse-Prevented Broad Neural Architecture SearchabstractIn this article, we propose BNAS-v2 to further improve the efficiency of broad neural architecture search (BNAS), which employs a broad convolutional neural network (BCNN) as the search space. In BNAS, the single-path sampling-updating strategy of an overparameterized BCNN leads to terrible unfair training issue, which restricts the efficiency improvement. To mitigate the unfair training issue, we employ a continuous relaxation strategy to optimize all paths of the overparameterized BCNN simultaneously. However, continuous relaxation leads to a performance collapse issue that leads to the unsatisfactory performance of the learned BCNN. For that, we propose the confident learning rate (CLR) and introduce the combination of partial channel connections and edge normalization. Experimental results show that 1) BNAS-v2 delivers state-of-the-art search efficiency on both CIFAR-10 (0.05 GPU days, which is$4\times $faster than BNAS) and ImageNet (0.19 GPU days) with better or competitive performance; 2) the above two solutions are effectively alleviating the performance collapse issue; and 3) BNAS-v2 achieves powerful generalization ability on multiple transfer tasks, e.g., MNIST, FashionMNIST, NORB, and SVHN. The code is available athttps://github.com/zixiangding/BNASv2. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | MGRL: Graph neural network based inference in a Markov network with reinforcement learning for visual navigation
Yaran Chen, Dongbin Zhao, Dong Li 0016 |
Neurocomputing | 3 |
| 2021 | Optimal Feedback Control of Pedestrian Flow in Heterogeneous CorridorsabstractMaintaining the orderliness and efficiency of pedestrian flow through an architectural area is critical for the evacuation process. Especially, clogs and jams are easily triggered in width-changing areas. In this article, we consider pedestrian movement in heterogeneous corridors and design an optimal feedback control to regulate pedestrian flow. Flow characteristics are first studied based on microscopic social-force simulations. A Gaussian process describes the relationship between flow variables with the observation data. The macroscopic model for flow in heterogeneous corridors is developed. To avoid jams, discharges among these corridors are balanced with the narrowest corridor as the primary concern. At the equilibrium, a continuous-time nonlinear control system is formulated, and the adaptive dynamic programming learns the optimal feedback controller. Policy iteration (PI) and neural networks are combined together, and the convergence of neural-network-based PI is demonstrated by analyzing its equivalence to the Gauss–Newton method. Batch normalization is introduced to stabilize the learning process. Simulated experiments demonstrate that the control design can effectively regulate pedestrian flow for both macroscopic and microscopic models.Note to Practitioners—The development of video-processing techniques provides a powerful tool to detect human behavior in real time. In crowd events, the pedestrian movement must be regulated; otherwise, it is easy to fall into the faster-is-slower effect. It is especially important for evacuation routes with different widths. In this article, the optimal feedback control is studied to regulate pedestrian flow in heterogeneous corridors. It takes flow densities as state and produces commands that are composed of entrance influx and free-flow velocities. These commands can be executed with the support of speakers, displays, or the recently developed interactive robots. To avoid congestion, discharges of different corridors are balanced, and the system is optimally stabilized at equilibrium. Based on our work, engineers are able to design pedestrian flow control and achieve optimal evacuation in arbitrary heterogeneous corridors. Yuanheng Zhu, Dongbin Zhao, Haibo He |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Shift-Invariant Convolutional Network SearchabstractThe development of Neural Architecture Search (NAS) makes Convolutional Neural Networks (CNN) more diverse and effective. But previous NAS approaches don't pay attention to the shift-invariant of CNN. Without the shift-invariant, convolutional network is not robust enough when input data is disturbed or damaged. Besides, taking accuracy as the only optimization goal of NAS cannot meet the increasingly diverse needs. In this paper, we propose Shift-Invariant Convolutional Network Search (SICNS). It uses one-shot NAS to search for shift-invariant convolutional network by incorporating the low-pass filter into the one-shot model. Furthermore, SICNS optimizes multiple indicators simultaneously through the multi-objective evolutionary algorithm. Through training one-shot model and evolving the architecture, we obtain convolutional networks which are robust and powerful on image classification task. Especially, our work can achieve 4.52% test error on CIFAR-10 with 0.7M parameters. And in case the input data are disturbed, the accuracy of searched network is 2.96% higher than network without low-pass filter. Nannan Li 0003, Yaran Chen, Zixiang Ding, Dongbin Zhao |
IJCNN | 4 |
| 2020 | RailNet: An Information Aggregation Network for Rail Track SegmentationabstractAs the basis of scenes understanding for the track inspection task, track segmentation is challenging due to the various illumination conditions, track crossing, and plant coverage. Since the rail has a strong shape prior, strict rail spacing and special distribution in the image, making full use of the spatial information of the rail features becomes an important factor to improve the accuracy of rail segmentation. In this paper, an information aggregation module is proposed to enhance the spatial relationship between pixels of the rail features. In other words, this module expands the receptive field. Furthermore, we build an information aggregation network based on this module, which is called as RailNet. Finally, the RailNet is evaluated in an open train track dataset. Experimental results show that RailNet can achieve the best performance so far in the dataset of trains. Haoran Li 0010, Dongbin Zhao, Yaran Chen |
IJCNN | 3 |
| 2020 | An Improved Minimax-Q Algorithm Based on Generalized Policy Iteration to Solve a Chaser-Invader GameabstractIn this paper, we use reinforcement learning and zero-sum games to solve a Chaser-Invader game, which is actually a Markov game (MG). Different from the single agent Markov Decision Process (MDP), MG can realize the interaction of multiple agents, which is an extension of game theory to a MDP environment. This paper proposes an improved algorithm based on the classical Minimax-Q algorithm. First, in order to solve the problem where Minimax-Q algorithm can only be applied for discrete and simple environment, we use Deep Q-network instead of traditional Q-learning. Second, we propose a generalized policy iteration to solve the zero-sum game. This method makes the agent use linear programming method to solve the Nash equilibrium action at each moment. Finally, through comparative experiments, we prove that the improved algorithm can perform as well as Monte Carlo Tree Search in simple environments and better than Monte Carlo Tree Search in complex environments. Minsong Liu, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 3 |
| 2020 | Cooperative Multi-Agent Deep Reinforcement Learning with Counterfactual RewardabstractIn partially observable fully cooperative games, agents generally tend to maximize global rewards with joint actions, so it is difficult for each agent to deduce their own contribution. To address this credit assignment problem, we propose a multi-agent reinforcement learning algorithm with counterfactual reward mechanism, which is termed as CoRe algorithm. CoRe computes the global reward difference in condition that the agent does not take its actual action but takes other actions, while other agents fix their actual actions. This approach can determine each agent's contribution for the global reward. We evaluate CoRe in a simplified Pig Chase game with a decentralised Deep Q Network (DQN) framework. The proposed method helps agents learn end-to-end collaborative behaviors. Compared with other DQN variants with global reward, CoRe significantly improves learning efficiency and achieves better results. In addition, CoRe shows excellent performances in various size game environments. Kun Shao, Yuanheng Zhu, Zhentao Tang, Dongbin Zhao |
IJCNN | 4 |
| 2020 | ContourRend: A Segmentation Method for Improving Contours by Rendering
Yaran Chen, Dongbin Zhao, Zhong-Hua Pang |
ISNN | 4 |
| 2020 | Dynamically Weighted Model Predictive Control of Affine Nonlinear Systems Based on Two-Timescale Neurodynamic Optimization
Jiasen Wang, Jun Wang 0002, Dongbin Zhao |
ISNN | 3 |
| 2020 | Advances in deep neural information processing
Dongbin Zhao, Shukai Duan 0001, Zheng Yan 0001, Cesare Alippi |
Neurocomputing | 1 |
| 2020 | Hierarchical optimal control for input-affine nonlinear systems through the formulation of Stackelberg game
Chaoxu Mu, Ke Wang 0037, Dongbin Zhao |
Inf. Sci. | 4 |
| 2020 | LMI-Based Synthesis of String-Stable Controller for Cooperative Adaptive Cruise ControlabstractController synthesis is a challenging problem in cooperative adaptive cruise control (CACC). Especially the requirement of string stability makes it even harder to choose appropriate control parameters. This paper applies a time-domain definition to string stability and converts the problem to the H∞control of a time-delay system. Based on the proposed control structure, the H∞norm and stability criteria of CACC are satisfied by a set of constraints in terms of a Lyapunov-Krasovskii functional candidate. These constraints are further reduced to linear matrix inequalities so that feasible solutions can be easily and efficiently computed. Simulations on an identified model validate the performance of our method in both frequency and time domains. Yuanheng Zhu, Haibo He, Dongbin Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Deep Reinforcement Learning-Based Automatic Exploration for Navigation in Unknown EnvironmentabstractThis paper investigates the automatic exploration problem under the unknown environment, which is the key point of applying the robotic system to some social tasks. The solution to this problem via stacking decision rules is impossible to cover various environments and sensor properties. Learning-based control methods are adaptive for these scenarios. However, these methods are damaged by low learning efficiency and awkward transferability from simulation to reality. In this paper, we construct a general exploration framework via decomposing the exploration process into the decision, planning, and mapping modules, which increases the modularity of the robotic system. Based on this framework, we propose a deep reinforcement learning-based decision algorithm that uses a deep neural network to learning exploration strategy from the partial map. The results show that this proposed algorithm has better learning efficiency and adaptability for unknown environments. In addition, we conduct the experiments on the physical robot, and the results suggest that the learned policy can be well transferred from simulation to the real robot. Haoran Li 0010, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Invariant Adaptive Dynamic Programming for Discrete-Time Optimal ControlabstractFor systems that can only be locally stabilized, control laws and their effective regions are both important. In this paper, invariant policy iteration is proposed to solve the optimal control of discrete-time systems. At each iteration, a given policy is evaluated in its invariantly admissible region, and a new policy and a new region are updated for the next iteration. Theoretical analysis shows the method is regionally convergent to the optimal value and the optimal policy. Combined with sum-of-squares polynomials, the method is able to achieve the near-optimal control of a class of discrete-time systems. An invariant adaptive dynamic programming algorithm is developed to extend the method to scenarios where system dynamics is not available. Online data are utilized to learn the near-optimal policy and the invariantly admissible region. Simulated experiments verify the effectiveness of our method. Yuanheng Zhu, Dongbin Zhao, Haibo He |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Auto-encoder based Graph Convolutional Networks for Online Financial Anti-fraudabstractMany practical problems can be formulated as graph-based semi-supervised classification problems. For example, online finance anti-fraud. Recently, many researchers attempt using deep learning methods to solve such problems. In this paper, we propose a novel neural network architecture to perform semi-supervised classification on graph-structured data. We improve the graph convolutional network (GCN) by replacing the graph convolution matrix with auto-encoder module. The proposed neural network is trained by a multi-task objective function. Except the classification task, we train the auto-encoder module to reconstruct the graph convolution matrix. It can be seen as an adaptive spectral convolution on graph. It can increase the depth of neural network without causing over-smooth effect. Additionally, the introduction of reconstruction task can mitigate the cold-start problem. Even the graph topological structure is extreme sparse, our method can learn expressive latent features for vertices. The experimental results show that our method can achieve the state of art performance. Le Lv, Jianbo Cheng, Nanbo Peng, Dongbin Zhao |
CIFEr | 5 |
| 2019 | Lane Change Decision-making through Deep Reinforcement Learning with Rule-based ConstraintsabstractAutonomous driving decision-making is a great challenge due to the complexity and uncertainty of the traffic environment. Combined with the rule-based constraints, a Deep Q-Network (DQN) based method is applied for autonomous driving lane change decision-making task in this study. Through the combination of high-level lateral decision-making and low-level rule-based trajectory modification, a safe and efficient lane change behavior can be achieved. With the setting of our state representation and reward function, the trained agent is able to take appropriate actions in a real-world-like simulator. The generated policy is evaluated on the simulator for 10 times, and the results demonstrate that the proposed rule-based DQN method outperforms the rule-based approach and the DQN method. Dongbin Zhao, Yaran Chen |
IJCNN | 3 |
| 2019 | Model-Free Reinforcement Learning based Lateral Control for Lane KeepingabstractIn this paper, the lateral control strategy for lane keeping task, which is an important module in the advanced assistant driver systems, is proposed based on the model-free reinforcement learning. Different from the model-based methods, our method only requires the generated data rather than the accurate system model. Furthermore, the lateral control strategy for driver model lane keeping is given, where driver controller and direct yaw controller (DYC) are working at the same time to maintain the vehicle stability. Note that the dynamic game theory is considered for this task, where the steering wheel controller for driver and the DYC compensated controller are obtained based on Nash game theory. Finally, we give simulation examples to prove the validity of the proposed schemes. Dongbin Zhao, Chaomin Luo, Dianwei Qian |
IJCNN | 3 |
| 2019 | Optimal Pedestrian Evacuation in Building with Consecutive Differential Dynamic ProgrammingabstractFast and efficient evacuation of pedestrians from an enclosed area is a difficult but crucial issue in modern society. In this paper, the optimization of evacuation from a building is studied. A graph is adopted to describe the building layout with nodes representing areas and edges representing connections. The dynamics of the evacuation process in the graph is formulated by a nonlinear discrete-time model at a macroscopic level. To find the optimal evacuation plan, a consecutive differential dynamic programming is developed. It inherits the differential dynamic programming property that solves the value and optimal policy locally. Additionally, it consecutively executes actions for multiple steps in the trajectory, which is beneficial to reduce computational burden and lower optimization difficulty. Simulations on a four-storey building layout demonstrates our method is efficient and suitable for on-site evacuation plan making. Yuanheng Zhu, Haibo He, Dongbin Zhao, Zhongsheng Hou |
IJCNN | 3 |
| 2019 | Graph-FCN for Image Semantic Segmentation
Yaran Chen, Dongbin Zhao |
ISNN (1) | 3 |
| 2019 | Deep Kalman Filter with Optical Flow for Multiple Object TrackingabstractDeep matching and Kalman filter-based multiple object tracking (DK-tracking) have been demonstrated to be promising. However, most of existing DK-tracking trackers assume that objects are slow-varying movement with a constant velocity. The assumption is hard to be satisfied in the real world, especially in the image space due to the sight distance. In this paper, we propose a novel multiple object tracking method combining deep feature matching, Kalman filter and flow information, which is called DK-flow-tracking, to improve tracking performance. In DK-flow-tracking, optical flow in consecutive frames is used to provide accurate object motion information for guiding Kalman filter to track objects. Experiments are performed on public datasets: MOT2016, MOT2017, and the proposed method achieves better performances compared to the DK-tracking with the assumption of a constant velocity movement. Yaran Chen, Dongbin Zhao, Haoran Li 0010 |
SMC | 2 |
| 2019 | Deep sparse representation-based mid-level visual elements discovery in fine-grained classification
Le Lv, Dongbin Zhao, Kun Shao |
Soft Comput. | 2 |
| 2019 | Adaptive cruise control via adaptive dynamic programming with experience replay
Bin Wang 0034, Dongbin Zhao, Jin Cheng 0004 |
Soft Comput. | 2 |
| 2019 | Data-Based Reinforcement Learning for Nonzero-Sum Games With Unknown Drift DynamicsabstractThis paper is concerned about the nonlinear optimization problem of nonzero-sum (NZS) games with unknown drift dynamics. The data-based integral reinforcement learning (IRL) method is proposed to approximate the Nash equilibrium of NZS games iteratively. Furthermore, we prove that the data-based IRL method is equivalent to the model-based policy iteration algorithm, which guarantees the convergence of the proposed method. For the implementation purpose, a single-critic neural network structure for the NZS games is given. To enhance the application capability of the data-based IRL method, we design the updating laws of critic weights based on the offline and online iterative learning methods, respectively. Note that the experience replay technique is introduced in the online iterative learning, which can improve the convergence rate of critic weights during the learning process. The uniform ultimate boundedness of the critic weights are guaranteed using the Lyapunov method. Finally, the numerical results demonstrate the effectiveness of the data-based IRL algorithm for nonlinear NZS games with unknown drift dynamics. Dongbin Zhao |
IEEE Trans. Cybern. | 2 |
| 2018 | Value Iteration Algorithm for Optimal Consensus Control of Multi-agent Systems
Dongbin Zhao |
ICONIP (7) | 2 |
| 2018 | Driving Control with Deep and Reinforcement Learning in The Open Racing Car Simulator
Yuanheng Zhu, Dongbin Zhao |
ICONIP (3) | 2 |
| 2018 | A temporal-based deep learning method for multiple objects detection in autonomous drivingabstractThis paper proposes a novel vision-based object detection method in autonomous driving, which introduces the temporal information into the deep learning-based detection method for moving object detection. Vision-based object detection is a critical technology for autonomous driving. The objects in the real world such as driving cars, don't have great changes in their positions and velocities. So the position change of objects between two consecutive frames is not large. This is usually ignored by traditional works, which usually use object detection methods on still-images to detect moving objects. Considering the relationship among consecutive frames (temporal information), we present a robust and real-time tracking method following image detection to refine the object detection results. Based on the three key attributes (distances, sizes and positions), the tracking method aims to build the association between the detected objects on the current frame and those in previous frames. The proposed object detection with temporal information dramatically improves the performance of existing object detection algorithms based on stillimage. With the proposed method, we won the champion in the preceding vehicle detection task in 2017 intelligent vehicle future challenge(2017 IVFC)1. Yaran Chen, Dongbin Zhao, Haoran Li 0010, Dong Li 0016, Ping Guo 0002 |
IJCNN | 2 |
| 2018 | DeepSign: Deep Learning based Traffic Sign RecognitionabstractThis paper investigates the traffic sign recognition task with deep learning methods. The proposed algorithm which is called DeepSign includes three modules: a detection module (PosNet) for locating the traffic sign in a static image, a classification module (PatchNet) for classifying the detected image patch, and a temporal filter for correcting the recognition results. The PosNet is a binary object detection convolution neural network which regards all traffic signs as one class and the background as the other class. Different from the traditional works which recognize the traffic sign on the static image, the proposed temporal filter exploits the contextual information to recover the missed detection region and correct the false classification. The experiments validate the effectiveness of the proposed algorithm. It achieved the third place on the traffic sign recognition task in 2017 China intelligent vehicle future challenge (2017 CIVFC). Dong Li 0016, Dongbin Zhao, Yaran Chen |
IJCNN | 2 |
| 2018 | Visual Navigation with Actor-Critic Deep Reinforcement LearningabstractVisual navigation in complex environments is crucial for intelligent agents. In this paper, we propose an efficient deep reinforcement learning (DRL) method to tackle visual navigation tasks. We present the synchronous advantage actor-critic (A2C) with generalized advantage estimator (GAE) algorithm. The A2C enables agents to learn from multiple processes, which significantly reduces the training time. The GAE used to estimate the advantage function improves the policy gradient estimates. We focus on visual navigation tasks in ViZDoom, and train agents in two health gathering scenarios. The experimental results show this method successfully teaches our agents to navigate in these scenarios. The A2C with GAE agent reaches the highest score in the first task, and a competitive score in the second task. In addition, this agent has better average scores and lower variances in both tasks. Kun Shao, Dongbin Zhao, Yuanheng Zhu |
IJCNN | 2 |
| 2018 | Model-Free Reinforcement Learning for Fully Cooperative Multi-Agent Graphical GamesabstractIn this paper, the optimal coordinated control problem for the homogeneous multi-agent graphical games with completely unknown dynamics is investigated. The off-policy reinforcement learning is proposed to approach the solution of the Hamilton-Jacobi equation under the framework of centralized training and decentralized execution. The actor-critic structure is adopted to learn the optimal control policies. Note that the critic network is centralized using the information from all the agents, and the parameter sharing scheme is adopted for the single actor network during the training process. For the execution process, the centralized critic network is not required, and only the trained actor network is used for each agent to obtain the control input based on its individual observation. For the implementation purpose, the neural network approximators with the actor-critic structure are constructed to approach the optimal centralized value function and the optimal policies for the multiagent graphical games. Finally, a simulation example is provided to demonstrate the effectiveness of the proposed algorithm. Dongbin Zhao, Frank L. Lewis |
IJCNN | 2 |
| 2018 | Multi-task learning for dangerous object detection in autonomous driving
Yaran Chen, Dongbin Zhao, Le Lv |
Inf. Sci. | 2 |
| 2018 | Policy Iteration for H∞ Optimal Control of Polynomial Nonlinear Systems via Sum of Squares ProgrammingabstractSum of squares (SOS) polynomials have provided a computationally tractable way to deal with inequality constraints appearing in many control problems. It can also act as an approximator in the framework of adaptive dynamic programming. In this paper, an approximate solution to the optimal control of polynomial nonlinear systems is proposed. Under a given attenuation coefficient, the Hamilton-Jacobi-Isaacs equation is relaxed to an optimization problem with a set of inequalities. After applying the policy iteration technique and constraining inequalities to SOS, the optimization problem is divided into a sequence of feasible semidefinite programming problems. With the converged solution, the attenuation coefficient is further minimized to a lower value. After iterations, approximate solutions to the smallest -gain and the associated optimal controller are obtained. Four examples are employed to verify the effectiveness of the proposed algorithm. Yuanheng Zhu, Dongbin Zhao, Xiong Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | A pdf-Free Change Detection Test Based on Density Difference EstimationabstractThe ability to detect online changes in stationarity or time variance in a data stream is a hot research topic with striking implications. In this paper, we propose a novel probability density function-free change detection test, which is based on the least squares density-difference estimation method and operates online on multidimensional inputs. The test does not require any assumption about the underlying data distribution, and is able to operate immediately after having been configured by adopting a reservoir sampling mechanism. Thresholds requested to detect a change are automatically derived once a false positive rate is set by the application designer. Comprehensive experiments validate the effectiveness in detection of the proposed method both in terms of detection promptness and accuracy. Li Bu, Cesare Alippi, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Event-Based Robust Control for Uncertain Nonlinear Systems Using Adaptive Dynamic ProgrammingabstractIn this paper, the robust control problem for a class of continuous-time nonlinear system with unmatched uncertainties is investigated using an event-based control method. First, the robust control problem is transformed into a corresponding optimal control problem with an augmented control and an appropriate cost function. Under the event-based mechanism, we prove that the solution of the optimal control problem can asymptotically stabilize the uncertain system with an adaptive triggering condition. That is, the designed event-based controller is robust to the original uncertain system. Note that the event-based controller is updated only when the triggering condition is satisfied, which can save the communication resources between the plant and the controller. Then, a single network adaptive dynamic programming structure with experience replay technique is constructed to approach the optimal control policies. The stability of the closed-loop system with the event-based control policy and the augmented control policy is analyzed using the Lyapunov approach. Furthermore, we prove that the minimal intersample time is bounded by a nonzero positive constant, which excludes Zeno behavior during the learning process. Finally, two simulation examples are provided to demonstrate the effectiveness of the proposed control scheme. Dongbin Zhao, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Special Issue on Deep Reinforcement Learning and Adaptive Dynamic ProgrammingabstractThe sixteen papers in this special section focus on deep reinforcement learning and adaptive dynamic programming (deep RL/ADP). Deep RL is able to output control signal directly based on input images, which incorporates both the advantages of the perception of deep learning (DL) and the decision making of RL or adaptive dynamic programming (ADP). This mechanism makes the artificial intelligence much closer to human thinking modes. Deep RL/ADP has achieved remarkable success in terms of theory and applications since it was proposed. Successful applications cover video games, Go, robotics, smart driving, healthcare, and so on. However, it is still an open problem to perform the theoretical analysis on deep RL/ADP, e.g., the convergence, stability, and optimality analyses. The learning efficiency needs to be improved by proposing new algorithms or combined with other methods. More practical demonstrations are encouraged to be presented. Therefore, the aim of this special issue is to call for the most advanced research and state-of-the-art works in the field of deep RL/ADP. Dongbin Zhao, Derong Liu 0001, Frank L. Lewis, José C. Príncipe, Stefano Squartini |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | FMR-GA - A Cooperative Multi-agent Reinforcement Learning Algorithm Based on Gradient Ascent
Zhen Zhang 0009, Dongqing Wang, Dongbin Zhao |
ICONIP (1) | 3 |
| 2017 | Off-Policy Reinforcement Learning for Partially Unknown Nonzero-Sum Games
Dongbin Zhao |
ICONIP (1) | 2 |
| 2017 | Policy gradient methods with Gaussian process modelling accelerationabstractPolicy gradient algorithm is often used to deal with the continuous control problems. But as a model-free algorithm, it suffers from the low data efficiency and long learning phase. In this paper, a policy gradient with Gaussian process modelling (PGGPM) algorithm is proposed to accelerate learning process. The system model is approximated by Gaussian process in an incremental way, which is used to explore state action space virtually by generating imaginary samples. Both the real and imaginary samples are used to train the actor and critic networks. Finally, we apply our algorithm to two experiments to verify that Gaussian process can accurately fit system model and the supplementary imaginary samples can speed up the learning phase. Dong Li 0016, Dongbin Zhao, Chaomin Luo |
IJCNN | 2 |
| 2017 | Multi-task Learning with Cartesian Product-Based Multi-objective Combination for Dangerous Object Detection
Yaran Chen, Dongbin Zhao |
ISNN (1) | 2 |
| 2017 | Data-driven adaptive dynamic programming for continuous-time fully cooperative games with partially constrained inputs
Dongbin Zhao, Yuanheng Zhu |
Neurocomputing | 2 |
| 2017 | FMRQ - A Multiagent Reinforcement Learning Algorithm for Fully Cooperative TasksabstractIn this paper, we propose a multiagent reinforcement learning algorithm dealing with fully cooperative tasks. The algorithm is called frequency of the maximum reward Q-learning (FMRQ). FMRQ aims to achieve one of the optimal Nash equilibria so as to optimize the performance index in multiagent systems. The frequency of obtaining the highest global immediate reward instead of immediate reward is used as the reinforcement signal. With FMRQ each agent does not need the observation of the other agents' actions and only shares its state and reward at each step. We validate FMRQ through case studies of repeated games: four cases of two-player two-action and one case of three-player two-action. It is demonstrated that FMRQ can converge to one of the optimal Nash equilibria in these cases. Moreover, comparison experiments on tasks with multiple states and finite steps are conducted. One is box-pushing and the other one is distributed sensor network problem. Experimental results show that the proposed algorithm outperforms others with higher performance. Zhen Zhang 0009, Dongbin Zhao, Junwei Gao, Dongqing Wang, Yujie Dai |
IEEE Trans. Cybern. | 2 |
| 2017 | Guest Editorial Special Issue on New Developments in Neural Network Structures for Signal Processing, Autonomous Decision, and Adaptive ControlabstractThere has been continuously increasing interest in applying neural networks (NNs) to identification and adaptive control of practical systems that are characterized by nonlinearity, uncertainty, communication constraints, and complexity. The past few years have witnessed a variety of new developments in NN-based approaches for behavior learning, information processing, autonomous decision, and system control. Biologically inspired NN structures can significantly enhance the capabilities of information processing, control, and computational performance. New discoveries in neurocognitive psychology, sociology, and elsewhere reveal new neurological learning structures with more powerful capabilities in complex problem solving and fast decision in dynamic environments. The goal of the special issue is to consolidate recent new developments in NN structures for signal processing, autonomous decision, and adaptive control with application to complex systems. It includes contributions from a wide range of research aspects relevant to the topic, ranging from neural computing, adaptive control, cooperative control, autonomous decision systems, mathematical and computational models, neuropsychology decision and control, algorithms and simulation, to applications and/or case studies. This issue contains 24 papers and the contents of which are summarized below. Yongduan Song 0001, Frank L. Lewis, Marios M. Polycarpou, Danil V. Prokhorov, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Iterative Adaptive Dynamic Programming for Solving Unknown Nonlinear Zero-Sum Game Based on Online Dataabstractcontrol is a powerful method to solve the disturbance attenuation problems that occur in some control systems. The design of such controllers relies on solving the zero-sum game (ZSG). But in practical applications, the exact dynamics is mostly unknown. Identification of dynamics also produces errors that are detrimental to the control performance. To overcome this problem, an iterative adaptive dynamic programming algorithm is proposed in this paper to solve the continuous-time, unknown nonlinear ZSG with only online data. A model-free approach to the Hamilton-Jacobi-Isaacs equation is developed based on the policy iteration method. Control and disturbance policies and value are approximated by neural networks (NNs) under the critic-actor-disturber structure. The NN weights are solved by the least-squares method. According to the theoretical analysis, our algorithm is equivalent to a Gauss-Newton method solving an optimization problem, and it converges uniformly to the optimal solution. The online data can also be used repeatedly, which is highly efficient. Simulation results demonstrate its feasibility to solve the unknown nonlinear ZSG. When compared with other algorithms, it saves a significant amount of online measurement time. Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | An Incremental Change Detection Test Based on Density Difference EstimationabstractWe propose incremental least squares density difference (LSDD) change detection method, an incremental test to detect changes in stationarity based on the difference between the unknown prechange and the post-change probability density functions (pdfs). The method is computationally light and, hence, adequate to process continuous data streams, as those emerging from the Internet of Things and the big data framework. The incremental change detection test operates on two nonoverlapping data windows to estimate the LSDD between the two pdfs. We construct a theoretical framework that shows how the distribution of LSDD values follows a linear combination of χ2distributions and provides thresholds to control false positive rates. The proposed test can operate online, with needed estimates and thresholds computed incrementally as fresh samples come. Comprehensive experiments validate the effectiveness of the test both in detecting abrupt and drift types of changes. Li Bu, Dongbin Zhao, Cesare Alippi |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2017 | Event-Triggered H∞ Control for Continuous-Time Nonlinear System via Concurrent LearningabstractIn this paper, the H∞optimal control problem for a class of continuous-time nonlinear systems is investigated using event-triggered method. First, the H∞optimal control problem is formulated as a two-player zero-sum (ZS) differential game. Then, an adaptive triggering condition is derived for the ZS game with an event-triggered control policy and a time-triggered disturbance policy. The event-triggered controller is updated only when the triggering condition is not satisfied. Therefore, the communication between the plant and the controller is reduced. Furthermore, a positive lower bound on the minimal intersample time is provided to avoid Zeno behavior. For implementation purpose, the event-triggered concurrent learning algorithm is proposed, where only one critic neural network (NN) is used to approximate the value function, the control policy and the disturbance policy. During the learning process, the traditional persistence of excitation condition is relaxed using the recorded data and instantaneous data together. Meanwhile, the stability of closed-loop system and the uniform ultimate boundedness (UUB) of the critic NN's parameters are proved by using Lyapunov technique. Finally, simulation results verify the feasibility to the ZS game and the corresponding H∞control problem. Dongbin Zhao, Yuanheng Zhu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | Ensemble LSDD-based change detection testsabstractThe least squares density difference change detection test (LSDD-CDT) has proven to be an effective method in detecting concept drift by inspecting features derived from the discrepancy between two probability density functions (pdfs). The first pdf is associated with the concept drift free case, the second to the possible post change one. Interestingly, the method permits to control the ratio of false positives. This paper introduces and investigates the performance of a family of LSDD methods constructed by exploring different ensemble options applied to the basic CDT procedure. Experiments show that most of proposed methods are characterized by improved performance in change detection once compared with the direct ensemble-free counterpart. Li Bu, Cesare Alippi, Dongbin Zhao |
IJCNN | 3 |
| 2016 | A general adaptive dynamic programming approach with experience replayabstractExperience replay is a promising approach to improve the learning efficiency of adaptive dynamic programming. A general model-free adaptive dynamic programming (ADP) approach with the experience replay technology is investigated in this paper to solve the optimal control problems in continuous state and action spaces. Both the critic network and action network are modeled with a feedforward neural network with one hidden layer. During the learning process, a number of recently observed data samples are recorded in a database. When updating the parameters of the neural networks, the data in the sample database are repeatedly used to update the weights of the action network and the critic network. Implementation details of the algorithm are given, and simulation experiments are utilized to verify the learning efficiency of the proposed approach. Bin Wang 0034, Dongbin Zhao, Jin Cheng 0004, Yuan Xu 0003, Yueyang Li 0001 |
IJCNN | 2 |
| 2016 | Model-free reinforcement learning for nonlinear zero-sum games with simultaneous explorationsabstractIn this paper, the continuous-time unknown nonlinear zero-sum game is investigated using a model-free online learning method. First, motivated by model-based policy iteration, an iterative equation without any knowledge of system dynamics is derived by introducing simultaneous explorations. Then, the model-free reinforcement learning based on the derived iterative equation is developed to approach the solution of the Hamilton-Jacobi-Isaacs equation. For the online implementation purpose, three neural networks are constructed to approach the value function, control and disturbance policies, respectively. Finally, a simulation example is provided to demonstrate the effectiveness of the proposed scheme. Dongbin Zhao, Yuanheng Zhu |
IJCNN | 2 |
| 2016 | Convolutional fitted Q iteration for vision-based control problemsabstractIn this paper a deep reinforcement learning (DRL) method is proposed to solve the control problem which takes raw image pixels as input states. A convolutional neural network (CNN) is used to approximate Q functions, termed as Q-CNN. A pretrained network, which is the result of a classification challenge on a vast set of natural images, initializes the parameters of Q-CNN. Such initialization assigns Q-CNN with the features of image representation, so it is more concentrated on the control tasks. The weights are tuned under the scheme of fitted Q iteration (FQI), which is an offline reinforcement learning method with the stable convergence property. To demonstrate the performance, a modified Food-Poison problem is simulated. The agent determines its movements based on its forward view. In the end the algorithm successfully learns a satisfied policy which has better performance than the results of previous researches. Dongbin Zhao, Yuanheng Zhu, Le Lv, Yaran Chen |
IJCNN | 1 |
| 2016 | Experience Replay for Optimal Control of Nonzero-Sum Game Systems With Unknown DynamicsabstractIn this paper, an approximate online equilibrium solution is developed for an N -player nonzero-sum (NZS) game systems with completely unknown dynamics. First, a model identifier based on a three-layer neural network (NN) is established to reconstruct the unknown NZS games systems. Moreover, the identifier weight vector is updated based on experience replay technique which can relax the traditional persistence of excitation condition to a simplified condition on recorded data. Then, the single-network adaptive dynamic programming (ADP) with experience replay algorithm is proposed for each player to solve the coupled nonlinear Hamilton- (HJ) equations, where only the critic NN weight vectors are required to tune for each player. The feedback Nash equilibrium is provided by the solution of the coupled HJ equations. Based on the experience replay technique, a novel critic NN weights tuning law is proposed to guarantee the stability of the closed-loop system and the convergence of the value functions. Furthermore, a Lyapunov-based stability analysis shows that the uniform ultimate boundedness of the closed-loop system is achieved. Finally, two simulation examples are given to verify the effectiveness of the proposed control scheme. Dongbin Zhao, Ding Wang 0001, Yuanheng Zhu |
IEEE Trans. Cybern. | 1 |
| 2016 | Fuzzy-Based Goal Representation Adaptive Dynamic ProgrammingabstractIn this paper, a novel nonlinear learning controller called fuzzy-based goal representation adaptive dynamic programming (Fuzzy-GrADP) is proposed. In the proposed GrADP method, a goal representation network is introduced to generate an adaptive internal reinforcement signal to the critic network to help the controller provide a general mapping between the input and output actions. Moreover, in the proposed architecture, the action network in the GrADP is improved by using the fuzzy hyperbolic model, which combines the merits of the fuzzy model and the neural network model. Based on the back-propagation technique, the parameters in the membership functions and the fuzzy rules are all undergo training and online adapting. The proposed controller is tested on two numerical benchmarks, and the simulation results show that the proposed controller outperforms the original adaptive dynamic fuzzy controller and the pure neural network-based GrADP controller. In addition, the proposed controller is further applied on a large multimachine power system for static var compensator damping control, where simulation results demonstrate the effectiveness of the proposed approach on real applications. Furthermore, in order to demonstrate the theoretical guarantee of the proposed method, Lyapunov stability analysis to support the proposed Fuzzy-GrADP approach has also been carried out. Yufei Tang, Haibo He, Zhen Ni, Xiangnan Zhong, Dongbin Zhao, Xin Xu 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2016 | Data-Based Adaptive Critic Designs for Nonlinear Robust Optimal Control With Uncertain DynamicsabstractIn this paper, the infinite-horizon robust optimal control problem for a class of continuous-time uncertain nonlinear systems is investigated by using data-based adaptive critic designs. The neural network identification scheme is combined with the traditional adaptive critic technique, in order to design the nonlinear robust optimal control under uncertain environment. First, the robust optimal controller of the original uncertain system with a specified cost function is established by adding a feedback gain to the optimal controller of the nominal system. Then, a neural network identifier is employed to reconstruct the unknown dynamics of the nominal system with stability analysis. Hence, the data-based adaptive critic designs can be developed to solve the Hamilton-Jacobi-Bellman equation corresponding to the transformed optimal control problem. The uniform ultimate boundedness of the closed-loop system is also proved by using the Lyapunov approach. Finally, two simulation examples are presented to illustrate the effectiveness of the developed control strategy. Ding Wang 0001, Derong Liu 0001, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2015 | Thermal comfort control based on MEC algorithm for HVAC systemsabstractThis paper combines an efficient reinforcement learning algorithm named Multisamples in Each Cell (MEC) with a building thermal comfort control problem. It implements the efficient exploration rule and makes high use of observed samples. A grid is utilized to partition the continuous state into cells that are used to store samples. A near-upper Q function is obtained based on the samples in each cell. The value iteration technique is designed to derive the near optimal control policy. The algorithm can efficiently balance exploration and exploitation. The entire implementation process needs no model of systems. The thermal comfort criterion, predicted mean vote, is introduced to evaluate zone thermal comfort status. A two story, multi-zone small office building equipped with a variable air volume direct expansion cooling system is built in EnergyPlus to establish an EnergyPlus-MATLAB co-simulation platform. A MEC thermal comfort control simulation is implemented to validate the high performance property compared with Q-learning. Dong Li 0016, Dongbin Zhao, Yuanheng Zhu, Zhongpu Xia |
IJCNN | 2 |
| 2015 | Online reinforcement learning by Bayesian inferenceabstractPolicy evaluation has long been one of the core issues of the online reinforcement learning, especially in the continuous state domain. In this paper, the issue is addressed by employing Gaussian processes to represent the action value function from the probability perspective. By modeling the return as a stochastic variable, the action value function can sequentially update according to observed variables such as state and reward by Bayesian inference during the policy evaluation. The update rule shows that it is a temporal difference learning method with the learning rate determined by the uncertainty of a collected sample. Incorporating the policy evaluation method with the ∈-greedy action selection method, we propose an online reinforcement learning algorithm referred as to Bayesian-SARSA. It is tested on some benchmark problems and the empirical results verifies its effectiveness. Zhongpu Xia, Dongbin Zhao |
IJCNN | 2 |
| 2015 | Event-Triggered H ∞ Control for Continuous-Time Nonlinear SystemabstractIn this paper, the H ∞ optimal control for a class of continuous-time nonlinear systems is investigated using event-triggered method. First, the H ∞ optimal control problem is formulated as a two-player zero-sum differential game. Then, an adaptive triggering condition is derived for the closed loop system with an event-triggered control policy and a time-triggered disturbance policy. For implementation purpose, the event-triggered concurrent learning algorithm is proposed, where only one critic neural network is required. Finally, an illustrated example is provided to demonstrate the effectiveness of the proposed scheme. Dongbin Zhao, Lingda Kong |
ISNN | 1 |
| 2015 | Computational Energy Management in Smart Grids
Stefano Squartini, Derong Liu 0001, Francesco Piazza, Dongbin Zhao, Haibo He |
Neurocomputing | 4 |
| 2015 | Convergence analysis and application of fuzzy-HDP for nonlinear discrete-time HJB systems
Yuanheng Zhu, Dongbin Zhao, Derong Liu 0001 |
Neurocomputing | 2 |
| 2015 | A data-based online reinforcement learning algorithm satisfying probably approximately correct principle
Yuanheng Zhu, Dongbin Zhao |
Neural Comput. Appl. | 2 |
| 2015 | Model-Free Optimal Control for Affine Nonlinear Systems With Convergence AnalysisabstractIn this paper, a self-learning control scheme is proposed for the infinite horizon optimal control of affine nonlinear systems based on the action dependent heuristic dynamic programming algorithm. The policy iteration technique is introduced to derive the optimal control policy with feasibility and convergence analysis. It shows that the “greedy” control action for each state is uniquely existent, the learned control policy after each policy iteration is admissible, and the optimal control policy is able to be obtained. Two three-layer perceptron neural networks are employed to implement the scheme. The critic network is trained by a novel rule to conform to the Bellman equation, and the action network is trained to yield a better control policy. Both training processes alternate until the optimal control policy is achieved. Two simulation examples are provided to validate the effectiveness of the approach. Note to Practitioners - The objective of designing optimal controllers without mathematical models is sought by control practitioners, whereas existing approaches usually derive optimal controllers by accessing the mathematical models or identified models. This paper proposes a new approach which derives optimal controllers by numerical iteration method without accessing any knowledge of the mathematical models. It gives evaluation for every state-action pair in the whole state-action space through the collected data of the underlying system, and then selects the action with the best evaluation for each state. What is required initial admissible control policy. Theorems show that optimal controllers can be acquired and simulation studies verify effectiveness. Further research will extend this approach to online self-learning optimal control approach, thus it can adapt the variation of underlying systems. Dongbin Zhao, Zhongpu Xia, Ding Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2015 | GrDHP: A General Utility Function Representation for Dual Heuristic Dynamic ProgrammingabstractA general utility function representation is proposed to provide the required derivable and adjustable utility function for the dual heuristic dynamic programming (DHP) design. Goal representation DHP (GrDHP) is presented with a goal network being on top of the traditional DHP design. This goal network provides a general mapping between the system states and the derivatives of the utility function. With this proposed architecture, we can obtain the required derivatives of the utility function directly from the goal network. In addition, instead of a fixed predefined utility function in literature, we conduct an online learning process for the goal network so that the derivatives of the utility function can be adaptively tuned over time. We provide the control performance of both the proposed GrDHP and the traditional DHP approaches under the same environment and parameter settings. The statistical simulation results and the snapshot of the system variables are presented to demonstrate the improved learning and controlling performance. We also apply both approaches to a power system example to further demonstrate the control capabilities of the GrDHP approach. Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001, Danil V. Prokhorov |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | MEC - A Near-Optimal Online Reinforcement Learning Algorithm for Continuous Deterministic SystemsabstractIn this paper, the first probably approximately correct (PAC) algorithm for continuous deterministic systems without relying on any system dynamics is proposed. It combines the state aggregation technique and the efficient exploration principle, and makes high utilization of online observed samples. We use a grid to partition the continuous state space into different cells to save samples. A near-upper Q operator is defined to produce a near-upper Q function using samples in each cell. The corresponding greedy policy effectively balances between exploration and exploitation. With the rigorous analysis, we prove that there is a polynomial time bound of executing nonoptimal actions in our algorithm. After finite steps, the final policy reaches near optimal in the framework of PAC. The implementation requires no knowledge of systems and has less computation complexity. Simulation studies confirm that it is a better performance than other similar PAC algorithms. Dongbin Zhao, Yuanheng Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | A data-based online reinforcement learning algorithm with high-efficient explorationabstractAn online reinforcement learning algorithm is proposed in this paper to directly utilizes online data efficiently for continuous deterministic systems without system parameters. The dependence on some specific approximation structures is crucial to limit the wide application of online reinforcement learning algorithms. We utilize the online data directly with the kd-tree technique to remove this limitation. Moreover, we design the algorithm in the Probably Approximately Correct principle. Two examples are simulated to verify its good performance. Yuanheng Zhu, Dongbin Zhao |
ADPRL | 2 |
| 2014 | A hierarchical classification algorithm for evaluating energy consumption behaviorsabstractResearches on office building energy consumption have been hot in these years, but few researchers consider the classification of office energy consumption performance which can evaluate user behaviors in order to offer a clear analysis of energy consumption and improve their energy saving consciousness. In this paper, we propose a novel hierarchical classification algorithm for evaluating energy consumption behaviors at a real energy management system, which combines fuzzy c-means clustering with GA (genetic algorithm)-based SVM (support vector machine) to fully utilize collected samples. The experiment results with real energy consumption data show that the proposed algorithm works well to distinguish the abnormal behaviors and classify energy consumption behaviors accurately on normal offices. Li Bu, Dongbin Zhao, Yu Liu 0005, Qiang Guan |
IJCNN | 2 |
| 2014 | A Kaiman filter-based actor-critic learning approachabstractKalman fiter is an efficient way to estimate the parameters of the value function in reinforcement learning. In order to solve Markov Decision Process (MDP) problems in both continuous state and action space, a new online reinforcement learning algorithm using Kalman filter technique, which is called Kalman filter-based actor-critic (KAC) learning is proposed in this paper. To implement the KAC algorithm, Cerebellar Model Articulation Controller (CMAC) neural networks are used to approximate the value function and the policy function respectively. Kalman filter is used to estimate the weights of the critic network. Two benchmark problems, namely the cart-pole balancing problem and the acrobot swing-up problem are provided to verify the effectiveness of the KAC approach. Experimental results demonstrate that the proposed KAC algorithm is more efficient than other similar algorithms. Bin Wang 0034, Dongbin Zhao |
IJCNN | 2 |
| 2014 | Event-triggered reinforcement learning approach for unknown nonlinear continuous-time systemabstractThis paper provides an adaptive event-triggered method using adaptive dynamic programming (ADP) for the nonlinear continuous-time system. Comparing to the traditional method with fixed sampling period, the event-triggered method samples the state only when an event is triggered and therefore the computational cost is reduced. We demonstrate the theoretical analysis on the stability of the event-triggered method, and integrate it with the ADP approach. The system dynamics are assumed unknown. The corresponding ADP algorithm is given and the neural network techniques are applied to implement this method. The simulation results verify the theoretical analysis and justify the efficiency of the proposed event-triggered technique using the ADP approach. Xiangnan Zhong, Zhen Ni, Haibo He, Xin Xu 0001, Dongbin Zhao |
IJCNN | 5 |
| 2014 | Dual Heuristic dynamic Programming for nonlinear discrete-time uncertain systems with state delay
Bin Wang 0034, Dongbin Zhao, Cesare Alippi, Derong Liu 0001 |
Neurocomputing | 2 |
| 2014 | Full-range adaptive cruise control based on supervised adaptive dynamic programming
Dongbin Zhao, Zhaohui Hu, Zhongpu Xia, Cesare Alippi, Yuanheng Zhu, Ding Wang 0001 |
Neurocomputing | 1 |
| 2014 | Detecting and Reacting to Changes in Sensing Units: The Active Classifier CaseabstractThe ability to detect concept drift, i.e., a structural change in the acquired datastream, and react accordingly is a major achievement for intelligent sensing units. This ability allows the unit, for actively tuning the application, to maintain high performance, changing online the operational strategy, detecting and isolating possible occurring faults to name a few tasks. In the paper, we consider a just-in-time strategy for adaptation; the sensing unit reacts exactly when needed, i.e., when concept drift is detected. Change detection tests (CDTs), designed to inspect structural changes in industrial and environmental data, are coupled here with adaptive k-nearest neighbor and support vector machine classifiers, and suitably retrained when the change is detected. Computational complexity and memory requirements of the CDT and the classifier, due to precious limited resources in embedded sensing, are taken into account in the application design. We show that a hierarchical CDT coupled with an adaptive resource-aware classifier is a suitable tool for processing and classifying sequential streams of data. Cesare Alippi, Derong Liu 0001, Dongbin Zhao, Li Bu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Real-time tracking on adaptive critic design with uniformly ultimately bounded conditionabstractIn this paper, we proposed a new nonlinear tracking controller based on heuristic dynamic programming (HDP) with the tracking filter. Specifically, we integrate a goal network into the regular HDP design and provide the critic network with detailed internal reward signal to help the value function approximation. The architecture is explicitly explained with the tracking filter, goal network, critic network and action network, respectively. We provide the stability analysis of our proposed controller with Lyapunov approach. It is shown that the filtered tracking errors and the weights estimation errors in neural networks are all uniformly ultimately bounded (UUB) under certain conditions. Finally, we compare our proposed approach with regular HDP approach in virtual reality (VR)/Simulink environment to justify the improved control performance. Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001 |
ADPRL | 4 |
| 2013 | Online Model-Free RLSPI Algorithm for Nonlinear Discrete-Time Non-affine Systems
Yuanheng Zhu, Dongbin Zhao |
ICONIP (2) | 2 |
| 2013 | A prior-free encode-decode change detection test to inspect datastreams for concept driftabstractOnline change detection in datastreams has attracted many researchers and is becoming a very hot topic whose relevance will further increase with research on Big Data. Concept drift is induced by changes in stationarity of the process generating the data caused by faults, time variance of the environment and inaccuracy of the change detection mechanism. Here, we propose a recurrent auto-associative Encode-Decode machine trained to reconstruct input data. The generated residual is then inspected for structural changes with a Change Detection Test (CDT). Although any CDT can be used, in the paper we focus the attention on the Hierarchical Intersection of Confidence Intervals change detection test for its capability of controlling false positives with a two layered test and an online version of the Lepage Change Point Model. Once concept drift is detected, the designed Encode-Decode machine, globally acting as an Encode-Decode CDT, is retrained on new data to detect subsequent changes. Cesare Alippi, Li Bu, Dongbin Zhao |
IJCNN | 3 |
| 2013 | Neural sliding-mode load frequency controller design of power systems
Dianwei Qian, Dongbin Zhao, Jianqiang Yi, Xiangjie Liu |
Neural Comput. Appl. | 2 |
| 2013 | A neural-network-based iterative GDHP approach for solving a class of nonlinear optimal control problems with control constraints
Ding Wang 0001, Derong Liu 0001, Dongbin Zhao, Yuzhu Huang, Dehua Zhang |
Neural Comput. Appl. | 3 |
| 2013 | Data-based control, optimization, modeling and applications
Dongbin Zhao, Yi Shen 0002, Xiaolin Hu 0001 |
Neural Comput. Appl. | 1 |
| 2013 | Special issue on intelligent control and information processing
Dongbin Zhao, Cesare Alippi, Derong Liu 0001, Huaguang Zhang |
Soft Comput. | 1 |
| 2013 | A supervised Actor-Critic approach for adaptive cruise control
Dongbin Zhao, Bin Wang 0034, Derong Liu 0001 |
Soft Comput. | 1 |
| 2012 | SVM-Based Just-in-Time Adaptive Classifiers
Cesare Alippi, Li Bu, Dongbin Zhao |
ICONIP (2) | 3 |
| 2012 | The Optimal Control of Discrete-Time Delay Nonlinear System with Dual Heuristic Dynamic Programming
Bin Wang 0034, Dongbin Zhao |
ICONIP (1) | 2 |
| 2012 | Reinforcement learning control based on multi-goal representation using hierarchical heuristic dynamic programmingabstractWe are interested in developing a multi-goal generator to provide detailed goal representations that help to improve the performance of the adaptive critic design (ACD). In this paper we propose a hierarchical structure of goal generator networks to cascade external reinforcement into more informative internal goal representations in the ACD. This is in contrast with previous designs in which the external reward signal is assigned to the critic network directly. The ACD control system performance is evaluated on the ball-and-beam balancing benchmark under noise-free and various noisy conditions. Simulation results in the form of a comparative study demonstrate effectiveness of our approach. Zhen Ni, Haibo He, Dongbin Zhao, Danil V. Prokhorov |
IJCNN | 3 |
| 2012 | Neural and fuzzy dynamic programming for under-actuated systemsabstractThis paper aims to integrate the fuzzy control with adaptive dynamic programming (ADP) scheme, to provide an optimized fuzzy control performance, together with faster convergence of ADP for the help of the fuzzy prior knowledge. ADP usually consists of two neural networks, one is the Actor as the controller, the other is the Critic as the performance evaluator. A fuzzy controller applied in many fields can be used instead as the Actor to speed up the learning convergence, because of its simplicity and prior information on fuzzy membership and rules. The parameters of the fuzzy rules are learned by ADP scheme to approach optimal control performance. The feature of fuzzy controller makes the system steady and robust to system states and uncertainties. Simulations on under-actuated systems, a cart-pole plant and a pendubot plant, are implemented. It is verified that the proposed scheme is capable of balancing under-actuated systems and has a wider control zone. Dongbin Zhao, Yuanheng Zhu, Haibo He |
IJCNN | 1 |
| 2012 | A Hierarchical Neural Network Architecture for Classification
Haibo He, Dongbin Zhao |
ISNN (1) | 5 |
| 2012 | Data-driven optimal algorithms and their applications to pattern recognition
Huaguang Zhang, Cesare Alippi, Dongbin Zhao |
Neurocomputing | 3 |
| 2012 | Self-teaching adaptive dynamic programming for Gomoku
Dongbin Zhao, Zhen Zhang 0009, Yujie Dai |
Neurocomputing | 1 |
| 2012 | Neural-Network-Based Optimal Control for a Class of Unknown Discrete-Time Nonlinear Systems Using Globalized Dual Heuristic ProgrammingabstractIn this paper, a neuro-optimal control scheme for a class of unknown discrete-time nonlinear systems with discount factor in the cost function is developed. The iterative adaptive dynamic programming algorithm using globalized dual heuristic programming technique is introduced to obtain the optimal controller with convergence analysis in terms of cost function and control law. In order to carry out the iterative algorithm, a neural network is constructed first to identify the unknown controlled system. Then, based on the learned system model, two other neural networks are employed as parametric structures to facilitate the implementation of the iterative algorithm, which aims at approximating at each iteration the cost function and its derivatives and the control law, respectively. Finally, a simulation example is provided to verify the effectiveness of the proposed optimal control approach. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2012 | Computational Intelligence in Urban Traffic Signal Control: A SurveyabstractUrban transportation system is a large complex nonlinear system. It consists of surface-way networks, freeway networks, and ramps with a mixed traffic flow of vehicles, bicycles, and pedestrians. Traffic congestions occur frequently, which affect daily life and pose all kinds of problems and challenges. Alleviation of traffic congestions not only improves travel safety and efficiencies but also reduces environmental pollution. Among all the solutions, traffic signal control (TSC) is commonly thought as the most important and effective method. TSC algorithms have evolved quickly, especially over the past several decades. As a result, several TSC systems have been widely implemented in the world, making TSC a major component of intelligent transportation system (ITS). In TSC and ITS, many new technologies can be adopted. Computational intelligence (CI), which mainly includes artificial neural networks, fuzzy systems, and evolutionary computation algorithms, brings flexibility, autonomy, and robustness to overcome nonlinearity and randomness of traffic systems. This paper surveys some commonly used CI paradigms, analyzes their applications in TSC systems for urban surface-way and freeway networks, and introduces current and potential issues of control and management of recurrent and nonrecurrent congestions in traffic networks, in order to provide valuable references for further research and development. Dongbin Zhao, Yujie Dai, Zhen Zhang 0009 |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2011 | Adaptive dynamic programming for optimal control of unknown nonlinear discrete-time systemsabstractAn intelligent optimal control scheme for unknown nonlinear discrete-time systems with discount factor in the cost function is proposed in this paper. An iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to obtain the optimal controller with convergence analysis. Three neural networks are used as parametric structures to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the unknown nonlinear system, respectively. Two simulation examples are provided to verify the effectiveness of the presented optimal control approach. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao |
ADPRL | 3 |
| 2011 | Supervised adaptive dynamic programming based adaptive cruise controlabstractThis paper proposes a supervised adaptive dynamic programming (SADP) algorithm for the full range Adaptive cruise control (ACC) system. The full range ACC system considers both the ACC situation in highway system and the stop and go (SG) situation in urban street way system. It can autonomously drive the host vehicle with desired speed and distance to the preceding vehicle in both situations. A traditional adaptive dynamic programming (ADP) algorithm is suited for this problem, but it suffers from the low learning efficiency. We propose the concept of inducing range to construct the supervisor and finally formulate the SADP algorithm, which greatly speeds up the learning efficiency. Several driving scenarios are designed and tested with the trained controller compared to traditional ones by simulation results, showing that trained SADP performs very well in all the scenarios, so that it provides an effective approach for the full range ACC problem. Dongbin Zhao, Zhaohui Hu |
ADPRL | 1 |
| 2011 | Neural-network-based optimal control for a class of nonlinear cdiscrete-time systems with control constraints using the citerative GDHP algorithmabstractIn this paper, a neural-network-based optimal control scheme for a class of nonlinear discrete-time systems with control constraints is proposed. The iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to design the optimal controller with convergence proof. Three neural networks are used to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the controlled nonlinear discrete-time system, respectively. A simulation study is carried out to demonstrate the effectiveness of the present approach in dealing with the nonlinear constrained optimal control problem. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao |
IJCNN | 3 |
| 2011 | DHP Method for Ramp Metering of Freeway TrafficabstractThis paper presents the design of dual heuristic programming (DHP) for the optimal coordination of ramp metering in freeway systems. Specifically, we implement the DHP method to solve both recurrent and nonrecurrent congestions with queuing consideration. A coordinated neural network controller is achieved by the DHP method with traffic models. Then, it is used for verifications with different traffic scenarios. Simulation studies performed on a hypothetical freeway indicate that the achieved neural controller maintains good control performance when compared with the classical ramp metering algorithm ALINEA. We emphasize that these neural controllers can be developed offline by using approximate traffic models. This offline mechanism avoids the risks of instability that incur during continual online training. We also discuss some real-time implementation issues. Dongbin Zhao, Xuerui Bai, Fei-Yue Wang 0001, Wensheng Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | Guest Editorial Data-Based Control, Modeling, and OptimizationabstractThe 21 papers in this special section focus on data-based control, modeling, and optimization. Tianyou Chai, Zhongsheng Hou, Frank L. Lewis, Amir Hussain 0001, Dongbin Zhao |
IEEE Trans. Neural Networks | 5 |
| 2010 | A comparative study of urban traffic signal control with reinforcement learning and Adaptive Dynamic ProgrammingabstractThis paper proposes a new algorithm that employs Adaptive Dynamic Programming(ADP) to solve the distributed control problem of urban traffic with an infinite horizon. Urban traffic congestions lead to a lot of time consumption and exhaust emissions. So alleviating congested situation will have a good impact on both economy and environment. The signal control at urban intersections is an effective and most important way to reduce the traffic jams and collisions. A lot of control theories including traditional mathematical ways and modern artificial intelligent ways have been exploited. ADP is an effective and amiable intelligent control method. We proposed an algorithm to adjust the signal time plan at urban traffic intersections based on ADP theory. Simulations are taken under a microscopic traffic simulation software, TSIS(Traffic Software Integrated System). Several criteria named MOEs(Measures of Effectiveness) are collected to compare with the widely used pre-timed control, actuated control, also with a machine learning method Q-learning control. Results show that ADP control method have a better adaptability to the various traffic simulating real traffic flows. Yujie Dai, Dongbin Zhao, Jianqiang Yi |
IJCNN | 2 |
| 2009 | ADHDP(λ) strategies based coordinated ramps metering with queuing considerationabstractRamp metering has been developed as a traffic management strategy to alleviate congestion on freeways. Most ramp metering control algorithms are concerned without queuing consideration, because its still a tough job to deal with the problems of coordinated multiple ramps metering with queuing consideration. In this paper, on the basis of our previous studies, we use action-dependent heuristic dynamic programming based on eligibility traces (ADHDP(lambda)) to solve local ramp metering and multiple ramps metering problems with queuing consideration. First, for the local ramp metering problem, we establish a comprehensive performance index which considers both traffic density and on-ramp queue length. Second, for the multiple ramps metering problem, based on ADHDP(lambda), the coordinated ramps metering and regulating queue lengths are achieved at the same time. Simulation studies on a hypothetical freeway are reported. It is shown that the proposed control scheme is efficient. Xuerui Bai, Dongbin Zhao, Jianqiang Yi |
ADPRL | 2 |
| 2009 | Control of the TORA system using SIRMs based type-2 fuzzy logicabstractThe translational oscillations with a rotational proof-mass actuator (TORA) is a well-known benchmark for examining the advantages and limitations of different nonlinear control design techniques. In this paper, a single-input-rule-modules (SIRMs) based type-2 fuzzy logic control scheme is proposed for this nonlinear multivariable system. And, genetic algorithms (GAs) are adopted to determine the parameters and to improve the performance of the SIRMs based type-2 fuzzy logic controller (SIRM-T2FLC). At last, simulations and comparisons are given to demonstrate the effectiveness, robustness and superiority of the proposed controller under three circumstances: normal case, the disturbance existing case, and the parameter varying case. From the design process and comparisons, it can be seen that: 1) this SIRMs based type-2 fuzzy control scheme can alleviate the difficulty to design conventional type-2 fuzzy logic controllers (T2FLCs) for this multivariable TORA system, 2) the SIRM-T2FLC is much easier to design and understand compared with conventional nonlinear control strategies for the TORA system, 3) better performance can be achieved. Chengdong Li, Jianqiang Yi, Dongbin Zhao |
FUZZ-IEEE | 3 |
| 2009 | Analysis and design of monotonic type-2 fuzzy inference systemsabstractThe prior knowledge-monotonicity property-is helpful for system analysis, modeling and design, especially when no specific physical structure knowledge about systems is available. This paper presents how to use interval type-2 fuzzy logic systems (IT2FLSs) to incorporate the monotonicity property into system design. First, we present sufficient conditions on the parameters of IT2FLSs to ensure the monotonicity between the inputs and outputs of IT2FLSs. Then, we transform the design of monotonic IT2FLSs to the least squares problem with linear-inequality constraints. At last, simulations are given to show the usefulness of the monotonicity property and the advantages of monotonic IT2FLSs under noisy circumstances. Chengdong Li, Jianqiang Yi, Dongbin Zhao |
FUZZ-IEEE | 3 |
| 2009 | Coordinated multiple ramps metering based on neuro-fuzzy adaptive dynamic programmingabstractThis paper aims to efficiently deal with the problems of multiple ramps metering. A new method which is called neuro-fuzzy adaptive dynamic programming with eligibility traces (NFADP(lambda)) is proposed. With the introduction of neuro-fuzzy and eligibility traces, the performance of ADP is greatly enhanced. First of all, the expert experience is introduced to ADP, therefore the convergence of ADP is greatly reinforced. Second, with the learning strategy revised, the training of action network is accelerated. In order to achieve multiple ramps metering control, special performance index function is established in NFADP(lambda). Extensive simulation on a hypothetical freeway are carried out with NFADP(lambda), compared to ALINEA as a stand-alone strategy. Simulation results indicate that NFADP(lambda) have good performances in both alleviating stochastic variations of the traffic demand and congestion situations. Xuerui Bai, Dongbin Zhao, Jianqiang Yi |
IJCNN | 2 |
| 2009 | Fuzzy logic based adjustment control of a cable-driven auto-leveling parallel robotabstractTo solve the level-adjusting and force-tuning problems of high accurate and costly payloads when loading and unloading, a cable-driven auto-leveling parallel robot is developed. A hierarchical fuzzy controller, which has the ability to deal with the rule explosion problem, is proposed in this paper. After a brief introduction of the architecture of the closed-loop control system for the cable-driven auto-leveling parallel robot, the construction of the hierarchical fuzzy controller is set up, in which the force offsets of the four cables and the angle deviations of the two diagonal inclinations are chosen as input variables, and the output variables are the position changes of the four linear motion units. The hierarchical fuzzy controller contains two layers - the low level layer which generates two outputs for leveling adjustment and force tuning, and the high level layer which is used to coordinate the two outputs from the low level layer. Experimental results have demonstrated that the hierarchical fuzzy controller can achieve the control objectives with high regulation accuracy and short adjusting time, and can be easily applied to practical systems. Jianqiang Yi, Chengdong Li, Dongbin Zhao |
IROS | 4 |
| 2009 | Trajectory Tracking Control of Omnidirectional Wheeled Mobile Manipulators: Robust Neural Network-Based Sliding Mode ApproachabstractThis paper addresses the robust trajectory tracking problem for a redundantly actuated omnidirectional mobile manipulator in the presence of uncertainties and disturbances. The development of control algorithms is based on sliding mode control (SMC) technique. First, a dynamic model is derived based on the practical omnidirectional mobile manipulator system. Then, a SMC scheme, based on the fixed large upper boundedness of the system dynamics (FLUBSMC), is designed to ensure trajectory tracking of the closed-loop system. However, the FLUBSMC scheme has inherent deficiency, which needs computing the upper boundedness of the system dynamics, and may cause high noise amplification and high control cost, particularly for the complex dynamics of the omnidirectional mobile manipulator system. Therefore, a robust neural network (NN)-based sliding mode controller (NNSMC), which uses an NN to identify the unstructured system dynamics directly, is further proposed to overcome the disadvantages of FLUBSMC and reduce the online computing burden of conventional NN adaptive controllers. Using learning ability of NN, NNSMC can coordinately control the omnidirectional mobile platform and the mounted manipulator with different dynamics effectively. The stability of the closed-loop system, the convergence of the NN weight-updating process, and the boundedness of the NN weight estimation errors are all strictly guaranteed. Then, in order to accelerate the NN learning efficiency, a partitioned NN structure is applied. Finally, simulation examples are given to demonstrate the proposed NNSMC approach can guarantee the whole system's convergence to the desired manifold with prescribed performance. Dongbin Zhao, Jianqiang Yi, Xiang-min Tan |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Control of a class of under-actuated systems with saturation using hierarchical sliding modeabstractThis paper presents a control scheme of a class of under-actuated systems with saturation using hierarchical sliding mode. This class with a single input and multiple outputs is made up of several subsystems. Based on this physical structure, the hierarchical structure of the sliding mode surfaces is developed as follows. The sliding surface of every subsystem is defined. Then the sliding surface of one subsystem is selected as the first layer sliding surface. The first layer sliding surface is used to construct the second layer sliding surface with the sliding surface of another subsystem. This process continues till all the subsystem sliding surfaces are included. The hierarchical sliding mode control law is deduced by using Lyapunov theorem. On account of saturation nonlinearity of the single input, asymptotic stability of the control system is proven by nonlinear small gain theorem. Parameter ranges of the subsystem sliding surfaces are also given. In practice, simulation and experimental results show the validity of this control method. Dianwei Qian, Jianqiang Yi, Dongbin Zhao |
ICRA | 3 |
| 2008 | Trajectory tracking control of omnidirecitonal wheeled mobile manipulators: Robust neural network based sliding mode approachabstractThis paper focuses on developing a robust neural network (NN) based sliding mode controller (NNSMC) to solve the trajectory tracking problem of a redundantly-actuated omnidirectional mobile manipulator. The SMC is designed to be robust to disturbances assuring the stability of the system. The NN is used to identify the unstructured uncertainty of system dynamics. The stability of the closed-loop system, the convergence of the NN weight-updating process, and the boundedness of the NN weight estimation errors are all strictly guaranteed. Through theories analysis, we know the controller is also capable of disturbance-rejection in the presence of time varying disturbances. Finally, simulation results demonstrate the proposed NNSMC approach can guarantee the whole system’s convergence to the desired manifold with prescribed performance. Dongbin Zhao, Jianqiang Yi, Xiang-min Tan, Zonghai Chen |
ICRA | 2 |
| 2008 | Ramp metering based on on-line ADHDP (lambda) controllerabstractIncreasing dependence on car-based travel has led to the daily occurrence of freeway congestions around the world. In order to improve the worse and worse traffic congestion situation and solve the problems brought with it, a new kind of effective, fast, and robust method should be presented. Ramp metering has been developed as a traffic management strategy to alleviate congestion on freeways. But, it doesnpsilat work well in uncertainty situations. In this paper, in order to solve the problems in uncertainty conditions, an on-line learning control method based on the fundamental principle of reinforcement learning is proposed. The method is ADP (adaptive dynamic programming) and in order to expedite the learning rate, the concept about eligibility traces is introduced here. Then eligibility trace and ADP is combined to present a new kind of traffic responsive control method. The new method is called action-dependent heuristic dynamic programming based on eligibility traces (ADHDP (lambda)). ADHDP (lambda) is an approximate optimal ramp metering method. Simulation studies on a hypothetical freeway indicate good control performance of the proposed real-time traffic controller. Xuerui Bai, Dongbin Zhao, Jianqiang Yi |
IJCNN | 2 |
| 2008 | Adaptive dynamic neuro-fuzzy system for traffic signal controlabstractThis paper aims at developing near optimal traffic signal control for multi-intersection in city. Fuzzy control is widely used in traffic signal control. For improving fuzzy controlpsilas adaptability in fluctuate states, a controller combined with neuro-fuzzy system and adaptive dynamic programming (ADP) is designed. This controller can be used for cooperative control of multi-intersection. The adaptive dynamic programming gives reinforcement for good neuro-fuzzy system behavior and punishment for poor behavior. The neuro-fuzzy system adjusts its parameters according to the reinforcement and punishment. Then, those actions leading to better results tend to be chosen preferentially in the future. Comparing with traditional ADP, this controller uses neuro-fuzzy system as the action network. The neuro-fuzzy system offers some existing knowledge and reduces the randomness of traditional ADP. In this paper, the objective of the controller is to minimize the average vehicular delay. The controller can be trained to adapt fluctuant traffic states by real-time traffic data, and achieves a near optimal control result in a long run. Simulation results show that the trained controller achieves shorter average vehicular delay than the controller with initial membership function. Dongbin Zhao, Jianqiang Yi |
IJCNN | 2 |
| 2007 | Robust Control Using Sliding Mode for a Class of Under-Actuated Systems With Mismatched UncertaintiesabstractBased on the methodology of sliding mode, this paper presents a robust controller for a class of under-actuated systems with mismatched uncertainties. Such a system consists of a nominal system and the mismatched uncertainties. The structural characteristic of the nominal system is that it is made up of several subsystems. Based on this characteristic, the hierarchical structure of the sliding mode surfaces is designed for the nominal system as follows. Firstly, the nominal system is divided into several subsystems and the sliding mode surface of every subsystem is defined. Secondly, the sliding mode surface of one subsystem is selected as the first layer sliding mode surface. The first layer sliding mode surface is then to construct the second layer sliding mode surface with the sliding mode surface of another subsystem. This process continues till the sliding mode surfaces of all the subsystems are included. For dealing with the mismatched uncertainties, a lumped sliding mode compensator is designed at the last layer sliding mode surface. The asymptotic stability of every layer sliding mode surface and the sliding mode surface of each subsystem is proven theoretically by Barbalat's lemma. Simulation results show the validity of this robust control method through stabilization control of a double inverted pendulums system with mismatched uncertainties. Dianwei Qian, Jianqiang Yi, Dongbin Zhao |
ICRA | 3 |
| 2007 | Robust adaptive tracking control of omnidirecitonal wheeled mobile manipulatorsabstractThis paper addresses the trajectory tracking problem for an redundantly-actuated omnidirectional mobile manipulator system with uncertainties and disturbances. The proposed algorithm is robust adaptive control strategy and the parameter estimates are tuned online. First, for designing controller, the conservative upper-bounded function of dynamic model of omnidirectional mobile manipulator system is derived based on the dynamic structure properties. Then, a robust adaptive control scheme is presented to ensure trajectory tracking effect of this closed-loop system. The asymptotical stability is verified a Lyapunov method. Finally, simulation examples are given to demonstrate the proposed approach can guarantee the whole system converge to the desired manifold with prescribed performance. Dongbin Zhao, Jianqiang Yi, Xiang-min Tan |
IROS | 2 |
| 2007 | Motion regulation of redundantly actuated omni-directional Wheeled Mobile Robots with internal force controlabstractBecause of the complexity of the mechanisms of redundantly actuated omni-directional Wheeled Mobile Robots (WMR), its motion regulation is a challenging problem, especially for consideration of the interaction force between the redundantly actuated wheels. The interaction force can be decomposed into motion-induced force and internal force, which are orthogonal between each other. Only the motion-induced force contributes to the motion of the robot, while the internal force abrades the wheels components, and causes the reduction of their life span. So the internal force should be eliminated or minimized. In this paper, kinematic model and dynamic model of redundantly actuated omni-directional WMR considering the interaction force is first established. A proportional differential plus motion regulator is presented. An integral feedback internal force controller is applied to minimize the internal force. Simulation results verify the effectiveness of the proposed control scheme. The robot is regulated successfully, and the internal force is reduced efficiently. Dongbin Zhao, Jianqiang Yi, Xuyue Deng |
IROS | 1 |
| 2007 | Approximate Dynamic Programming for Ship Course Control
Xuerui Bai, Jianqiang Yi, Dongbin Zhao |
ISNN (1) | 3 |
| 2007 | A Comparison of Four Data Mining Models: Bayes, Neural Network, SVM and Decision Trees in Identifying Syndromes in Coronary Heart Disease
Yanwei Xing, Guangcheng Xi, Jianqiang Yi, Dongbin Zhao, Jie Wang 0107 |
ISNN (1) | 6 |
| 2007 | Application of ADP to Intersection Signal Control
Dongbin Zhao, Jianqiang Yi |
ISNN (1) | 2 |
| 2007 | Multiple Approximate Dynamic Programming Controllers for Congestion Control
Yanping Xiang, Jianqiang Yi, Dongbin Zhao |
ISNN (1) | 3 |
| 2006 | A New Fuzzy Autopilot for Way-point Tracking Control of ShipsabstractA new fuzzy control design for way-point tracking control problem of ship autopilot is proposed. To effectively control ships in a designed trajectory is always an important task for ship manipulators. The paper gives the design method of a kind of fuzzy autopilot for way-point tracking control system. The fuzzy control rules are constructed based on human operator's manipulating experience. This control design approach greatly simplifies the control design process and the control algorithm, and is easily applied to practical ship tracking system. Simulation results show that the proposed fuzzy autopilot has desired performance with high accuracy of course keeping, short time of rudder actions and less frequency of changing rudder direction. Jin Cheng 0004, Jianqiang Yi, Dongbin Zhao |
FUZZ-IEEE | 3 |
| 2006 | Exponential Convergence Flow Control Model for Congestion Control
Jianqiang Yi, Dongbin Zhao, John T. Wen |
ICIC (1) | 3 |
| 2006 | Time Based Congestion Control (TBCC) for High Speed High Delay Networks
Yanping Xiang, Jianqiang Yi, Dongbin Zhao, John T. Wen |
ICIC (1) | 3 |
| 2006 | Hierarchical Sliding Mode Control for Series Double Inverted Pendulums SystemabstractThis paper proposes a hierarchical sliding mode controller for series double inverted pendulums system. This provides a simple method to control a class of under-actuated systems with three subsystems by sliding mode control. Firstly, the given system is divided into three subsystems according to its structure characteristic. Then, the 1st-level sliding mode surface is defined for every subsystem and the 2nd-level sliding mode surface is constituted by them. Based on the two levels structure, the equivalent control of each subsystem is deduced and the total control law is derived by the Lyapunov stability theorem. The asymptotical stability of the entire sliding mode surfaces is proved theoretically. Finally, simulation results show the validity of this control strategy. And the influence of the controller parameter changes for the performances is also discussed Dianwei Qian, Jianqiang Yi, Dongbin Zhao, Yinxing Hao |
IROS | 3 |
| 2006 | A Particle Swarm Optimized Fuzzy Neural Network Control for Acrobot
Dongbin Zhao, Jianqiang Yi |
ISNN (2) | 1 |
| 2005 | Pose Estimation and Structure Recovery from Point PairsabstractThis paper presents a new feature point pairs based technique for object pose estimation and structure recovery from a single view. It first estimates rotational matrix independently, then computes translation vector and recovers the 3D structure of the object directly. Linear and nonlinear strategies are presented to estimate the rotational matrix. One is for small rotational motion and the other is used to estimate large rotational parameters. When the nonlinear technique is applied, its initial guesses are given automatically by the proposed linear estimation method. On the other hand, the presented structure recovery method is not sensitive to the rotational matrix estimation results. The proposed method is applicable to three, four or more feature points and has no constraints, such as collinear or coplanar, on their relative positions. As the number of feature points increases, the estimation results are improved while the computation cost is almost unchanged. Many experiments are performed on synthetic data and real images to demonstrate the presented technique. Zhiguang Zhong, Jianqiang Yi, Dongbin Zhao |
ICRA | 3 |
| 2005 | Tracking control of mobile manipulator with dynamical uncertaintiesabstractTracking control problem of mobile manipulators with dynamical uncertainties is addressed in this paper. The controller is designed based on model of mobile manipulators consisting of two cascaded subsystems: a chained-like kinematical model without uncertainties and a dynamical model with uncertainties. The proposed control law can ensure that full states of closed-loop system can track given trajectories in presence of dynamical uncertainties. A globally asymptotic stability is obtained in Lyapunov sense. Simulation studies show feasibility and effectiveness of the proposed approach. Zuoshi Song, Dongbin Zhao, Jianqiang Yi, Xinchun Li |
IROS | 2 |
| 2005 | Double layer sliding mode control for second-order underactuated mechanical systemsabstractA new stable sliding mode control method for a class of underactuated mechanical systems is proposed in this paper. The controller has the double-layer structure. Firstly, the system states are divided into several different subsystems. For each of these subsystems, a first-layer sliding plane is constructed. From these first-layer sliding planes, then we further construct a second-layer sliding plane. By analyzing the features of the mathematical model of the underactuated mechanical systems, we derive the sliding-mode control law and indicate the ranges of the controller parameters. Using Lyapunov law, the paper proves the stability of all the sliding planes theoretically. The simulation results show the validity of this method for this class of underactuated mechanical systems. Wei Wang 0115, Jianqiang Yi, Dongbin Zhao |
IROS | 3 |
| 2005 | Cascade sliding-mode controller for large-scale underactuated systemsabstractOn the basis of sliding mode control, a new cascade sliding-mode controller (CSMC) for a class of large-scale underactuated systems is proposed. The large-scale underactuated systems include several subsystems. Firstly, two states are chosen to construct the first-layer sliding surface. Secondly, the first-layer sliding surface and one of the left states are used to construct the second-layer sliding surface. This process continues till the last-layer sliding surface is obtained. By theoretical analysis, the cascade sliding-mode controller is proved to be globally stable in the sense that all signals involved are bounded. The simulation results show the validity of this method. Jianqiang Yi, Wei Wang 0115, Dongbin Zhao |
IROS | 3 |
| 2005 | A Reinforcement Learning Based Radial-Bassis Function Network Control System
Jianqiang Yi, Dongbin Zhao, Guangcheng Xi |
ISNN (1) | 3 |
| 2005 | Adaptive Inverse Control System Based on Least Squares Support Vector Machines
Jianqiang Yi, Dongbin Zhao |
ISNN (3) | 3 |
| 2005 | A computed torque controller for uncertain robotic manipulator systems: Fuzzy approach
Zuoshi Song, Jianqiang Yi, Dongbin Zhao, Xinchun Li |
Fuzzy Sets Syst. | 3 |
| 2005 | Effective pose estimation from point pairs
Zhiguang Zhong, Jianqiang Yi, Dongbin Zhao, Yiping Hong |
Image Vis. Comput. | 3 |
| 2004 | A robust doorplate recognition systemabstractIn real applications, captured doorplate images often contain unexpected noise from irregular illumination conditions, various imaging angles, different imaging distances, etc. In this paper, a robust doorplate recognition system is presented. First, an efficient method based on Sobel operator, region splitting and merging is applied to extract doorplate. Then, according to the doorplate shape and the character position in the doorplate, the character region is determined. If the candidate of some doorplate characters extracted from the character region cannot be determined at the segmentation step, a speculation based on known knowledge about geometrical relations between each doorplate character is executed. The threshold for character extraction from candidates is adjusted when the corresponding character is rejected after classification. The data set and the images in experiments all come from practical applications. Experimental results indicate that the proposed doorplate recognition system effectively improves the recognition result. Yiping Hong, Jianqiang Yi, Dongbin Zhao, Xinzheng Li |
ICARCV | 3 |
| 2004 | Smooth time-varying regulation of nonholonomic chained systemsabstractThe problem of regulating a nonholonomic system in chained form to an equilibrium state is addressed and solved by a smooth time-varying control law. The proposed scheme based on Lyapunov analysis can guarantee asymptotic stabilization and exponential stabilization, respectively, by simply changing a term in the controller expression. Therefore, the controller can be used for theoretical analysis and practical application. The additional novelty of the proposed approach is that the parameters in our controller have clear physical meanings and are easily implemented and tuned. Simulation results based on a numerical example and a tricycle-type mobile robot are presented to demonstrate the effectiveness of the developed method. Zuoshi Song, Jianqiang Yi, Dongbin Zhao, Xinchun Li |
ICARCV | 3 |
| 2004 | Motion vision for mobile robot localizationabstractThis paper presents a localization method using motion vision. The proposed method locates a mobile robot relative to the object to which the robot moves. Two points are selected from the object as feature points. Consecutive two images containing the feature points are taken before and after the robot moves. Then, the poses of the robot can be determined according to the image coordinates of the feature points. A search algorithm is also presented to enhance the localization precision. It can find more correct image coordinates of the feature points based on the real detected image coordinates. Many experiments are performed on real images to justify this search algorithm. The results show that it is effective. Zhiguang Zhong, Jianqiang Yi, Dongbin Zhao, Yiping Hong, Xinzheng Li |
ICARCV | 3 |
| 2004 | Passive Adaptive Grasp Multi-fingered Humanoid Robot Hand with High Under-actuated FunctionabstractThis paper proposed a design idea of a novel under-actuated finger mechanism, and designed the finger mechanism. The finger has no actuator in itself, is only driven by the other finger joints and object grasped. The finger is similar to a human finger and can be easily arranged in series to realize a finger with super under-actuation and high integration. It can be mounted in humanoid robot hand to make the hand obtain more DOFs with less actuators, and good grasping function of shape adaptation, decrease the requirement of control system. This paper analyzed the relationship between the grasping force of the finger and its design parameters, proposed the design principle of structure optimization of the finger. Based on the finger, a multi-fingered humanoid robot hand: TH-2 Hand has been designed. TH-2 Hand has many excellent features: high personification, super under-actuation and be very compact, easy to real-time control, small volume, light in weight, strong grasping function, etc. Wenzeng Zhang, Qiang Chen 0009, Zhenguo Sun, Dongbin Zhao |
ICRA | 4 |
| 2003 | Under-actuated passive adaptive grasp humanoid robot hand with control of grasping forceabstractConventional dexterous hands have too many DOFs, their driver systems are too big to be installed in a humanoid robot arm, and their controls are tool complex. This paper develops an under-actuated passive adaptive grasp humanoid robot hand named TH-1 hand with control of grasping force. With humanoid appearance and size, TH-1 hand is light, fewer DOFs, and can be easily controlled. Its motors and driver circuit boards are embedded in itself. These features make it fit to be installed in a humanoid robot arm. In addition, for stably grasping operation, a mechanical finger with control of grasping force is designed and applied in TH-1 hand's index. To get more DOFs with fewer drivers, a novel under-actuated passive adaptive grasp mechanical finger is design and applied in TH-1 hand's thumb. Wenzeng Zhang, Qiang Chen 0009, Zhenguo Sun, Dongbin Zhao |
ICRA | 4 |