Bo Zhao 0015

dblp:94/4810-15 · DBLP profile ↗
← Back
83ranked-venue papers
16as first author
60since 2021 · last 2026
0000-0002-7684-7342ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 12 first-author · 46 since 2021Human-computer interaction and ubiquitous computing · 14 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Compensating Distribution Drifts in Continual Learning with Pre-trained Vision Transformers
abstract
Recent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class features, can be an effective strategy for class-incremental learning (CIL). However, this approach is susceptible to distribution drift, caused by the sequential optimization of shared backbone parameters. This results in a mismatch between the distributions of the previously learned classes and that of the updated model, ultimately degrading the effectiveness of classifier performance over time. To address this issue, we introduce a latent space transition operator and propose Sequential Learning with Drift Compensation (SLDC). SLDC aims to align feature distributions across tasks to mitigate the impact of drift. First, we present a linear variant of SLDC, which learns a linear operator by solving a regularized least-squares problem that maps features before and after fine-tuning. Next, we extend this with a weakly nonlinear SLDC variant, which assumes that the ideal transition operator lies between purely linear and fully nonlinear transformations. This is implemented using learnable, weakly nonlinear mappings that balance flexibility and generalization. To further reduce representation drift, we apply knowledge distillation (KD) in both algorithmic variants. Extensive experiments on standard CIL benchmarks demonstrate that SLDC significantly improves the performance of SeqFT. Notably, by combining KD to address representation drift with SLDC to compensate distribution drift, SeqFT achieves performance comparable to joint training across all evaluated datasets.
Xuan Rao, Simian Xu, Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Cesare Alippi
AAAI4
2026 Computation-aware Transformer-based encoding for efficient latent spatial neural architecture search
Jiamin Xiao, Bo Zhao 0015, Derong Liu 0001, Yonghua Wang 0001, Jiacai Huang
Neurocomputing2
2026 Control barrier function-based self-learning robust control of safety-critical nonlinear systems
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001
Neurocomputing2
2026 Event-triggered decentralized adaptive critic learning control for interconnected systems with nonlinear inequality state constraints
Wenqian Du 0006, Mingduo Lin, Guoling Yuan, Bo Zhao 0015
Neural Networks4
2026 Dynamic Self-Triggered Prescribed-Time Optimal Tracking Control for Helicopters via Adaptive Dynamic Programming
Shunchao Zhang, Bo Zhao 0015
IEEE Trans Autom. Sci. Eng.2
2026 Active Dataset Distillation via Dual-Space Informative Matching
abstract
Dataset distillation improves neural network training efficiency by compressing large real datasets into compact synthetic datasets. Existing methods typically optimize matching objectives, such as aligning gradients, features, and trajectories between the synthetic and original datasets to ensure the distilled data retains essential properties for model training. However, many of these approaches rely on predefined distillation pools to streamline the process or treat all real data points equally, overlooking the dynamic nature of the synthetic dataset's training requirements during optimization. To address these limitations, we propose Active Dataset Distillation via Dual-Space Informative Matching (ACDD), an active learning-based algorithm that dynamically selects the most informative real data subset to align with the synthetic dataset's evolving needs. By adaptively refining the distillation pool, ACDD enhances training efficiency and generalization while ensuring the synthetic dataset effectively captures the original data's key characteristics. ACDD operates through two interconnected loops: the dual-space active loop (DAL) and the distillation loop. DAL plays a key role by dynamically selecting samples that balance diversity and uncertainty, adding them to the target distillation pool to meet the evolving informational needs of the current distillation loop. As a result, ACDD enables the synthetic dataset to achieve superior performance compared to SOTA methods across multiple benchmarks, including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset. Moreover, ACDD reduces the required real dataset to just 20%-40% of the original, demonstrating its efficiency and effectiveness in data distillation.
Ding Qi, Jian Li 0062, Shuguang Dou, Junyao Gao 0002, Yabiao Wang, Bo Zhao 0015, Cairong Zhao
IEEE Trans. Image Process.6
2026 Fixed-Time Formation Hunting Control of Multi-Marine Surface Vehicle System Based on a Novel Deep Reinforcement Learning
abstract
In this article, a fixed-time deep reinforcement learning (DRL) formation hunting control problem is investigated for a multi-marine surface vehicle (MSV) system. First, considering the lack of dynamic adaptability caused by the conventional deep neural network (DNN) framework, an online adaptive DNN method is proposed for the high-dimensional multi-MSV system. Second, a novel DRL framework is developed for designing fixed-time formation hunting controllers, which integrates the online adaptive DNNs method with the actor–critic-based reinforcement learning (RL) algorithm. Finally, a nonsmooth fixed-time stability analysis is established for the nonsmooth closed-loop system induced by the DRL-based structure, which rigorously demonstrates that all signals converge within a fixed-time interval independent of initial states. The simulation example demonstrates the practical viability of the presented scheme.
Weiwei Bai, Yuanhao Wang 0015, Bo Zhao 0015, Dewang Chen, Andrea D'Ariano
IEEE Trans. Syst. Man Cybern. Syst.3
2026 Fractional-Order Online Policy Iteration-Based Approximate Optimal Control for Fractional-Order Nonlinear Systems
abstract
In this article, a fractional-order online policy iteration (FOOPI) algorithm-based approximate optimal control scheme is developed for fractional-order nonlinear systems (FONSs). Using the fractional Taylor series and the property of Hadamard product, the fractional Hamilton–Jacobi–Bellman (FHJB) equation corresponding to FONSs is formulated. Motivated by the procedure of solving the optimal control problem of integer-order nonlinear systems (IONSs), the design of the fractional-order (FO) controller for FONSs is transformed into solving the FHJB equation utilizing a FOOPI algorithm. By constructing a critic neural network (NN) and using the fractional calculus theory, the FO derivatives of the value function are established, followed by the approximate fractional Hamiltonian, which is obtained via the FOOPI algorithm. Hereafter, the FO control policy is obtained approximately to guarantee the stability of closed-loop FONSs via the FO Lyapunov’s direct method. Simulation results guarantee the effectiveness of the proposed FOOPI-based approximate optimal control scheme.
Guoling Yuan, Bo Zhao 0015
IEEE Trans. Syst. Man Cybern. Syst.3
2026 Reinforcement Learning-Based Dynamic Event-Triggered Control for Unknown Nonaffine Systems Using Dynamic Feedback
abstract
In this article, a dynamic feedback (DF)-based dynamic event-triggered (DET) control method for unknown nonaffine systems (UNSs) is developed by using reinforcement learning (RL). Through introducing a DF signal as a virtual control input, the UNS is augmented into a partially unknown affine system (PUAS). Subsequently, by designing a novel cost function that reflects the system states, and the actual and virtual control inputs, the DET optimal control (OC) problem of UNS is transformed into a DET OC problem of PUAS. To relax the requirement of PUAS dynamics, a neural network (NN)-based observer is established by using the measured system data. Moreover, a novel DET condition is established based on the static event-triggered (SET) rule, and the relationship of the triggering interval between SET and DET is revealed. In order to solve the DET Hamilton–Jacobi–Bellman equation (HJBE), a critic NN is constructed with the concurrent learning method to release the persistence of excitation (PE) condition. Furthermore, according to Lyapunov’s direct method, the stability of the closed-loop system is guaranteed under the developed DF-based DET control strategy. Finally, simulation results of two examples demonstrate the effectiveness of the present DF-based DET method.
Jinquan Lin, Bo Zhao 0015, Yonghua Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2025 MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval
abstract
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70\times more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our code, synthesized dataset, and pre-trained models are publicly available at https://github.com/VectorSpaceLab/MegaPairs.
Junjie Zhou 0001, Yongping Xiong, Zheng Liu 0011, Shitao Xiao, Yueze Wang, Bo Zhao 0015, Chen Zhang 0013, Defu Lian
ACL (1)7
2025 MLVU: Benchmarking Multi-task Long Video Understanding
abstract
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU.
Junjie Zhou 0001, Bo Zhao 0015, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Yongping Xiong, Tiejun Huang 0001, Zheng Liu 0011
CVPR3
2025 Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
abstract
Multimodal Large Language Models (MLLMs) have displayed remarkable performance in multi-modal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a benchmark with 12 categories and design evaluation metrics that assess the degree of error in MLLM responses even when the visual content is seemingly understood. Based on this benchmark, we test 15 leading MLLMs and analyze the distribution of attention maps and logits of some MLLMs. Our investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that "directly" relate to the visual content, leading to a bias in MLLMs’ responses to other indirect questions, and 2) MLLMs’ attention to visual tokens is notably lower than to system and question tokens. We further observe that attention scores between questions and visual tokens as well as the model’s confidence in the answers are lower in response to misleading questions than to straightforward ones. To address the first challenge, we introduce a paired positive and negative data construction pipeline to diversify the dataset. For the second challenge, we propose to enhance the model’s focus on visual content during decoding by refining the text and visual prompt. For the text prompt, we propose a content guided refinement strategy that performs preliminary visual content analysis to generate structured information before answering the question. Additionally, we employ a visual attention refinement strategy that highlights question-relevant visual tokens to increase the model’s attention to visual content that aligns with the question. Extensive experiments demonstrate that these challenges can be significantly mitigated with our proposed dataset and techniques.
Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu, Muyang He, Jian Li 0062, Zheng Liu 0011, Harry Yang, Ser-Nam Lim, Bo Zhao 0015
CVPR11
2025 Towards Universal Dataset Distillation via Task-Driven Diffusion
abstract
Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, most recent research has primarily focused on image classification tasks, with limited exploration in detection and segmentation. Two key challenges remain: (i) Task Optimization Heterogeneity, where existing methods focus on class-level information but fail to address the diverse needs of detection and segmentation, and (ii) Inflexible Image Generation, where current generation methods rely on global updates for single-class targets and lack localized optimization for specific object regions. To address these challenges, we propose UniDD, a universal dataset distillation framework built on a task-driven diffusion model for diverse DD tasks, as shown in Fig. 1. Our approach operates in two stages: Universal Task Knowledge Mining, which captures task-relevant information through task-specific proxy model training, and Universal Task-Driven Diffusion, where these proxies guide the diffusion process to generate task-specific synthetic images. Extensive experiments across ImageNet-1K, Pascal VOC, and MS COCO demonstrate that UniDD consistently outperforms state-of-the-art methods. In particular, on ImageNet-1K with IPC-10, UniDD surpasses previous diffusion-based methods by 6.1%, while also reducing deployment costs.
Ding Qi, Jian Li 0062, Junyao Gao 0002, Shuguang Dou, Ying Tai, Jianlong Hu, Bo Zhao 0015, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
CVPR7
2025 Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
abstract
Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens, their strategies encounter severe bottleneck, restricting MLLMs’ ability to perceive fine-grained visual details. In this work, we propose Video-XL, a novel approach that leverages MLLMs’ inherent key-value (KV) sparsification capacity to condense the visual input. Specifically, we introduce a new special token, the Visual Summarization Token (VST), for each interval of the video, which summarizes the visual information within the interval as its associated KV. The VST module is trained by instruction fine-tuning, where two optimizing strategies are offered. 1. Curriculum learning, where VST learns to make small (easy) and large compression (hard) progressively. 2. Composite data curation, which integrates single-image, multi-image, and synthetic data to overcome the scarcity of long-video instruction data. The compression quality is further improved by dynamic compression, which customizes compression granularity based on the information density of different video intervals. Video-XL’s effectiveness is verified from three aspects. First, it achieves a superior long-video understanding capability, outperforming state-of-the-art models of comparable sizes across multiple popular benchmarks. Second, it effectively preserves video information, with minimal compression loss even at 16 × compression ratio. Third, it realizes outstanding cost-effectiveness, enabling high-quality processing of thousands of frames on a single A100 GPU.
Zheng Liu 0011, Peitian Zhang, Minghao Qin, Junjie Zhou 0001, Zhengyang Liang, Tiejun Huang 0001, Bo Zhao 0015
CVPR8
2025 Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation
abstract
Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision Language Model (VLM) based navigation agent taking the lead in explicitly showcasing the capabilities of spatial intelligence in the real world. Given natural language instructions, NaVid-4D requires only egocentric RGB-D video streams as observations to perform spatial understanding and reasoning for generating precise instruction-following robotic actions. NaVid-4D learns navigation policies using the data from simulation environments and is endowed with precise spatial understanding and reasoning capabilities using web data. Without the need to pre-train an RGB-D foundation model, we propose a method capable of directly injecting the depth features into the visual encoder of a VLM. We further compare the use of factually captured depth information with the monocularly estimated one and find NaVid-4D works well with both while using estimated depth offers greater gener-alization capability and better mitigates the sim-to-real gap. Extensive experiments demonstrate that NaVid-4D achieves state-of-the-art performance in simulation environment and makes impressive VLN performance with spatial intelligence happen in the real world.
Weikang Wan, Xiqian Yu, Jiazhao Zhang, Bo Zhao 0015, Zhibo Chen 0001, Zhongyuan Wang 0006, Zhizheng Zhang 0004, He Wang 0010
ICRA6
2025 Distributed Fault-Tolerant Consensus Control Based on Zero-Sum Differential Games for Nonlinear Multi-agent Systems
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001
ISNN3
2025 MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
abstract
Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LVMR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1,200 seconds in duration, and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, and object-level, covering common tasks like action recognition, object localization, causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker to facilitate future research in this area.
Huaying Yuan, Jian Ni, Zheng Liu 0011, Yueze Wang, Junjie Zhou 0001, Zhengyang Liang, Bo Zhao 0015, Zhao Cao, Ji-Rong Wen, Zhicheng Dou
NeurIPS7
2025 Event-triggered neuro-optimal fault tolerant control for uncertain macro-micro composite stage system with actuator faults
Shunchao Zhang, Bo Zhao 0015, Yongwei Zhang 0002
Eng. Appl. Artif. Intell.2
2025 Image Captions are Natural Prompts for Training Data Synthesis
Shiye Lei, Hao Chen 0011, Sen Zhang 0006, Bo Zhao 0015, Dacheng Tao
Int. J. Comput. Vis.4
2025 Improved cost function-based fault tolerant control for nonlinear systems with simultaneous faults
Chujian Zeng, Bo Zhao 0015, Derong Liu 0001
Neurocomputing2
2025 Event-triggered robust hierarchical control for uncertain multiplayer Stackelberg games via adaptive dynamic programming
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Marios M. Polycarpou, Shiguo Peng, Shunchao Zhang
Neurocomputing2
2025 Neuro-dynamic programming-based event-triggered fault tolerant control for nonlinear systems with multiple faults
Haowei Lin, Weifeng Su, Runlin Huang, Bo Zhao 0015, Wentao Fan 0001
Neural Networks4
2025 On robust learning of memory attractors with noisy deep associative memory networks
Xuan Rao, Bo Zhao 0015, Derong Liu 0001
Neural Networks2
2025 Event-triggered control for input-constrained nonzero-sum games through particle swarm optimized neural networks
Qiuye Wu, Bo Zhao 0015, Derong Liu 0001
Neural Networks2
2025 Optimal Learning Output Tracking Control: A Model-Free Policy Optimization Method With Convergence Analysis
abstract
Optimal learning output tracking control (OLOTC) in a model-free manner has received increasing attention in both the intelligent control and the reinforcement learning (RL) communities. Although the model-free tracking control has been achieved via off-policy learning and Q-learning, another popular RL idea of direct policy learning, with its easy-to-implement feature, is still rarely considered. To fill this gap, this article aims to develop a novel model-free policy optimization (PO) algorithm to achieve the OLOTC for unknown linear discrete-time (DT) systems. The iterative control policy is parameterized to directly improve the discounted value function of the augmented system via the gradient-based method. To implement this algorithm in a model-free manner, a model-free two-point policy gradient (PG) algorithm is designed to approximate the gradient of discounted value function by virtue of the sampled states and the reference trajectories. The global convergence of model-free PO algorithm to the optimal value function is demonstrated with the sufficient quantity of samples and proper conditions. Finally, numerical simulation results are provided to validate the effectiveness of the present method.
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 FX-DARTS: Designing Topology-Unconstrained Architectures With Differentiable Architecture Search and Entropy-BasedSuper-Network Shrinking
abstract
Strong priors are imposed on the search space of differentiable architecture search (DARTS), such that cells of the same type share the same topological structure and each intermediate node retains two operators from distinct nodes. While these priors reduce optimization difficulties and improve the applicability of searched architectures, they hinder the subsequent development of automated machine learning (auto-ML) and prevent the optimization algorithm from exploring more powerful neural networks through improved architectural flexibility. This article aims to reduce these prior constraints by eliminating restrictions on cell topology and modifying the discretization mechanism for super-networks. Specifically, the flexible DARTS (FX-DARTS) method, which leverages an entropy-based super-network shrinking (ESS) framework, is presented to address the challenges arising from the elimination of prior constraints. Notably, FX-DARTS enables the derivation of neural architectures without strict prior rules while maintaining the stability in the enlarged search space. Experimental results on image classification benchmarks demonstrate that FX-DARTS is capable of exploring a set of neural architectures with competitive trade-offs between performance and computational complexity within a single search procedure.
Xuan Rao, Bo Zhao 0015, Derong Liu 0001, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.2
2025 Self-Triggered Approximate Optimal Neuro-Control for Nonlinear Systems Through Adaptive Dynamic Programming
abstract
In this article, a novel self-triggered approximate optimal neuro-control scheme is presented for nonlinear systems by utilizing adaptive dynamic programming (ADP). According to the Bellman principle of optimality, the cost function of the general nonlinear system is approximated by building a critic neural network with a nested updating weight vector. Thus, the Hamilton-Jacobi-Bellman equation is solved to indirectly obtain the approximate optimal neuro-control input. In order to reduce the computation, the communication bandwidth, and the energy consumption, an appropriate self-triggering condition is designed as an alternative way to predict the updating time instants of the approximate optimal neuro-control policy. On the basis of Lyapunov's direct method, the stability of the closed-loop nonlinear system is analyzed and guaranteed to be uniformly ultimately bounded. Simulation results of two practical systems illustrate the present ADP-based self-triggered approximate optimal neuro-control scheme to be reasonable and effective.
Bo Zhao 0015, Shunchao Zhang, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Distributed Optimal Containment Control of Wheeled Mobile Robots via Adaptive Dynamic Programming
abstract
In this article, the distributed optimal containment (DOC) control of wheeled mobile robots (WMRs) is investigated via adaptive dynamic programming. To begin with, a novel performance index function which contains containment errors and their derivatives is designed for each following WMR without requiring the discount factor, which simplifies the controller design process and enhances the practicality of the control method. Subsequently, the DOC control of WMRs is formulated as a differential graphical game whose Nash equilibrium can be formed by using the optimal responses of all following WMRs. Moreover, a critic-only structure is built to obtain an approximate DOC control law, which provides a solution for the coupled Hamilton–Jacobi–Bellman equation of each following WMR. Stability analysis demonstrates that the containment error of each following WMR is uniformly ultimately bounded. Finally, a group of WMRs are utilized to verify the effectiveness of the present DOC control scheme.
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
abstract
Multi-modal retrieval becomes increasingly popular in practice.However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information.Despite the presence of vision-language models like CLIP, the current methods are severely limited in representing the text-only and imageonly data.In this work, we present a new embedding model VISTA for universal multimodal retrieval.Our work brings forth threefold technical contributions.Firstly, we introduce a flexible architecture which extends a powerful text encoder with the image understanding capability by introducing visual token embeddings.Secondly, we develop two data generation strategies, which bring highquality composed image-text to facilitate the training of the embedding model.Thirdly, we introduce a multi-stage training algorithm, which first aligns the visual token embedding with the text encoder using massive weakly labeled data, and then develops multi-modal representation capability using the generated composed image-text data.In our experiments, VISTA achieves superior performances across a variety of multi-modal retrieval tasks in both zero-shot and supervised settings.Our model, data, and source code are available at https://github.com/FlagOpen/FlagEmbedding.
Junjie Zhou 0001, Zheng Liu 0011, Shitao Xiao, Bo Zhao 0015, Yongping Xiong
ACL (1)4
2024 Massive Multi-agent Mean-Field Game Using Online Federated Adaptive Critic-Density Learning
Mingduo Lin, Guoling Yuan, Bo Zhao 0015, Derong Liu 0001
ICONIP (4)3
2024 ENAO: Evolutionary Neural Architecture Optimization in the Approximate Continuous Latent Space of a Deep Generative Model
abstract
Neural architecture search (NAS) has emerged as a transformative approach for automating the design of neural networks, demonstrating exceptional performance across a variety of tasks. Numerous NAS methods aim to optimize neural architectures within discrete or continuous search spaces, but each method possesses its own inherent limitations. Additionally, the search efficiency is notably impeded by suboptimal encoding methods, presenting an ongoing challenge. In response to these obstacles, this paper introduces a novel approach, evolutionary neural architecture optimization (ENAO), which optimizes architectures in an approximate continuous search space. ENAO begins with training a deep generative model to embed discrete architectures into a condensed latent space, leveraging unsupervised representation learning. Subsequently, evolutionary algorithm is employed to refine neural architectures within this approximate continuous latent space. Empirical comparisons against several NAS benchmarks underscore the effectiveness of the ENAO method. Thanks to its foundation in deep unsupervised representation learning, ENAO demonstrates a distinguished ability to identify high-quality architectures with fewer evaluations and achieve state-of-the-art result in NAS-Bench-201 dataset. Overall, the ENAO method is a promising approach for optimizing neural network architectures in an approximate continuous search space with evolutionary algorithms and may be a useful tool for researchers and practitioners in the field of NAS.
Xuan Rao, Shaojie Liu, Bo Zhao 0015, Derong Liu 0001
IJCNN4
2024 Temporal Normalization Flow for Probabilistic Time Series Forecasting
abstract
Time series data has the characteristics of strong randomness, complex data structures, and high non-stationarity. Time series probabilistic forecasting is of great significance for quantifying the uncertainty in time series. This study presents a novel probabilistic forecasting approach that is distribution-free and relies on the normalization flow technology. The developed method utilizes normalization flow to transform complex target distributions into simpler ones, which is convenient for probabilistic modeling of the target distributions. Additionally, the method we have developed employs both convolution and attention mechanisms to identify temporal patterns across both short and extended timeframes within the time series. Comprehensive empirical assessments across various real-world time series datasets confirm that the proposed method surpasses standard models in predictive accuracy.
Jiarui Ye, Bo Zhao 0015, Derong Liu 0001
INDIN2
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS36
2024 SegVol: Universal and Interactive Volumetric Medical Image Segmentation
abstract
Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence of a 3D foundation segmentation model that can segment a wide range of anatomical categories with easy user interaction. In this paper, we propose a 3D foundation segmentation model, named SegVol, supporting universal and interactive volumetric medical image segmentation. By scaling up training data to 90K unlabeled Computed Tomography (CT) volumes and 6K labeled CT volumes, this foundation model supports the segmentation of over 200 anatomical categories using semantic and spatial prompts. To facilitate efficient and precise inference on volumetric images, we design a zoom-out-zoom-in mechanism. Extensive experiments on 22 anatomical segmentation tasks verify that SegVol outperforms the competitors in 19 tasks, with improvements up to 37.24\% compared to the runner-up methods. We demonstrate the effectiveness and importance of specific designs by ablation study. We expect this foundation model can promote the development of volumetric medical image analysis. The model and code are publicly available at https://github.com/BAAI-DCAI/SegVol.
Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015
NeurIPS4
2024 Fetch and Forge: Efficient Dataset Condensation for Object Detection
abstract
Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses on image classification, with less exploration of object detection. This is primarily due to two challenges: (i) the multitasking nature of object detection complicates the condensation process, and (ii) Object detection datasets are characterized by large-scale and high-resolution data, which are difficult for existing DC methods to handle. As a remedy, we propose DCOD, the first dataset condensation framework for object detection. It operates in two stages: Fetch and Forge, initially storing key localization and classification information into model parameters, and then reconstructing synthetic images via model inversion. For the complex of multiple objects in an image, we propose Foreground Background Decoupling to centrally update the foreground of multiple instances and Incremental PatchExpand to further enhance the diversity of foregrounds. Extensive experiments on various detection datasets demonstrate the superiority of DCOD. Even at an extremely low compression rate of 1\%, we achieve 46.4\% and 24.7\% $\text{AP}_{50}$ on the VOC and COCO, respectively, significantly reducing detector training duration.
Ding Qi, Jian Li 0062, Jinlong Peng, Bo Zhao 0015, Shuguang Dou, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
NeurIPS4
2024 Uncertainty-based Continual Learning for Neural Networks with Low-rank Variance Matrices
abstract
Bayesian inference has provided the continual learning (CL) with an elegant framework where past experiences and new knowledge are consolidated into the posterior constantly. Typical approaches rely on Bayesian neural networks whose parameters are updated by variational inference, namely, maximizing the evidence lower bound of log-likelihood. In this paper, we discuss the effects of local reparameterization on the optimization of such networks in the context of CL. The empirical results show that it does not only increase the inference speed of neural networks, but also enhance the CL performance in some scenarios. Additionally, motivated by the observation that variance matrices have low-rank structures, we propose the d-tied variational continual learning (d-tied-VCL) to improve the parameter efficiency of variational continual learning (VCL). Experiments on random classification, per-muted MNIST, and split CIFAR100 show that even VCL with rank-1 variance matrices achieves competitive performance.
Xuan Rao, Bo Zhao 0015, Derong Liu 0001
SMC2
2024 Dynamic compensator-based near-optimal control for unknown nonaffine systems via integral reinforcement learning
Jinquan Lin, Bo Zhao 0015, Derong Liu 0001, Yonghua Wang 0001
Neurocomputing2
2024 Fault-tolerant tracking control for nonlinear systems with multiplicative actuator faults in view of zero-sum differential games
Mingduo Lin, Hongbing Xia, Bo Zhao 0015
Neurocomputing4
2024 Semi-supervised accuracy predictor-based multi-objective neural architecture search
Songyi Xiao, Bo Zhao 0015, Derong Liu 0001
Neurocomputing2
2024 Event-Triggered Robust Adaptive Dynamic Programming for Multiplayer Stackelberg-Nash Games of Uncertain Nonlinear Systems
abstract
In this article, an event-triggered robust adaptive dynamic programming (ETRADP) algorithm is developed to solve a class of multiplayer Stackelberg-Nash games (MSNGs) for uncertain nonlinear continuous-time systems. Considering the different roles of players in the MSNG, the hierarchical decision-making process is described as the designed value functions for the leader and all followers, which assist to transform the robust control problem of the uncertain nonlinear system into an optimal regulation problem of the nominal system. Then, an online policy iteration algorithm is formulated to solve the derived coupled Hamilton-Jacobi equation. Meanwhile, an event-triggered mechanism is designed to alleviate computational and communication burdens. Moreover, critic neural networks (NNs) are constructed to obtain the event-triggered approximate optimal control polices for all players, which constitute the Stackelberg-Nash equilibrium of the MSNG. By using Lyapunov's direct method, the stability of the closed-loop uncertain nonlinear system is guaranteed under the ETRADP-based control scheme in the sense of uniform ultimate boundedness. Finally, a numerical simulation is provided to demonstrate the effectiveness of the present ETRADP-based control scheme.
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Cybern.2
2024 Synchronization of Delayed Memristor-Based Neural Networks via Pinning Control With Local Information
abstract
In this article, a novel pinning control method, only requiring information from partial nodes, is developed to synchronize drive-response memristor-based neural networks (MNNs) with time delay. An improved mathematical model of MNNs is established to describe the dynamic behaviors of MNNs accurately. In the existing literature, pinning controllers for synchronization of drive-response systems were designed based on information of all nodes, but in some specific situations, the control gains may be very large and challenging to realize in practice. To overcome this problem, a novel pinning control policy is developed to achieve synchronization of delayed MNNs, which depends only on local information of MNNs, for reducing communication and calculation burdens. Furthermore, sufficient conditions for synchronization of delayed MNNs are provided. Finally, numerical simulation and comparative experiments are conducted to verify the effectiveness and superiority of the proposed pinning control method.
Zhanyu Yang, Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Distributed Fault Tolerant Consensus Control of Nonlinear Multiagent Systems via Adaptive Dynamic Programming
abstract
This article develops a distributed fault-tolerant consensus control (DFTCC) approach for multiagent systems by using adaptive dynamic programming. By establishing a local fault observer, the potential actuator faults of each agent are estimated. Subsequently, the DFTCC problem is transformed into an optimal consensus control problem by designing a novel local value function for each agent which contains the estimated fault, the consensus errors, and the control laws of the local agent and its neighbors. In order to solve the coupled Hamilton-Jacobi-Bellman equation of each agent, a critic-only structure is established to obtain the approximate local optimal consensus control law of each agent. Moreover, by using Lyapunov's direct method, it is proven that the approximate local optimal consensus control law guarantees the uniform ultimate boundedness of the consensus error of all agents, which means that all following agents with potential actuator faults synchronize to the leader. Finally, two simulation examples are provided to validate the effectiveness of the present DFTCC scheme.
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang
IEEE Trans. Neural Networks Learn. Syst.2
2024 Event-Triggered Decentralized Integral Sliding Mode Control for Input-Constrained Nonlinear Large-Scale Systems With Actuator Failures
abstract
In this article, an event-triggered decentralized integral sliding mode control (ETDISMC) method is investigated for a class of input-constrained nonlinear large-scale systems with actuator failures based on adaptive dynamic programming (ADP). An integral sliding mode control method is developed to maintain the subsystem trajectories on the sliding mode surface, eliminate the effect of actuator failures, and obtain the sliding mode dynamics (SMDs). Then, the control problem is transformed into an optimal control (OC) problem for the nominal form of the SMDs by constructing a modified local value function. To obtain the event-triggered OC law, a critic-only structure is applied to approximate the local optimal value function of the nominal subsystem for solving the event-triggered Hamilton–Jacobi–Bellman equation. An event-triggered ADP control method is developed to decrease the updating frequency of the OC law and to reduce the computational burden. In addition, an experience replay-based weight updating policy is presented to relax the persistence of excitation condition. Furthermore, we prove that the developed method can guarantee the closed-loop system to be asymptotically stable by using Lyapunov’s direct method. Finally, a numerical example and a practical system are employed for simulation to demonstrate the effectiveness of the proposed ETDISMC scheme.
Shunchao Zhang, Bo Zhao 0015, Derong Liu 0001, Yongwei Zhang 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Fault tolerant control for a class of nonlinear systems with multiple faults using neuro-dynamic programming
Chujian Zeng, Bo Zhao 0015, Derong Liu 0001
Neurocomputing2
2023 Event-triggered adaptive dynamic programming for decentralized tracking control of input constrained unknown nonlinear interconnected systems
Qiuye Wu, Bo Zhao 0015, Derong Liu 0001, Marios M. Polycarpou
Neural Networks2
2023 Policy gradient adaptive dynamic programming for nonlinear discrete-time zero-sum games with unknown dynamics
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001
Soft Comput.2
2023 Adaptive Dynamic Programming-Based Event-Triggered Robust Control for Multiplayer Nonzero-Sum Games With Unknown Dynamics
abstract
In this article, the event-triggered robust control of unknown multiplayer nonlinear systems with constrained inputs and uncertainties is investigated by using adaptive dynamic programming. To relax the requirement of system dynamics, a neural network-based identifier is constructed by using the system input-output data. Subsequently, by designing a nonquadratic value function, which contains the bounded functions, the system states, and the control inputs of all players, the event-triggered robust stabilization problem is converted into an event-triggered constrained optimal control problem. To obtain the approximate solution of the event-triggered Hamilton-Jacobi (HJ) equation, a critic network for each player is established with a novel weight updating law to relax the persistence of excitation condition based on the experience replay technique. Furthermore, according to the Lyapunov stability theorem, the present event-triggered robust optimal control ensures the multiplayer system to be uniformly ultimately bounded. Finally, two simulation examples are employed to show the effectiveness of the present method.
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang
IEEE Trans. Cybern.2
2023 Adaptive Dynamic Programming-Based Cooperative Motion/Force Control for Modular Reconfigurable Manipulators: A Joint Task Assignment Approach
abstract
This article develops a cooperative motion/force control (CMFC) scheme based on adaptive dynamic programming (ADP) for modular reconfigurable manipulators (MRMs) with the joint task assignment approach. By separating terms depending on local variables only, the dynamic model of the entire MRM system can be regarded as a set of joint modules interconnected by coupling torque. In addition, the Jacobian matrix, which reflects the interaction force of the MRM end-effector, can be mapped into each joint. Using this approach, both the motion and force tasks on the end-effector of the entire MRM system can be assigned to each joint module cooperatively. Then, by substituting the actual states of coupled joint modules with their desired ones, the norm-boundedness assumption on the interconnection of joint module can be relaxed. By using the measured input-output data of each joint module, a neural network (NN)-based robust decentralized observer, which guarantees the observation error to be asymptotically stable is established. An improved local value function is constructed for each joint module to reflect the interconnection. Then, the local Hamilton-Jacobi-Bellman equation is solved by constructing a local critic NN with a nested learning structure. Hereafter, the ADP-based CMFC is obtained by the assistance of force feedback compensation. Based on the Lyapunov stability analysis, the closed-loop MRM system is guaranteed to be uniformly ultimately bounded under the present ADP-based CMFC scheme. The simulation on a two-degree of freedom MRM system demonstrates the effectiveness of the present control approach.
Bo Zhao 0015, Yongwei Zhang 0002, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Policy Optimization Adaptive Dynamic Programming for Optimal Control of Input-Affine Discrete-Time Nonlinear Systems
abstract
In this article, a policy optimization adaptive dynamic programming (POADP) method is developed for optimal control of discrete-time unknown nonlinear systems, where the iterative control policy is parameterized to optimize the iterative$Q$-function directly. The relaxed condition for the learning rate is given to guarantee the convergence of the present algorithm. Furthermore, the Polyak– ojasiewicz inequality is introduced to analyze the optimality, i.e., the iterative$Q$-function converges to the optimum within a given computational threshold under a finite iteration, and the rate of convergence (i.e., the required minimum number of iterations) for the developed POADP method is also illustrated. To ease real implementations, the iterative$Q$-function and the iterative control policy are approximated by employing an actor–critic structure. Then, an experiment-based method is developed to obtain the initial weights of actor–critic structure. Finally, numerical simulation results of two examples are provided to validate the effectiveness of the POADP algorithm.
Mingduo Lin, Bo Zhao 0015
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Event-Triggered Local Control for Nonlinear Interconnected Systems Through Particle Swarm Optimization-Based Adaptive Dynamic Programming
abstract
This article investigates local control problems for nonlinear interconnected systems by using adaptive dynamic programming (ADP) with particle swarm optimization (PSO). Through constructing a proper local value function, a local critic neural network, whose weight vector is tuned via the PSO algorithm, is employed to solve the local Hamilton–Jacobi–Bellman equation. By introducing the event-triggering mechanism, the sampling time instants of each interconnected subsystem are determined by establishing a proper event-triggering condition. Then, the ADP-based event-triggered local control policy can be derived indirectly to ensure the closed-loop nonlinear interconnected system to be asymptotically stable through the Lyapunov stability analysis. The positive lower bound on the minimal intersampling instant for each interconnected subsystem is provided to exclude the Zeno behavior. Simulation results of a practical system and a numerical example demonstrate the effectiveness of the present event-triggered local control scheme.
Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2022 DMPP: Differentiable multi-pruner and predictor for neural network pruning
Bo Zhao 0015, Derong Liu 0001
Neural Networks2
2022 Synergetic learning structure-based neuro-optimal fault tolerant control for unknown nonlinear systems
Hongbing Xia, Bo Zhao 0015, Ping Guo 0002
Neural Networks2
2022 Policy Gradient Adaptive Critic Designs for Model-Free Optimal Tracking Control With Experience Replay
abstract
A model-free optimal tracking controller is designed for discrete-time nonlinear systems through policy gradient adaptive critic designs (PGACDs) with experience replay (ER). By using system transformation, optimal tracking control problems are converted into optimal regulation problems. An off-policy PGACD algorithm is developed to minimize the iterative$Q$-function and improve the tracking control performance. The proposed method is realized based on the critic network and the actor network (AN), which are applied to approximate the iterative$Q$-function and the iterative control policy, respectively. Then, the policy gradient technique is introduced to derive a novel weight updating law of the AN explicitly by using measured system data only. The convergence of the iteration is established through theoretical analysis, and the uniform ultimate boundedness is demonstrated for the closed-loop system under the PGACD-based controller by using Lyapunov’s direct method. To guarantee the stability and increase the data usage efficiency of the learning process, an ER-based learning framework is designed to improve the realizability of the proposed method. Finally, simulation results of two examples are provided to demonstrate the performance of the off-policy PGACD algorithm.
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Event-Triggered Control of Discrete-Time Zero-Sum Games via Deterministic Policy Gradient Adaptive Dynamic Programming
abstract
In order to address zero-sum game problems for discrete-time (DT) nonlinear systems, this article develops a novel event-triggered control (ETC) approach based on the deterministic policy gradient (PG) adaptive dynamic programming (ADP) algorithm. By adopting the input and output data, the proposed ETC method updates the control law and the disturbance law with a gradient descent algorithm. Compared with the conventional PG ADP-based control scheme, the present controller is updated aperiodically to reduce the computational and communication burden. Then, the actor-critic-disturbance framework is adopted to obtain the optimal control law and the worst disturbance law, which guarantee the input-to-state stability of the closed-loop system. Moreover, a novel neural network weight updating law which guarantees the uniform ultimate boundedness of weight estimation errors is provided based on the experience replay technique. Finally, the validity of the present method is verified by simulation of two DT nonlinear systems.
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Event-triggered control for input constrained non-affine nonlinear systems based on neuro-dynamic programming
Shunchao Zhang, Bo Zhao 0015, Yongwei Zhang 0002
Neurocomputing2
2021 Observer-based event-triggered control for zero-sum games of input constrained multi-player nonlinear systems
Shunchao Zhang, Bo Zhao 0015, Derong Liu 0001, Yongwei Zhang 0002
Neural Networks2
2021 Particle swarm optimized neural networks based local tracking control scheme of unknown nonlinear interconnected systems
Bo Zhao 0015, Fangchao Luo, Haowei Lin, Derong Liu 0001
Neural Networks1
2021 Event-triggered adaptive dynamic programming for multi-player zero-sum games with unknown dynamics
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001
Soft Comput.2
2021 Sliding-Mode Surface-Based Approximate Optimal Control for Uncertain Nonlinear Systems With Asymptotically Stable Critic Structure
abstract
This article develops a novel sliding-mode surface (SMS)-based approximate optimal control scheme for a large class of nonlinear systems affected by unknown mismatched perturbations. The observer-based perturbation estimation procedure is employed to establish the online updated value function. The solution to the Hamilton-Jacobi-Bellman equation is approximated by an SMS-based critic neural network whose weights error dynamics is designed to be asymptotically stable by nested update laws. The sliding-mode control strategy is combined with the approximate optimal control design procedure to obtain a faster control action. The stability is proved based on the Lyapunov's direct method. The simulation results show the effectiveness of the developed control scheme.
Bo Zhao 0015, Derong Liu 0001, Cesare Alippi
IEEE Trans. Cybern.1
2021 Adaptive Dynamic Programming for Control: A Survey and Recent Advances
abstract
This article reviews the recent development of adaptive dynamic programming (ADP) with applications in control. First, its applications in optimal regulation are introduced, and some skilled and efficient algorithms are presented. Next, the use of ADP to solve game problems, mainly nonzero-sum game problems, is elaborated. It is followed by applications in large-scale systems. Note that although the functions presented in this article are based on continuous-time systems, various applications of ADP in discrete-time systems are also analyzed. Moreover, in each section, not only some existing techniques are discussed, but also possible directions for future work are pointed out. Finally, some overall prospects for the future are given, followed by conclusions of this article. Through a comprehensive and complete investigation of its applications in many existing fields, this article fully demonstrates that the ADP intelligent control method is promising in today's artificial intelligence era. Furthermore, it also plays a significant role in promoting economic and social development.
Derong Liu 0001, Shan Xue 0004, Bo Zhao 0015, Biao Luo 0001, Qinglai Wei
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Augmented Bi-path Network for Few-shot Learning
abstract
Few-shot Learning (FSL) which aims to learn from few labeled training data is becoming a popular research topic, due to the expensive labeling cost in many real-world applications. One kind of successful FSL method learns to compare the testing (query) image and training (support) image by simply concatenating the features of two images and feeding it into the neural network. However, with few labeled data in each class, the neural network has difficulty in learning or comparing the local features of two images. Such simple image-level comparison may cause serious mis-classification. To solve this problem, we propose Augmented Bi-path Network (ABNet) for learning to compare both global and local features on multi-scales. Specifically, the salient patches are extracted and embedded as the local features for every image. Then, the model learns to augment the features for better robustness. Finally, the model learns to compare global and local features separately, i.e., in two paths, before merging the similarities. Extensive experiments show that the proposed ABNet outperforms the state-of-the-art methods. Both quantitative and visual ablation studies are provided to verify that the proposed modules lead to more precise comparison results.
Baoming Yan, Bo Zhao 0015, Kan Guo, Ming Zhang 0004, Yizhou Wang 0001
ICPR3
2020 Boosting and Residual Learning Scheme with Pseudoinverse Learners
abstract
The traditional gradient descent based optimization algorithms for neural network are subjected too many vulnerabilities, such as slow convergent rate, gradient vanishing and falling into local minima. Therefore, the alternative non-gradient descent learning algorithm was proposed and prevalently applied in kinds of domains, such as pseudoinverse learning algorithm (PIL). However, when a special variant of the PIL, taking the random configuration of weight parameters, is adopted, the generalization ability needs further improvement although it has excellent training efficiency. Thus, on consideration of integrating the idea of ensemble learning, we proposes two methods to enhance basic PIL. One method is equivalent to an additive model, which can raise the network's performance by introducing boosting mechanism, and the other is to adopt a recursive way to rectify the hidden layer output of the neural network, then the relative better model is used in the subsequent prediction. Comprehensive evaluating experiments are conducted on several datasets, and the experimental results illustrate that the our proposed methods are effective on the classification accuracy.
Xiaoxuan Sun, Rundong Shi, Bo Zhao 0015, Ping Guo 0002
SMC3
2020 Deterministic policy gradient adaptive dynamic programming for model-free optimal control
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001
Neurocomputing2
2020 Asymptotically stable critic designs for approximate optimal stabilization of nonlinear systems subject to mismatched external disturbances
Bo Zhao 0015, Ding Wang 0001
Neurocomputing1
2020 Reinforcement Learning-Based Optimal Stabilization for Unknown Nonlinear Systems Subject to Inputs With Uncertain Constraints
abstract
This article presents a novel reinforcement learning strategy that addresses an optimal stabilizing problem for unknown nonlinear systems subject to uncertain input constraints. The control algorithm is composed of two parts, i.e., online learning optimal control for the nominal system and feedforward neural networks (NNs) compensation for handling uncertain input constraints, which are considered as the saturation nonlinearities. Integrating the input-output data and recurrent NN, a Luenberger observer is established to approximate the unknown system dynamics. For nominal systems without input constraints, the online learning optimal control policy is derived by solving Hamilton-Jacobi-Bellman equation via a critic NN alone. By transforming the uncertain input constraints to saturation nonlinearities, the uncertain input constraints can be compensated by employing a feedforward NN compensator. The convergence of the closed-loop system is guaranteed to be uniformly ultimately bounded by using the Lyapunov stability analysis. Finally, the effectiveness of the developed stabilization scheme is illustrated by simulation studies.
Bo Zhao 0015, Derong Liu 0001, Chaomin Luo
IEEE Trans. Neural Networks Learn. Syst.1
2019 Large-Scale Datasets for Going Deeper in Image Understanding
abstract
Recently, extensive efforts have been devoted to computer vision and machine learning by exploiting big data to explore many practical applications. However, these research fields are still quite limited not only by the sheer volume, but also the versatility and diversity, of the available datasets. In this paper, we target at four challenging and yet important computer vision tasks, namely, human-centered scene classification, attribute based zero-shot learning (recognition), human keypoint detection and image Chinese captioning. Four novel large-scale datasets are collected and annotated to facilitate these tasks of deeper image understanding. Labels, bounding boxes, attributes, keypoints and captions are annotated in corresponding datasets. These rich annotations bridge the semantic gap between low-level images and high-level concepts. Extensive experiments on baseline methods have been implemented and compared, which show that these learning tasks on our datasets are still challenging.
Jiahong Wu 0006, He Zheng, Bo Zhao 0015, Baoming Yan, Shipei Zhou, Guosen Lin, Yanwei Fu 0001, Yizhou Wang 0001
ICME3
2019 Local Near-Optimal Control for Interconnected Systems with Time-Varying Delays
Qiuye Wu, Haowei Lin, Bo Zhao 0015, Derong Liu 0001
ICONIP (2)3
2019 Distributed Adaptive Dynamic Programming Algorithm for Office Energy Control with Multiple Batteries
Chao Li 0024, Bo Zhao 0015, Qinglai Wei, Derong Liu 0001
IJCNN3
2019 A Solution of Two-Person Zero Sum Differential Games with Incomplete State Information
Kanghao Du, Ruizhuo Song, Qinglai Wei, Bo Zhao 0015
ISNN (1)4
2019 An Ensemble Model for Error Modeling with Pseudoinverse Learning Algorithm
abstract
In Bayesian theory, the maximum posterior estimator uses prior information to estimate the noise in the machine learning model by adding the regularization term. The regularization terms L1and L2correspond to Laplacian prior and Guassian prior, respectively. In existing deep learning models, in order to use the gradient descent optimization algorithm and achieve good results, most models take L2regularization as the regularization term of the network model to fit the complex Guassian noise. However in practice, the Laplace noise and the Guassian noise are both considered as data noise. For multi-layer perceptrons, the difficulty caused by adding L1and L2into the optimization function of the network is solved by proposing an ensemble model for error modeling through adopting the divide and conquer strategy. First, several base learners are trained to fit different noise distributions of data, then the final results can be obtained by taking the results of each base leaner as new data to train a meta leaner, and get the final results. Among them, coordinate regression method is used to solve L1loss, while the pseudo-inverse learning algorithm is employed to solve L2loss. Both methods are nongradient optimization algorithms. The comparison results of the model on several data sets show that the proposed ensemble model achieves better performance.
Sibo Feng, Xiaodan Deng, Ping Guo 0002, Bo Zhao 0015, Qian Yin 0001
SMC4
2019 Zero-Shot Learning Via Recurrent Knowledge Transfer
abstract
Zero-shot learning (ZSL) which aims to learn new concepts without any labeled training data is a promising solution to large-scale concept learning. Recently, many works implement zero-shot learning by transferring structural knowledge from the semantic embedding space to the image feature space. However, we observe that such direct knowledge transfer may suffer from the space shift problem in the form of the inconsistency of geometric structures in the training and testing spaces. To alleviate this problem, we propose a novel method which actualizes recurrent knowledge transfer (RecKT) between the two spaces. Specifically, we unite the two spaces into the joint embedding space in which unseen image data are missing. The proposed method provides a synthesis-refinement mechanism to learn the shared subspace structure (SSS) and synthesize missing data simultaneously in the joint embedding space. The synthesized unseen image data are utilized to construct the classifier for unseen classes. Experimental results show that our method outperforms the state-of-the-art on three popular datasets. The ablation experiment and visualization of the learning process illustrate how our method can alleviate the space shift problem. By product, our method provides a perspective to interpret the ZSL performance by implementing subspace clustering on the learned SSS.
Bo Zhao 0015, Xinwei Sun 0001, Xiaopeng Hong, Yuan Yao 0011, Yizhou Wang 0001
WACV1
2019 An echo state network based approach to room classification of office buildings
Bo Zhao 0015, Chao Li 0024, Qinglai Wei, Derong Liu 0001
Neurocomputing2
2018 MSplit LBI: Realizing Feature Selection and Dense Estimation Simultaneously in Few-shot and Zero-shot Learning
abstract
It is one typical and general topic of learning a good embedding model to efficiently learn the representation coefficients between two spaces/subspaces. To solve this task, $L_{1}$ regularization is widely used for the pursuit of feature selection and avoiding overfitting, and yet the sparse estimation of features in $L_{1}$ regularization may cause the underfitting of training data. $L_{2}$ regularization is also frequently used, but it is a biased estimator. In this paper, we propose the idea that the features consist of three orthogonal parts, namely sparse strong signals, dense weak signals and random noise, in which both strong and weak signals contribute to the fitting of data. To facilitate such novel decomposition, MSplit LBI is for the first time proposed to realize feature selection and dense estimation simultaneously. We provide theoretical and simulational verification that our method exceeds $L_{1}$ and $L_{2}$ regularization, and extensive experimental results show that our method achieves state-of-the-art performance in the few-shot and zero-shot learning.
Bo Zhao 0015, Xinwei Sun 0001, Yanwei Fu 0001, Yuan Yao 0011, Yizhou Wang 0001
ICML1
2018 Local Tracking Control for Unknown Interconnected Systems via Neuro-Dynamic Programming
Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Ding Wang 0001, Yancai Xu, Qinglai Wei
ICONIP (7)1
2018 Decentralized Control for Large-Scale Nonlinear Systems With Unknown Mismatched Interconnections via Policy Iteration
abstract
In this paper, the decentralized control problem is solved based on a policy iteration algorithm for large-scale nonlinear systems with unknown mismatched interconnections. The unknown interconnection is approximated by a neural network with local states of isolated subsystem and substituted reference states of coupled subsystems. Then, an adaptive estimation term is utilized to construct the improved local performance index function that reflects the substitution error. Hereafter, the closed-loop large-scale nonlinear system is guaranteed to be ultimately uniformly bounded by the implementation of a set of developed decentralized optimal control policies. Two simulation examples are given to verify the effectiveness of the presented scheme. The significant contribution of this scheme lies in that it removes the common assumptions on satisfying matching condition and upper boundedness of interconnections, when designing the decentralized optimal control for large-scale nonlinear systems.
Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Neuro-control of Nonlinear Systems with Unknown Input Constraints
Bo Zhao 0015, Derong Liu 0001
ICONIP (1)1
2017 A Generalized Policy Iteration Adaptive Dynamic Programming Algorithm for Optimal Control of Discrete-Time Nonlinear Systems with Actuator Saturation
Qiao Lin 0003, Qinglai Wei, Bo Zhao 0015
ISNN (2)3
2017 Self-tuned local feedback gain based decentralized fault tolerant control for a class of large-scale nonlinear systems
Bo Zhao 0015, Derong Liu 0001
Neurocomputing1
2017 Observer based adaptive dynamic programming for fault tolerant control of a class of nonlinear systems
Bo Zhao 0015, Derong Liu 0001
Inf. Sci.1
2016 Decentralized Stabilization for Nonlinear Systems with Unknown Mismatched Interconnections
Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001
ICONIP (3)1
2016 Decentralized adaptive neural network sliding mode control for reconfigurable manipulators with data-based modeling
abstract
In this paper, a decentralized adaptive neural network sliding mode control scheme is proposed for trajectory tracking control problem of reconfigurable manipulators based on data-based modeling. This method can be implemented to reconfigurable manipulators with different configurations and degrees of freedom without modifying any control parameters. Different from the previous works, the proposed control strategy is applied to mechanism model and data-based model of reconfigurable manipulators, respectively. The data-based model which is more comprehensive and precise is trained by the BP neural network with sampled input-output data. The gradient descent method is used to attain higher identification precision. Then the asymptotical stability of the system is proved using the Lyapunov theorem. Simulations are presented to not only illustrate the effectiveness of the proposed decentralized control scheme, but make a detailed comparison for control performance of the two modeling methods.
Guibin Ding, Bo Zhao 0015, Bo Dong 0002
IJCNN3
2016 Asteroid landing via onboard optimal guidance based on bidirectional extreme learning machine
abstract
In order to autonomously design the optimal descending trajectory for spacecraft soft landing on an asteroid, an onboard guidance based on the bidirectional extreme learning machine (B-ELM) is proposed. The optimization problem is formulated and transformed into a two-point boundary value problem (TPBVP). And then, based on the sample trajectories obtained off-line, a single-hidden layer feed-forward neural network (SLFN) trained by B-ELM is employed to design the optimal descending trajectory onboard. Finally, Monte Carlo simulations are performed to verify the effectiveness of the proposed guidance. Also, the learning process of the B-ELM is compared with traditional algorithms in simulations. Simulation results show that the guidance via the B-ELM trained SLFN meets the requirement of the soft landing in a lower learning and implementation cost.
Xiaosong Liu, Bo Zhao 0015, Mujun Xie, Keping Liu
IJCNN3
2016 Adaptive dynamic programming based fault compensation control for nonlinear systems with actuator failures
abstract
This paper develops a novel fault compensation control scheme based on adaptive dynamic programming for nonlinear systems with actuator failures. The control scheme consists of a policy iteration algorithm and a fault compensation. For fault-free dynamic models, the Hamilton-Jacobi-Bellman equation is solved by policy iteration algorithm via constructing a critic neural network, and then the approximate optimal control policy can be derived directly. On the other hand, the online fault compensation is achieved without the fault detection and isolation mechanism by reconstructing the actuator failure. The closed-loop system is guaranteed to be asymptotically stable based on Lyapunov stability theorem. Two simulation examples are given to demonstrate the effectiveness of the present fault compensation control scheme.
Bo Zhao 0015, Derong Liu 0001
IJCNN1