VLDB 2026 Research / reviewers in the wild / expert
Wei Pan 0004
dblp:69/4450-4
· DBLP profile ↗
37ranked-venue papers
0as first author
33since 2021 · last 2026
0000-0003-1121-9879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 24 since 2021Systems, architecture and hardware · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vehicle Dynamics Embedded World Models for Autonomous DrivingabstractWorld models have gained significant attention as a promising approach for autonomous driving. By emulating human-like perception and decision-making processes, these models can predict and adapt to dynamic environments. Existing methods typically map high-dimensional observations into compact latent spaces and learn optimal policies within these latent representations. However, prior work usually jointly learns ego-vehicle dynamics and environmental transition dynamics from the image input, leading to inefficiencies and a lack of robustness to variations in vehicle dynamics. To address these issues, we propose the Vehicle Dynamics embedded Dreamer (VDD) method, which decouples the modeling of ego-vehicle dynamics from environmental transition dynamics. This separation allows the world model to generalize effectively across vehicles with diverse parameters. Additionally, we introduce two strategies to further enhance the robustness of the learned policy: Policy Adjustment during Deployment (PAD) and Policy Augmentation during Training (PAT). Comprehensive experiments in simulated environments demonstrate that the proposed model significantly improves both driving performance and robustness to variations in vehicle dynamics, outperforming existing approaches. Huiqian Li, Wei Pan 0004, Jin Huang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | Object-Reconstruction-Aware Whole-Body Control of Mobile ManipulatorsabstractObject reconstruction and inspection tasks play a crucial role in various robotics applications. Identifying paths that reveal the most unknown areas of the object is paramount in this context, as it directly affects reconstruction efficiency. This problem is known as the view path planning problem. Current methods often use sampling-based path planning techniques, evaluating potential views along the path to enhance reconstruction performance. However, these methods are computationally expensive as they require evaluating several candidate views on the path. To this end, we propose a computationally efficient solution that relies on calculating a focus point in the most informative (unknown) region and having the robot maintain this point in the camera field of view along the path. In this way, object reconstruction-related information is incorporated into the whole-body control of a mobile manipulator employing a visibility constraint without the need for an additional path planner. We conducted comprehensive and realistic simulations using a large dataset of 114 diverse objects of varying sizes from 57 categories to compare our method with a sampling-based planning strategy and a strategy that does not employ informative paths using Bayesian data analysis. Furthermore, to demonstrate the applicability and generality of the proposed approach, we conducted real-world experiments with an 8-DoF omnidirectional mobile manipulator and a legged manipulator. Our results suggest that, when compared to a sampling based strategy, there is no statistically significant difference in object reconstruction entropy, and there is a 52.3% probability that they are practically equivalent in terms of coverage. In contrast, our method is 6.2 to 19.36 times faster in terms of computation time and reduces the total time the robot spends between views by 13.76% to 27.9%, depending on the camera field of view and model resolution. When compared with strategies that do not exploit informative paths, our method improves, on average, coverage by 4.9% and entropy by 9.72% at the expense of spending 8.72% more time in the reconstruction process. Fatih Dursun, Bruno Vilhena Adorno, Simon Watson 0001, Wei Pan 0004 |
IEEE Trans. Robotics | 4 |
| 2025 | Multi-objective Sequential Decision Making for Holistic Supply Chain Optimization
Rifny Rachman, Josh C. Tingey, Richard Allmendinger 0001, Pradyumn Kumar Shukla, Wei Pan 0004 |
EMO (1) | 5 |
| 2025 | Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman FilteringabstractVision-language pretrained models, such as CLIP, have established new benchmarks in multimodal data mining. In such models, few-shot fine-tuning is a major challenge to achieve optimal performance on both in-distribution (ID) and out-of-distribution (OOD) datasets, especially when labeled data is scarce. Most existing fine-tuning approaches rely on first-order gradient-based optimizers, which typically suffer from slow convergence, sensitivity to step-size hyperparameters, and poor generalization in OOD settings. In contrast, second-order methods utilize local curvature information of the loss landscape to adjust the update step size. This is particularly beneficial for CLIP models, whose non-convex loss functions often contain sharp critical points. In such cases, natural gradient direction can offer more substantial and efficient per-iteration updates when fine-tuning with limited data. Natural Gradient Descent (NGD) is obtained by preconditioning the standard gradient with the inverse Fisher Information Matrix (FIM), which is computationally expensive for large models. To address this, we propose a Bayesian approximation of NGD using a Kalman filter for CLIP models. Our method combines the benefits of second-order optimization with Bayesian inference, which enhances generalization while providing uncertainty quantification. Extensive experiments conducted on diverse image classification datasets demonstrate that our algorithm consistently achieves superior-or comparable-ID performance and improved OOD robustness compared to state-of-the-art baselines. To the best of our knowledge, this work represents the first successful application of Kalman filtering to fine-tuning CLIP-based models, which enables more robust and efficient learning in vision-language tasks. Hossein Abdi, Mingfei Sun 0001, Wei Pan 0004 |
ICDM | 3 |
| 2025 | DroneDiffusion: Robust Quadrotor Dynamics Learning with Diffusion ModelsabstractAn inherent fragility of quadrotor systems stems from model inaccuracies and external disturbances. These factors hinder performance and compromise the stability of the system, making precise control challenging. Existing model-based approaches either make deterministic assumptions, utilize Gaussian-based representations of uncertainty, or rely on nominal models, all of which often fall short in capturing the complex, multimodal nature of real-world dynamics. This work introduces DroneDiffusion, a novel framework that leverages conditional diffusion models to learn quadrotor dynamics, formulated as a sequence generation task. DroneDiffusion achieves superior generalization to unseen, complex scenarios by capturing the temporal nature of uncertainties and mitigating error propagation. We integrate the learned dynamics with an adaptive controller for trajectory tracking with stability guarantees. Extensive experiments in both simulation and real-world flights demonstrate the robustness of the framework across a range of scenarios, including unfamiliar flight paths and varying payloads, velocities, and wind disturbances. Project page: https://sites.google.com/view/dronediffusion. Avirup Das, Rishabh Dev Yadav, Sihao Sun, Mingfei Sun 0001, Samuel Kaski, Wei Pan 0004 |
ICRA | 6 |
| 2025 | Explosive Jumping with Rigid and Articulated Soft Quadrupeds via Example Guided Reinforcement LearningabstractAchieving controlled jumping behaviour for a quadruped robot is a challenging task, especially when introducing passive compliance in mechanical design. This study addresses this challenge via imitation-based deep reinforcement learning with a progressive training process. To start, we learn the jumping skill by mimicking a coarse jumping example generated by model-based trajectory optimization. Subsequently, we generalize the learned policy to broader situations, including various distances in both forward and lateral directions, and then pursue robust jumping in unknown ground unevenness. In addition, without tuning the reward much, we learn the jumping policy for a quadruped with parallel elasticity. Results show that using the proposed method, i) the robot learns versatile jumps by learning only from a single demonstration, ii) the robot with parallel compliance reduces the landing error by 11.1%, saves energy cost by 15.2% and reduces the peak torque by 15.8%, compared to the rigid robot without parallel elasticity, iii) the robot can perform jumps of variable distances with robustness against ground unevenness (maximal ±4cm height perturbations) using only proprioceptive perception. Georgios Apostolides, Wei Pan 0004, Jens Kober, Cosimo Della Santina, Jiatao Ding |
IROS | 2 |
| 2025 | TAR: Teacher-Aligned Representations via Contrastive Learning for Quadrupedal LocomotionabstractQuadrupedal locomotion via Reinforcement Learning (RL) is commonly addressed using the teacher-student paradigm, where a privileged teacher guides a proprioceptive student policy. However, key challenges such as representation misalignment between privileged teacher and proprioceptive-only student, covariate shift due to behavioral cloning, and lack of deployable adaptation; lead to poor generalization in real-world scenarios. We propose Teacher-Aligned Representations via Contrastive Learning (TAR), a framework that leverages privileged information with self-supervised contrastive learning to bridge this gap. By aligning representations to a privileged teacher in simulation via contrastive objectives, our student policy learns structured latent spaces and exhibits robust generalization to Out-of-Distribution (OOD) scenarios, surpassing the fully privileged “Teacher”. Results showed accelerated training by 2× compared to state-of-the-art baselines to achieve peak performance. OOD scenarios showed better generalization by 40% on average compared to existing methods. Moreover, TAR transitions seamlessly into learning during deployment without requiring privileged states, setting a new benchmark in sample-efficient, adaptive locomotion and enabling continual fine-tuning in real-world scenarios. Open-source code and videos are available at https://amrmousa.com/TARLoco/. Amr Mousa, Neil Karavis, Michele Caprio, Wei Pan 0004, Richard Allmendinger 0001 |
IROS | 4 |
| 2025 | Local Path Optimization in The Latent Space Using Learned Distance GradientabstractConstrained motion planning is a common but challenging problem in robotic manipulation. In recent years, data-driven constrained motion planning algorithms have shown impressive planning speed and success rate. Among them, the latent motion method based on manifold approximation is the most efficient planning algorithm. Due to errors in manifold approximation and the difficulty in accurately identifying collision conflicts within the latent space, time-consuming path validity checks and path replanning are required. In this paper, we propose a method that trains a neural network to predict the minimum distance between the robot and obstacles using latent vectors as inputs. The learned distance gradient is then used to calculate the direction of movement in the latent space to move the robot away from obstacles. Based on this, a local path optimization algorithm in the latent space is proposed, and it is integrated with the path validity checking process to reduce the time of replanning. The proposed method is compared with state-of-the-art algorithms in multiple planning scenarios, demonstrating the fastest planning speed. Jiawei Zhang 0014, Chengchao Bai, Wei Pan 0004, Tianhang Liu, Jifeng Guo 0004 |
IROS | 3 |
| 2025 | Spatial-Aware Decision-Making with Ring Attractors in Reinforcement Learning SystemsabstractRing attractors, mathematical models inspired by neural circuit dynamics, provide a biologically plausible mechanism to improve learning speed and accuracy in Reinforcement Learning (RL). Serving as specialized brain-inspired structures that encode spatial information and uncertainty, ring attractors explicitly encode the action space, facilitate the organization of neural activity, and enable the distribution of spatial representations across the neural network in the context of Deep Reinforcement Learning (DRL). These structures also provide temporal filtering that stabilizes action selection during exploration, for example, by preserving the continuity between rotation angles in robotic control or adjacency between tactical moves in game-like environments. The application of ring attractors in the action selection process involves mapping actions to specific locations on the ring and decoding the selected action based on neural activity. We investigate the application of ring attractors by both building an exogenous model and integrating them as part of DRL agents. Our approach significantly improves state-of-the-art performance on the Atari 100k benchmark, achieving a 53\% increase in performance over selected baselines. Marcos Negre Saura, Richard Allmendinger 0001, Wei Pan 0004, Theodore Papamarkou |
NeurIPS | 3 |
| 2025 | Hierarchical Physics-Informed Neural Network for Rotor System Health AssessmentabstractDue to coupled nonlinearities and complex measurement noise, assess the condition of the rotor system remains a challenge, particularly in cases where historical run-to-failure data is lacking. To this end, we proposed a hierarchical physics-informed neural network (HPINN) to identify/discover the ordinary differential equations (ODEs) of a healthy/faulty rotor system from noise measurements and then assess the rotor condition based on the discovered ODEs. Specifically, the ODEs of a healthy rotor system are first stably identified from noisy measurement through HPINN guided by rotor dynamics. Based on the identified healthy ODEs, the extra fault terms in the ODEs of the faulty rotor system are then sparsely regressed from the predefined library embedded in HPINN, in which the phase compensation and alternating training strategy are developed to guarantee training convergence. Moreover, with the mathematical terms of discovered fault, the potential fault and the health indicator (HI) are diagnosed and constructed to assess the condition of the rotor system, respectively. Finally, the effectiveness of the proposed method is verified with simulation and test bench datasets, showing the potential for practical industrial applications.Note to Practitioners—This paper investigates the health assessment problem (condition monitoring and fault diagnosis) of the rotor system, a critical component in large rotating machinery. The proposed HPINN provides a hierarchical framework to firstly identify the ODEs of healthy rotor system and then discover the ODEs of faulty rotor system with limited monitoring data (3-5 seconds data collected from sensor commonly, depending on the rotating speeds). With the mathematical terms of discovered fault, the fault can be diagnosed and a health indicator (HI) can be constructed to assess the condition of rotor system in a fully interpretative way. This approach is applicable to large rotating machinery in safety-critical industries, such as circulating water pumps. Wei Cheng 0007, Ji Xing, Xuefeng Chen 0002, Zhibin Zhao 0002, Rongyong Zhang, Hongpeng Zhou, Wei Xing Zheng 0001, Wei Pan 0004 |
IEEE Trans Autom. Sci. Eng. | 11 |
| 2025 | NiSNN-A: Noniterative Spiking Neural Network With Attention With Application to Motor Imagery EEG ClassificationabstractMotor imagery (MI), an important category in electroencephalogram (EEG) research, often intersects with scenarios demanding low energy consumption, such as portable medical devices and isolated environment operations. Traditional deep learning (DL) algorithms, despite their effectiveness, are characterized by significant computational demands accompanied by high energy usage. As an alternative, spiking neural networks (SNNs), inspired by the biological functions of the brain, emerge as a promising energy-efficient solution. However, SNNs typically exhibit lower accuracy than their counterpart convolutional neural networks (CNNs). Although attention mechanisms successfully increase network accuracy by focusing on relevant features, their integration in the SNN framework remains an open question. In this work, we combine the SNN and the attention mechanisms for the EEG classification, aiming to improve precision and reduce energy consumption. To this end, we first propose a noniterative leaky integrate-and-fire (NiLIF) neuron model, overcoming the gradient issues in traditional SNNs that use iterative LIF neurons for long time steps. Then, we introduce the sequence-based attention mechanisms to refine the feature map. We evaluated the proposed noniterative SNN with attention (NiSNN-A) model on two MI EEG datasets, OpenBMI and BCIC IV 2a. Experimental results demonstrate that: 1) our model outperforms other SNN models by achieving higher accuracy and 2) our model increases energy efficiency compared with the counterpart CNN models (i.e., by 2.13 times) while maintaining comparable accuracy. Wei Pan 0004, Cosimo Della Santina |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPCabstractDistributed model predictive control (DMPC) is promising in achieving optimal cooperative control in multirobot systems (MRS). However, real-time DMPC implementation relies on numerical optimization tools to periodically calculate local control sequences online. This process is computationally demanding and lacks scalability for large-scale, nonlinear MRS. This article proposes a novel distributed learning-based predictive control framework for scalable multirobot control. Unlike conventional DMPC methods that calculate open-loop control sequences, our approach centers around a computationally fast and efficient distributed policy learning algorithm that generates explicit closed-loop DMPC policies for MRS without using numerical solvers. The policy learning is executed incrementally and forward in time in each prediction interval through an online distributed actor–critic implementation. The control policies are successively updated in a receding-horizon manner, enabling fast and efficient policy learning with the closed-loop stability guarantee. The learned control policies could be deployed online to MRS with varying robot scales, enhancing scalability and transferability for large-scale MRS. Furthermore, we extend our methodology to address the multirobot safe learning challenge through a force field-inspired policy learning approach. We validate our approach's effectiveness, scalability, and efficiency through extensive experiments on cooperative tasks of large-scale wheeled robots and multirotor drones. Our results demonstrate the rapid learning and deployment of DMPC policies for MRS with scales up to 10 000 units. Wei Pan 0004, Cong Li 0015, Xin Xu 0001, Xiangke Wang, Dewen Hu |
IEEE Trans. Robotics | 2 |
| 2024 | Impact of Computation in Integral Reinforcement Learning for Continuous-Time ControlabstractIntegral reinforcement learning (IntRL) demands the precise computation of the utility function's integral at its policy evaluation (PEV) stage. This is achieved through quadrature rules, which are weighted sums of utility functions evaluated from state samples obtained in discrete time. Our research reveals a critical yet underexplored phenomenon: the choice of the computational method -- in this case, the quadrature rule -- can significantly impact control performance. This impact is traced back to the fact that computational errors introduced in the PEV stage can affect the policy iteration's convergence behavior, which in turn affects the learned controller. To elucidate how computation impacts control, we draw a parallel between IntRL's policy iteration and Newton's method applied to the Hamilton-Jacobi-Bellman equation. In this light, computational error in PEV manifests as an extra error term in each iteration of Newton's method, with its upper bound proportional to the computational error. Further, we demonstrate that when the utility function resides in a reproducing kernel Hilbert space (RKHS), the optimal quadrature is achievable by employing Bayesian quadrature with the RKHS-inducing kernel function. We prove that the local convergence rates for IntRL using the trapezoidal rule and Bayesian quadrature with a Matérn kernel to be $O(N^{-2})$ and $O(N^{-b})$, where $N$ is the number of evenly-spaced samples and $b$ is the Matérn kernel's smoothness parameter. These theoretical findings are finally validated by two canonical control tasks. Wenhan Cao, Wei Pan 0004 |
ICLR | 2 |
| 2024 | Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningabstractIn offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy regularization. However, these methods often suffer from the issue of unnecessary conservativeness, hampering policy improvement. This occurs due to the indiscriminate use of all actions from the behavior policy that generates the offline dataset as constraints. The problem becomes particularly noticeable when the quality of the dataset is suboptimal. Thus, we propose Adaptive Advantage-guided Policy Regularization (A2PR), obtaining high-advantage actions from an augmented behavior policy combined with VAE to guide the learned policy. A2PR can select high-advantage actions that differ from those present in the dataset, while still effectively maintaining conservatism from OOD actions. This is achieved by harnessing the VAE capacity to generate samples matching the distribution of the data points. We theoretically prove that the improvement of the behavior policy is guaranteed. Besides, it effectively mitigates value overestimation with a bounded performance gap. Empirically, we conduct a series of experiments on the D4RL benchmark, where A2PR demonstrates state-of-the-art performance. Furthermore, experimental results on additional suboptimal mixed datasets reveal that A2PR exhibits superior performance. Code is available at https://github.com/ltlhuuu/A2PR. Tenglong Liu, Yang Li 0116, Yixing Lan, Hao Gao 0014, Wei Pan 0004, Xin Xu 0001 |
ICML | 5 |
| 2024 | Open Ad Hoc Teamwork with Cooperative Game TheoryabstractAd hoc teamwork poses a challenging problem, requiring the design of an agent to collaborate with teammates without prior coordination or joint training. Open ad hoc teamwork (OAHT) further complicates this challenge by considering environments with a changing number of teammates, referred to as open teams. One promising solution in practice to this problem is leveraging the generalizability of graph neural networks to handle an unrestricted number of agents with various agent-types, named graph-based policy learning (GPL). However, its joint Q-value representation over a coordination graph lacks convincing explanations. In this paper, we establish a new theory to understand the representation of the joint Q-value for OAHT and its learning paradigm, through the lens of cooperative game theory. Building on our theory, we propose a novel algorithm named CIAO, based on GPL's framework, with additional provable implementation tricks that can facilitate learning. The demos of experimental results are available on https://sites.google.com/view/ciao2024, and the code of experiments is published on https://github.com/hsvgbkhgbv/CIAO. Yang Li 0116, Yuan Zhang 0027, Wei Pan 0004, Samuel Kaski |
ICML | 4 |
| 2024 | Language and Sketching: An LLM-driven Interactive Multimodal Multitask Robot Navigation FrameworkabstractThe socially-aware navigation system has evolved to adeptly avoid various obstacles while performing multiple tasks, such as point-to-point navigation, human-following, and -guiding. However, a prominent gap persists: in Human-Robot Interaction (HRI), the procedure of communicating commands to robots demands intricate mathematical formulations. Furthermore, the transition between tasks does not quite possess the intuitive control and user-centric interactivity that one would desire. In this work, we propose an LLM-driven interactive multimodal multitask robot navigation framework, termed LIM2N, to solve the above new challenge in the navigation field. We achieve this by first introducing a multimodal interaction framework where language and hand-drawn inputs can serve as navigation constraints and control objectives. Next, a reinforcement learning agent is built to handle multiple tasks with the received information. Crucially, LIM2N creates smooth cooperation among the reasoning of multimodal input, multitask planning, and adaptation and processing of the intelligent sensing modules in the complicated system. Detailed experiments are conducted in both simulation and the real world demonstrating that LIM2N has solid user needs understanding, alongside an enhanced interactive experience. Weiqin Zu, Wenbin Song, Ruiqing Chen, Ze Guo, Fanglei Sun, Zheng Tian 0002, Wei Pan 0004, Jun Wang 0012 |
ICRA | 7 |
| 2024 | Aligning Individual and Collective Objectives in Multi-Agent CooperationabstractAmong the research topics in multi-agent learning, mixed-motive cooperation is one of the most prominent challenges, primarily due to the mismatch between individual and collective goals. The cutting-edge research is focused on incorporating domain knowledge into rewards and introducing additional mechanisms to incentivize cooperation. However, these approaches often face shortcomings such as the effort on manual design and the absence of theoretical groundings. To close this gap, we model the mixed-motive game as a differentiable game for the ease of illuminating the learning dynamics towards cooperation. More detailed, we introduce a novel optimization method named \textbf{\textit{A}}ltruistic \textbf{\textit{G}}radient \textbf{\textit{A}}djustment (\textbf{\textit{AgA}}) that employs gradient adjustments to progressively align individual and collective objectives. Furthermore, we theoretically prove that AgA effectively attracts gradients to stable fixed points of the collective objective while considering individual interests, and we validate these claims with empirical evidence. We evaluate the effectiveness of our algorithm AgA through benchmark environments for testing mixed-motive collaboration with small-scale agents such as the two-player public good game and the sequential social dilemma games, Cleanup and Harvest, as well as our self-developed large-scale environment in the game StarCraft II. Yang Li 0116, Shao Zhang, Yali Du 0001, Ying Wen 0001, Wei Pan 0004 |
NeurIPS | 7 |
| 2024 | Enhancing healthcare decision support through explainable AI models for risk predictionabstractElectronic health records (EHRs) are a valuable source of information that can aid in understanding a patient’s health condition and making informed healthcare decisions. However, modelling longitudinal EHRs with heterogeneous information is a challenging task. Although recurrent neural networks (RNNs), which are current artificial intelligence (AI) models, have the capability to capture longitudinal information, their explanatory power is limited. Predictive clustering is a recent development in this field, which provides cluster-level explainable evidence for disease risk prediction. Nonetheless, the challenge of determining the optimal number of clusters has put a brake on the widespread application of predictive clustering for disease risk prediction. In this paper, we introduce a novel non-parametric predictive clustering-based risk prediction model that integrates the Dirichlet Process Mixture Model (DPMM) with predictive clustering via neural networks. To enhance the model’s interpretability, we integrate attention mechanisms that enable the capture of local-level evidence in addition to the cluster-level evidence provided by predictive clustering. The outcome of this research is the development of a multi-level explainable artificial intelligence (AI) model. We evaluated the proposed model on two real-world datasets and demonstrated its effectiveness in capturing longitudinal EHR information for disease risk prediction. Additionally, the model was successful in generating explainable evidence to support its predictions. Qing Yin, Jing Ma 0004, Yunya Song, Liang Bai 0001, Wei Pan 0004, Xian Yang 0001 |
Decis. Support Syst. | 7 |
| 2024 | Tackling Cooperative Incompatibility for Zero-Shot Human-AI CoordinationabstractSecuring coordination between AI agent and teammates (human players or AI agents) in contexts involving unfamiliar humans continues to pose a significant challenge in Zero-Shot Coordination. The issue of cooperative incompatibility becomes particularly prominent when an AI agent is unsuccessful in synchronizing with certain previously unknown partners. Traditional algorithms have aimed to collaborate with partners by optimizing fixed objectives within a population, fostering diversity in strategies and behaviors. However, these techniques may lead to learning loss and an inability to cooperate with specific strategies within the population, a phenomenon named cooperative incompatibility in learning. In order to solve cooperative incompatibility in learning and effectively address the problem in the context of ZSC, we introduce the Cooperative Open-ended LEarning (COLE) framework, which formulates open-ended objectives in cooperative games with two players using perspectives of graph theory to evaluate and pinpoint the cooperative capacity of each strategy. We present two practical algorithms, specifically COLESV and COLER, which incorporate insights from game theory and graph theory. We also show that COLE could effectively overcome the cooperative incompatibility from theoretical and empirical analysis. Subsequently, we created an online Overcooked human-AI experiment platform, the COLE platform, which enables easy customization of questionnaires, model weights, and other aspects. Utilizing the COLE platform, we enlist 130 participants for human experiments. Our findings reveal a preference for our approach over state-of-the-art methods using a variety of subjective metrics. Moreover, objective experimental outcomes in the Overcooked game environment indicate that our method surpasses existing ones when coordinating with previously unencountered AI agents and the human proxy model. Our code and demo are publicly available at https://sites.google.com/view/cole-2023. Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004 |
J. Artif. Intell. Res. | 8 |
| 2024 | Cross-Utterance Conditioned VAE for Speech GenerationabstractSpeech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned Variational Autoencoder speech synthesis (CUC-VAE S2) framework to enhance prosody and ensure natural speech generation. This framework leverages the powerful representational capabilities of pre-trained language models and the re-expression abilities of variational autoencoders (VAEs). The core component of the CUC-VAE S2 framework is the cross-utterance CVAE, which extracts acoustic, speaker, and textual features from surrounding sentences to generate context-sensitive prosodic features, more accurately emulating human prosody generation. We further propose two practical algorithms tailored for distinct speech synthesis applications: CUC-VAE TTS for text-to-speech and CUC-VAE SE for speech editing. The CUC-VAE TTS is a direct application of the framework, designed to generate audio with contextual prosody derived from surrounding texts. On the other hand, the CUC-VAE SE algorithm leverages real mel spectrogram sampling conditioned on contextual information, producing audio that closely mirrors real sound and thereby facilitating flexible speech editing based on text such as deletion, insertion, and replacement. Experimental results on the LibriTTS datasets demonstrate that our proposed models significantly enhance speech synthesis and editing, producing more natural and expressive speech. Yang Li 0116, Guangzhi Sun, Weiqin Zu, Zheng Tian 0002, Ying Wen 0001, Wei Pan 0004, Chao Zhang 0031, Jun Wang 0012, Yang Yang 0001, Fanglei Sun |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2024 | Learning-Based Multi-UAV Flocking Control With Limited Visual Field and Instinctive RepulsionabstractThis article explores deep reinforcement learning (DRL) for the flocking control of unmanned aerial vehicle (UAV) swarms. The flocking control policy is trained using a centralized-learning-decentralized-execution (CTDE) paradigm, where a centralized critic network augmented with additional information about the entire UAV swarm is utilized to improve learning efficiency. Instead of learning inter-UAV collision avoidance capabilities, a repulsion function is encoded as an inner-UAV "instinct." In addition, the UAVs can obtain the states of other UAVs through onboard sensors in communication-denied environments, and the impact of varying visual fields on flocking control is analyzed. Through extensive simulations, it is shown that the proposed policy with the repulsion function and limited visual field has a success rate of 93.8% in training environments, 85.6% in environments with a high number of UAVs, 91.2% in environments with a high number of obstacles, and 82.2% in environments with dynamic obstacles. Furthermore, the results indicate that the proposed learning-based methods are more suitable than traditional methods in cluttered environments. Chengchao Bai, Peng Yan 0005, Haiyin Piao, Wei Pan 0004, Jifeng Guo 0004 |
IEEE Trans. Cybern. | 4 |
| 2024 | Data-Enabled Tire-Road Friction Estimation Based on Explainable Dynamics Mechanism Under Straight Stationary Driving ManeuversabstractThe tire-road friction coefficient (TRFC) is the critical parameter that significantly improves the control performance of distributed electric vehicles. Nonetheless, achieving precise TRFC estimation during straight stationary driving maneuvers, characterized by constant longitudinal speed (e.g., where the longitudinal acceleration is nearly zero) on a straight road, poses a particularly formidable challenge. In the paper, we propose a new learning strategy that leverages multi-domain fusion feature extraction in both the time domain and time-frequency domain to estimate the TRFC during straight stationary driving maneuvers. Specifically, the frequency response function of the in-wheel-motor-drive system first is inferred from the longitudinal dynamics model and single wheel dynamics model. Then, the input selection of learning strategy is determined through frequency response characteristics analysis and explainable dynamics mechanism. In addition, a parallel spatial-temporal convolutional neural network (PSTCNN) is built to extract features in both the time domain and in the time-frequency domain, respectively. Finally, the TRFC learning strategy is verified by experimental tests on different road surfaces. Our results demonstrate that the proposed methodology is capable of estimating the TRFC with a lower error than the traditional learning-based method and the classical slip-slope method. Zhaobo Qin, Manjiang Hu, Yougang Bian, Wei Pan 0004 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Cooperative Open-ended Learning Framework for Zero-Shot CoordinationabstractZero-shot coordination in cooperative artificial intelligence (AI) remains a significant challenge, which means effectively coordinating with a wide range of unseen partners. Previous algorithms have attempted to address this challenge by optimizing fixed objectives within a population to improve strategy or behaviour diversity. However, these approaches can result in a loss of learning and an inability to cooperate with certain strategies within the population, known as cooperative incompatibility. To address this issue, we propose the Cooperative Open-ended LEarning (COLE) framework, which constructs open-ended objectives in cooperative games with two players from the perspective of graph theory to assess and identify the cooperative ability of each strategy. We further specify the framework and propose a practical algorithm that leverages knowledge from game theory and graph theory. Furthermore, an analysis of the learning process of the algorithm shows that it can efficiently overcome cooperative incompatibility. The experimental results in the Overcooked game environment demonstrate that our method outperforms current state-of-the-art methods when coordinating with different-level partners. Our demo is available at https://sites.google.com/view/cole-2023. Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004 |
ICML | 7 |
| 2023 | Reinforcement Learning for Safe Robot Control using Control Lyapunov Barrier FunctionsabstractReinforcement learning (RL) exhibits impressive performance when managing complicated control tasks for robots. However, its wide application to physical robots is limited by the absence of strong safety guarantees. To overcome this challenge, this paper explores the control Lyapunov barrier function (CLBF) to analyze the safety and reachability solely based on data without explicitly employing a dynamic model. We also proposed the Lyapunov barrier actor-critic (LBAC), a model-free RL algorithm, to search for a controller that satisfies the data-based approximation of the safety and reachability conditions. The proposed approach is demonstrated through simulation and real-world robot control experiments, i.e., a 2D quadrotor navigation task. The experimental findings reveal this approach's effectiveness in reachability and safety, surpassing other model-free RL methods. Desong Du, Shaohang Han, Naiming Qi, Haitham Bou-Ammar, Jun Wang 0012, Wei Pan 0004 |
ICRA | 6 |
| 2023 | Sim-and-Real Reinforcement Learning for Manipulation: A Consensus-based ApproachabstractSim-and-real training is a promising alternative to sim-to-real training for robot manipulations. However, the current sim-and-real training is neither efficient, i.e., slow con-vergence to the optimal policy, nor effective, i.e., sizeable real-world robot data. Given limited time and hardware budgets, the performance of sim-and-real training is not satisfactory. In this paper, we propose a Consensus-based Sim-And-Real deep reinforcement learning algorithm (CSAR) for manipulator pick-and-place tasks, which shows comparable performance in both sim-and- real worlds. In this algorithm, we train the agents in simulators and the real world to get the optimal policies for both sim-and-real worlds. We found two interesting phenomenons: (1) Best policy in simulation is not the best for sim-and-real training. (2) The more simulation agents, the better sim-and-real training. The experimental video is available at: https://youtu.be/mcHJtNIsTEQ. Hanlin Niu, Wei Pan 0004, Guido Herrmann, Joaquín Carrasco |
ICRA | 3 |
| 2023 | Maintaining Visibility of Dynamic Objects in Cluttered Environments Using Mobile Manipulators and Vector Field InequalitiesabstractVision-based perception has become prevalent in robotic applications, especially in those where the control loop relies on visual data, such as visual servoing. For those applications, ensuring that the features or target object remain visible to the camera is critical, necessitating visibility-aware control. In this paper, we propose a method to guarantee the visibility of a dynamic object using a constrained kinematic controller and Vector Field Inequalities (VFIs) to include a linear visibility constraint. Unlike existing methods, we introduce constraints into the kinematic controller to ensure the target's visibility without needing a trajectory optimizer or local planner. Our method maintains the target object in the camera field of view (FoV) by representing the FoV with four infinite planes and maintaining the distance between the target object and each plane higher than a predefined distance. We evaluated the proposed approach using a mobile manipulator in two simulations involving cluttered environments: the first scenario involves a stationary target object, whereas the second scenario presents a more challenging workspace involving a moving target. Our results demonstrate that the proposed approach successfully maintains the target within the FoV while avoiding obstacles in the workspace, showing the potential of our method to improve the safety and reliability of visual-servoing-based robotic systems. Fatih Dursun, Bruno Vilhena Adorno, Simon Watson 0001, Wei Pan 0004 |
IROS | 4 |
| 2023 | A Secure Robot Learning Framework for Cyber Attack Scheduling and CountermeasureabstractThe problem of learning-based control for robots has been extensively studied, whereas the security issue under malicious adversaries has not been paid much attention to. Malicious adversaries can invade intelligent devices and communication networks used in robots, causing incidents, achieving illegal objectives, and even injuring people. This article first investigates the problems of optimal false data injection attack scheduling and countermeasure design for car-like robots in the framework of deep reinforcement learning. Using a state-of-the-art deep reinforcement learning approach, an optimal false data injection attack scheme is proposed to deteriorate the tracking performance of a robot, guaranteeing the tradeoff between the attack efficiency and the limited attack energy. Then, an optimal tracking control strategy is learned to mitigate attacks and recover the tracking performance. More importantly, a theoretical stability guarantee of a robot using the learning-based secure control scheme is achieved. Both simulated and real-world experiments are conducted to show the effectiveness of the proposed schemes. Chengwei Wu 0001, Weiran Yao, Wensheng Luo 0001, Wei Pan 0004, Guanghui Sun, Hui Xie 0003, Ligang Wu 0001 |
IEEE Trans. Robotics | 4 |
| 2022 | Barrier Function-based Safe Reinforcement Learning for Formation Control of Mobile RobotsabstractDistributed model predictive control (DMPC) concerns how to online control multiple robotic systems with constraints effectively. However, the nonlinearity, nonconvexity, and strong interconnections of dynamic system models and constraints can make the real-time and real-world DMPC implementations nontrivial. Reinforcement learning (RL) algorithms are promising for control policy design. However, how to ensure safety in terms of state constraints in RL remains a significant issue. This paper proposes a barrier function-based safe reinforcement learning algorithm for DMPC of nonlinear multi-robot systems under state constraints. The proposed approach is composed of several local learning-based MPC regulators. Each regulator, associated with a local system, learns and deploys the local control policy using a safe reinforcement learning algorithm in a distributed manner, i.e., with state information only among the neighbor agents. As a prominent feature of the proposed algorithm, we present a novel barrier-based policy structure to ensure safety, which has a clear mechanistic interpretation. Both simulated and real-world experiments on the formation control of mobile robots with collision avoidance show the effectiveness of the proposed safe reinforcement learning algorithm for DMPC. Yaoqian Peng, Wei Pan 0004, Xin Xu 0001, Haibin Xie |
ICRA | 3 |
| 2022 | Learning-Based Multi-Robot Formation Control With Obstacle AvoidanceabstractMulti-robot formation control has been intensively studied in recent years. In practical applications, the multi-robot system’s ability to independently change the formation to avoid collision among the robots or with obstacles is critical. In this study, a multi-robot adaptive formation control framework based on deep reinforcement learning is proposed. The framework consists of two layers, namely the execution layer and the decision-making layer. The execution layer enables the robot to approach its target position and avoid collision with other robots and obstacles through a deep network trained by a reinforcement learning method. The decision-making layer organizes all robots into a formation through a new leader–follower configuration and provides target positions to the leader and followers. The leader’s target position is kept unchanged, while the follower’s target position is changed according to the situation it encounters. In addition, to operate more effectively in environments with different levels of complexity, a hybrid switching control strategy is proposed. The simulation results demonstrate that our proposed formation control framework enables the robots to adjust formation independently to pass through obstacle areas and can be generalized to different scenarios with unknown obstacles and varying number of robots. Chengchao Bai, Peng Yan 0005, Wei Pan 0004, Jifeng Guo 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Model-Reference Reinforcement Learning for Collision-Free Tracking Control of Autonomous Surface VehiclesabstractThis paper presents a novel model-reference reinforcement learning algorithm for the intelligent tracking control of uncertain autonomous surface vehicles with collision avoidance. The proposed control algorithm combines a conventional control method with reinforcement learning to enhance control accuracy and intelligence. In the proposed control design, a nominal system is considered for the design of a baseline tracking controller using a conventional control approach. The nominal system also defines the desired behaviour of uncertain autonomous surface vehicles in an obstacle-free environment. Thanks to reinforcement learning, the overall tracking controller is capable of compensating for model uncertainties and achieving collision avoidance at the same time in environments with obstacles. In comparison to traditional deep reinforcement learning methods, our proposed learning-based control can provide stability guarantees and better sample efficiency. We demonstrate the performance of the new algorithm using an example of autonomous surface vehicles. Qingrui Zhang, Wei Pan 0004, Vasso Reppa |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Reinforcement Learning for Orientation Estimation Using Inertial Sensors with Performance GuaranteeabstractThis paper presents a deep reinforcement learning (DRL) algorithm for orientation estimation using inertial sensors combined with a magnetometer. Lyapunov’s method in control theory is employed to prove the convergence of orientation estimation errors. The estimator gains and a Lyapunov function are parametrised by deep neural networks and learned from samples based on the theoretical results. The DRL estimator is compared with three well-known orientation estimation methods on both numerical simulations and real dataset collected from commercially available sensors. The results show that the proposed algorithm is superior for arbitrary estimation initialisation and can adapt to a drastic angular velocity profile for which other algorithms can be hardly applicable. To the best of our knowledge, this is the first DRL-based orientation estimation method with an estimation error boundedness guarantee. Liang Hu 0002, Yujie Tang 0004, Wei Pan 0004 |
ICRA | 4 |
| 2021 | Reinforcement Learning Compensated Extended Kalman Filter for Attitude EstimationabstractInertial measurement units are widely used in different fields to estimate the attitude. Many algorithms have been proposed to improve estimation performance. However, most of them still suffer from 1) inaccurate initial estimation, 2) inaccurate initial filter gain, and 3) non-Gaussian process and/or measurement noise. This paper will leverage reinforcement learning to compensate for the classical extended Kalman filter estimation, i.e., to learn the filter gain from the sensor measurements. We also analyse the convergence of the estimate error. The effectiveness of the proposed algorithm is validated on both simulated data and real data. Yujie Tang 0004, Liang Hu 0002, Qingrui Zhang, Wei Pan 0004 |
IROS | 4 |
| 2021 | Learning Tracking Control for Cyber-Physical SystemsabstractThis article investigates the problem of optimal tracking control for cyber-physical systems (CPSs) when the cyber realm is attacked by Denial-of-Service (DoS) attacks which can prevent the control signal transmitting to the actuator. Attention is focused on how to design the optimal tracking control scheme without using the system dynamics and analyze the impact of DoS attacks on tracking performance. First, a Riccati equation for the augmented system, including the system model and the reference model is derived under the framework of dynamic programming. The existence and uniqueness of its solution are proved. Second, the impact of the successful DoS attack probability on tracking performance is analyzed. A critical value of the probability is given, beyond which the solution to the Riccati equation cannot converge. The tracking controller cannot be designed. Third, reinforcement learning is introduced to design the optimal tracking control schemes, in which the system dynamics are not necessary to be known. Finally, both a dc motor and an F16 aircraft are used to evaluate the proposed control schemes in this article. Chengwei Wu 0001, Wei Pan 0004, Guanghui Sun, Jianxing Liu, Ligang Wu 0001 |
IEEE Internet Things J. | 2 |
| 2020 | Lyapunov-Based Reinforcement Learning for Decentralized Multi-agent Control
Qingrui Zhang, Hao Dong 0003, Wei Pan 0004 |
DAI | 3 |
| 2020 | Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
Baozhou Zhu, Zaid Al-Ars, Wei Pan 0004 |
ECAI | 3 |
| 2019 | Probabilistic Recursive Reasoning for Multi-Agent Reinforcement Learning
Ying Wen 0001, Yaodong Yang 0001, Rui Luo 0001, Jun Wang 0012, Wei Pan 0004 |
ICLR (Poster) | 5 |
| 2019 | BayesNAS: A Bayesian Approach for Neural Architecture SearchabstractOne-Shot Neural Architecture Search (NAS) is a promising method to significantly reduce search time without any separate training. It can be treated as a Network Compression problem on the architecture parameters from an over-parameterized network. However, there are two issues associated with most one-shot NAS methods. First, dependencies between a node and its predecessors and successors are often disregarded which result in improper treatment over zero operations. Second, architecture parameters pruning based on their magnitude is questionable. In this paper, we employ the classic Bayesian learning approach to alleviate these two issues by modeling architecture parameters using hierarchical automatic relevance determination (HARD) priors. Unlike other NAS methods, we train the over-parameterized network for only one epoch then update the architecture. Impressively, this enabled us to find the architecture in both proxy and proxyless tasks on CIFAR-10 within only 0.2 GPU days using a single GPU. As a byproduct, our approach can be transferred directly to compress convolutional neural networks by enforcing structural sparsity which achieves extremely sparse networks without accuracy deterioration. Hongpeng Zhou, Jun Wang 0012, Wei Pan 0004 |
ICML | 4 |