EDBT 2026 Demo / reviewers in the wild / expert
Xiaoming Duan
dblp:131/1732
· DBLP profile ↗
12ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-2655-3987ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Online Informative Motion Planning for Active Information Gathering of a Non-Stationary Gaussian ProcessabstractInformation gathering focuses on designing strategies for a robot to collect data about a physical process, aiming for accurate field reconstruction. While many recent methods have been proposed to address this problem, they often assume the model of the physical process is a priori known and stationary-assumptions that rarely hold in practice. This paper presents a novel informative motion planning approach for online information gathering of a non-stationary Gaussian process. Our approach comprises two key components: an informative path planner that explores the physical field and an adaptive velocity planner that adjusts the robot's velocity profile exploiting the field's spatial variability. Additionally, we propose a path smoothing and tracking strategy to ensure continuous robot motion. Extensive simulations on a bathymetric mapping task demonstrate the effectiveness of our approach, showing superior performance in reconstructing non-stationary physical fields compared to several baseline methods. Kexiang Mao, Jianping He 0001, Xiaoming Duan |
ICRA | 3 |
| 2025 | CDPMM-DMP: Conditional Dirichlet Process Mixture Model-Based Dynamic Movement PrimitivesabstractMovement Primitives (MPs) are compact generators for representation and generalization of modular movements, which are usually used to implement learning from demonstration tasks in robotics. Existing works on MPs mostly utilize combinations of basis functions to represent diverse movements, whether employing probabilistic or dynamic approaches. However, applying these approaches requires manual specification of hyperparameters related to basis functions, resulting in inconvenience and a reliance on specific expertise. In this paper, we develop a Conditional Dirichlet Process Mixture Model-based Dynamic Movement Primitive (CDPMM-DMP) to achieve a non-parametric improvement for the Dynamic Movement Primitive (DMP). First, inspired by Bayesian nonparametric theory, we explore the use of the Dirichlet Process Mixture Model (DPMM) to replace the original radial basis functions in the DMP, and construct the required training set from demonstrations. Then, we study the output generation mechanism driven by the DPMM, particularly by employing conditional sampling to avoid the anomalous outputs caused by direct sampling from the DPMM. Finally, we provide analyses of the various properties brought by our nonparametric transformation of DMP. The analyses and validation results show that the proposed CDPMM-DMP can significantly reduce the parameter tuning burden in usage with its nonparametric learning property. Besides, our method still retains the inherent properties of DMP, while also incorporating some properties of probabilistic MPs, such as multi-sample learning and co-activation. Hao Jiang 0027, Jianping He 0001, Xiaoming Duan |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Learning-Based Motion Planning with Mixture Density NetworksabstractThe trade-off between computation time and path optimality is a key consideration in motion planning algorithms. While classical sampling based algorithms fall short of computational efficiency in high dimensional planning, learning based methods have shown great potential in achieving time efficient and optimal motion planning. The SOTA learning based motion planning algorithms utilize paths generated by sampling based methods as expert supervision data and train networks via regression techniques. However, these methods often overlook the important multimodal property of the optimal paths in the training set, making them incapable of finding good paths in some scenarios. In this paper, we propose a Multimodal Neuron Planner (MNP) based on the mixture density networks that explicitly takes into account the multimodality of the training data and simultaneously achieves time efficiency and path optimality. For environments represented by point clouds, MNP first efficiently compresses point clouds into a latent vector by encoding networks that are suitable for processing point clouds. We then design multimodal planning networks which enables MNP to learn and predict multiple optimal solutions. Simulation results show that our method outperforms SOTA learning based method MPNet and advanced sampling based methods IRRT* and BIT*. Yinghan Wang, Xiaoming Duan, Jianping He 0001 |
ICRA | 2 |
| 2024 | Collaboration Strategies for Two Heterogeneous Pursuers in A Pursuit-Evasion Game Using Deep Reinforcement LearningabstractWe investigate a pursuit-evasion game taking place in an unbounded three-dimensional space, where a flexible pursuer with hybrid dynamics collaborates with a fast pursuer and aims to capture a flexible evader within a finite time. The key feature of this problem lies in the hybrid dynamics of the flexible pursuer, which can change its dynamics once during the game and switch to a fast pursuer with increased speed but lower maneuverability. To address this challenge, we devise a hybrid strategy based on the soft actor-critic framework, tailored specifically for the flexible pursuer, which encompasses both maneuvering and switch tactics. We introduce a switch factor to the input of the actor network and incorporate switch actions to further expand the action space. These additions enable the flexible pursuer to execute maneuvering actions and determine a moment to switch to a fast pursuer. The reward function is designed to account for related angle, altitude, speed, and sparse reward. Through extensive ablation experiments conducted in a simulated environment, we demonstrate the efficacy of our algorithm in facilitating the learning of hybrid strategies for the flexible pursuer, resulting in significantly improved capture rates compared to alternative methods. Zhanping Zhong, Zhuoning Dong, Xiaoming Duan, Jianping He 0001 |
IROS | 3 |
| 2024 | Differentially Private No-regret Exploration in Adversarial Markov Decision ProcessesabstractWe study learning adversarial Markov decision process (MDP) in the episodic setting under the constraint of differential privacy (DP). This is motivated by the widespread applications of reinforcement learning (RL) in non-stationary and even adversarial scenarios, where protecting users’ sensitive information is vital. We first propose two efficient frameworks for adversarial MDPs, spanning full-information and bandit settings. Within each framework, we consider both Joint DP (JDP), where a central agent is trusted to protect the sensitive data, and Local DP (LDP), where the information is protected directly on the user side. Then, we design novel privacy mechanisms to privatize the stochastic transition and adversarial losses. By instantiating such privacy mechanisms to satisfy JDP and LDP requirements, we obtain near-optimal regret guarantees for both frameworks. To our knowledge, these are the first algorithms to tackle the challenge of private learning in adversarial MDPs. Shaojie Bai, Lanting Zeng, Chengcheng Zhao, Xiaoming Duan, Mohammad Sadegh Talebi, Peng Cheng 0001, Jiming Chen 0001 |
UAI | 4 |
| 2023 | Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
Yash Paliwal, Rajarshi Roy 0002, Jean-Raphaël Gaglione, Nasim Baharisangari, Daniel Neider, Xiaoming Duan, Ufuk Topcu, Zhe Xu 0005 |
CD-MAKE | 6 |
| 2023 | Balancing Efficiency and Unpredictability in Multi-robot Patrolling: A MARL-Based ApproachabstractPatrolling with multiple robots is a challenging task. While the robots collaboratively and repeatedly cover the regions of interest in the environment, their routes should satisfy two often conflicting properties: i) (efficiency) the time intervals between two consecutive visits to the regions are small; ii) (unpredictability) the patrolling trajectories are random and unpredictable. We manage to strike a balance between the two goals by i) recasting the original patrolling problem as a Graph Deep Learning problem; ii) directly solving this problem on the graph in the framework of cooperative multi-agent reinforcement learning. Treating the decisions of a team of agents as a sequence input, our model outputs the agents' actions in order by an autoregressive mechanism. Extensive simulation studies show that our approach has comparable performance with existing algorithms in terms of efficiency and outperforms them in terms of unpredictability. To our knowledge, this is the first work that successfully solves the patrolling problem with reinforcement learning on a graph. Lingxiao Guo, Haoxuan Pan, Xiaoming Duan, Jianping He 0001 |
ICRA | 3 |
| 2023 | Evaluation and learning in two-player symmetric games via best and better responses
Rui Yan 0002, Weixian Zhang, Ruiliang Deng, Xiaoming Duan, Zongying Shi, Yisheng Zhong |
Inf. Sci. | 4 |
| 2022 | Finite-horizon equilibria for neuro-symbolic concurrent stochastic gamesabstractWe present novel techniques for neuro-symbolic concurrent stochastic games, a recently proposed modelling formalism to represent a set of probabilistic agents operating in a continuous-space environment using a combination of neural network based perception mechanisms and traditional symbolic methods. To date, only zero-sum variants of the model were studied, which is too restrictive when agents have distinct objectives. We formalise notions of equilibria for these models and present algorithms to synthesise them. Focusing on the finite-horizon setting, and (global) social welfare subgame-perfect optimality, we consider two distinct types: Nash equilibria and correlated equilibria. We first show that an exact solution based on backward induction may yield arbitrarily bad equilibria. We then propose an approximation algorithm called frozen subgame improvement, which proceeds through iterative solution of nonlinear programs. We develop a prototype implementation and demonstrate the benefits of our approach on two case studies: an automated car-parking system and an aircraft collision avoidance system. Rui Yan 0002, Gabriel Santos, Xiaoming Duan, David Parker 0001, Marta Z. Kwiatkowska |
UAI | 3 |
| 2021 | Multi-robot Target Search under Multi-peak Distribution: A Dynamic Approach based on High Confidence AreaabstractTarget search with multiple robots has attracted widespread attention for its numerous applications,$eg$., surveillance, reconnaissance and environmental exploration. In this paper, we propose a heuristic target search scheme for multi-robot, considering the prior information of targets is unreliable. To tactfully capture the multi-peak characteristics of the probability distribution map (PDM) of each target, we introduce the concept of high confidence area (HCA) based on the Gaussian mixture model. Then, a coordinated search method consisting of task allocation and path planning is designed to achieve efficient search performance. The main novelty of our method is twofold. First, the probability information of multi-peak is sufficiently captured by HCA and evaluated by reliability degree. Second, target allocation and path planning are designed coordinately, which dynamically update with real-time status and alleviate misleading effects even when PDM is not reliable, thereby largely reducing search time. Extensive contrastive simulations demonstrate that the HCA-based search method outperforms two other existing methods. Qing Jiao, Yushan Li 0001, Xiaoming Duan, Jianping He 0001, Qing-Guo Wang |
VTC Fall | 3 |
| 2020 | Robotic Surveillance Based on the Meeting Time of Random WalksabstractThis article analyzes the meeting time between a pair of pursuer and evader performing random walks on digraphs. The existing bounds on the meeting time usually work only for certain classes of walks and cannot be used to formulate optimization problems and design robotic strategies. First, by analyzing multiple random walks on a common graph as a single random walk on the Kronecker product graph, we provide the first closed-form expression for the expected meeting time in terms of the transition matrices of the moving agents. This novel expression leads to necessary and sufficient conditions for the meeting time to be finite and to insightful graph-theoretic interpretations. Second, based on the closed-form expression, we set up and study the minimization problem for the expected capture time for a pursuer/evader pair. We report theoretical and numerical results on basic case studies to show the effectiveness of the design. Xiaoming Duan, Mishel George, Rushabh Patel, Francesco Bullo |
IEEE Trans. Robotics | 1 |
| 2020 | Analysis of Consensus-Based Economic Dispatch Algorithm Under Time DelaysabstractUnder consensus-based economic dispatch (ED) algorithm, multiple agents, which control local generation units, cooperatively minimize the total generation cost subject to the balance of the generation and expected demand in smart grids. As ubiquitous time delays on communication links exist in communication networks, studying the effect of delays on the dispatch performance is of both theoretical merit and practical value for the efficient and stable operation of smart grids. In this paper, we consider a well-developed consensus-based ED protocol under constant time delays. We find that there always exists a sufficiently small learning gain parameter under finite constant delays such that the convergence of the consensus-based algorithm is guaranteed. Further, an analytical expression of the upper bound is established for the learning gain parameter, which is determined by the largest delay, the weight matrix and the parameters of generation cost functions. In order to guarantee the optimality of the final solution, we propose the updating rule for iterations when initial states are not received by their neighbors due to time delays. The optimality of the final solution under the proposed updating rule is analyzed. We validate our theoretical results through extensive simulation studies. Chengcheng Zhao, Xiaoming Duan, Yang Shi 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |