VLDB 2026 Research / reviewers in the wild / expert
Zhe Xu 0005
dblp:97/3701-5
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-0440-0912ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Logic-Based Approach to Causal Discovery: Signal Temporal Logic PerspectiveabstractCausal discovery in time-series datasets is critical for understanding complex systems, especially when the \textit{effectiveness} of causal relationships depends on both the \textit{duration} and \textit{magnitude} of the cause. We introduce a novel framework for causal discovery based on \textbf{Signal Temporal Logic (STL)}, enabling the extraction of interpretable causal diagrams (STL-CD) that explicitly capture these temporal dynamics. Our method first identifies statistically meaningful time intervals, then infers STL formulas that classify system behaviors, and finally employs transfer entropy to determine direct causal relationships among the formulas. This approach not only uncovers causal structure but also identifies the temporal persistence required for causal influence—an insight missed by existing methods. Experimental results on synthetic and real-world datasets demonstrate that our method achieves superior structural accuracy over state-of-the-art baselines, providing more informative and temporally precise causal models. Nasim Baharisangari, Yucheng Ruan, Chengcheng Zhao, Zhe Xu 0005 |
IJCAI | 4 |
| 2025 | Successor Features for Transfer in Alternating Markov GamesabstractThis paper explores successor features for knowledge transfer in zero-sum, complete-information, and turn-based games. Prior research in single-agent systems has shown that successor features can provide a "jump start" for agents when facing new tasks with varying reward structures. However, knowledge transfer in games typically relies on value and equilibrium transfers, which heavily depends on the similarity between tasks. This reliance can lead to failures when the tasks differ significantly. To address this issue, this paper presents an application of successor features to games and presents a novel algorithm called Game Generalized Policy Improvement (GGPI), designed to address Markov games in multi-agent reinforcement learning. The proposed algorithm enables the transfer of learning values and policies across games. An upper bound of the errors for transfer is derived as a function the similarity of the task. Through experiments with a turn-based pursuer-evader game, we demonstrate that the GGPI algorithm can generate high-reward interactions and one-shot policy transfer. When further tested in a wider set of initial conditions, the GGPI algorithm achieves higher success rates with improved path efficiency compared to those of the baseline algorithms. Sunny Amatya, Zhe Xu 0005 |
IROS | 3 |
| 2025 | Decentralizing Multi-agent Reinforcement Learning with Temporal Causal Information
Jan Corazza, Hadi Partovi Aria, Hyohun Kim, Daniel Neider, Zhe Xu 0005 |
ECML/PKDD (6) | 5 |
| 2025 | Non-Parametric Neuro-Adaptive Formation ControlabstractWe develop a learning-based algorithm for the distributed formation control of networked multi-agent systems governed by unknown, nonlinear dynamics. Most existing algorithms either assume certain parametric forms for the unknown dynamic terms or resort to unnecessarily large control inputs in order to provide theoretical guarantees. The proposed algorithm avoids these drawbacks by integrating neural network-based learning with adaptive control in a two-step procedure. In the first step of the algorithm, each agent learns a controller, represented as a neural network, using training data that correspond to a collection of formation tasks and agent parameters. These parameters and tasks are derived by varying the nominal agent parameters and a user-defined formation task to be achieved, respectively. In the second step of the algorithm, each agent incorporates the trained neural network into an online and adaptive control policy in such a way that the behavior of the multi-agent closed-loop system satisfies the user-defined formation task. Both the learning phase and the adaptive control policy are distributed, in the sense that each agent computes its own actions using only local information from its neighboring agents. The proposed algorithm does not use any a priori information on the agents’ unknown dynamic terms or any approximation schemes. We provide formal theoretical guarantees on the achievement of the formation task. Note to Practitioners—This paper is motivated by control of multi-agent systems, such as teams of robots, smart grids, or wireless sensor networks, with uncertain dynamic models. Existing works develop controllers that rely on unrealistic or impractical assumptions on these models. We propose an algorithm that integrates offline learning with neural networks and real-time feedback control to accomplish a multi-agent task. The task consists of the formation of a pre-defined geometric pattern by the multi-agent team. The learning module of the proposed algorithm aims to learn stabilizing controllers that accomplish the task from data that are obtained from offline runs of the system. However, the learned controller might result in poor performance owing to potential data inaccuracies and the fact that learning algorithms can only approximate the stabilizing controllers. Therefore, we complement the learned controller with a real-time feedback-control module that adapts on the fly to such discrepancies. In practise, the data can be collected from pre-recorded trajectories of the multi-agent system, but these trajectories do need to accomplish the task at hand. The real-time feedback-control is a closed-form function of the states of each agent and its neighbours and the trained neural networks and can be straightforwardly implemented. The experimental results show that the proposed algorithm achieves greater performance than algorithms that use only the trained neural networks or only the real-time feedback-control policy. Our future research will address the sensitivity of the algorithm to the quality and quantity of the employed data as well as to the learning performance of the neural networks. Christos K. Verginis, Zhe Xu 0005, Ufuk Topcu |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | State-Constrained Zero-Sum Differential Games with One-Sided InformationabstractWe study zero-sum differential games with state constraints and one-sided information, where the informed player (Player 1) has a categorical payoff type unknown to the uninformed player (Player 2). The goal of Player 1 is to minimize his payoff without violating the constraints, while that of Player 2 is to violate the state constraints if possible, or to maximize the payoff otherwise. One example of the game is a man-to-man matchup in football. Without state constraints, Cardaliaguet (2007) showed that the value of such a game exists and is convex to the common belief of players. Our theoretical contribution is an extension of this result to games with state constraints and the derivation of the primal and dual subdynamic principles necessary for computing behavioral strategies. Different from existing works that are concerned about the scalability of no-regret learning in games with discrete dynamics, our study reveals the underlying structure of strategies for belief manipulation resulting from information asymmetry and state constraints. This structure will be necessary for scalable learning on games with continuous actions and long time windows. We use a simplified football game to demonstrate the utility of this work, where we reveal player positions and belief states in which the attacker should (or should not) play specific random deceptive moves to take advantage of information asymmetry, and compute how the defender should respond. Mukesh Ghimire, Zhe Xu 0005 |
ICML | 3 |
| 2024 | Decentralized graph-based multi-agent reinforcement learning using reward machines
Jueming Hu, Zhe Xu 0005, Weichang Wang, Guannan Qu, Yutian Pang, Yongming Liu |
Neurocomputing | 2 |
| 2024 | Value Approximation for Two-Player General-Sum Differential Games With State ConstraintsabstractSolving Hamilton–Jacobi–Isaacs (HJI) PDEs numerically enables equilibrial feedback control in two-player differential games, yet faces the curse of dimensionality (CoD). While physics-informed neural networks (PINNs) have shown promise in alleviating CoD in solving PDEs, vanilla PINNs fall short in learning discontinuous solutions due to their sampling nature, leading to poor safety performance of the resulting policies when values are discontinuous due to state or temporal logic constraints. In this study, we explore three potential solutions to this challenge: 1) a hybrid learning method that is guided by both supervisory equilibria and the HJI PDE, 2) a value-hardening method where a sequence of HJIs are solved with increasing Lipschitz constant on the constraint violation penalty, and 3) the epigraphical technique that lifts the value to a higher dimensional state space where it becomes continuous. Evaluations through 5-D and 9-D vehicle and 13-D drone simulations reveal that the hybrid method outperforms others in terms of generalization and safety performance by taking advantage of both the supervisory equilibrium values and co-states, and the low cost of PINN loss gradients. Mukesh Ghimire, Zhe Xu 0005 |
IEEE Trans. Robotics | 4 |
| 2023 | Learning Interpretable Temporal Properties from Positive Examples OnlyabstractWe consider the problem of explaining the temporal behavior of black-box systems using human-interpretable models. Following recent research trends, we rely on the fundamental yet interpretable models of deterministic finite automata (DFAs) and linear temporal logic (LTL_f) formulas. In contrast to most existing works for learning DFAs and LTL_f formulas, we consider learning from only positive examples. Our motivation is that negative examples are generally difficult to observe, in particular, from black-box systems. To learn meaningful models from positive examples only, we design algorithms that rely on conciseness and language minimality of models as regularizers. Our learning algorithms are based on two approaches: a symbolic and a counterexample-guided one. The symbolic approach exploits an efficient encoding of language minimality as a constraint satisfaction problem, whereas the counterexample-guided one relies on generating suitable negative examples to guide the learning. Both approaches provide us with effective algorithms with minimality guarantees on the learned models. To assess the effectiveness of our algorithms, we evaluate them on a few practical case studies. Rajarshi Roy 0002, Jean-Raphaël Gaglione, Nasim Baharisangari, Daniel Neider, Zhe Xu 0005, Ufuk Topcu |
AAAI | 5 |
| 2023 | Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
Yash Paliwal, Rajarshi Roy 0002, Jean-Raphaël Gaglione, Nasim Baharisangari, Daniel Neider, Xiaoming Duan, Ufuk Topcu, Zhe Xu 0005 |
CD-MAKE | 8 |
| 2023 | Reinforcement Learning with Reward Machines in Stochastic GamesabstractWe investigate multi-agent reinforcement learning for stochastic games with complex tasks, where the reward functions are non-Markovian. We utilize reward machines to incorporate high-level knowledge of complex tasks. We develop an algorithm called Q-learning with reward machines for stochastic games (QRM-SG), to learn the best-response strategy at Nash equilibrium for each agent. In QRM-SG, we define the Q-function at a Nash equilibrium in augmented state space. The augmented state space integrates the state of the stochastic game and the state of reward machines. Each agent learns the Q-functions of all agents in the system. We prove that Q-functions learned in QRM-SG converge to the Q-functions at a Nash equilibrium if the stage game at each time step during learning has a global optimum point or a saddle point, and the agents update Q-functions based on the best-response strategy at this point. We use the Lemke-Howson method to derive the best-response strategy given current Q-functions. The three case studies show that QRM-SG can learn the best-response strategies effectively. QRM-SG learns the best-response strategies after around 7500 episodes in Case Study I, 1000 episodes in Case Study II, and 1500 episodes in Case Study III, while baseline methods such as Nash Q-learning and MADDPG fail to converge to the Nash equilibrium in all three case studies. Jueming Hu, Jean-Raphaël Gaglione, Zhe Xu 0005, Ufuk Topcu, Yongming Liu |
ECAI | 4 |
| 2023 | Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential GamesabstractFinding Nash equilibrial policies for two-player differential games requires solving Hamilton-Jacobi-Isaacs (HJI) PDEs. Self-supervised learning has been used to approximate solutions of such PDEs while circumventing the curse of dimensionality. However, this method fails to learn discontinuous PDE solutions due to its sampling nature, leading to poor safety performance of the resulting controllers in robotics applications when player rewards are discontinuous. This paper investigates two potential solutions to this problem: a hybrid method that leverages both supervised Nash equilibria and the HJI PDE, and a value-hardening method where a sequence of HJIs are solved with a gradually hardening reward. We compare these solutions using the resulting generalization and safety performance in two vehicle interaction simulation studies with 5D and 9D state spaces, respectively. Results show that with informative supervision (e.g., collision and near-collision demonstrations) and the low cost of self-supervised learning, the hybrid method achieves better safety performance than the supervised, self-supervised, and value hardening approaches on equal computational budget. Value hardening fails to generalize in the higher-dimensional case without informative supervision. Lastly, we show that the neural activation function needs to be continuously differentiable for learning PDEs and its choice can be case dependent. Mukesh Ghimire, Zhe Xu 0005 |
ICRA | 4 |
| 2021 | Adaptive Teaching of Temporal Logic Formulas to Preference-based Learners
Zhe Xu 0005, Yuxin Chen 0001, Ufuk Topcu |
AAAI | 1 |
| 2021 | Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentabstractWe study a class of reinforcement learning tasks in which the agent receives its reward for complex, temporally-extended behaviors sparsely. For such tasks, the problem is how to augment the state-space so as to make the reward function Markovian in an efficient way. While some existing solutions assume that the reward function is explicitly provided to the learning algorithm (e.g., in the form of a reward machine), the others learn the reward function from the interactions with the environment, assuming no prior knowledge provided by the user. In this paper, we generalize both approaches and enable the user to give advice to the agent, representing the user’s best knowledge about the reward function, potentially fragmented, partial, or even incorrect. We formalize advice as a set of DFAs and present a reinforcement learning algorithm that takes advantage of such advice, with optimal con- vergence guarantee. The experiments show that using well- chosen advice can reduce the number of training steps needed for convergence to optimal policy, and can decrease the computation time to learn the reward function by up to two orders of magnitude. Daniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu, Bo Wu 0005, Zhe Xu 0005 |
AAAI | 6 |
| 2021 | Learning Linear Temporal Properties from Noisy Data: A MaxSAT-Based Approach
Jean-Raphaël Gaglione, Daniel Neider, Rajarshi Roy 0002, Ufuk Topcu, Zhe Xu 0005 |
ATVA | 5 |
| 2021 | Active Finite Reward Automaton Inference and Reinforcement Learning Using Queries and Counterexamples
Zhe Xu 0005, Bo Wu 0005, Aditya Ojha, Daniel Neider, Ufuk Topcu |
CD-MAKE | 1 |
| 2019 | Transfer of Temporal Logic Formulas in Reinforcement LearningabstractTransferring high-level knowledge from a source task to a target task is an effective way to expedite reinforcement learning (RL). For example, propositional logic and first-order logic have been used as representations of such knowledge. We study the transfer of knowledge between tasks in which the timing of the events matters. We call such tasks temporal tasks. We concretize similarity between temporal tasks through a notion of logical transferability, and develop a transfer learning approach between different yet similar temporal tasks. We first propose an inference technique to extract metric interval temporal logic (MITL) formulas in sequential disjunctive normal form from labeled trajectories collected in RL of the two tasks. If logical transferability is identified through this inference, we construct a timed automaton for each sequential conjunctive subformula of the inferred MITL formulas from both tasks. We perform RL on the extended state which includes the locations and clock valuations of the timed automata for the source task. We then establish mappings between the corresponding components (clocks, locations, etc.) of the timed automata from the two tasks, and transfer the extended Q-functions based on the established mappings. Finally, we perform RL on the extended state for the target task, starting with the transferred extended Q-functions. Our implementation results show, depending on how similar the source task and the target task are, that the sampling efficiency for the target task can be improved by up to one order of magnitude by performing RL in the extended state space, and further improved by up to another order of magnitude using the transferred extended Q-functions. Zhe Xu 0005, Ufuk Topcu |
IJCAI | 1 |
| 2019 | Advisory Temporal Logic Inference and Controller Design for Semiautonomous RobotsabstractIn this paper, we present a method to learn (infer) and refine a set of advices from the trajectories generated in the successful and failed attempts in a task or game, in the form of advisory signal temporal logic (STL) formulas. Each advice consists of an advisory motion STL formula that characterizes the spatial-temporal pattern of the motion as a feature of success and an advisory selection STL formula as a criterion for the environment to select the advice. For the inference of advisory STL formulas, we provide a theoretical framework of perfect classification with a labeled set of trajectories with different time lengths. We design an advisory controller that can drive the robots to satisfy an advisory motion STL formula based on the advice selected according to the advisory selection STL formula. The advisory controller can advise or guide the human operators or the robots for better performance with the shared autonomy between the human operator and the controller. We provide two case studies to test the effectiveness of the advisory controller, one with a Baxter-On-Wheels simulator and the other with two quadrotors in an experimental testbed in iteratively improving the success rates of completing the tasks with the help of the designed advisory controller. Zhe Xu 0005, Botao Hu, Sandipan Mishra, A. Agung Julius |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2018 | Census Signal Temporal Logic Inference for Multiagent Group Behavior AnalysisabstractIn this paper, we define a novel census signal temporal logic (CensusSTL) that focuses on the number of agents in different subsets of a group that complete a certain task specified by the STL. CensusSTL consists of an “inner logic” STL formula and an “outer logic” STL formula. We present a new inference algorithm to infer CensusSTL formulas from the trajectory data of a group of agents. We first identify the “inner logic” STL formula and then infer the subgroups based on whether the agents' behaviors satisfy the “inner logic” formula at each time point. We use two different approaches to infer the subgroups based on similarity and complementarity, respectively. The “outer logic” CensusSTL formula is inferred from the census trajectories of different subgroups. We apply the algorithm in analyzing data from a soccer match by inferring the CensusSTL formula for different subgroups of a soccer team. Zhe Xu 0005, A. Agung Julius |
IEEE Trans Autom. Sci. Eng. | 1 |