EDBT 2026 Demo / reviewers in the wild / expert
Zijian Guo 0002
dblp:11/4679-2
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-9791-6749ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 64% Trustworthy machine learning · 13% Robot navigation and mapping · 6% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
safe reinforcement learning |
3.7 | 5 | 2025 | One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning · NeurIPS 2025 Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning |
2.3 | 3 | 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning · ICML 2024 Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Machine learning › Trustworthy machine learning
robustness |
1.3 | 2 | 2023 | Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023 On the Robustness of Safe Reinforcement Learning under Observational Perturbations · ICLR 2023 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic in computer science › logical foundations › non-classical logics › temporal logic
linear temporal logic |
0.9 | 1 | 2025 | One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems · NeurIPS 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation
safe navigation |
0.9 | 1 | 2025 | HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › temporal planning
temporal logic planning |
0.9 | 1 | 2025 | One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
signal temporal logic |
0.8 | 1 | 2024 | Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning · ICML 2024 |
Visual content generation and editing › video generation
controllable video generation |
0.8 | 1 | 2024 | TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs · ICRA 2024 |
Visual content generation and editing
video generation |
0.8 | 1 | 2024 | TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs · ICRA 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.7 | 1 | 2023 | Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy learning |
0.7 | 1 | 2023 | Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer |
0.7 | 1 | 2023 | Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.7 | 1 | 2023 | Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Automata and formal languages › omega-automata
büchi automata |
0.3 | 1 | 2025 | One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
0.2 | 1 | 2024 | TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs · ICRA 2024 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.2 | 1 | 2023 | Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023 |
Mathematical optimization
multi-objective optimization |
0.2 | 1 | 2023 | Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
decision transformer · 2.1subgoal observation reduction · 1.7reach-avoid decomposition · 1.7neural ordinary differential equation · 1.5dynamical system modeling · 1.5out-of-distribution regularization · 0.9hierarchical reinforcement learning · 0.9control barrier functions · 0.9constraint conditioning · 0.9signal temporal logic · 0.8zero-shot adaptation · 0.7multi-objective optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning (OSRL) aims to learn policies with high rewards while satisfying safety constraints solely from data collected offline. However, the learned policies often struggle to handle states and actions that are not present or out-of-distribution (OOD) from the offline dataset, which can result in violation of the safety constraints or overly conservative behaviors during their online deployment. Moreover, many existing methods are unable to learn policies that can adapt to varying constraint thresholds. To address these challenges, we propose constraint-conditioned actor-critic (CCAC), a novel OSRL method that models the relationship between state-action distributions and safety constraints, and leverages this relationship to regularize critics and policy learning. CCAC learns policies that can effectively handle OOD data and adapt to varying constraint thresholds. Empirical evaluations on the $\texttt{DSRL}$ benchmarks show that CCAC significantly outperforms existing methods for learning adaptive, safe, and high-reward policies. Zijian Guo 0002, Weichao Zhou, Shengao Wang, Wenchao Li 0001 |
ICLR | 1 |
| 2025 | HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsabstractWe address the problem of safe policy learning in multi-agent safety-critical autonomous systems.
In such systems, it is necessary for each agent to meet the safety requirements at all times while also cooperating with other agents to accomplish the task. Toward this end, we propose a safe Hierarchical Multi-Agent Reinforcement Learning (HMARL) approach based on Control Barrier Functions (CBFs). Our proposed hierarchical approach decomposes the overall reinforcement learning problem into two levels –- learning joint cooperative behavior at the higher level and learning safe individual behavior at the lower or agent
level conditioned on the high-level policy. Specifically, we propose a skill-based HMARL-CBF algorithm in which the higher-level problem involves learning a joint policy over the skills for all the agents and the lower-level problem involves
learning policies to execute the skills safely with CBFs. We validate our approach on challenging environment scenarios whereby a large number of agents have to safely navigate through conflicting road networks. Compared with existing state-of-the-art methods, our approach significantly improves the safety achieving near perfect (within $5\%$) success/safety rate while also improving performance across all the environments. H. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja, Zijian Guo 0002, Songyuan Zhang, Chuchu Fan, Christos G. Cassandras, Wenchao Li 0001 |
NeurIPS | 5 |
| 2025 | One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement LearningabstractGeneralizing to complex and temporally extended task objectives and safety constraints remains a critical challenge in reinforcement learning (RL). Linear temporal logic (LTL) offers a unified formalism to specify such requirements, yet existing methods are limited in their abilities to handle nested long-horizon tasks and safety constraints, and cannot identify situations when a subgoal is not satisfiable and an alternative should be sought. In this paper, we introduce GenZ-LTL, a method that enables zero-shot generalization to arbitrary LTL specifications. GenZ-LTL leverages the structure of Büchi automata to decompose an LTL task specification into sequences of reach-avoid subgoals. Contrary to the current state-of-the-art method that conditions on subgoal sequences, we show that it is more effective to achieve zero-shot generalization by solving these reach-avoid problems $\textit{one subgoal at a time}$ through proper safe RL formulations. In addition, we introduce a novel subgoal-induced observation reduction technique that can mitigate the exponential complexity of subgoal-state combinations under realistic assumptions. Empirical results show that GenZ-LTL substantially outperforms existing methods in zero-shot generalization to unseen LTL specifications. Zijian Guo 0002, Ilker Isik, H. M. Sabbir Ahmad, Wenchao Li 0001 |
NeurIPS | 1 |
| 2024 | Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning (RL) aims to train a constraint satisfaction policy from a fixed dataset. Current state-of-the-art approaches are based on supervised learning with a conditioned policy. However, these approaches fall short in real-world applications that involve complex tasks with rich temporal and logical structures. In this paper, we propose temporal logic Specification-conditioned Decision Transformer (SDT), a novel framework that harnesses the expressive power of signal temporal logic (STL) to specify complex temporal rules that an agent should follow and the sequential modeling capability of Decision Transformer (DT). Empirical evaluations on the DSRL benchmarks demonstrate the better capacity of SDT in learning safe and high-reward policies compared with existing approaches. In addition, SDT shows good alignment with respect to different desired degrees of satisfaction of the STL specification that it is conditioned on. Zijian Guo 0002, Weichao Zhou, Wenchao Li 0001 |
ICML | 1 |
| 2024 | TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image PairsabstractVideos capture the evolution of continuous dynamical systems over time in the form of discrete image sequences. Recently, video generation models have been widely used in robotic research. However, generating controllable videos from image-text pairs is an important yet underexplored research topic in both robotic and computer vision communities. This paper introduces an innovative and elegant framework named TiV-ODE, formulating this task as modeling the dynamical system in a continuous space. Specifically, our framework leverages the ability of Neural Ordinary Differential Equations (Neural ODEs) to model the complex dynamical system depicted by videos as a nonlinear ordinary differential equation. The resulting framework offers control over the generated videos’ dynamics, content, and frame rate, a feature not provided by previous methods. Experiments demonstrate the ability of the proposed method to generate highly controllable and visually consistent videos and its capability of modeling dynamical systems. Overall, this work is a significant step towards developing advanced controllable video generation models that can handle complex and dynamic scenes. Nanbo Li, Arushi Goel, Zonghai Yao, Zijian Guo 0002, Hamidreza Kasaei 0001, Seyed Mohammadreza Mohades Kasaei, Zhibin Li 0001 |
ICRA | 5 |
| 2023 | On the Robustness of Safe Reinforcement Learning under Observational Perturbations
Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Jie Tan 0001, Bo Li 0026, Ding Zhao |
ICLR | 2 |
| 2023 | Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataabstractPrevious work demonstrates that the optimal safe reinforcement learning policy in a noise-free environment is vulnerable and could be unsafe under observational attacks. While adversarial training effectively improves robustness and safety, collecting samples by attacking the behavior agent online could be expensive or prohibitively dangerous in many applications. We propose the robuSt vAriational ofF-policy lEaRning (SAFER) approach, which only requires benign training data without attacking the agent. SAFER obtains an optimal non-parametric variational policy distribution via convex optimization and then uses it to improve the parameterized policy robustly via supervised learning. The two-stage policy optimization facilitates robust training, and extensive experiments on multiple robot platforms show the efficiency of SAFER in learning a robust and safe policy: achieving the same reward with much fewer constraint violations during training than on-policy baselines. Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Yihang Yao, Hanjiang Hu, Ding Zhao |
ICML | 2 |
| 2023 | Constrained Decision Transformer for Offline Safe Reinforcement LearningabstractSafe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment. We aim to tackle a more challenging problem: learning a safe policy from an offline dataset. We study the offline safe RL problem from a novel multi-objective optimization perspective and propose the $\epsilon$-reducible concept to characterize problem difficulties. The inherent trade-offs between safety and task performance inspire us to propose the constrained decision transformer (CDT) approach, which can dynamically adjust the trade-offs during deployment. Extensive experiments show the advantages of the proposed method in learning an adaptive, safe, robust, and high-reward policy. CDT outperforms its variants and strong offline safe RL baselines by a large margin with the same hyperparameters across all tasks, while keeping the zero-shot adaptation capability to different constraint thresholds, making our approach more suitable for real-world RL under constraints. Zuxin Liu, Zijian Guo 0002, Yihang Yao, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao |
ICML | 2 |