VLDB 2026 Research / reviewers in the wild / expert
Weichao Zhou
dblp:207/8077
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 60% Trustworthy machine learning · 31% Motion planning and robot control · 5% | |
| Theoretical computer science
2 papers |
Automated reasoning and model checking · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
2.2 | 4 | 2024 | Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment · NeurIPS 2024 A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines · ICML 2022 Programmatic Reward Design by Example · AAAI 2022 |
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning |
1.6 | 2 | 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning · ICML 2024 |
Machine learning › Trustworthy machine learning
robustness |
1.5 | 2 | 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 REGLO: Provable Neural Network Repair for Global Robustness Properties · AAAI 2024 |
Machine learning › Reinforcement learning
reward design |
1.1 | 2 | 2022 | A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines · ICML 2022 Programmatic Reward Design by Example · AAAI 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.9 | 1 | 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | REGLO: Provable Neural Network Repair for Global Robustness Properties · AAAI 2024 |
Machine learning › Trustworthy machine learning › verification
formal verification of neural networks |
0.8 | 1 | 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.8 | 1 | 2024 | Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › fairness
individual fairness |
0.8 | 1 | 2024 | REGLO: Provable Neural Network Repair for Global Robustness Properties · AAAI 2024 |
Machine learning › Trustworthy machine learning › robustness
neural network repair |
0.8 | 1 | 2024 | REGLO: Provable Neural Network Repair for Global Robustness Properties · AAAI 2024 |
Robotics › Motion planning and robot control
signal temporal logic |
0.8 | 1 | 2024 | Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning · ICML 2024 |
Automated reasoning and model checking › neural network verification
neural network control system verification |
0.8 | 1 | 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Automated reasoning and model checking
reachability |
0.8 | 1 | 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Machine learning › Reinforcement learning › reward learning
reward learning from demonstrations |
0.6 | 1 | 2022 | Programmatic Reward Design by Example · AAAI 2022 |
Machine learning › Reinforcement learning › reward design
reward machine |
0.6 | 1 | 2022 | A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines · ICML 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
symbolic reasoning |
0.6 | 1 | 2022 | A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines · ICML 2022 |
Automated reasoning and model checking › model checking
probabilistic model checking |
0.3 | 1 | 2018 | Safety-Aware Apprenticeship Learning · CAV (1) 2018 |
Embedded and real-time systems › cyber-physical systems
cyber-physical system safety |
0.2 | 1 | 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Methods — techniques the papers use, named apart from their topics
parallel computation · 2.3taylor model arithmetic · 1.5over-approximation · 1.5out-of-distribution regularization · 0.9constraint conditioning · 0.9signal temporal logic · 0.8semi-supervised learning · 0.8robust convex optimization · 0.8overapproximation · 0.8gradient analysis · 0.8decision transformer · 0.8adversarial training · 0.8probabilistic model checking · 0.3counterexample-guided abstraction refinement · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Constraint-Conditioned Actor-Critic for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning (OSRL) aims to learn policies with high rewards while satisfying safety constraints solely from data collected offline. However, the learned policies often struggle to handle states and actions that are not present or out-of-distribution (OOD) from the offline dataset, which can result in violation of the safety constraints or overly conservative behaviors during their online deployment. Moreover, many existing methods are unable to learn policies that can adapt to varying constraint thresholds. To address these challenges, we propose constraint-conditioned actor-critic (CCAC), a novel OSRL method that models the relationship between state-action distributions and safety constraints, and leverages this relationship to regularize critics and policy learning. CCAC learns policies that can effectively handle OOD data and adapt to varying constraint thresholds. Empirical evaluations on the $\texttt{DSRL}$ benchmarks show that CCAC significantly outperforms existing methods for learning adaptive, safe, and high-reward policies. Zijian Guo 0002, Weichao Zhou, Shengao Wang, Wenchao Li 0001 |
ICLR | 2 |
| 2024 | REGLO: Provable Neural Network Repair for Global Robustness PropertiesabstractWe present REGLO, a novel methodology for repairing pretrained neural networks to satisfy global robustness and individual fairness properties. A neural network is said to be globally robust with respect to a given input region if and only if all the input points in the region are locally robust. This notion of global robustness also captures the notion of individual fairness as a special case. We prove that any counterexample to a global robustness property must exhibit a corresponding large gradient. For ReLU networks, this result allows us to efficiently identify the linear regions that violate a given global robustness property. By formulating and solving a suitable robust convex optimization problem, REGLO then computes a minimal weight change that will provably repair these violating linear regions. Feisi Fu, Zhilu Wang, Weichao Zhou, Yixuan Wang 0001, Jiameng Fan, Chao Huang 0015, Qi Zhu 0002, Xin Chen 0002, Wenchao Li 0001 |
AAAI | 3 |
| 2024 | Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning (RL) aims to train a constraint satisfaction policy from a fixed dataset. Current state-of-the-art approaches are based on supervised learning with a conditioned policy. However, these approaches fall short in real-world applications that involve complex tasks with rich temporal and logical structures. In this paper, we propose temporal logic Specification-conditioned Decision Transformer (SDT), a novel framework that harnesses the expressive power of signal temporal logic (STL) to specify complex temporal rules that an agent should follow and the sequential modeling capability of Decision Transformer (DT). Empirical evaluations on the DSRL benchmarks demonstrate the better capacity of SDT in learning safe and high-reward policies compared with existing approaches. In addition, SDT shows good alignment with respect to different desired degrees of satisfaction of the STL specification that it is conditioned on. Zijian Guo 0002, Weichao Zhou, Wenchao Li 0001 |
ICML | 2 |
| 2024 | Rethinking Inverse Reinforcement Learning: from Data Alignment to Task AlignmentabstractMany imitation learning (IL) algorithms use inverse reinforcement learning (IRL) to infer a reward function that aligns with the demonstration.
However, the inferred reward functions often fail to capture the underlying task objectives.
In this paper, we propose a novel framework for IRL-based IL that prioritizes task alignment over conventional data alignment. Our framework is a semi-supervised approach that leverages expert demonstrations as weak supervision to derive a set of candidate reward functions that align with the task rather than only with the data. It then adopts an adversarial mechanism to train a policy with this set of reward functions to gain a collective validation of the policy's ability to accomplish the task. We provide theoretical insights into this framework's ability to mitigate task-reward misalignment and present a practical implementation. Our experimental results show that our framework outperforms conventional IL baselines in complex and transfer learning scenarios. Weichao Zhou, Wenchao Li 0001 |
NeurIPS | 1 |
| 2024 | POLAR-Express: Efficient and Precise Formal Reachability Analysis of Neural-Network Controlled SystemsabstractNeural networks (NNs) playing the role of controllers have demonstrated impressive empirical performance on challenging control problems. However, the potential adoption of NN controllers in real-life applications has been significantly impeded by the growing concerns over the safety of these NN-controlled systems (NNCSs). In this work, we present POLAR-Express, an efficient and precise formal reachability analysis tool for verifying the safety of NNCSs. POLAR-Express uses Taylor model (TM) arithmetic to propagate TMs layer-by-layer across an NN to compute an overapproximation of the NN. It can be applied to analyze any feedforward NNs with continuous activation functions, such as ReLU, Sigmoid, and Tanh activation functions that cover the common benchmarks for NNCS reachability analysis. Compared with its earlier prototype POLAR, we develop a novel approach in POLAR-Express to propagate TMs more efficiently and precisely across ReLU activation functions, and provide parallel computation support for TM propagation, thus significantly improving the efficiency and scalability. Across the comparison with six other state-of-the-art tools on a diverse set of common benchmarks, POLAR-Express achieves the best verification efficiency and tightness in the reachable set analysis. POLAR-Express is publicly available athttps://github.com/ChaoHuang2018/POLAR_Tool. Yixuan Wang 0001, Weichao Zhou, Jiameng Fan, Zhilu Wang, Xin Chen 0002, Chao Huang 0015, Wenchao Li 0001, Qi Zhu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Programmatic Reward Design by ExampleabstractReward design is a fundamental problem in reinforcement learning (RL). A misspecified or poorly designed reward can result in low sample efficiency and undesired behaviors. In this paper, we propose the idea of programmatic reward design, i.e. using programs to specify the reward functions in RL environments. Programs allow human engineers to express sub-goals and complex task scenarios in a structured and interpretable way. The challenge of programmatic reward design, however, is that while humans can provide the high-level structures, properly setting the low-level details, such as the right amount of reward for a specific sub-task, remains difficult. A major contribution of this paper is a probabilistic framework that can infer the best candidate programmatic reward function from expert demonstrations. Inspired by recent generative-adversarial approaches, our framework searches for themost likely programmatic reward function under whichthe optimally generated trajectories cannot be differen-tiated from the demonstrated trajectories. Experimental results show that programmatic reward functions learned using this framework can significantly outperform those learned using existing reward learning algorithms, and enable RL agents to achieve state-of-the-art performance on highly complex tasks. Weichao Zhou, Wenchao Li 0001 |
AAAI | 1 |
| 2022 | A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward MachinesabstractA misspecified reward can degrade sample efficiency and induce undesired behaviors in reinforcement learning (RL) problems. We propose symbolic reward machines for incorporating high-level task knowledge when specifying the reward signals. Symbolic reward machines augment existing reward machine formalism by allowing transitions to carry predicates and symbolic reward outputs. This formalism lends itself well to inverse reinforcement learning, whereby the key challenge is determining appropriate assignments to the symbolic values from a few expert demonstrations. We propose a hierarchical Bayesian approach for inferring the most likely assignments such that the concretized reward machine can discriminate expert demonstrated trajectories from other trajectories with high accuracy. Experimental results show that learned reward machines can significantly improve training efficiency for complex RL tasks and generalize well across different task environment configurations. Weichao Zhou, Wenchao Li 0001 |
ICML | 1 |
| 2020 | Runtime-Safety-Guided Policy Repair
Weichao Zhou, Ruihan Gao, BaekGyu Kim, Eunsuk Kang, Wenchao Li 0001 |
RV | 1 |
| 2018 | Safety-Aware Apprenticeship LearningabstractApprenticeship learning (AL) is a kind of Learning from Demonstration techniques where the reward function of a Markov Decision Process (MDP) is unknown to the learning agent and the agent has to derive a good policy by observing an expert’s demonstrations. In this paper, we study the problem of how to make AL algorithms inherently safe while still meeting its learning objective. We consider a setting where the unknown reward function is assumed to be a linear combination of a set of state features, and the safety property is specified in Probabilistic Computation Tree Logic (PCTL). By embedding probabilistic model checking inside AL, we propose a novel counterexample-guided approach that can ensure safety while retaining performance of the learnt policy. We demonstrate the effectiveness of our approach on several challenging AL scenarios where safety is essential. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Weichao Zhou, Wenchao Li 0001 |
CAV (1) | 1 |