EDBT 2026 Demo / reviewers in the wild / expert
Hongpeng Cao
dblp:285/4627
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4717-8714ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NLBAC: A neural ODE-based algorithm for state-wise stable and safe reinforcement learningabstractEnsuring safety and stability is critical when using reinforcement learning (RL) to control safety-critical systems. However, model-free RL algorithms usually suffer from low sample efficiency, and employing widely-used methods like dual ascent to solve constrained RL problems may be challenging due to their sensitivity to hyperparameters. To address these difficulties, in this work, we first propose an augmented Lagrangian-based method to maintain safety and stability through state-wise control Lyapunov function (CLF) and pre-defined control barrier function (CBFs) constraints in non-constrained Markov decision process (non-CMDP) settings. To handle tasks without pre-defined CBFs, we extend this method by training a barrier certificate jointly with the control policy, supported by theoretical guarantees to ensure monotonically improved control performance. Moreover, we investigate the issue of infeasibility arising from the presence of multiple state-wise constraints. A practical algorithm, Neural ordinary differential equations-based Lyapunov-Barrier Actor-Critic (NLBAC), is further designed by integrating the proposed method with the Soft Actor-Critic (SAC) and leveraging neural ordinary differential equations (NODEs) for system modeling. Comparisons with baselines and ablation experiments demonstrate that our algorithm achieves superior performance in terms of safety and driving the system towards the desired state with higher sample efficiency. Liqun Zhao, Keyan Miao, Hongpeng Cao, Konstantinos Gatsis, Antonis Papachristodoulou |
Neurocomputing | 3 |
| 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsabstractThis paper proposes the Phy-DRL: a physics-regulated deep reinforcement learning (DRL) framework for safety-critical autonomous systems. The Phy-DRL has three distinguished invariant-embedding designs: i) residual action policy (i.e., integrating data-driven-DRL action policy and physics-model-based action policy), ii) automatically constructed safety-embedded reward, and iii) physics-model-guided neural network (NN) editing, including link editing and activation editing. Theoretically, the Phy-DRL exhibits 1) a mathematically provable safety guarantee and 2) strict compliance of critic and actor networks with physics knowledge about the action-value function and action policy. Finally, we evaluate the Phy-DRL on a cart-pole system and a quadruped robot. The experiments validate our theoretical results and demonstrate that Phy-DRL features guaranteed safety compared to purely data-driven DRL and solely model-based design while offering remarkably fewer learning parameters and fast training towards safety guarantee. Hongpeng Cao, Yanbing Mao, Lui Sha, Marco Caccamo |
ICLR | 1 |
| 2024 | Equivariant Ensembles and Regularization for Reinforcement Learning in Map-based Path PlanningabstractIn reinforcement learning (RL), exploiting environmental symmetries can significantly enhance efficiency, robustness, and performance. However, ensuring that the deep RL policy and value networks are respectively equivariant and invariant to exploit these symmetries is a substantial challenge. Related works try to design networks that are equivariant and invariant by construction, limiting them to a very restricted library of components, which in turn hampers the expressiveness of the networks. This paper proposes a method to construct equivariant policies and invariant value functions without specialized neural network components, which we term equivariant ensembles. We further add a regularization term for adding inductive bias during training. In a map-based path planning case study, we show how equivariant ensembles and regularization benefit sample efficiency and performance. Mirco Theile, Hongpeng Cao, Marco Caccamo, Alberto L. Sangiovanni-Vincentelli |
IROS | 2 |
| 2023 | Towards Safe AI: Sandboxing DNNs-Based Controllers in Stochastic GamesabstractNowadays, AI-based techniques, such as deep neural networks (DNNs), are widely deployed in autonomous systems for complex mission requirements (e.g., motion planning in robotics). However, DNNs-based controllers are typically very complex, and it is very hard to formally verify their correctness, potentially causing severe risks for safety-critical autonomous systems. In this paper, we propose a construction scheme for a so-called Safe-visor architecture to sandbox DNNs-based controllers. Particularly, we consider the construction under a stochastic game framework to provide a system-level safety guarantee which is robust to noises and disturbances. A supervisor is built to check the control inputs provided by a DNNs-based controller and decide whether to accept them. Meanwhile, a safety advisor is running in parallel to provide fallback control inputs in case the DNN-based controller is rejected. We demonstrate the proposed approaches on a quadrotor employing an unverified DNNs-based controller. Bingzhuo Zhong, Hongpeng Cao, Majid Zamani 0001, Marco Caccamo |
AAAI | 2 |
| 2023 | Flexible Gear Assembly with Visual Servoing and Force FeedbackabstractThis paper presents a vision-guided two-stage approach with force feedback to achieve high-precision and flexible gear assembly. The proposed approach integrates YOLO to coarsely localize the target workpiece in a searching phase and deep reinforcement learning (DRL) to complete the insertion. Specifically, DRL addresses the challenge of partial visibility when the on-wrist camera is too close to the workpiece of a small size. Moreover, we use force feedback to improve the robustness of the vision-guided assembly process. To reduce the effort of collecting training data on real robots, we use synthetic RGB images for training YOLO and construct an offline interaction environment leveraging sampled real-world data for training DRL agents. The proposed approach was evaluated in an industrial gear assembly experiment, which requires an assembly clearance of 0.3 mm, demonstrating high robustness and efficiency in gear searching and insertion from arbitrary positions. Junjie Ming, Daniel Bargmann, Hongpeng Cao, Marco Caccamo |
IROS | 3 |
| 2022 | Cloud-Edge Training Architecture for Sim-to-Real Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) is a promising approach to solve complex control tasks by learning policies through interactions with the environment. However, the training of DRL policies requires large amounts of training experiences, making it impractical to learn the policy directly on physical systems. Sim-to-real approaches leverage simulations to pretrain DRL policies and then deploy them in the real world. Unfortunately, the direct real-world deployment of pretrained policies usually suffers from performance deterioration due to the different dynamics, known as the reality gap. Recent sim-to-real methods, such as domain randomization and domain adaptation, focus on improving the robustness of the pretrained agents. Nevertheless, the simulation-trained policies often need to be tuned with real-world data to reach optimal performance, which is challenging due to the high cost of real-world samples. This work proposes a distributed cloud-edge architecture to train DRL agents in the real world in real-time. In the architecture, the inference and training are assigned to the edge and cloud, separating the real-time control loop from the computationally expensive training loop. To overcome the reality gap, our architecture exploits sim-to-real transfer strategies to continue the training of simulation-pretrained agents on a physical system. We demonstrate its applicability on a physical inverted-pendulum control system, analyzing critical parameters. The real-world experiments show that our architecture can adapt the pretrained DRL agents to unseen dynamics consistently and efficiently.11A video showing a real-world training process under the proposed method can be found from https://youtu.be/hMY9-c0SST0. Hongpeng Cao, Mirco Theile, Federico G. Wyrwal, Marco Caccamo |
IROS | 1 |