Zengjie Zhang

dblp:211/5735 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1875-1032ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Sample-efficient reinforcement learning with symmetry-guided demonstrations for robotic manipulation
Amir M. Soufi Enayati, Zengjie Zhang, Kashish Gupta, Homayoun Najjaran
Neural Comput. Appl.2
2025 Risk-Aware Autonomous Driving with Linear Temporal Logic Specifications
abstract
Human drivers naturally balance the risks of different concerns while driving, including traffic rule violations, minor accidents, and fatalities. However, achieving the same behavior in autonomous driving systems remains an open problem. This paper extends a risk metric that has been verified in human-like driving studies to encompass more complex driving scenarios specified by linear temporal logic (LTL) that go beyond just collision risks. This extension incorporates the timing and severity of events into LTL specifications, thereby reflecting a human-like risk awareness. Without sacrificing expressivity for traffic rules, we adopt LTL specifications composed of safety and co-safety formulas, allowing the control synthesis problem to be reformulated as a reachability problem. By leveraging occupation measures, we further formulate a linear programming (LP) problem for this LTL-based risk metric. Consequently, the synthesized policy balances different types of driving risks, including both collision risks and traffic rule violations. The effectiveness of the proposed approach is validated by three typical traffic scenarios in Carla simulator.
Shuhao Qi, Zengjie Zhang, Zhiyong Sun 0001, Sofie Haesaert
IROS2
2025 Distributed Coverage Control of Constrained Constant-Speed Unicycle Multi-Agent Systems
abstract
This paper proposes a novel distributed coverage controller for a multi-agent system with constant-speed unicycle robots (CSUR). The work is motivated by the limitation of the conventional method that does not ensure the satisfaction of hard state-and input-dependent constraints and leads to feasibility issues for multi-CSUR systems. In this paper, we solve these problems by designing a novel coverage cost function and a saturated gradient-search-based control law. Theoretical proofs are provided to guarantee that the CSURs ultimately move to the optimal coverage configuration without moving out of the covered domain. The controller is implemented in a distributed manner based on a novel communication standard among the agents. A series of simulation studies are conducted to validate the correctness of our theory by showing the efficacy of the proposed coverage controller in different initial conditions and with various control parameters. A comparison study in simulation reveals the advantage of the proposed method over the conventional method in terms of avoiding infeasibility. The experimental study verifies the applicability of the method to real robots. The development procedure of the method from theoretical analysis to experimental validation provides a novel framework for multi-agent system coordinate control with complex dynamics.Note to Practitioners—This paper gives a novel method to effectively cover a polygonal area using multiple constant-speed unicycle robots (CSUR) like wheeled robots and fixed-wing unmanned aerial vehicles (fUAV). Compared to the conventional approaches, our method allows these robots to cover a target region using circular orbits without departing the covered region. Also, the method satisfies common control saturation constraints in practice and can be implemented in a reliable distributed scheme. While the efficacy and correctness of the proposed method are rigorously proved using control theory, we also provide necessary interpretive elucidations to explain its underlying mechanism and selection rationale. The method is validated to be effective for wheeled robots in experimental studies, although it can also be applied to fUAVs in theory.
Qingchen Liu, Zengjie Zhang, Nhan Khanh Le, Jiahu Qin, Fangzhou Liu 0001, Sandra Hirche
IEEE Trans Autom. Sci. Eng.2
2024 Self-supervised graph autoencoder with redundancy reduction for community detection
Xiaofeng Wang 0004, Guodong Shen, Zengjie Zhang, Shuaiming Lai, Shuailei Zhu, Yuntao Chen, Daying Quan
Neurocomputing3
2024 Using Implicit Behavior Cloning and Dynamic Movement Primitive to Facilitate Reinforcement Learning for Robot Motion Planning
abstract
Reinforcement learning (RL) for motion planning of multi-degree-of-freedom robots still suffers from low efficiency in terms of slow training speed and poor generalizability. In this article, we propose a novel RL-based robot motion planning framework that uses implicit behavior cloning (IBC) and dynamic movement primitive (DMP) to improve the training speed and generalizability of an off-policy RL agent. IBC utilizes human demonstration data to leverage the training speed of RL, and DMP serves as a heuristic model that transfers motion planning into a simpler planning space. To support this, we also create a human demonstration dataset using a pick-and-place experiment that can be used for similar studies. Comparison studies reveal the advantage of the proposed method over the conventional RL agents with faster training speed and higher scores. A real-robot experiment indicates the applicability of the proposed method to a simple assembly task. Our work provides a novel perspective on using motion primitives and human demonstration to leverage the performance of RL for robot applications.
Zengjie Zhang, Jayden Hong, Amir M. Soufi Enayati, Homayoun Najjaran
IEEE Trans. Robotics1
2023 Risk-Aware Reward Shaping of Reinforcement Learning Agents for Autonomous Driving
abstract
Reinforcement learning (RL) is an effective approach to motion planning in autonomous driving, where an optimal driving policy can be automatically learned using the interaction data with the environment. Nevertheless, the reward function for an RL agent, which is significant to its performance, is challenging to determine. The conventional work mainly focuses on rewarding safe driving states but does not incorporate the awareness of risky driving behaviors of the vehicles. In this paper, we investigate how to use risk-aware reward shaping to leverage the training and test performance of RL agents in autonomous driving. Based on the essential requirements that prescribe the safety specifications for general autonomous driving in practice, we propose additional reshaped reward terms that encourage exploration and penalize risky driving behaviors. A simulation study in OpenAI Gym indicates the advantage of risk-aware reward shaping for various RL agents. Also, we point out that proximal policy optimization (PPO) is likely to be the best RL method that works with risk-aware reward shaping.
Lin-Chi Wu, Zengjie Zhang, Sofie Haesaert, Zhiqiang Ma 0001, Zhiyong Sun 0001
IECON2
2023 A Persistent-Excitation-Free Method for System Disturbance Estimation Using Concurrent Learning
abstract
Observer-based methods are widely used to estimate the disturbances of different dynamic systems. However, a drawback of the conventional disturbance observers is that they all assume persistent excitation (PE) of the systems. As a result, they may lead to poor estimation precision when PE is not ensured, for instance, when the disturbance gain of the system is close to the singularity. In this paper, we propose a novel disturbance observer based on concurrent learning (CL) with time-variant history stacks, which ensures high estimation precision even in PE-free cases. The disturbance observer is designed in both continuous and discrete time. The estimation errors of the proposed method are proved to converge to a bounded set using the Lyapunov method. A history-sample-selection procedure is proposed to reduce the estimation error caused by the accumulation of old history samples. A simulation study on epidemic control shows that the proposed method produces higher estimation precision than the conventional disturbance observer when PE is not satisfied. This justifies the correctness of the proposed CL-based disturbance observer and verifies its applicability to solving practical problems.
Zengjie Zhang, Fangzhou Liu 0001, Tong Liu 0031, Jianbin Qiu, Martin Buss
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A methodical interpretation of adaptive robotics: Study and reformulation
Amir M. Soufi Enayati, Zengjie Zhang, Homayoun Najjaran
Neurocomputing2
2021 An Online Robot Collision Detection and Identification Scheme by Supervised Learning and Bayesian Decision Theory
abstract
This article is dedicated to developing an online collision detection and identification (CDI) scheme for human-collaborative robots. The scheme is composed of a signal classifier and an online diagnosor, which monitors the sensory signals of the robot system, detects the occurrence of a physical human–robot interaction, and identifies its type within a short period. In the beginning, we conduct an experiment to construct a data set that contains the segmented physical interaction signals with ground truth. Then, we develop the signal classifier on the data set with the paradigm of supervised learning. To adapt the classifier to the online application with requirements on response time, an auxiliary online diagnosor is designed using the Bayesian decision theory. The diagnosor provides not only a collision identification result but also a confidence index which represents the reliability of the result. Compared to the previous works, the proposed scheme ensures rapid and accurate CDI even in the early stage of a physical interaction. As a result, safety mechanisms can be triggered before further injuries are caused, which is quite valuable and important toward a safe human–robot collaboration. In the end, the proposed scheme is validated on a robot manipulator and applied to a demonstration task with collision reaction strategies. The experimental results reveal that the collisions are detected and classified within 20 ms with an overall accuracy of 99.6%, which confirms the applicability of the scheme to collaborative robots in practice.Note to Practitioners—This article is intended to provide a novel online collision event handling scheme for robots in industrial environments. This scheme is designed to quickly and accurately detect an accidental collision and distinguish it from the intentional human–robot interaction. The method takes the raw signals from external torque sensors and provides a collision diagnosis result with a reliability index. The simple structure makes it easy to be implemented as a regular fault monitoring routine for collaborative robots. Different from the conventional methods, the proposed collision identification scheme in this article especially focuses on overcoming the following two challenges in practice: first, to timely and accurately report a collision within its early stage, and second, to ensure a high identification accuracy in a complicated environment, where ubiquitous disturbance and noise are unneglectable. The experimental validation at the end of this article confirms its promising application value in human–robot collaboration.
Zengjie Zhang, Kun Qian 0003, Björn W. Schuller, Dirk Wollherr
IEEE Trans Autom. Sci. Eng.1