Ze Ji

dblp:44/1965 · DBLP profile ↗
← Back
41ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0002-8968-9902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Physics-Informed Demonstration-Guided Learning Framework for Granular Material Manipulation
abstract
Due to the complex physical properties of granular materials, research on robot learning for manipulating such materials predominantly either disregards the consideration of their physical characteristics or uses surrogate models to approximate their physical properties. Learning to manipulate granular materials based on physical information obtained through precise modeling remains an unsolved problem. In this article, we propose to address this challenge by constructing a differentiable physics-based simulator for granular materials using the Taichi programming language and developing a learning framework accelerated by demonstrations generated through gradient-based optimization on nongranular materials within our simulator, eliminating the costly data collection and model training of prior methods. Experimental results show that our method, with its flexible design, trains robust policies that are capable of executing the task of transporting granular materials in both simulated and real-world environments, beyond the capabilities of standard reinforcement learning (RL), imitation learning (IL), and prior task-specific granular manipulation methods.
Minglun Wei, Xintong Yang, Yukun Lai, Seyed Amir Tafrishi, Ze Ji
IEEE Trans. Neural Networks Learn. Syst.5
2026 DDBot: Differentiable Physics-Based Digging Robot for Unknown Granular Materials
abstract
Automating the manipulation of granular materials poses significant challenges due to complex contact dynamics, unpredictable material properties, and intricate system states. Existing approaches often fail to achieve efficiency and accuracy in such tasks. To fill the research gap, this paper studies the small-scale and high-precision granular material digging task with unknown physical properties. A key scientific problem addressed is the feasibility of applying first-order gradient based optimisation to complex differentiable granular material simulation and overcoming associated numerical instability. A new framework, named differentiable digging robot (DDBot), is proposed to manipulate granular materials, including sand and soil. Specifically, we equip DDBot with a differentiable physics based simulator, tailored for granular material manipulation, powered by GPU-accelerated parallel computing and automatic differentiation. DDBot can perform efficient differentiable system identification and high-precision digging skill optimisation for unknown granular materials, which is enabled by a differentiable skill-to-action mapping, a task-oriented demonstration method, gradient clipping and line search-based gradient descent. Experimental results show that DDBot can efficiently (converge within 5 to 20 minutes) identify unknown granular material dynamics and optimise digging skills, with high-precision results in zero-shot real-world deployments, highlighting its practicality. Benchmark results against state-of-the-art baselines also confirm the robustness and efficiency of DDBot in such digging tasks.
Xintong Yang, Minglun Wei, Yukun Lai, Ze Ji
IEEE Trans. Robotics4
2025 Multi-Class Part Parsing Based on Multi-Class Boundaries
abstract
Multi-class part parsing is a dense prediction task that segments objects into semantic components with multi-level abstractions. Despite its significance, this task remains challenging due to ambiguities at both part and class levels. In this paper, we propose a network that incorporates multi-class boundaries to precisely identify and emphasize the spatial boundaries of part classes, thereby improving segmentation quality. Additionally, we employ a weighted multi-label cross-entropy loss function to ensure balanced and effective learning from all parts. Experimental results validate the effectiveness of the proposed method, demonstrating its ability to enhance baseline performance on benchmark datasets.
Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji
ICIP4
2025 Celebi's Choice: Causality-Guided Skill Optimisation for Granular Manipulation via Differentiable Simulation
abstract
Robotic soil manipulation is essential for automated farming, particularly in excavation and levelling tasks. However, the nonlinear dynamics of granular materials challenge traditional control methods, limiting stability and efficiency. We propose Celebi, a causality-enhanced optimisation method that integrates differentiable physics simulation with adaptive step-size adjustments based on causal inference. To enable gradient-based optimisation, we construct a differentiable simulation environment for granular material interactions. We further define skill parameters with a differentiable mapping to end-effector motions, facilitating efficient trajectory optimisation. By modelling causal effects between task-relevant features extracted from point cloud observations and skill parameters, Celebi selectively adjusts update step sizes to enhance optimisation stability and convergence efficiency. Experiments in both simulated and real-world environments validate Celebi’s effectiveness, demonstrating robust and reliable performance in robotic excavation and levelling tasks.
Minglun Wei, Xintong Yang, Junyu Yan, Yukun Lai, Ze Ji
IROS5
2025 Skeleton-Guided Rolling-Contact Kinematics for Arbitrary Point Clouds via Locally Controllable Parameterized Curve Fitting
abstract
Rolling contact kinematics plays a vital role in dexterous manipulation and rolling-based locomotion. Yet, in practical applications, the environments and objects involved are often captured as discrete point clouds, creating substantial difficulties for traditional motion control and planning frameworks that rely on continuous surface representations. In this work, we propose a differential geometry-based framework that models point cloud data for continuous rolling contact using locally parameterized representations. Our approach leverages skeletonization to define a rotational reference structure for rolling interactions and applies a Fourier-based curve fitting technique to extract and represent meaningful controllable local geometric structure. We further introduce a novel 2D manifold coordinate system tailored to arbitrary surface curves, enabling local parameterization of complex shapes. The governing kinematic equations for rolling contact are then derived, and we demonstrate the effectiveness of our method through simulations on various object examples.
Qingmeng Wen, Ze Ji, Yukun Lai, Mikhail M. Svinin, Seyed Amir Tafrishi
IROS2
2025 RGB-D Video Mirror Detection
abstract
Mirror detection aims to identify mirror areas in a scene, with recent methods either integrating depth information (RGB-D) or making use of temporal information (video). However, utilizing both data is still under-explored due to the lack of a high-quality dataset and an effective method for the RGB-D Video Mirror Detection (DVMD) problem. To the best of our knowledge, this is the first work to address the DVMD problem. To exploit depth and temporal information in mirror segmentation, we first construct a large-scale RGB-D Video Mirror Detection Dataset (DVMD-D), which contains 17977 RGB-D images from 273 diverse videos. We further develop a novel model, named DVMDNet, which can first locate the mirrors based on triple consistencies: local consistency, cross-modality consistency and global consistency, and then refine the mirror boundaries through content discontinuity, taking the temporal information within videos into account. We conduct a comparative study on the DVMD dataset, evaluating 12 state-of-the-art models (including single-image mirror detection, single-image glass detection, RGB-D mirror detection, video shadow detection, video glass detection, and video mirror detection methods). Code is available from https://github.com/UpChen/2025_DVMDNet.
Mingchen Xu, Peter Herbert, Yukun Lai, Ze Ji, Jing Wu 0004
WACV4
2025 FBSM: Foveabox-based boundary-aware segmentation method for green apples in natural orchards
Weikuan Jia, Zhifen Wang, Ruina Zhao, Ze Ji
Expert Syst. Appl.4
2025 Quantifying the degree of scientific innovation breakthrough: Considering knowledge trajectory change and impact
Runhui Lin, Ze Ji, Qiqi Xie, Chen Xiaoyu
Inf. Process. Manag.3
2025 Deep Reinforcement Learning With Multiple Unrelated Rewards for AGV Mapless Navigation
abstract
Mapless navigation for Automated Guided Vehicles (AGV) via Deep Reinforcement Learning (DRL) algorithms has attracted significantly rising attention in recent years. Collision avoidance from dynamic obstacles in unstructured environments, such as pedestrians and other vehicles, is one of the key challenges for mapless navigation. Autonomous navigation requires a policy to make decisions to optimize the path distance towards the goal but also to reduce the probability of collisions with obstacles. Mostly, the reward for AGV navigation is calculated by combining multiple reward functions for different purposes, such as encouraging the robot to move towards the goal or avoiding collisions, as a state-conditioned function. The combined reward, however, may lead to biased behaviours due to the empirically chosen weights when multiple rewards are combined and dangerous situations are misjudged. Therefore, this paper proposes a learning-based method with multiple unrelated rewards, which represent the evaluation of different behaviours respectively. The policy network, named Multi-Feature Policy Gradients (MFPG), is conducted by two separate Q networks that are constructed by two individual rewards, corresponding to goal distance shortening and collision avoidance, respectively. In addition, we also propose an auto-tuning method, named Ada-MFPG, that allows the MFPG algorithm to automatically adjust the weights for the two separate policy gradients. For collision avoidance, we present a new social norm-oriented continuous biased reward for performing specific social norm so as to reduce the probabilities of AGV collisions. By adding an offset gain to one of the reward functions, vehicles conducted by the proposed algorithm exhibited the predetermined features. The work was tested in different simulation environments under multiple scenarios with a single robot or multiple robots. The proposed MFPG method is compared with standard Deep Deterministic Policy Gradient (DDPG), the modified DDPG, SAC and TD3 with a social norm mechanism. MFPG significantly increases the success rate in robot navigation tasks compared with the DDPG. Besides, among all the benchmarking algorithms, the MFPG-based algorithms have the optimal task completion duration and lower variance compared with the baselines. The work has also been tested on real robots. Experiments on the real robots demonstrate the viability of the trained model for the real world scenarios. The learned model can be used for multi-robot mapless navigation in complex environments, such as a warehouse, that need multi-robot cooperation. Our source code and supplementary material is available athttps://github.com/dornenkrone/MFPGNote to Practitioners—Autonomous navigation for AGVs in complex and large-scale environments, such as factories and warehouses, is challenging. AGVs are usually centrally controlled and depend on reliable communications. However, centralized control is not always reliable due to poor signal strengths or crashes of the server, and hence unsuitable due to the requirements of accurate information of the dynamic environments and fast responses of decision making. Therefore, it is necessary for the vehicles to perform reliable decision making based on only onboard sensors and processors, for efficient and safe autonomous navigation. Existing methods, such as simultaneous localization and mapping (SLAM) and motion planning algorithms, have been widely used. However, they are neither flexible nor generalizable enough. This paper proposes a method for autonomous navigation based on reinforcement learning (RL), which allows vehicles to gain experience through cumulative rewards by continuously interacting with the environment. The RL-based controller is designed for optimising its performance in two independent aspects, namely collision avoidance and navigation, which are quantified as separate rewards. Instead of carefully hand-crafting a combined reward, our proposed approach trains the agent using the two rewards separately to obtain one optimal policy. It is clearly easier and more practical to design the individual rewards than manually combining them. Besides, the algorithm includes a mechanism for incorporating social norms to encourage the vehicles to follow the right-hand rule, such that they can avoid pedestrians or other vehicles in a socially acceptable manner. This is achieved by adding a continuous bias on the collision avoidance reward. Experiments using simulation environments and real robots suggest that the method is generalizable to multi-robot systems, while guaranteeing safety. In future research, we will focus on incorporating uncertainties of sensor readings for safe and reliable autonomous navigation.
Boliang Cai, Changyun Wei, Ze Ji
IEEE Trans Autom. Sci. Eng.3
2025 A Survey of Object Goal Navigation
abstract
Object Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. Though it has drawn increasing attention from researchers in the Embodied AI community, there has not been a contemporary and comprehensive survey of ObjectNav. In this survey, we give an overview of this field by summarizing more than 70 recent papers. First, we give the preliminaries of the ObjectNav: the definition, the simulator, and the metrics. Then, we group the existing works into three categories: 1) end-to-end methods that directly map the observations to actions, 2) modular methods that consist of a mapping module, a policy module, and a path planning module, and 3) zero-shot methods that use zero-shot learning to do navigation. Finally, we summarize the performance of existing works and the main failure modes and discuss the challenges of ObjectNav. This survey would provide comprehensive information for researchers in this field to have a better understanding of ObjectNav.Note to Practitioners—This work was motivated by the increased interest in real-world applications of mobile robots. Object Goal Navigation (ObjectNav), which is an important task in these applications, requires an agent to find an object in an unseen environment. To accomplish that, the agent needs to be equipped with the capability to move in the environment, decide where to go, and recognize the object categories. So far, most works on ObjectNav have been done in a simulation environment. We present an overview of the existing works in ObjectNav and introduce them in three categories. Additionally, we analyze the current performance of ObjectNav and the challenges for future research. This paper provides researchers and practitioners with a comprehensive overview of the developed methods in ObjectNav, which can help them to have a good understanding of this task and develop suitable solutions for applications in the real world.
Jing Wu 0004, Ze Ji, Yukun Lai
IEEE Trans Autom. Sci. Eng.3
2025 Cognitive UAV Tracking: Leveraging DRL and Hybrid Curriculum Learning for Target Reacquisition
abstract
Tracking a moving unmanned ground vehicle (UGV) with an autonomous Unmanned Aerial Vehicle (UAV) is challenging, particularly in GNSS-denied indoor environments where reacquiring the UGV after losing track poses a significant obstacle. This paper presents a novel learning framework designed to address these challenges, enabling a quadrotor UAV to effectively chase a moving UGV and regain tracking in an indoor environment. The proposed framework encompasses two primary components: the Track-HCL and the Tracking Vision System (TVS). The TVS leverages a lightweight tracker to offer real-time recognition and localization of the UGV. Additionally, the Chronological Ghosting (CG) method is employed to describe the UGV’s motion trend within a single frame. The Track-HCL component introduces a hybrid curriculum strategy to guide policy learning for the Deep Reinforcement Learning (DRL) agent. The Track-HCL enables the agent to learn the tracking policy conducive to target chasing and proficient reacquisition. We demonstrate the effectiveness of the proposed method in both simulation and field experiments.
Jiaqing Wang, Baichuan Zeng, Lan Deng, Ze Ji, Changyun Wei, Zheng Zeng 0003
IEEE Trans Autom. Sci. Eng.4
2025 Deep Reinforcement Learning With Explicit Context Representation
abstract
Though reinforcement learning (RL) has shown an outstanding capability for solving complex computational problems, most RL algorithms lack an explicit method that would allow learning from contextual information. On the other hand, humans often use context to identify patterns and relations among elements in the environment, along with how to avoid making wrong actions. However, what may seem like an obviously wrong decision from a human perspective could take hundreds of steps for an RL agent to learn to avoid. This article proposes a framework for discrete environments called Iota explicit context representation (IECR). The framework involves representing each state using contextual key frames (CKFs), which can then be used to extract a function that represents the affordances of the state; in addition, two loss functions are introduced with respect to the affordances of the state. The novelty of the IECR framework lies in its capacity to extract contextual information from the environment and learn from the CKFs' representation. We validate the framework by developing four new algorithms that learn using context: Iota deep Q-network (IDQN), Iota double deep Q-network (IDDQN), Iota dueling deep Q-network (IDuDQN), and Iota dueling double deep Q-network (IDDDQN). Furthermore, we evaluate the framework and the new algorithms in five discrete environments. We show that all the algorithms, which use contextual information, converge in around 40000 training steps of the neural networks, significantly outperforming their state-of-the-art equivalents.
Francisco Munguia-Galeano, Ah-Hwee Tan, Ze Ji
IEEE Trans. Neural Networks Learn. Syst.3
2024 GRPSNET: Multi-Class Part Parsing Based on Graph Reasoning
abstract
Multi-class part parsing is a dense prediction task that decomposes objects into semantic components with multi-level abstractions. Despite the importance of this problem, it remains challenging due to the presence of both part-level and class-level ambiguities. In this paper, we propose GRPSNet network which integrates graph reasoning to capture relationships between parts for part segmentation. These captured relationships help to enhance the recognition and localization of parts. We also propose to exploit the relationships of part boundaries to further enhance the accuracy of part segmentation. The experimental results demonstrate the effectiveness of the proposed method and show that it achieves state-of-the-art performance on the benchmark datasets.
Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji
ICME4
2024 Fusion of Short-term and Long-term Attention for Video Mirror Detection
abstract
Techniques for detecting mirrors from static images have witnessed rapid growth in recent years. However, these methods detect mirrors from single input images. Detecting mirrors from video requires further consideration of temporal consistency between frames. We observe that humans can recognize mirror candidates, from just one or two frames, based on their appearance (e.g. shape, color). However, to ensure that the candidate is indeed a mirror (not a picture or a window), we often need to observe more frames for a global view. This observation motivates us to detect mirrors by fusing appearance features extracted from a short-term attention module and context information extracted from a long-term attention module. To evaluate the performance, we build a challenging benchmark dataset of 19,255 frames from 281 videos. Experimental results demonstrate that our method achieves state-of-the-art performance on the benchmark dataset.
Mingchen Xu, Jing Wu 0004, Yukun Lai, Ze Ji
ICME4
2024 VO-Safe Reinforcement Learning for Drone Navigation
abstract
This work is focused on reinforcement learning (RL)-based navigation for drones, whose localisation is based on visual odometry (VO). Such drones should avoid flying into areas with poor visual features, as this can lead to deteriorated localization or complete loss of tracking. To achieve this, we propose a hierarchical control scheme, which uses an RL-trained policy as the high-level controller to generate waypoints for the next control step and a low-level controller to guide the drone to reach subsequent waypoints. For the high-level policy training, unlike other RL-based navigation approaches, we incorporate awareness of VO performance into our policy by introducing pose estimation-related punishment. To aid robots in distinguishing between perception-friendly areas and unfavoured zones, we instead provide semantic scenes, as input for decision-making instead of raw images. This approach also helps minimise the sim-to-real application gap.
Feiqiang Lin, Changyun Wei, Raphael Grech, Ze Ji
ICRA4
2024 SCaR: Refining Skill Chaining for Long-Horizon Robotic Manipulation via Dual Regularization
abstract
Long-horizon robotic manipulation tasks typically involve a series of interrelated sub-tasks spanning multiple execution stages. Skill chaining offers a feasible solution for these tasks by pre-training the skills for each sub-task and linking them sequentially. However, imperfections in skill learning or disturbances during execution can lead to the accumulation of errors in skill chaining process, resulting in execution failures. In this paper, we investigate how to achieve stable and smooth skill chaining for long-horizon robotic manipulation tasks. Specifically, we propose a novel skill chaining framework called Skill Chaining via Dual Regularization (SCaR). This framework applies dual regularization to sub-task skill pre-training and fine-tuning, which not only enhances the intra-skill dependencies within each sub-task skill but also reinforces the inter-skill dependencies between sequential sub-task skills, thus ensuring smooth skill chaining and stable long-horizon execution. We evaluate the SCaR framework on two representative long-horizon robotic manipulation simulation benchmarks: IKEA furniture assembly and kitchen organization. Additionally, we conduct a simple real-world validation in tabletop robot pick-and-place tasks. The experimental results show that, with the support of SCaR, the robot achieves a higher success rate in long-horizon tasks compared to relevant baselines and demonstrates greater robustness to perturbations.
Ze Ji, Jing Huo, Yang Gao 0001
NeurIPS2
2024 Evaluating Human-Robot Interaction User Experiences in Manufacturing: An Initial Assessment Framework
abstract
In the manufacturing sector, enhancing user experiences (UX) in Human-Robot Interaction (HRI) and Human-robot collaboration (HRC) are becoming increasingly essential. The core contribution of this research is the development of a UX assessment framework tailored for evaluating HRI in manufacturing, including five facets of UX. Through qualitative semi-structured interviews, we focus on identifying key factors that constitute UX in the manufacturing context. This framework is derived from an in-depth analysis of user feedback, providing a structured approach to understanding user interactions with robots. Our work highlights the importance of fostering intuitive and productive human-robot relationships in manufacturing. Future research should explore diverse manufacturing environments and integrate emerging technologies to further refine and validate the framework.
Yanzhang Tong, Ze Ji
RO-MAN3
2024 RSMPNet: Relationship Guided Semantic Map Prediction
abstract
In semantic navigation, a top-down map with accurate and complete semantic information is vital to subsequent decision-making. However, due to occlusions and limitations of the robot’s field of view (FOV), there are often unobserved areas in the top-down maps. To address this problem, recent works have studied semantic map prediction to complete the top-down maps. In this work, we propose to improve map prediction by integrating relational information. We propose RSMPNet, a relationship-guided semantic map prediction network, which makes use of semantic and spatial relationships to predict unobserved areas from accumulated semantic maps. Specifically, we propose a Relationship Reasoning Layer that includes two modules, namely 1) the Semantic Relationship Graph Reasoning Module (SeGRM) to capture the semantic relationship and 2) the Spatial Relationship Graph Reasoning Module (SpGRM) to utilize the spatial relationship. We also design a semantic relationship enhanced loss to enhance our model to learn semantic relationship information. Experiments show the effectiveness of our proposed network which achieves state-of-the-art performance in semantic map prediction. Our code and dataset are publicly available at https://github.com/jws39/semantic-map-prediction
Jing Wu 0004, Ze Ji, Yukun Lai
WACV3
2024 Sparse Convolutional Networks for Surface Reconstruction from Noisy Point Clouds
abstract
Reconstructing accurate 3D surfaces from noisy point clouds is a fundamental problem in computer vision. Among different approaches, neural implicit methods that map 3D coordinates to occupancy values benefit from the learning capabilities of deep neural networks and the flexible topology of implicit representations, achieving promising reconstruction results. However, existing methods utilize standard (dense) 3D convolutional neural networks for feature extraction and occupancy prediction, which significantly restricts their capability to reconstruct details. In this paper, we propose a neural implicit method based on sparse convolutions, where features and network calculations only focus on grid points close to the surface to be reconstructed. This allows us to build significantly higher resolution 3D grids and reconstruct high-fidelity details. We further build a 3D residual UNet to extract features which are robust to noise, while ensuring details are retained. A 3D position along with features extracted at the position are fed into the occupancy probability predictor network to obtain occupancy. As features at nearby grid points to the query position may not exist due to the sparse nature, we propose a normalized weight interpolation approach to obtain smooth interpolation with sparse data. Experimental results demonstrate that our method achieves promising results, both qualitatively and quantitatively, outperforming existing methods.
Jing Wu 0004, Ze Ji, Yukun Lai
WACV3
2024 Benchmarking visual SLAM methods in mirror environments
abstract
Visual simultaneous localisation and mapping (vSLAM) finds applications for indoor and outdoor navigation that routinely subjects it to visual complexities, particularly mirror reflections. The effect of mirror presence (time visible and its average size in the frame) was hypothesised to impact localisation and mapping performance, with systems using direct techniques expected to perform worse. Thus, a dataset, MirrEnv, of image sequences recorded in mirror environments, was collected, and used to evaluate the performance of existing representative methods. RGBD ORB-SLAM3 and BundleFusion appear to show moderate degradation of absolute trajectory error with increasing mirror duration, whilst the remaining results did not show significantly degraded localisation performance. The mesh maps generated proved to be very inaccurate, with real and virtual reflections colliding in the reconstructions. A discussion is given of the likely sources of error and robustness in mirror environments, outlining future directions for validating and improving vSLAM performance in the presence of planar mirrors. The MirrEnv dataset is available at https://doi.org/10.17035/d.2023.0292477898 .
Peter Herbert, Jing Wu 0004, Ze Ji, Yukun Lai
Comput. Vis. Media3
2024 GAM: General affordance-based manipulation for contact-rich object disentangling tasks
abstract
Picking up an entangled object is a difficult manipulation task due to its rich contact dynamics. Most existing solutions fail to produce grasp poses to enable reliable manipulation due to the dependence on simplified assumptions for the motion policies. Grasps generated by these methods tend to drop objects or cause undesired movements of non-grasped objects. To improve such object-disentangling tasks, we propose to extend the concept of reinforcement learning (RL)-based affordance to include arbitrary action consequences and implement a general affordance-based manipulation (GAM) framework. In the GAM, we train an RL agent that uses more fine-grained actions and outperforms previous methods with a smaller chance of dropping objects and making contact with non-grasped hooks. Then, a manipulation affordance prediction (MAP) model is trained to estimate the performances of the RL agent. Finally, the manipulation affordance-based grasp filter (MAGF) selects grasp poses that afford the desired manipulation performances, showing substantial improvements in five challenging hook disentangling tasks in simulation. The experiments show (1) the limitation of TAG generators, (2) the effectiveness of filtering TAGs with predicted manipulation performances based on the general affordance theory, and (3) the importance of avoiding contact with non-grasped objects in contact-rich manipulation.
Xintong Yang, Jing Wu 0004, Yukun Lai, Ze Ji
Neurocomputing4
2024 Efficient Hierarchical Reinforcement Learning for Mapless Navigation With Predictive Neighbouring Space Scoring
abstract
Solving reinforcement learning (RL)-based mapless navigation tasks is challenging due to their sparse reward and long decision horizon nature. Hierarchical reinforcement learning (HRL) has the ability to leverage knowledge at different abstract levels and is thus preferred in complex mapless navigation tasks. However, it is computationally expensive and inefficient to learn navigation end-to-end from raw high-dimensional sensor data, such as Lidar or RGB cameras. The use of subgoals based on a compact intermediate representation is therefore preferred for dimension reduction. This work proposes an efficient HRL-based framework to achieve this with a novel scoring method, named Predictive Neighbouring Space Scoring (PNSS). The PNSS model estimates the explorable space for a given position of interest based on the current robot observation. The PNSS values for a few candidate positions around the robot provide a compact and informative state representation for subgoal selection. We study the effects of different candidate position layouts and demonstrate that our layout design facilitates higher performances in longer-range tasks. Moreover, a penalty term is introduced in the reward function for the high-level (HL) policy, so that the subgoal selection process takes the performance of the low-level (LL) policy into consideration. Comprehensive evaluations demonstrate that using the proposed PNSS module consistently improves performances over the use of Lidar only or Lidar and encoded RGB featuresNote to Practitioners—This paper seeks to improve robot mapless navigation capabilities where the robot is expected to navigate to a goal location without knowing the map of the environment. This ability is highly demanded in many applications that require autonomous operations in unstructured environments, including both indoor and outdoor scenarios, involving tasks such as service robots for domestic and public environments, logistics in industrial warehouses, urban search and rescue missions, and disaster relief efforts, where detailed and accurate maps are difficult to obtain in advance. In this work, we focus on reinforcement learning-based mapless navigation. It is known that such methods struggle in complex long-range tasks, e.g. stuck in a local region by multiple objects. Therefore, this paper proposes a novel mapless navigation method inspired by human navigation behaviours. We enable a robot to split a long-range navigation task into multiple segments, by selecting and navigating to short-term goals. These subgoals are selected each time from a number of candidate positions located around the robot. The process stops when the robot reaches the final target location. When selecting a short-term goal, we use a deep neural network to predict the openness around each candidate subgoal position, named the Predictive Neighbouring Space Scoring (PNSS), from raw images and Lidar scans. In addition, we study the effects of different arrangements of candidate subgoal locations and select the optimal one. Experiments conducted in photo-realistic simulation environments demonstrate the effectiveness of our method, showcasing superior performance over baselines. It is worth noting that our agent is only trained in domestic environments using the iGibson simulator. For applications in other environments, additional training in more representative settings specific to corresponding scenarios will be necessary. In the future, our intention is to validate our methods in complex real-world environments and narrow the simulation-to-reality gap for long-horizon navigation tasks.
Yan Gao 0021, Jing Wu 0004, Xintong Yang, Ze Ji
IEEE Trans Autom. Sci. Eng.4
2023 AFPSNet: Multi-Class Part Parsing based on Scaled Attention and Feature Fusion
abstract
Multi-class part parsing is a dense prediction task that seeks to simultaneously detect multiple objects and the semantic parts within these objects in the scene. This problem is important in providing detailed object understanding, but is challenging due to the existence of both class-level and part-level ambiguities. In this paper, we propose to integrate an attention refinement module and a feature fusion module to tackle the part-level ambiguity. The attention refinement module aims to enhance the feature representations by focusing on important features. The feature fusion module aims to improve the fusion operation for different scales of features. We also propose an object-to-part training strategy to tackle the class-level ambiguity, which improves the localization of parts by exploiting prior knowledge of objects. The experimental results demonstrated the effectiveness of the proposed modules and the training strategy, and showed that our proposed method achieved state-of-the-art performance on the benchmark datasets.
Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji
WACV4
2023 PPLC: Data-driven offline learning approach for excavating control of cutter suction dredgers
abstract
Cutter suction dredgers (CSDs) play a very important role in the construction of ports, waterways and navigational channels. Currently, most of CSDs are mainly manipulated by human operators, and a large amount of instrument data needs to be monitored in real time in case of unforeseen accidents. In order to reduce the heavy workload of the operators, we propose a data-driven offline learning approach, named Preprocessing-Prediction-Learning Control (PPLC), for obtaining the optimal control policy of the excavating operation of CSDs. The proposed framework consists of three modules, i.e., a data preprocessing module, a dynamics prediction module realized by a Convolutional Neural Network (CNN), and a deep reinforcement learning based control module. The first module is responsible for filtering out irrelevant variables through correlation analysis and dimensionality reduction of raw data. The second module works as a state transition function that provides the dynamics prediction of the excavating operation of a CSD. To realize the learning control, the third module employs the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to control the swing speed during the excavating operation. The simulation results show that the proposed framework can provide an effective and reliable solution to the automated excavating control of a CSD.
Changyun Wei, Haonan Bai, Ze Ji, Zenghui Liu
Eng. Appl. Artif. Intell.4
2022 Safe distance prediction for braking control of bridge cranes considering anti-swing
abstract
Cranes are widely deployed for lifting and moving heavy objects in dynamic environments with human coexistence. Suddenly appeared workers, vehicles, and robots can affect the safety of the cranes. To avoid possible collisions, the cranes must have prediction ability to know how dangerous the situation is. In this paper, we address the safety issues of bridge cranes based on its online physical states and control model. Due to the swing of the payload, the safe braking distance cannot be a constant value. Therefore, we here propose a model prediction control (MPC)-based anti-swing method for non-zero initial states, where a new reference trajectory and a new cost function for optimization are proposed, such that the proposed MPC method can control the crane to follow the proposed reference trajectory and achieve a stable stop state with anti-swing. Furthermore, an offline learning mechanism is introduced to learn a statistical model between the velocity of the crane and the safe braking distance achieved by using the proposed MPC braking control method. In this way, we can predict how far the crane would require to safely stop without swing based on its current velocity, which is the safe distance prediction to evaluate the dangerous level of the dynamic obstacle. Experiments using both a simulated crane and a real crane demonstrate that the proposed safe braking distance prediction method is effective for safe braking control of the bridge cranes.
Huili Chen, Guohui Tian, Jianhua Zhang 0010, Ze Ji
Int. J. Intell. Syst.5
2022 Anisotropic GPMP2: A Fast Continuous-Time Gaussian Processes Based Motion Planner for Unmanned Surface Vehicles in Environments With Ocean Currents
abstract
In the past decade, there is an increasing interest in the deployment of unmanned surface vehicles (USVs) for undertaking ocean missions in dynamic, complex maritime environments. The success of these missions largely relies on motion planning algorithms that can generate optimal navigational trajectories to guide a USV. Apart from minimising the distance of a path, when deployed a USVs’ motion planning algorithms also need to consider other constraints such as energy consumption, the affected of ocean currents as well as the fast collision avoidance capability. In this paper, we propose a new algorithm named anisotropic GPMP2 to revolutionise motion planning for USVs based upon the fundamentals of GP (Gaussian process) motion planning (GPMP, or its updated version GPMP2). Firstly, we integrated the anisotropy into GPMP2 to make the generated trajectories follow ocean currents where necessary to reduce energy consumption on resisting ocean currents. Secondly, to further improve the computational speed and trajectory quality, a dynamic fast GP interpolation is integrated in the algorithm. Finally, the new algorithm has been validated on a WAM-V 20 USV in a ROS environment to show the practicability of anisotropic GPMP2. Note to Practitioners—The work reported in this article will be significant for USVs to conduct missions in complex, dynamic maritime environments where various obstacles and time-varying ocean currents exit. We develop this novel motion planning algorithm based on Gaussian process and optimise the trajectory using probabilistic inferences. The new algorithm can generate collision free trajectories that also minimise the influences caused by adverse ocean currents in a highly efficient way. In addition, the planning has been undertaken in a continuous-time domain making the generated trajectory have a guaranteed smoothness and readily feasible for autopilots to track. We use a coastal area with time-varying vortexes to present a challenging practical maritime environment. The presented algorithm integrates the available information about a fluid field regarding energy consumption and hazard level, along with the density of obstacles to plan a navigational route efficiently. To increase the practical performance of the proposed method, diverse models for generating ocean currents need to be developed in the future to tackle unpredictable situations.
Jiawei Meng, Yuanchang Liu, Richard Bucknall, Weihong Grace Guo, Ze Ji
IEEE Trans Autom. Sci. Eng.5
2022 Hierarchical Reinforcement Learning With Universal Policies for Multistep Robotic Manipulation
abstract
Multistep tasks, such as block stacking or parts (dis)assembly, are complex for autonomous robotic manipulation. A robotic system for such tasks would need to hierarchically combine motion control at a lower level and symbolic planning at a higher level. Recently, reinforcement learning (RL)-based methods have been shown to handle robotic motion control with better flexibility and generalizability. However, these methods have limited capability to handle such complex tasks involving planning and control with many intermediate steps over a long time horizon. First, current RL systems cannot achieve varied outcomes by planning over intermediate steps (e.g., stacking blocks in different orders). Second, the exploration efficiency of learning multistep tasks is low, especially when rewards are sparse. To address these limitations, we develop a unified hierarchical reinforcement learning framework, named Universal Option Framework (UOF), to enable the agent to learn varied outcomes in multistep tasks. To improve learning efficiency, we train both symbolic planning and kinematic control policies in parallel, aided by two proposed techniques: 1) an auto-adjusting exploration strategy (AAES) at the low level to stabilize the parallel training, and 2) abstract demonstrations at the high level to accelerate convergence. To evaluate its performance, we performed experiments on various multistep block-stacking tasks with blocks of different shapes and combinations and with different degrees of freedom for robot control. The results demonstrate that our method can accomplish multistep manipulation tasks more efficiently and stably, and with significantly less memory consumption.
Xintong Yang, Ze Ji, Jing Wu 0004, Yukun Lai, Changyun Wei, Rossitza Setchi
IEEE Trans. Neural Networks Learn. Syst.2
2021 ShorelineNet: An Efficient Deep Learning Approach for Shoreline Semantic Segmentation for Unmanned Surface Vehicles
abstract
This paper introduces a novel deep learning approach to semantic segmentation of the shoreline environments with a high frames-per-second (fps) performance, making the approach readily applicable to autonomous navigation for Unmanned Surface Vehicles (USV). The proposed ShorelineNet is an efficient deep neural network of high performance relying only on visual input. ShorelineNet uses monocular visual input to produce accurate shoreline separation and obstacle detection compared to the state-of-the-art, and achieves this with real-time performance. Experimental validation on a challenging multi-modal maritime obstacle detection dataset, the MODD2 dataset, achieves a much faster inference (25fps on an NVIDIA Tesla K80 and 6fps on a CPU) with respect to the recent state-of-the-art methods, while keeping the performance equally high (73.1% F-score). This makes ShorelineNet a robust and effective model to be used for reliable USV navigation that require real-time and high-performance semantic segmentation of maritime environments.
Linghong Yao, Dimitrios Kanoulas, Ze Ji, Yuanchang Liu
IROS3
2021 Online human action recognition with spatial and temporal skeleton features using a distributed camera network
abstract
Online action recognition is an important task for human-centered intelligent services. However, it remains a highly challenging problem due to the high varieties and uncertainties of spatial and temporal scales of human actions. In this paper, the following core ideas are proposed to deal with the online action recognition problem. First, we combine spatial and temporal skeleton features to represent human actions, which include not only geometrical features, but also multiscale motion features, such that both spatial and temporal information of the actions are covered. We use an efficient one-dimensional convolutional neural network to fuse spatial and temporal features and train them for action recognition. Second, we propose a group sampling method to combine the previous action frames and current action frames, which are based on the hypothesis that the neighboring frames are largely redundant, and the sampling mechanism ensures that the long-term contextual information is also considered. Third, the skeletons from multiview cameras are fused in a distributed manner, which can improve the human pose accuracy in the case of occlusions. Finally, we propose a Restful style based client-server service architecture to deploy the proposed online action recognition module on the remote server as a public service, such that camera networks for online action recognition can benefit from this architecture due to the limited onboard computational resources. We evaluated our model on the data sets of JHMDB and UT-Kinect, which achieved highly promising accuracy levels of 80.1% and 96.9%, respectively. Our online experiments show that our memory group sampling mechanism is far superior to the traditional sliding window.
Yichao Cao, Guohui Tian, Ze Ji
Int. J. Intell. Syst.5
2021 Output Feedback NCS of DoS Attacks Triggered by Double-Ended Events
abstract
In recent years, the research of the network control system under the event triggering mechanism subjected to network attacks has attracted foreign and domestic scholars’ wide attention. Among all kinds of network attacks, denial-of-service (DoS) attack is considered the most likely to impact the performance of NCS significantly. The existing results on event triggering do not assess the occurrence of DoS attacks and controller changes, which will reduce the control performance of the addressed system. Aiming at the network control system attacked by DoS, this paper combines double-ended elastic event trigger control, DoS attack, and quantitative feedback control to study the stability of NCS with quantitative feedback of DoS attack triggered by a double-ended elastic event. Simulation examples show that this method can meet the requirements of control performance and counteract the known periodic DoS attacks, which save limited resources and improve the system’s antijamming ability.
Xinzhi Feng, Yang Yang 0105, Xiaozhong Qi, Ze Ji
Secur. Commun. Networks5
2020 Automated Robot-based Large-Scale 3D Surface Imaging
abstract
This work develops a robot-based automated 3D imaging system for large-scale surface measurement at high resolution. The system has the advantages of allowing 1) high-resolution 3D surface imaging based on photometric stereo, and 2) automatic stitching of multiple images collected by a robot for large-scale surface measurement. We developed a dome-shaped image acquisition system with 16 individually controlled lights, mounted on a robot (Kuka iiwa lbr). A photometric stereo with a lighting selection mechanism is used for the reconstruction of local surface regions. To allow image stitching for large-scale surface measurement, one challenge arises from the robot arm’s limited encoder precision and accuracy, which is about ±150 µm for its repeatability and even lower for its nominal accuracy. This is unsuitable for the applications of surface metrology or inspection. To compensate the errors introduced by the chained robot arm’s encoders, for image stitching, we experimented with two feature descriptors extracted from the normal and the curvature space respectively, and performed comparative studies with standard feature descriptors from the standard grey-scale intensity space. The normal-based feature descriptor demonstrated advantages of illumination invariance while the curvature-based feature descriptor demonstrated clear advantages of rotation invariance, and feasibility of aligning multiple images with high accuracy.
Jingjing Wen, Jing Wu 0004, Ze Ji
KES4
2020 Machine Learning-enabled feedback loops for metal powder bed fusion additive manufacturing
abstract
Metal Powder Bed Fusion (PBF) has been attracting an increasing attention as an emerging metal Additive Manufacturing (AM) technology. Despite its distinctive advantages compared to traditional subtractive manufacturing such as high design flexibility, short development time, low tooling cost, and low production waste, the inconsistent part quality caused by inappropriate product design, non-optimal process plan and inadequate process control has significantly hindered its wide acceptance in the industry. To improve the part quality control in metal PBF process, this paper proposes a novel Machine Learning (ML)-enabled approach for developing feedback loops throughout the entire metal PBF process. A categorisation of metal PBF feedback loops is proposed along with a summary of the critical PBF manufacturing data in each process stage. A generic framework of ML-enabled metal PBF feedback loops is proposed with detailed explanations and examples. The opportunities and challenges of the proposed approach are also discussed. The applications of ML techniques in metal PBF process allow efficient and effective decision-makings to be achieved in each PBF process stage, and hence have a great potential in reducing the number of experiments needed, thus saving a significant amount of time and cost in metal PBF production.
Chao Liu 0031, Léopold Le Roux, Ze Ji, Pierre Kerfriden, Franck Lacan, Samuel Bigot
KES3
2019 Low-cost Measurement of Industrial Shock Signals via Deep Learning Calibration
abstract
Special high-end sensors with expensive hardware are usually needed to measure shock signals with high accuracy. In this paper, we show that cheap low-end sensors calibrated by deep neural networks are also capable to measure high-g shocks accurately. Firstly we perform drop shock tests to collect a dataset of shock signals measured by sensors of different fidelity. Secondly, we propose a novel network to effectively learn both the signal peak and overall shape. The results show that the proposed network is capable to map low-end shock signals to its high-end counterparts with satisfactory accuracy. To the best of our knowledge, this is the first work to apply deep learning techniques to calibrate shock sensors.
Houpu Yao, Jingjing Wen, Ze Ji
ICASSP5
2013 Integrating Robot Task Planner with Common-sense Knowledge Base to Improve the Efficiency of Planning
abstract
This paper presents a developed approach for intelligently generating symbolic plans by mobile robots acting in domestic environments, such as offices and houses. The significance of the approach lies in developing a new framework that consists of the new modeling of high-level robot actions and then their integration with common-sense knowledge in order to support a robotic task planner. This framework will enable interactions between the task planner and the semantic knowledge base directly. By using common-sense domain knowledge, the task planner will take into consideration the properties and relations of objects and places in its environment, before creating semantically related actions that will represent a plan. This plan will accomplish the user order. The robot task planner will use the available domain knowledge to check the next related actions to the current one and the action's conditions met will be chosen. Then the robot will use the immediately available knowledge information to check whether the plan outcomes are met or violated.
Ahmed Abdulhadi Al-Moadhen, Renxi Qiu, Michael S. Packianather, Ze Ji, Rossitza Setchi
KES4
2012 Towards automated task planning for service robots using semantic knowledge representation
abstract
Automated task planning for service robots faces great challenges in handling dynamic domestic environments. Classical methods in the Artificial Intelligence (AI) area mostly focus on relatively structured environments with fewer uncertainties. This work proposes a method to combine semantic knowledge representation with classical approaches in AI to build a flexible framework that can assist service robots in task planning at the high symbolic level. A semantic knowledge ontology is constructed for representing two main types of information: environmental description and robot primitive actions. Environmental knowledge is used to handle spatial uncertainties of particular objects. Primitive actions, which the robot can execute, are constructed based on a STRIPS-style structure, allowing a feasible solution (an action sequence) for a particular task to be created. With the Care-O-Bot (CoB) robot as the platform, we explain this work with a simple, but still challenging, scenario named “get a milk box”. A recursive back-trace search algorithm is introduced for task planning, where three main components are involved, namely primitive actions, world states, and mental actions. The feasibility of the work is demonstrated with the CoB in a simulated environment.
Ze Ji, Renxi Qiu, Alexandre Noyvirt, Anthony Soroka, Michael S. Packianather, Rossitza Setchi, Dayou Li
INDIN1
2012 Challenges for service robots operating in non-industrial environments
abstract
The concept of service robotics has grown considerably over the past two decades with many robots being used in non-industrial environments such homes, hospitals and airports. Many of these environments were never designed to have mobile service robots deployed within them. This paper describes some of the challenges that are faced and need to be overcome in order for robots to successfully work in non-industrial environments (specifically homes and hospitals). These include the problems caused by an environment not having been designed to be robot-friendly, the unstructured nature of the environment and finally the challenges presented by certain user populations who may have difficulties interacting with a robot.
Anthony Soroka, Renxi Qiu, Alexandre Noyvirt, Ze Ji
INDIN4
2012 Towards robust personal assistant robots: Experience gained in the SRS project
abstract
SRS is a European research project for building robust personal assistant robots using ROS (Robotic Operating System) and Care-O-bot (COB) 3 as the initial demonstration platform. In this paper, experience gained while building the SRS system is presented. A main contribution of the paper is the SRS autonomous control framework. The framework is divided into two parts. First, it has an automatic task planner, which initialises actions on the symbolic level. The planner produces proactive robotic behaviours based on updated semantic knowledge. Second, it has an action executive for coordination actions at the level of sensing and actuation. The executive produces reactive behaviours in well-defined domains. The two parts are integrated by fuzzy logic based symbolic grounding. As a whole, they represent the framework for autonomous control. Based on the framework, several new components and user interfaces are integrated on top of COB's existing capabilities to enable robust fetch and carry in unstructured environments. The implementation strategy and results are discussed at the end of the paper.
Renxi Qiu, Ze Ji, Alexandre Noyvirt, Anthony Soroka, Rossitza Setchi, Duc Truong Pham, Nayden Shivarov, Lucia Pigini, Georg Arbeiter, Florian Weisshardt, Birgit Graf, Marcus Mast, Lorenzo Blasi, David Facal, Martijn Rooker, Rafa López, Dayou Li, Beisheng Liu, Gernot Kronreif, Pavel Smrz
IROS2
2011 Histogram based classification of tactile patterns on periodically distributed skin sensors for a humanoid robot
abstract
The main target of this work is to improve human-robot interaction capabilities, by adding a new modality of sense, touch, to KASPAR, a humanoid robot. Large scale distributed skin-like sensors are designed and integrated on the robot, covering KASPAR at various locations. One of the challenges is to classify different types of touch. Unlike digital images represented by grids of pixels, the geometrical structure of the sensor array limits the capability of straightforward application of well-established approaches for image patterns. This paper introduces a novel histogram-based classification algorithm, transforming tactile data into histograms of local features termed as codebook. Tactile pattern can be invariant at periodical locations, allowing tactile pattern classification using a smaller number of training data, instead of using training data from everywhere on the large scale skin sensors. To generate the codebook, this method uses a two-layer approach, namely local neighbourhood structures and encodings of pressure distribution of the local neighbourhood. Classification is performed based on the constructed features using Support Vector Machine (SVM) with the intersection kernel. Real experimental data are used for experiment to classify different patterns and have shown promising accuracy. To evaluate the performance, it is also compared with the SVM using the Radial Basis Function (RBF) kernel and results are discussed from both aspects of accuracy and the location invariance property.
Ze Ji, Farshid Amirabdollahian, Daniel Polani, Kerstin Dautenhahn
RO-MAN1
2011 Adaptive Bees Algorithm - Bioinspiration from Honeybee Foraging to Optimize Fuel Economy of a Semi-Track Air-Cushion Vehicle
abstract
This interdisciplinary study covers bionics, optimization and vehicle engineering. Semi-track air-cushion vehicle (STACV) provides a solution to transportation on soft terrain, whereas it also brings a new problem of excessive fuel consumption. By mimicking the foraging behaviour of honeybees, the bioinspired adaptive bees algorithm (ABA) is proposed to calculate its running parameters for fuel economy optimization. Inherited from the basic algorithm prototype, it involves parallel-operated global search and local search, which undertake exploration and exploitation, respectively. The innovation of this improved algorithm lies in the adaptive adjustment mechanism of the range of local search (called ‘patch size’) according to the source and the rate of change of the current optimum. Three gradually in-depth experiments are implemented for 143 kinds of soils. First, the two optimal STACV running parameters present the same increasing or decreasing trend with soil parameters. This result is consistent with the terramechanics-based theoretical analysis. Second, the comparisons with four alternative algorithms exhibit the ABA's effectiveness and efficiency, and accordingly highlight the advantage of the novel adaptive patch size adjustment mechanism. Third, the impacts of two selected optimizer parameters to optimization accuracy and efficiency are investigated and their recommended values are thus proposed.
Ze Ji, Duc Truong Pham, Renxi Qiu
Comput. J.4
2010 Tactile interaction with a humanoid robot for children with autism: A case study analysis involving user requirements and results of an initial implementation
abstract
The work presented in this paper is part of our investigation in the ROBOSKIN project. The project aims to develop and demonstrate a range of new robot capabilities based on the tactile feedback provided by a robotic skin. One of the project's objectives is to improve human-robot interaction capabilities in the application domain of robot-assisted play. This paper presents design challenges in augmenting a humanoid robot with tactile sensors specifically for interaction with children with autism. It reports on a preliminary study that includes requirements analysis based on a case study evaluation of interactions of children with autism with the child-sized, minimally expressive robot KASPAR. This is followed by the implementation of initial sensory capabilities on the robot that were then used in experimental investigations of tactile interaction with children with autism.
Ben Robins, Farshid Amirabdollahian, Ze Ji, Kerstin Dautenhahn
RO-MAN3
2009 A new computer interface based on in-solid acoustic source localization
abstract
Designing ergonomic interfaces for man-machine interaction is a major task for today's computer design engineers. A new type of tangible man-machine communication interface which is based on acoustics is presented in this paper. The proposed new approach is the use of pattern recognition to match the pattern of a received signal's feature with a template, acquired during a learning stage, associated with a predefined location. Theoretically this method, referred as location patten matching (LPM) can work on heterogeneous medium of any shape or material using one or two sensors. Therefore they overcome the limitations of the more widely used approaches based on time delay of arrival (TDOA).
Duc Truong Pham, Mostafa Al-Kutubi, Ze Ji, Zuobin Wang
INDIN4