EDBT 2026 Demo / reviewers in the wild / expert
Jiyu Cheng
dblp:205/3847
· DBLP profile ↗
16ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0002-8063-2547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning packing-and-unpacking synergistic policy via LLM-guided DRL for robust online robotic packing
Shuai Song, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021 |
Adv. Eng. Informatics | 4 |
| 2025 | C2P-Net: Comprehensive Depth Map to Planar Depth Conversion for Room Layout EstimationabstractRoom layout estimation seeks to infer the overall spatial configuration of indoor scenes using perspective or panoramic images. As the layout is determined by the dominant indoor planes, this problem inherently requires the reconstruction of these planes. Some studies reconstruct indoor planes from perspective images by learning pixel-level or instance-level plane parameters. However, directly learning these parameters has the problems of susceptibility to occlusions and position dependency. In this paper, we introduce the Comprehensive depth map to Planar depth (C2P) conversion, which reformulates planar depth reconstruction into the prediction of a comprehensive depth map and planar visibility confidence. Based on the parametric representation of planar depth we propose, the C2P conversion is applicable to both panoramic and perspective images. Accordingly, we present an effective framework for room layout estimation that jointly learns the comprehensive depth map and planar visibility confidence. Due to the differentiability of the C2P conversion, our network autonomously learns planar visibility confidence by constraining the estimated plane parameters and reconstructed planar depth map. We further propose a novel approach for 3D layout generation through sequential planar depth map integration. Experimental results demonstrate the superiority of our method across all evaluated panoramic and perspective datasets. Weidong Zhang 0005, Mengjie Zhou, Jiyu Cheng, Ying Liu 0026, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Autonomous and Adaptive Role Selection for Multi-Robot Collaborative Area Search Based on Deep Reinforcement LearningabstractIn the tasks of multi-robot collaborative area search, we propose the unified approach for simultaneous mapping for sensing more targets (exploration) while searching and locating the targets (coverage). Specifically, we implement a hierarchical multi-agent reinforcement learning algorithm to decouple task planning from task execution. The role concept is integrated into the upper-level task planning for role selection, which enables robots to learn the role based on the state status from the upper-view. Besides, an intelligent role switching mechanism enables the role selection module to function between two timesteps, promoting both exploration and coverage interchangeably. Then, the primitive policy learns how to plan based on their assigned roles and local observation for sub-task execution. The well-designed experiments show the scalability and generalization of our method compared with state-of-the-art approaches in scenes with varying complexity and numbers of robots. Our code is released at https://github.com/linaug/Role_selection. Jiyu Cheng, Hao Zhang 0113, Zhichao Cui, Wei Zhang 0021, Yuehu Liu |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | A Deep Reinforcement Learning Approach Using Asymmetric Self-Play for Robust Multirobot FlockingabstractFlocking control, as an essential approach for survivable navigation of multirobot systems, has been widely applied in fields, such as logistics, service delivery, and search and rescue. However, realistic environments are typically complex, dynamic, and even aggressive, posing considerable threats to the safety of flocking robots. In this article, based on deep reinforcement learning, anAsymmetricSelf-play-empoweredFlockingControl framework is proposed to address this concern. Specifically, the flocking robots are trained concurrently with learnable adversarial interferers to stimulate the intelligence of the flocking strategy. A two-stage self-play training paradigm is developed to improve the robustness and generalization of the model. Furthermore, an auxiliary training module regarding the learning of transition dynamics is designed, dramatically enhancing the adaptability to environmental uncertainties. Feature-level and agent-level attention are implemented for action and value generation, respectively. Both extensive comparative experiments and real-world deployment demonstrate the superiority and practicality of the proposed framework. Yunjie Jia, Yong Song 0005, Jiyu Cheng, Jiong Jin, Wei Zhang 0021, Simon X. Yang, Sam Kwong |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Empowering Multirobot Flocking in Complex Environments via Effective Communication: A Deep Reinforcement Learning ApproachabstractMultirobot flocking is crucial for safe and cooperative navigation, with wide applications in logistics, service delivery, and mobile surveillance. Despite significant progress, developing effective flocking strategies under complex conditions remains challenging. Communication is a vital technique for multirobot coordination. In this article, we propose refinement and enhancement of communication information (REIN), a novel deep reinforcement learning-based framework designed to improve communication effectiveness in leader–follower flocking systems through the REIN. First, regarding information refinement, a graph-based information refiner, integrating directed graph-structured communication with an innovative edge filter, is developed for selective multirobot interaction. It helps robots adaptively focus on relevant neighbors, considerably alleviating information overload. Second, for information enhancement, a cognition-aligned information enhancer is designed that boosts information expressiveness by encouraging team consensus. It utilizes two cascaded leader-related objectives to optimize information towards cognitive alignment among decentralized followers. Extensive comparisons with state-of-the-art approaches and ablation versions demonstrate the superiority of our framework. Physical experiments are also conducted to validate its practicality. Yunjie Jia, Yong Song 0005, Jiyu Cheng, Heteng Zhang, Wei Zhang 0021, Rui Song 0002, Simon X. Yang, Sam Kwong |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement LearningabstractDespite significant advancements in multirobot technologies, efficiently and collaboratively exploring an unknown environment remains a major challenge. In this paper, we propose AIM-Mapping, an Asymmetric InforMation enhanced Mapping framework based on deep reinforcement learning. The framework fully leverages the privileged information to help construct the environmental representation as well as the supervised signal in an asymmetric actor-critic training framework. Specifically, privileged information is used to evaluate exploration performance through an asymmetric feature representation module and a mutual information evaluation module. The decision-making network employs the trained feature encoder to extract structural information of the environment and integrates it with a topological map constructed based on geometric distance. By leveraging this topological map representation, we apply topological graph matching to assign corresponding boundary points to each robot as long-term goal points. We conduct experiments in both iGibson simulation environments and real-world scenarios. The results demonstrate that the proposed method achieves significant performance improvements compared to existing approaches. Jiyu Cheng, Junhui Fan, Xiaolei Li 0003, Paul L. Rosin, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans. Robotics | 1 |
| 2024 | Heuristics Integrated Deep Reinforcement Learning for Online 3D Bin PackingabstractOnline 3D Bin Packing Problem (3D-BPP) has a wide range of industrial applications and there is an emerging research interest in learning optimal bin packing policy and deploying it for real logistics applications. From the heuristic methods to the deep reinforcement learning (DRL) methods, the previous works have proposed many solutions to solve the online 3D-BPP. However, none of them have studied what and how heuristics can be modelled into DRL to build a more effective and practical bin packing pipeline. In this work, we thoroughly investigate what heuristics can be used in online 3D-BPP and how to effectively integrate the heuristics with the DRL. First, we design 3 different heuristics based on the physical rules of the real world and the experiences of the human packers, including the Physics-Heuristics, the Packing-Heuristics and the Unpacking-Heuristics. Second, we model the 3 types of heuristics into the DRL framework and propose a novel heuristic DRL method to solve the online 3D-BPP. Extensive experimental results show that our method achieves state-of-the-art bin packing performance and the resulting real-world system is able to reliably finish the bin packing task in real logistics scenarios. Supplementary video is available athttps://www.youtube.com/watch?v=x8GpmEELq18. Note to Practitioners—The rapid growth of e-commerce has significantly increased the burden of human packers in logistic warehouses, where the workers need to pick the products from a conveyor and pack them into bins (i.e. the online 3D bin packing). Thus it is of great importance to develop intelligent robotic systems to replace human labor, which is a long-standing topic in the field of control and automation science. This paper makes a substantial contribution to the related field by studying the online 3D bin packing in terms of both the theory and practice. On the one hand, the simulated experiments suggest that the presented algorithm significantly improves the space utilization of bin packing. On the other hand, the robotic system developed based on the proposed method can favourably finish the bin packing task in real logistics scenarios, demonstrating the practical use of our approach. Consequently, the approach proposed in this paper is totally applicable in logistic warehouses and is promising to drastically improve the working efficiency of the product packing in real warehouses. In the future, we will extend the presented approach to pack irregular-shaped objects and then facilitate more logistics applications. Shuai Song, Shilei Chu, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Hierarchical Perception-Improving for Decentralized Multi-Robot Motion Planning in Complex ScenariosabstractMulti-robot cooperative navigation is an important task, which has been widely studied in many fields like logistics, transportation, and disaster rescue. However, most of the existing methods either require some strong assumptions or are validated in simple scenarios, which greatly hinders their implementation in the real world. In this paper, more complex environments are considered in which robots can only acquire local observations from their own sensors and have only limited communication capabilities for mapless collaborative navigation. To address this challenging task, we propose a hierarchical framework, by fusing bothSensor-wise andAgent-wise features forPerception-Improving (SAPI), which can adaptively integrate features from different information sources to improve perception capabilities. Specifically, to facilitate scene understanding, we assign prior knowledge to the visual coder to generate efficient embeddings. For effective feature representation, an attention-based sensor fusion network is designed to fuse sensor-level information of visual and LiDAR sensors, while graph convolution with multi-head attention mechanism is applied to aggregate agent-level information from an arbitrary number of neighbors. In addition, reinforcement learning is used to optimize the policy, where a novel compound reward function is introduced to guide training. Extensive experiments demonstrate that our method has excellent generalization ability in different scenarios and scalability for large-scale systems. Yunjie Jia, Yong Song 0005, Bo Xiong 0001, Jiyu Cheng, Wei Zhang 0021, Simon X. Yang, Sam Kwong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Multi-Robot Environmental Coverage With a Two-Stage Coordination Strategy via Deep Reinforcement LearningabstractMulti-robot environmental coverage can be widely used in many applications like search and rescue. However, it is challenging to coordinate the robot team for high coverage efficiency. In this paper, we propose a Two-Stage Coordination (TSC) strategy, which consists of a high-level leader module and a low-level action executor. The former provides the robots with the topology and geometry of the environment, which are crucial for robots to learn “where” they should go and avoid invalid coverage. Based on the observed information and the environmental topology, the latter module takes primitive action to reach the sub-goal. To facilitate cooperation among the robots, we aggregate local perception information of neighbors from different hops based on graph neural networks. We compare our method with state-of-the-art multi-robot coverage approaches. Experiments and supporting ablation studies show the superior efficiency, scalability, and generalization of our algorithm especially in unseen style and scale of scenes, and an unseen number of robots. Jiyu Cheng, Hao Zhang 0113, Wei Zhang 0021, Yuehu Liu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Watch and Act: Learning Robotic Manipulation From Visual DemonstrationabstractLearning from demonstration holds the promise of enabling robots to learn diverse actions from expert experience. In contrast to learning from observation-action pairs, humans learn to imitate in a more flexible and efficient manner: learning behaviors by simply “watching.” In this article, we propose a “watch-and-act” imitation learning pipeline that endows a robot with the ability of learning diverse manipulations from visual demonstrations. Specifically, we address this problem by intuitively casting it as two subtasks: 1) understanding the demonstration video and 2) learning the demonstrated manipulations. First, a captioning module based on visual change is presented to understand the demonstration by translating the demonstration video into a command sentence. Then, to execute the captioning command, a manipulation module that learns the demonstrated manipulations is built upon an instance segmentation model and a manipulation affordance prediction model. We validate the superiority of the two modules over existing methods separately via extensive experiments and demonstrate the whole robotic imitation system developed based on the two modules in diverse scenarios using a real robotic arm. Supplementary video is available athttps://vsislab.github.io/watch-and-act/. Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Hesheng Wang 0001, Yibin Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Effective Safety Strategy for Mobile Robots Based on Laser-Visual Fusion in Home EnvironmentsabstractThe proven efficacy of safety strategies based on 2-D laser rangefinder (LRF) strongly stimulates their application to mobile robots operating in the home environment. However, it remains a challenge for the robot to avoid collisions with all obstacles in the environment. Since LRF can only scan a horizontal slice of the world, some objects cannot be fully observed, such as tables and chairs. In this article, an effective solution based on laser-visual fusion is presented to enhance the safety of the robot. First, a vision sensor is adopted to help detect obstacles that are not fully visible to LRF. Then we propose a method to convert the depth information of the visual image into 2-Dpseudo-laser datarepresentation. With this representation, a strategy for 2-D mapping is developed. On this basis, a novel map fusion algorithm is proposed to generate an improved grid map that amends the incorrect representation of obstacles on the traditional 2-D grid map. We further investigate a robot autonomous navigation strategy that considers LRF data and pseudo-laser data to avoid all obstacles. Experimental results show that the improved grid map together with the presented navigation strategy allows the robot not only to plan a “real” collision-free path, but also to navigate safely in both static and dynamic scenarios, and the proposed strategies can significantly enhance the performance of robot navigation in terms of safety, reliability and robustness. Ying Zhang 0043, Guohui Tian, Xuyang Shao 0002, Jiyu Cheng |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Autonomous Multi-View Navigation via Deep Reinforcement LearningabstractIn this paper, we propose a novel deep reinforcement learning (DRL) system for the autonomous navigation of mobile robots that consists of three modules: map navigation, multi-view perception and multi-branch control. Our DRL system takes as the input a routed map provided by a global planner and three RGB images captured by a multi-camera setup to gather global and local information, respectively. In particular, we present a multi-view perception module based on an attention mechanism to filter out redundant information caused by multi-camera sensing. We also replace raw RGB images with low-dimensional representations via a specifically designed network, which benefits a more robust sim2real transfer learning. Extensive experiments in both simulated and real-world scenarios demonstrate that our system outperforms state-of-the-art approaches. Xueqin Huang, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001 |
ICRA | 5 |
| 2021 | PackerBot: Variable-Sized Product Packing with Heuristic Deep Reinforcement LearningabstractProduct packing is a typical application in ware-house automation that aims to pick objects from unstructured piles and place them into bins with optimized placing policy. However, it still remains a significant challenge to finish the product packing tasks in general logistics scenarios where the objects are variable-sized and the configurations are complex. In this work, we present the PackerBot, a complete robotic pipeline for performing variable-sized product packing in unstructured scenes. First, by leveraging the imperfect experience of human packer, we propose a heuristic DRL framework for learning optimal online 3D bin packing policy. Then we integrate it with a 6-DoF suction-based picking module and a product size estimation module, leading to a complete product packing system, namely the PackerBot. Extensive experimental results show that our method achieves the state-of-the-art performance in both simulated and real-world tests. The video demonstration is available at: https://vsislab.github.io/packerbot. Zifei Yang, Shuai Song, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001 |
IROS | 6 |
| 2021 | Semantic-Aware Informative Path Planning for Efficient Object Search Using Mobile RobotabstractIn this article, a novel informative path planning (IPP) framework is proposed for efficient robotic object search. We innovatively reformulate the object search into an IPP problem, which takes account of the knowledge of possible target object locations. To model the target object distribution knowledge, the semantic information of the focused environment is utilized to obtain the probabilities of finding the target object at possible locations. Then, the probability distribution is modeled by Gaussian mixture model (GMM) to generate an information map. Based on the map, a sampling-based IPP method is proposed to minimize the object search cost. It is worth noting that the object search path is planned with a tree structure and evaluated by a utility function that concerns both search information gain and path cost. Moreover, to improve the quality of the search path, a novel informative sampling strategy and a rewire mechanism are conceived. The performance of the proposed object search framework is fully evaluated through both simulation experiments and real-world tests with a mobile robot platform. Results demonstrated that our method can find the target object efficiently and robustly with shorter path length than three comparative methods in the literature and the mobile robot shows human-like behavior when searching for the target object. Chaoqun Wang 0009, Jiyu Cheng, Wenzheng Chi, Tingfang Yan, Max Q.-H. Meng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Improving Visual Localization Accuracy in Dynamic Environments Based on Dynamic Region RemovalabstractVisual localization is a fundamental capability in robotics and has been well studied for recent decades. Although many state-of-the-art algorithms have been proposed, great success usually builds on the assumption that the working environment is static. In most of the real scenes, the assumption cannot hold because there are inevitably moving objects, especially humans, which significantly degrade the localization accuracy. To address this problem, we propose a robust visual localization system building on top of a feature-based visual simultaneous localization and mapping algorithm. We design a dynamic region detection method and use it to preprocess the input frame. The detection process is achieved in a Bayesian framework which considers both the prior knowledge generated from an object detection process and observation information. After getting the detection result, feature points extracted from only the static regions will be used for further visual localization. We performed the experiments on the public TUM data set and our recorded data set, which shows the daily dynamic scenarios. Both qualitative and quantitative results are provided to show the feasibility and effectiveness of the proposed method. Jiyu Cheng, Hong Zhang 0013, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2020 | Robust Visual Localization in Dynamic Environments Based on Sparse Motion RemovalabstractVisual localization has been well studied in recent decades and applied in many fields as a fundamental capability in robotics. However, the success of the state of the arts usually builds on the assumption that the environment is static. In dynamic scenarios where moving objects are present, the performance of the existing visual localization systems degrades a lot due to the disturbance of the dynamic factors. To address this problem, we propose a novel sparse motion removal (SMR) model that detects the dynamic and static regions for an input frame based on a Bayesian framework. The similarity between the consecutive frames and the difference between the current frame and the reference frame are both considered to reduce the detection uncertainty. After the detection process is finished, the dynamic regions are eliminated while the static ones are fed into a feature-based visual simultaneous localization and mapping (SLAM) system for further visual localization. To verify the proposed method, both qualitative and quantitative experiments are performed and the experimental results have demonstrated that the proposed model can significantly improve the accuracy and robustness for visual localization in dynamic environments.Note to Practitioners-This article was motivated by the visual localization problem in dynamic environments. Visual localization is well applied in many robotic fields such as path planning and exploration as the basic capability for a mobile robot. In the GPS-denied environments, one robot needs to localize itself through perceiving the unknown environment based on a visual sensor. In real-world scenes, the existence of the moving objects will significantly degrade the localization accuracy, which makes the robot implementation unreliable. In this article, an SMR model is designed to handle this problem. Once receiving a frame, the proposed model divides it into dynamic and static regions through a Bayesian framework. The dynamic regions are eliminated, while the static ones are maintained and fed into a feature-based visual SLAM system for further visual localization. The proposed method greatly improves the localization accuracy in dynamic environments and guarantees the robustness for robotic implementation. Jiyu Cheng, Chaoqun Wang 0009, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 1 |