EDBT 2026 Demo / reviewers in the wild / expert
Hui Cheng 0002
dblp:87/3633-2
· DBLP profile ↗
17ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0003-2579-7004ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsabstractRecent advancements in diffusion-based imitation learning, which shows impressive performance in modeling multimodal distributions and training stability, have led to substantial progress in various robot learning tasks. In visual navigation, previous diffusion-based policies typically generate action sequences by initiating from denoising Gaussian noise. However, the target action distribution often diverges significantly from Gaussian noise, leading to redundant denoising steps and increased learning complexity. Additionally, the sparsity of effective action distributions makes it challenging for the policy to generate accurate actions without guidance. To address these issues, we propose a novel, unified visual navigation framework leveraging the denoising diffusion bridge models named NaviBridger. This approach enables action generation by initiating from any informative prior actions, enhancing guidance and efficiency in the denoising process. We explore how diffusion bridges can enhance imitation learning in visual navigation tasks and further examine three source policies for generating prior actions. Extensive experiments in both simulated and real-world indoor and outdoor scenarios demonstrate that NaviBridger accelerates policy inference and outperforms the baselines in generating target action sequences. Code is available at https: //github.com/hren20/NaiviBridger. Hao Ren 0006, Yiming Zeng 0008, Zetong Bi, Zhaoliang Wan, Junlong Huang, Hui Cheng 0002 |
CVPR | 6 |
| 2025 | NaviDiffusor: Cost-Guided Diffusion Model for Visual NavigationabstractVisual navigation, a fundamental challenge in mobile robotics, demands versatile policies to handle diverse environments. Classical methods leverage geometric solutions to minimize specific costs, offering adaptability to new scenarios but are prone to system errors due to their multi-modular design and reliance on hand-crafted rules. Learning-based methods, while achieving high planning success rates, face difficulties in generalizing to unseen environments beyond the training data and often require extensive training. To address these limitations, we propose a hybrid approach that combines the strengths of learning-based methods and classical approaches for RGB-only visual navigation. Our method first trains a conditional diffusion model on diverse path-RGB observation pairs. During inference, it integrates the gradients of differentiable scene-specific and task-level costs, guiding the diffusion model to generate valid paths that meet the constraints. This approach alleviates the need for retraining, offering a plug-and-play solution. Extensive experiments in both indoor and outdoor settings, across simulated and real-world scenarios, demonstrate zero-shot transfer capability of our approach, achieving higher success rates and fewer collisions compared to baseline methods. Code will be released at https://github.com/SYSU-RoboticsLab/NaviD. Yiming Zeng 0008, Hao Ren 0006, Shuhang Wang, Junlong Huang, Hui Cheng 0002 |
ICRA | 5 |
| 2025 | SFExplorer: A Surface-Frontier-based Efficient UAV Exploration Method for Large-Scale Unknown EnvironmentsabstractAutonomous exploration in unknown environments is a crucial challenge for various applications of unmanned aerial vehicles (UAVs). However, in large-scale scenarios, existing methods suffer from inefficient environmental information acquisition, computationally expensive exploration planning, and inconsistent motion. In this work, we present a novel method for rapid UAV autonomous exploration in large-scale environments. We develop a surface frontier guided viewpoints generation strategy that supports efficient coverage of scenario. Besides, we introduce an incremental viewpoint clustering method to approximate distant viewpoints using fewer anchor points, decreasing the computational costs of exploration tour planning. Building upon this, we propose a history-informed tour planning method that incorporates information from previous tour into the optimization process, maintaining motion consistency. Extensive simulation experiments validate that our method outperforms existing state-of-the-art methods in terms of exploration time, travel distance, and run time. Various real-world experiments are conducted to indicate the practicality of our approach. The source code will be released to benefit the community1. Peiming Duan, Xiaoxun Zhang, Lanxiang Zheng, Junlong Huang, Jiahui Liang, Hui Cheng 0002 |
IROS | 6 |
| 2025 | FASTEX: Fast UAV Exploration in Large-Scale Environments Using Dynamically Expanding Grids and Coverage PathsabstractAutonomous exploration is essential for the effective deployment of quadrotors in various applications. However, existing approaches face significant challenges in large-scale environments, particularly in balancing global coverage efficiency and computational overhead. These limitations often result in poor adaptability to environmental changes and redundant revisits to previously explored areas, reducing overall exploration efficiency. To address these issues, we propose FASTEX, a fast UAV exploration framework designed for large-scale environments, using dynamically expanding grids and coverage paths to improve exploration efficiency. To support efficient exploration planning in large-scale scenarios, we introduce an efficient environment preprocessing method, including a dynamic grid expansion mechanism and a sparse roadmap. Furthermore, we present a hierarchical exploration planning framework that integrates an incremental global planner with a local planner, ensuring high coverage and computational efficiency. Extensive simulation tests demonstrate the superior performance and robustness of the proposed method compared to the state-of-the-art methods. In addition, we conduct various real-world experiments to validate the feasibility of our autonomous exploration system. Xiaoxun Zhang, Peiming Duan, Lanxiang Zheng, Junlong Huang, Hui Cheng 0002 |
IROS | 5 |
| 2025 | RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Embodied IntelligenceabstractThis paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy.
To achieve it, we propose the RAPID Hand, a co-optimized hardware and software platform where the compact 20-DoF hand, robust whole-hand perception, and high-DoF teleoperation interface are jointly designed.
Specifically, RAPID Hand adopts a compact and practical hand ontology and a hardware-level perception framework that stably integrates wrist-mounted vision, fingertip tactile sensing, and proprioception with sub-7 ms latency and spatial alignment.
Collecting high-quality demonstrations on high-DoF hands is challenging, as existing teleoperation methods struggle with precision and stability on complex multi-fingered systems.
We address this by co-optimizing hand design, perception integration, and teleoperation interface through a universal actuation scheme, custom perception electronics, and two retargeting constraints. We evaluate the platform’s hardware, perception, and teleoperation interface. Training a diffusion policy on collected data shows superior performance over prior works, validating the system’s capability for reliable, high-quality data collection.
The platform is constructed from low-cost and off-the-shelf components and will be made public to ensure reproducibility and ease of adoption. Zhaoliang Wan, Zetong Bi, Zida Zhou, Hao Ren 0006, Yiming Zeng 0008, Lu Qi 0001, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002 |
NeurIPS | 10 |
| 2025 | Energy Efficient Scheduling for Position Reconfiguration of Swarm DronesabstractEnhancing the energy efficiency of drones, particularly in extending the flight lifetime, has emerged as a crucial area. Position reconfiguration has been explored as a mechanism to achieve this goal for swarm drones. Building on this concept, we investigate how position reconfiguration can be applied within urban wind environments to further extend the lifetime of drone swarms. Despite its potential, efficiently implementing position reconfiguration remains challenging. To address it, we propose an efficient position reconfiguration scheme that reduces the energy consumption imbalance of the swarm and prolongs the lifetime. The scheme includes: (1) a MIP (mixed integer programming)-based optimization method. (2) an approximation algorithm that runs in pseudo-polynomial time and without the need for an optimization solver. The scheme provides a complete position reconfiguration solution that determines (i) the number of position reconfiguration; (ii) when to perform reconfiguration; (iii) who to change positions. Simulation and experimental results demonstrate the effectiveness of our scheme. Note to Practitioners—In urban environments, the significant variation in wind speeds leads to an energy imbalance among swarm drones performing tasks. This paper addresses the practical issue of extending the lifetime of drones in such environments by optimizing position reconfiguration. Specifically, drones operating in high wind speed areas require more energy to maintain hovering, resulting in faster battery depletion. By allowing drones with more remaining energy to exchange positions with those experiencing higher energy consumption, the overall energy usage can be balanced, thus extending the mission duration. We propose an energy-efficient scheduling scheme to determine when and which drones should reconfigure their positions. The scheme strikes a balance between the benefits of reconfiguration and the associated energy costs, preventing unnecessary movement that could waste energy while ensuring drones do not deplete their batteries prematurely. This solution is particularly suited for drone swarms operating in urban environments. Future research could further explore the integration of this scheme into real-time drone fleet management systems. Mingxin Wei, Shuai Zhao 0004, Hui Cheng 0002, Kai Huang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Meta-Learning Enhanced Model Predictive Contouring Control for Agile and Precise Quadrotor FlightabstractIn agile quadrotor flight, accurately modeling the varying aerodynamic drag forces encountered at different speeds is critical. These drag forces significantly impact the performance and maneuverability of the UAV, especially during high-speed maneuvers. Traditional control models based on first principles struggle to capture these dynamics due to the complexity and variability of aerodynamic effects, which are challenging to model accurately. To address these challenges, this study proposes a meta-learning-based control strategy for accurately modeling quadrotor dynamics under varying speeds, treating each velocity condition as an independent learning task with a specifically trained neural network to ensure precise dynamic predictions. The meta-learning framework rapidly generates task-specific parameters adapted to speed variations by solving an optimization problem and employs an online incremental learning strategy to integrate real-time data for continuous model updates, enhancing system robustness. Regularization is introduced to prevent overfitting and improve generalizability. The integration of the meta-learned model into Model Predictive Contouring Control (MPCC) allows the system to achieve optimal control across different velocity levels, ensuring efficient and accurate flight control even during sharp turns and high-speed maneuvers. Extensive simulations and real-world experiments confirm that the proposed algorithm maintains a high level of control precision despite the nonlinear effects of rapid speed changes, complex flight trajectories and wind disturbances. The results highlight the advantages of combining meta-learning with adaptive control strategies, providing a robust framework for quadrotors operating in diverse and dynamic environments. Mingxin Wei, Lanxiang Zheng, Ying Wu 0011, Ruidong Mei, Hui Cheng 0002 |
IEEE Trans. Robotics | 5 |
| 2025 | AAGE: Air-Assisted Ground Robotic Autonomous Exploration in Large-Scale Unknown EnvironmentsabstractThe article presents an air-assisted ground robotic autonomous exploration framework, which leverages the high mobility and wide aerial perspective of unmanned aerial vehicles (UAVs) to assist unmanned ground vehicles (UGVs) in detailed exploration, enhancing exploration efficiency and improving the quality of point cloud collection in regions of interest in large-scale, unknown environments. In this framework, the UAV, equipped with an onboard RGB camera, rapidly surveys large unknown areas and generates a bird's eye view (BEV) to identify critical zones for UGV exploration. With prior information about the unexplored area's outline from the real-time shared BEV, the UGV can carry out more efficient and informed exploration from a global perspective. To maximize the utility of this prior information and optimize point cloud collection, a hierarchical exploration strategy and an attention mechanism are incorporated to guide the UGV's focus toward areas requiring detailed mapping, rather than broad, featureless regions. Real-world experiments validate the effectiveness of the framework, demonstrating significant improvements in exploration efficiency and point cloud collection compared to state-of-the-art methods. The results further show that even with a coarse BEV, the UGV's exploration efficiency is greatly enhanced. Lanxiang Zheng, Mingxin Wei, Ruidong Mei, Junlong Huang, Hui Cheng 0002 |
IEEE Trans. Robotics | 6 |
| 2024 | VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and ProprioceptionabstractThis paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real entries, collected via simulations in Mujoco and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/. Zhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi 0001, Wang Wei Lee, Minglei Lu, Xiao Teng, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002 |
ICML | 12 |
| 2024 | OPG-Policy: Occluded Push-Grasp Policy Learning with Amodal SegmentationabstractGoal-oriented grasping in dense clutter, a fundamental challenge in robotics, demands an adaptive policy to handle occluded target objects and diverse configurations. Previous methods typically learn policies based on partially observable segments of the occluded target to generate motions. However, these policies often struggle to generate optimal motions due to uncertainties regarding the invisible portions of different occluded target objects across various scenes, resulting in low motion efficiency. To this end, we propose OPG-Policy, a novel framework that leverages amodal segmentation to predict occluded portions of the target and develop an adaptive push-grasp policy for cluttered scenarios where the target object is partially observed. Specifically, our approach trains a dedicated amodal segmentation module for diverse target objects to generate amodal masks. These masks and scene observations are mapped to the future rewards of grasp and push motion primitives via deep Q-learning to learn the motion critic. Afterward, the push and grasp motion candidates predicted by the critic, along with the relevant domain knowledge, are fed into the coordinator to generate the optimal motion implemented by the robot. Extensive experiments conducted in both simulated and real-world environments demonstrate the effectiveness of our approach in generating motion sequences for retrieving occluded targets, outperforming other baseline methods in success rate and motion efficiency. Yiming Zeng 0008, Zhaoliang Wan, Hui Cheng 0002 |
IROS | 4 |
| 2024 | Efficient multi-agent cooperation: Scalable reinforcement learning with heterogeneous graph networks and limited communication
Zhi Li 0091, Yanjie Yang, Hui Cheng 0002 |
Knowl. Based Syst. | 3 |
| 2024 | Safe Learning-Based Control for Multiple UAVs Under Uncertain DisturbancesabstractThis paper presents a safe learning control strategy aimed at ensuring the accurate tracking of multiple unmanned aerial vehicles (UAVs) along their predetermined trajectories while also guaranteeing safety under uncertain environments such as trajectory conflict, airflow interference between UAVs, and external disturbances. The proposed control framework employs a high-level learning-based feedback linearization control combined with model predictive control (LB-FBL-MPC), coupled with a low-level safety barrier certificates and control Lyapunov function-based quadratic programs (SC), for nonlinear multiple-UAV systems. The high-level LB-FBL-MPC uses incremental Gaussian processes (IGPs) to learn uncertain disturbances online, and feedback linearization is applied to approximate the linear system. The MPC optimizes the reference trajectory based on the linearized dynamical model to enhance the adaptivity of the system. Furthermore, the low-level SC guarantees the safety and asymptotic stability of the multi-UAV system by using the prediction distribution of the IGPs. Ablation and benchmark comparison experiments demonstrate the efficacy of the proposed tracking control strategy.Note to Practitioners—Controlling multiple unmanned systems to achieve precision and safety in complex environmental disturbances is a challenge. Existing machine learning-based control frameworks are mostly limited by low learning efficiency and poor interpretability, making it difficult to deploy them in practical robot systems. This article introduces a machine learning and control theory combined framework for the safe control of multiple unmanned aerial vehicles. On the one hand, the framework enables UAVs to quickly learn uncertain environmental disturbances without the need for any pre-collected data. The use of a linearized system model greatly reduces computation time and provides a new approach for practical engineering deployment. On the other hand, by designing separate quadratic programming, we prove the stability and safety of the system. Extensive experiments demonstrate that the designed control strategy significantly improves system control performance while ensuring system safety. Mingxin Wei, Lanxiang Zheng, Ying Wu 0011, Hui Cheng 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Multitarget Pursuit-Evasion Based on Distributed and Competitive MechanismsabstractThe pursuit-evasion game is a critical problem in artificial intelligence and draws a lot of attentions. In this article, we study the coordinated capture of multiple targets using multiple pursuers. A task allocation algorithm named distributed multitarget k-winners-take-all (DMK-WTA) is proposed for multiple evaders and multiple pursuers in this article, which is distributed and based on competition. In this algorithm, pursuers obtain hunting qualification through competition of task cost. After that, robots are controlled by predator-pack encirclement model (PPM), through which pursuers can automatically navigate to the target while avoiding collisions with obstacles and other robots. Combined with DMK-WTA and PPM, a distributed multitarget pursuit scheme in a dynamic environment has formed. By comparing with Kuhn-Munkres algorithm and genetic algorithm, we have evaluated the efficiency of DMK-WTA algorithm. Extensive simulations and physical experiments are conducted on a variety of robots to verify the viability and applicability of the proposed approach. Ning Tan 0003, Yang Liu 0395, Ruikun Hu, Hui Cheng 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Safe learning-based gradient-free model predictive control based on cross-entropy method
Lei Zheng 0007, Rui Yang 0025, Zhixuan Wu, Jiesen Pan, Hui Cheng 0002 |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | RDC-SLAM: A Real-Time Distributed Cooperative SLAM System Based on 3D LiDARabstractTo improve the accuracy and efficiency of 3D LiDAR mapping, real-time cooperative SLAM has been considered to explore large and complex areas. To merge the individual maps from multiple robots, it is crucial to identify the common areas and obtain alternative matches between them. However, data transmission, especially in sparse networks with narrow bandwidth and limited range, is a challenging issue for the above problem. Since the distribution manner is suitable for limited communication, we proposed a common framework of 3D real-time distributed cooperative SLAM to fill the community gap. Assuming that each robot can communicate with others, the presented framework consists of four key modules: place recognition, relative pose estimation, distributed graph optimization, and communication. Meanwhile, we developed a complete real-time distributed cooperative SLAM system, called RDC-SLAM, by integrating state-of-the-art components into the framework. For computation and data transmission efficiency, descriptor-based registration is used instead of the conventional point cloud matching. An intensity-based descriptor is developed to perform the place recognition and obtain the alternative matches, while an eigenvalue-based segment descriptor is applied to further refine the relative pose estimations between these alternative matches. A distributed graph optimization method is utilized to obtain the maximum likelihood of multi-trajectory estimation. A communication protocol is also designed to associate data among robots that are easy to deploy and have low network requirements. The RDC-SLAM is validated by real-world experiments and exhibits superior performance concerning accuracy, computation efficiency, and data efficiency. Yachen Zhang, Long Chen 0005, Hui Cheng 0002, Wei Tu 0001, Dongpu Cao, Qingquan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | Multi-Scale Guided Mask Refinement for Coarse-to-Fine RGB-D PerceptionabstractPixel-level object segmentation is highly desired in many vision applications. Although segmentation methods purely based on visual input have achieved great success in the past decade, their further improvement is still hindered by the intrinsic drawback of color camouflage. With the rapid development and wide deployment of depth sensors, depth assisted methods are increasingly popular in visual perception systems. It is expected that RGB-D based methods can lead to significant performance improvements because color and depth are naturally complementary. However, how to merge color and depth modalities for segmentation with both high efficiency and high accuracy remains an open problem to be addressed. In this letter, we propose to divide the segmentation process into “coarse” and “refining” stages because a coarse segmentation can be easily obtained by various light-weight methods. In this way, we can tackle this problem by focusing on the refinement of coarse segmentations. In particular, we propose a multi-scale approach that selectively inherits the effective features of both edge-preserving filtering and deep neural networks. The proposed approach is evaluated on several benchmark datasets, respectively, using the coarse segmentations from background subtraction and object detection as the input. Numerous results indicate that our approach can achieve significant accuracy improvements compared to other alternatives, demonstrating superior edge-preserving capability. Besides an effective method for merging RGB-D information, our study on the capability of coarse-to-fine refinement also brings new inspirations for designing light-weight perception systems. Chongyu Chen, Haoguang Huang, Chuangrong Chen, Zhuoqi Zheng, Hui Cheng 0002 |
IEEE Signal Process. Lett. | 5 |
| 2018 | Image-to-Video Person Re-Identification With Temporally Memorized Similarity LearningabstractWith the development of video surveillance in public safety field, there is an increasing research on person re-identification (re-id). In this paper, we address the image-to-video person re-id, in which the probe is an image and the gallery is consists of videos captured by nonoverlapping cameras. Compared with image, video sequence contains more temporal information that can be explored to improve the performance of re-identification system. However, it is challenging to model temporal information in the matching process of image-to-video person re-id. In this paper, we proposed a novel temporally memorized similarity learning neural network for this problem. In specific, the proposed network mainly consisted of two parts, including feature representation sub-network and similarity sub-network. In the first part, we adopted a convolutional neural network (CNN) to extract features from the input image. Given a video sequence of a person, features were first extracted from each its frame by using CNN and further forward to a long shot term memory (LSTM) network to encode the temporal information of video sequence. The outputs of LSTM were concatenated together as the feature vector of video sequences. Finally, the feature vectors of probe image and the video sequence were further forward to the similarity sub-network for distance metric learning. In the proposed framework, the feature representation and the similarity metric learning can be learned and optimized simultaneously. We evaluated the proposed framework on three public person re-id data sets, and the experimental results showed that the proposed approach is effective for the image-to-video person re-id. Dongyu Zhang 0002, Wenxi Wu, Hui Cheng 0002, Ruimao Zhang, Zhenjiang Dong, Zhaoquan Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |