VLDB 2026 Research / reviewers in the wild / expert
Zhe Liu 0022
dblp:70/1220-22
· DBLP profile ↗
81ranked-venue papers
13as first author
56since 2021 · last 2026
0000-0001-6753-0303ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 6 first-author · 36 since 2021Systems, architecture and hardware · 28 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DifFlow3D: Hierarchical Diffusion Models for Uncertainty-Aware 3D Scene Flow Estimationabstract3D scene flow represents the dense per-point motion field in dynamic scenes, playing a crucial role in various downstream tasks, including motion segmentation, dynamic scene reconstruction, 4D content generation, etc. However, previous regression-based works commonly suffer from unreliable correlations caused by locally constrained search ranges and struggle with the absence of timely feedback regarding the flow estimation uncertainty during training. To address these challenges, we propose a novel uncertainty-aware network for scene flow estimation, termed DifFlow3D, based on the conditional probabilistic diffusion model. Hierarchical diffusion-based flow estimation blocks are designed to enhance the correlation robustness and resilience to challenging cases, e.g., dynamics, noisy inputs, repetitive patterns, etc. To mitigate the generation diversity, three key flow-related features are leveraged as conditions in our diffusion model. Furthermore, we develop an uncertainty estimation module within diffusion to assess the reliability of estimated scene flow dynamically. A Hidden State Denoising strategy (HSD) is also introduced to further boost the stability of the reverse denoising process. Extensive experiments conducted on four scene flow datasets, including both synthetic and real-world datasets (FlyingThings3D, KITTI 2015, Argoverse, and Waymo Open), demonstrate the superiority of our proposed DifFlow3D. Compared to prior state-of-the-art methods, DifFlow3D has 26.0%, 36.4%, 35.3%, and 17.7% EPE3D reduction respectively across four datasets. Only trained on the synthetic FlyingThings3D dataset, our method achieves an unprecedented millimeter-level accuracy (0.0070 m EPE3D) on the real-scene KITTI dataset, highlighting its exceptional generalization capability. Additionally, our diffusion-based refinement paradigm can be seamlessly integrated as a plug-and-play module into existing scene flow networks, significantly enhancing their estimation accuracy. We also introduce our pre-trained scene flow estimator as explicit motion priors into the novel dynamic LiDAR view synthesis task, which validates its great potential for improving the 4D LiDAR reconstruction performance. Jiuming Liu, Weicai Ye, Guangming Wang 0001, Chaokang Jiang, Jinru Han, Zhe Liu 0022, Guofeng Zhang 0001, Hesheng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) requires agents to follow natural language instructions through environments, with memory-persistent variants demanding progressive improvement through accumulated experience. Existing approaches for memory-persistent VLN face critical limitations: they lack effective memory access mechanisms, instead relying on entire memory incorporation or fixed-horizon lookup, and predominantly store only environmental observations while neglecting navigation behavioral patterns that encode valuable decision-making strategies. We present Memoir, which employs imagination as a retrieval mechanism grounded by explicit memory: a world model imagines future navigation states as queries to selectively retrieve relevant environmental observations and behavioral histories. The approach comprises: 1) a language-conditioned world model that imagines future states serving dual purposes: encoding experiences for storage and generating retrieval queries; 2) Hybrid Viewpoint-Level Memory that anchors both observations and behavioral patterns to viewpoints, enabling hybrid retrieval; and 3) an experience-augmented navigation model that integrates retrieved knowledge through specialized encoders. Extensive evaluation across diverse memory-persistent VLN benchmarks with 10 distinct testing scenarios demonstrates Memoir's effectiveness: significant improvements across all scenarios, with 5.4% SPL gains on IR2R over the best memory-persistent baseline, accompanied by $8.3\times$8.3× training speedup and 74% inference memory reduction. The results validate that predictive retrieval of both environmental and behavioral memories enables more effective navigation, with analysis indicating substantial headroom (73.3% vs 93.4% upper bound) for this imagination-guided paradigm. Yunzhe Xu, Yiyuan Pan, Zhe Liu 0022 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Effective and Scalable Path Planning and Motion Coordination for Four-Way Shuttle Vehicles in Storage/Retrieval SystemsabstractRecently, Shuttle-Based Storage and Retrieval Systems (SBS/RSs) have emerged as a cornerstone of modern robotic handling systems. However, motion conflicts between four-way shuttle vehicles are introduced frequently due to the sparsely connected roadmap topology. Furthermore, the different time and energy costs between linear and direction-changing motions bring additional difficulties for classic solvers when applied directly. To address these issues, we propose Dynamic Graph-Driven A* (DGA*) for cooperative path planning in SBS/RSs. By fully exploiting the structured topology of SBS/RS roadmaps, the algorithm extracts sector-level topological graphs enriched with spatiotemporal attributes, effectively resolving conflicts while simultaneously optimizing motion costs. In addition, Recursive Preemption and Avoidance (RPA) is introduced to ensure motion coordination in congestion-prone areas commonly seen in SBS/RSs, through a combination of independent avoidance and recursive avoidance mechanisms. We implement our approach in a real warehouse with more than 240 vehicles and achieve a 17% improvement in system throughput while reducing runtime. Xingyao Han, Zhe Liu 0022, Jieshi Xu, Yuhong Tan, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Spatial-Aware and Viewpoint-Robust Vision-Language NavigationabstractVision-language navigation (VLN) requires an agent to follow visual observations and language instructions to navigate within a 3D environment. Most prior VLN approaches utilize recurrent units, topological graphs, or grid maps to represent the agent’s previously explored environment. However, these methods face two key limitations: (1) difficulty in accurately and adequately representing multi-level spatial information, and (2) a lack of robustness in semantic feature extraction against viewpoint variations. To overcome these challenges, we design a novel spatial-aware scene representation (SSR) as well as a viewpoint-robust feature extraction method. Specifically, for SSR, we first construct a multi-layer map that models the scene at different granularities, incorporating waypoints, objects, rooms, and floors. We then propose to extract spatial features from this multi-layer map based on a heterogeneous graph transformer and align them with the instructions, effectively guiding the agent’s planning process. As to viewpoint-robust feature extraction, we integrate 3D Gaussian splatting to capture the 3D geometry and visual texture of the environment, enabling novel view object observation and feature extraction. This improves observation coverage and enhances the viewpoint robustness of feature extraction. Extensive experiments demonstrated that our approach achieves state-of-the-art performance on two VLN benchmarks and shows strong performance in real-world scenarios. Source codes will be publicly available upon paper acceptance. Zhide Zhong, Xiangchen Liu, Xinhu Zheng, Zhe Liu 0022, Hesheng Wang 0001, Haoang Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Non-Communicative Decentralized Cooperative Navigation: Reinforcement Learning From Point CloudsabstractLast-mile transport is costly and operationally fragile due to frequent layout changes, mixed traffic with pedestrians and other robots, and unreliable connectivity that makes precise maps hard to maintain. Classical multi-agent planners rely on globally consistent maps or communication. Local BEV pipelines depend on tightly registered perception stacks. End-to-end vision models can be brittle across scene changes. While 3D LiDAR offers geometry that transfers well, its high dimensionality and occlusions complicate real-time, multi-agent control without messaging. We propose an end-to-end mapless, communication-free navigator that consumes a single onboard 3D LiDAR scan and outputs discrete actions for decentralized execution. The approach introduces a dual-channel LiDAR projection that preserves near-field free space while aligning goal guidance. A graph-attention interaction head infers neighbors’ motion tendencies from local observations only, which guides different coordination strategies for pedestrians and vehicles. A structure-aware dense reward stabilizes learning around occlusion-prone layouts (e.g., long walls, U-shaped bays). Our simulations demonstrate strong generalizability and scalability across layouts and sizes. The module trained only in random scenes with 10 agents transfers to crowd and warehouse styles, which to clusters with over$10^{3}$agents. Real-world trials validate sim-to-real transfer without retraining. We deploy the model on transport vehicles and small robots under changing environments. The results indicate robust, scalable, and deployment-friendly navigation for last-mile operations. Zhe Liu 0022, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language NavigationabstractHumans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of embodied agents in unseen environments. However, existing Vision-and-Language Navigation (VLN) agents lack a memory mechanism of this kind. To address this, we propose a novel architecture that equips agents with a reality-imagination hybrid memory system. This system enables agents to maintain and expand their memory through both imaginative mechanisms and navigation actions. Additionally, we design tailored pre-training tasks to develop the agent's imaginative capabilities. Our agent can imagine high-fidelity RGB images for future scenes, achieving state-of-the-art results in a Success rate weighted by Path Length (SPL). Yiyuan Pan, Yunzhe Xu, Zhe Liu 0022, Hesheng Wang 0001 |
AAAI | 3 |
| 2025 | FLAME: Learning to Navigate with Multimodal LLM in Urban EnvironmentsabstractLarge Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specialized VLN models. We introduce FLAME (FLAMingo-Architected Embodied Agent), a novel Multimodal LLM-based agent and architecture designed for urban VLN tasks that efficiently handles multiple observations. Our approach implements a three-phase tuning technique for effective adaptation to navigation tasks, including single perception tuning for street view description, multiple perception tuning for route summarization, and end-to-end training on VLN datasets. The augmented datasets are synthesized automatically. Experimental results demonstrate FLAME's superiority over existing methods, surpassing state-of-the-art methods by a 7.3% increase in task completion on Touchdown dataset. This work showcases the potential of Multimodal LLMs (MLLMs) in complex navigation tasks, representing an advancement towards applications of MLLMs in the field of embodied intelligence. Yunzhe Xu, Yiyuan Pan, Zhe Liu 0022, Hesheng Wang 0001 |
AAAI | 3 |
| 2025 | Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsabstractPoint cloud videos can faithfully capture real-world spatial geometries and temporal dynamics, which are essential for enabling intelligent agents to understand the dynamically changing world. However, designing an effective 4D backbone remains challenging, mainly due to the irregular and unordered distribution of points and temporal inconsistencies across frames. Also, recent transformer-based 4D backbones commonly suffer from large computational costs due to their quadratic complexity, particularly for long video sequences. To address these challenges, we propose a novel point cloud video understanding backbone purely based on the State Space Models (SSMs). Specifically, we first disentangle space and time in 4D video sequences and then establish the spatio-temporal correlation with the unified spatial-temporal Mamba blocks. The Intra-frame Spatial Mamba module is developed to encode locally similar geometric structures within a certain temporal stride. Subsequently, locally correlated tokens are delivered to the Inter-frame Temporal Mamba module, which integrates long-term point features across the entire video with linear complexity. Our proposed Mamba4D achieves competitive performance on the MSR-Action3D action recognition (+10.4% accuracy), HOI4D action segmentation (+0.7 F1 Score), and Synthia4D semantic segmentation (+0.19 mIoU) datasets. Mamba4D also has a significant efficiency improvement, especially for long video sequences, with 87.5% GPU memory reduction and × 5.36 speed-up. Codes are released at https://github.com/IRMVLab/Mamba4D. Jiuming Liu, Jinru Han, Angelica I. Avilés-Rivero, Chaokang Jiang, Zhe Liu 0022, Hesheng Wang 0001 |
CVPR | 6 |
| 2025 | Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic WarehousingabstractIn warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing framework aiming to foresee and act in advance, which consists of task flow prediction and hybrid task allocation. For task prediction, we design the spatio-temporal representations of the task flow and introduce a periodicity-decoupled mechanism tailored for the generation patterns of aggregated orders, and then further extract spatial features of task distribution with a novel combination of graph structures. In hybrid tasks allocation, we consider the known tasks and predicted future tasks simultaneously to optimize the task allocation. In addition, we consider factors such as predicted task uncertainty and sector-level efficiency to realize more balanced and rational allocations. We validate our task prediction model across datasets derived from factories, achieving SOTA performance. Furthermore, we implement our system in a real-world robotic warehouse, demonstrating more than 30% improvements in efficiency. Zhe Liu 0022, Xingyao Han, Shunbo Zhou, Hesheng Wang 0001 |
ICRA | 2 |
| 2025 | MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep ThinkingabstractMoving object segmentation plays a vital role in understanding dynamic visual environments. While existing methods rely on multi-frame image sequences to identify moving objects, single-image MOS is critical for applications like motion intention prediction and handling camera frame drops. However, segmenting moving objects from a single image remains challenging for existing methods due to the absence of temporal cues. To address this gap, we propose MovSAM, the first framework for single-image moving object segmentation. MovSAM leverages a Multimodal Large Language Model (MLLM) enhanced with Chain-of-Thought (CoT) prompting to search the moving object and generate text prompts based on deep thinking for segmentation. These prompts are cross-fused with visual features from the Segment Anything Model (SAM) and a Vision-Language Model (VLM), enabling logic-driven moving object segmentation. The segmentation results then undergo a deep thinking refinement loop, allowing MovSAM to iteratively improve its understanding of the scene context and inter-object relationships with logical reasoning. This innovative approach enables MovSAM to segment moving objects in single images by considering scene understanding. We implement MovSAM in the real world to validate its practical application and effectiveness for autonomous driving scenarios where the multi-frame methods fail. Furthermore, despite the inherent advantage of multi-frame methods in utilizing temporal information, MovSAM achieves state-of-the-art performance across public MOS benchmarks, reaching 92.5% on J&F. Our implementation will be available at https://github.com/IRMVLab/MovSAM. Chang Nie, Yiqing Xu, Guangming Wang 0001, Zhe Liu 0022, Yanzi Miao |
IROS | 4 |
| 2025 | GIPD: Global Intent Prediction and Decomposition of Cooperative Multi-Robot System in Non-Communication EnvironmentsabstractIn complex multi-robot application scenarios, particularly in dynamically adversarial, hazardous, or disaster environments, traditional cooperation paradigms face significant challenges due to unreliable or absent communication links. Achieving efficient cooperation in the absence of communication has become a key bottleneck limiting the performance of multirobot systems. In this paper, we propose a Global Intent Prediction and Decomposition (GIPD) framework that enables robots to perform cooperative behavior without relying on communication. Each robot independently infers a globally consistent intent based solely on its local observations, ensuring implicit alignment across the system. Given the inferred global intent, robots autonomously determine their responsibilities and select the most appropriate tasks. They then base their local decision-making on the global intent, selected tasks, and individual observations, thereby facilitating effective execution and cooperation. We validate our approach using the MPE and SMAC benchmarks. Additionally, real-world experiments involving multiple ships demonstrate the effectiveness and practical applicability of the proposed GIPD method. Zhe Liu 0022, Haoyu Wei, Duwen Zhai, Kefan Jin, Haibin Shao |
IROS | 2 |
| 2025 | Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual CalibrationabstractAutonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningful novelty. Moreover, existing curiosity mechanisms exhibit a uniform novelty bias, treating all unexpected observations equally. However, peer behavior novelty, which encode latent task dynamics, are often overlooked, resulting in suboptimal exploration in decentralized, communication-free MARL settings. To this end, inspired by how human children adaptively calibrate their own exploratory behaviors via observing peers, we propose a novel approach to enhance multi-agent exploration. We introduce CERMIC, a principled framework that empowers agents to robustly filter noisy surprise signals and guide exploration by dynamically calibrating their intrinsic curiosity with inferred multi-agent context. Additionally, CERMIC generates theoretically-grounded intrinsic rewards, encouraging agents to explore state transitions with high information gain. We evaluate CERMIC on benchmark suites including VMAS, Meltingpot, and SMACv2. Empirical results demonstrate that exploration with CERMIC significantly outperforms SoTA algorithms in sparse-reward environments. Yiyuan Pan, Zhe Liu 0022, Hesheng Wang 0001 |
NeurIPS | 2 |
| 2025 | Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual NavigationabstractVisual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-based agents typically increase architectural complexity that paradoxically become counterproductive in the small-sample regime. This paper introduce NeuRO, a integrated learning-to-optimize framework that tightly couples perception networks with downstream task-level robust optimization. Specifically, NeuRO addresses core difficulties in this integration: (i) it transforms noisy visual predictions under data scarcity into convex uncertainty sets using Partially Input Convex Neural Networks (PICNNs) with conformal calibration, which directly parameterize the optimization constraints; and (ii) it reformulates planning under partial observability as a robust optimization problem, enabling uncertainty-aware policies that transfer across environments. Extensive experiments on both unordered and sequential multi-object navigation tasks demonstrate that NeuRO establishes SoTA performance, particularly in generalization to unseen environments. Our work thus presents a significant advancement for developing robust, generalizable autonomous agents. Yiyuan Pan, Yunzhe Xu, Zhe Liu 0022, Hesheng Wang 0001 |
NeurIPS | 3 |
| 2025 | Memorize My Movement: Efficient Sensorimotor Navigation With Self-Motion-Based Spatial CognitionabstractNavigation is a fundamental capability for robots to operate in expansive spaces. Reliable navigation in unknown environments is crucial for deploying robots in areas such as disaster rescue and industrial inspection. In such scenarios, it is essential for robots to construct memories based on historical data to support long-term, optimized decision-making. However, many existing techniques focus on memorizing direct features from raw perceptions, often resulting in redundancy due to irrelevant textures and areas. This approach leads to inefficiencies in computation and storage, and produces a memory structure that lacks general applicability. We suggest that it may not be necessary to store specific scene features. Instead, recalling the robot’s episodic movements could provide sufficient cognitive cues for navigation. To address this, we introduce Memory Enhanced Navigation with Embedded Odometry (MENEO), a framework consisting of three steps: ego-motion estimation, memory aggregation, and adaptive policy generation. MENEO offers two main advantages: its streamlined architecture significantly boosts computational and storage efficiency, and its universal design supports various sensor modalities, adapts to multiple navigation tasks, and accommodates different scenarios. We test MENEO in two different environments: maze exploration using a lidar-IMU sensor, and image-goal visual navigation in photorealistic indoor scenes. In both cases, MENEO demonstrates competitive navigation performance, outperforming existing methods by reducing storage and computational requirements. Additionally, MENEO’s compact memory representation not only enhances adaptability across diverse environments but also shows flexibility in real-world applications.Note to Practitioners—In learning-based navigation, the memory mechanism is essential for long-term optimized policies. It allows intelligent robots to make informed decisions by utilizing a wide range of temporal and spatial cues derived from historical data. Traditional methods use various types of memory (such as internal, unstructured, or structured), but these often result in computational and storage inefficiencies due to the direct inclusion of complex and redundant raw scene features. Furthermore, because these memory systems are closely linked to specific scene features, they lack general applicability across different sensor configurations, scene types, and tasks. In this paper, we aim to eliminate the need to directly manage the redundant environmental features found in previous memory structures. Instead, we propose focusing solely on memorizing a robot’s self-movements. Since the pose sequence is streamlined and compact, our approach not only enhances computational and storage efficiency but also improves the interpretability and universality of the navigation system. This advantage enables MENEO to be seamlessly integrated into a wide variety of intelligent navigation systems. It is especially beneficial for small robots with limited computing power, such as those used in search and rescue operations, where enhanced memory can greatly enhance their autonomous navigation capabilities. Qiming Liu 0001, Dingbang Huang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Online Geometric Memory Generation and Maintenance for Visuomotor Navigation in Structural Dynamic EnvironmentsabstractAutonomous navigation substantially depends on memory mechanisms for enhancing path optimality. However, the dynamic nature of environments can cause inconsistencies between the stored scene memory and real-time perception, leading to potentially catastrophic navigation errors. Existing memory structures often fall short in addressing these challenges, as they typically only account for object-level dynamics and falter when faced with long-term alterations in environmental structure. To counter these constraints, this paper presents a learning-based framework designed to construct and maintain a geometric memory, thereby facilitating improved visuomotor navigation under structural dynamics. This proposed framework employs visual inputs to construct a geometric representation of the environment, and identifies structural changes by assessing the consistency between the established memory and instantaneous perception. To update the geometric memory efficiently, we introduce a memory updater grounded in a structure storage pool. Furthermore, a two-phase hierarchical planner is proposed to decompose navigation tasks and formulate smooth navigation strategies. Experimental results from photorealistic simulations underscore the efficacy of the proposed system in managing long-term dynamics and navigation control. The effectiveness of the proposed system is further corroborated through deployment and testing in real-world environments.Note to Practitioners—Classic geometry-based navigation methods hinge on the construction of a global map for spatial reasoning and optimized robot control. Within a learning-based pipeline, the dependence on memory information becomes critical for enabling robots to develop spatial and temporal awareness. While the maintenance of memory structures can significantly expand the field of the robot’s perception in time and space, discrepancies between historical memory and real-time perception in dynamic environments can lead to misguided decisions. Unlike most existing research that addresses short-term dynamic issues at the object level, this paper centers on long-term, large-scale dynamics precipitated by changes in environmental structure. We put forward a learning-based framework that incorporates dynamic perception, map maintenance, and hierarchical navigation. The experimental results highlight the efficiency and real-time processing capability of this method in handling structural dynamics, hence enhancing navigation efficiency. Qiming Liu 0001, Neng Xu, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Sample-Efficient Deep Reinforcement Learning of Mobile Manipulation for 6-DOF Trajectory FollowingabstractThe whole-body control of mobile manipulators for the 6-DOF trajectory following task is the basis of many continuous tasks. However, traditional control strategies rely on accurate models and expert knowledge for solving the trajectory following task. Deep reinforcement learning (DRL) provides a promising model-free solution, but it is sample-inefficient. To this end, we propose Trajectory Following Hindsight Experience Replay (TF-HER), a sample-efficient DRL algorithm for the whole-body coupled trajectory following task with dense rewards. TF-HER builds a multi-trajectory state space, and relabels the low-reward data to generate informative high-reward experiences. Also, the distributional shift caused by the relabeling is corrected by estimating the density ratio of relabeled experiences. Extensive demonstrations on both nonholonomic and holonomic bases in simulation validate that our algorithm can accelerate the model convergence and significantly improve the sample efficiency. Furthermore, we present real-world experiments to demonstrate the effectiveness of our approach. The code is available:https://github.com/IRMV-Manipulation-Group/TF-HER.Note to Practitioners—The whole-body 6-DOF trajectory tracking capability for mobile manipulators is crucial in industrial automation, serving as the foundation for a wide range of continuous and precise operations including automated assembly, welding, and material handling. This paper proposes a reinforcement learning approach to enhance the efficiency and effectiveness of mobile manipulators, requiring no prior model information. Beyond this specific application, this research has promising applications in other robotic systems because it is model-free. By utilizing a gradient-descent-based relabeling method and an adaptive density ratio estimator, we address distributional shifts and mitigate hindsight bias, guiding the mobile manipulator to accurately follow complex 6-DOF trajectories with dense rewards. The experimental results highlight our method’s superior performance over existing model-free techniques, with improved sample efficiency and reduced tracking error across diverse robotic platforms in both simulated and real-world settings. The successful policy transfer from simulation to physical robots with fine-tuning further validates the robustness and practical applicability. Qiyu Feng, Yixuan Zhou 0003, Jianghao Lin, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Visual Navigation for Embodied Agents Using Semantic-Based Multi-Modal Cognitive GraphabstractVisual navigation is fundamental for embodied agents operating in expansive workspaces. The cognitive abilities of these agents form the essential basis for creating intelligent behavioral patterns. Memory and reasoning are vital components among these abilities. The former enhances decision-making by preserving a wide array of episodic spatio-temporal perception cues, while the latter allows proactive and advanced probabilistic inference of task distributions based on long-term experiences. Despite individual studies on these two cognitive modalities, their integration for enhanced decision-making presents a considerable challenge due to their substantial differences in representation and behavioral characteristics. In this paper, we introduce Semantic-based Multi-modal Cognitive Graph (SMCG) for intelligent visual navigation. This framework is distinguished by its unified semantic-level representation of both memory and reasoning capabilities. Specifically, SMCG, rather than directly memorizing perceptual features as per previous methods, records observed object sequences. Simultaneously, reasoning is based on a semantic relation graph that represents correlations among objects. We additionally develop a hierarchical cognition extraction (HCE) pipeline and employ it to decode cognitive cues within SMCG and situation-aware subgraphs, thereby enhancing intelligent navigation behavior. Experimental results in image-goal navigation show pronounced performance improvements, credited to the effective induction and rational application of heterogeneous cognitive modalities. Qiming Liu 0001, Xinmin Du, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Visuomotor Navigation for Embodied Robots With Spatial Memory and Semantic Reasoning CognitionabstractThe fundamental prerequisite for embodied agents to make intelligent decisions lies in autonomous cognition. Typically, agents optimize decision-making by leveraging extensive spatiotemporal information from episodic memory. Concurrently, they utilize long-term experience for task reasoning and foster conscious behavioral tendencies. However, due to the significant disparities in the heterogeneous modalities of these two cognitive abilities, existing literature falls short in designing effective coupling mechanisms, thus failing to endow robots with comprehensive intelligence. This article introduces a navigation framework, the hierarchical topology-semantic cognitive navigation (HTSCN), which seamlessly integrates both memory and reasoning abilities within a singular end-to-end system. Specifically, we represent memory and reasoning abilities with a topological map and a semantic relation graph, respectively, within a unified dual-layer graph structure. Additionally, we incorporate a neural-based cognition extraction process to capture cross-modal relationships between hierarchical graphs. HTSCN forges a link between two different cognitive modalities, thus further enhancing decision-making performance and the overall level of intelligence. Experimental results demonstrate that in comparison to existing cognitive structures, HTSCN significantly enhances the performance and path efficiency of image-goal navigation. Visualization and interpretability experiments further corroborate the promoting role of memory, reasoning, as well as their online learned relationships, on intelligent behavioral patterns. Furthermore, we deploy HTSCN in real-world scenarios to further verify its feasibility and adaptability. Qiming Liu 0001, Guangzhan Wang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | End-to-End 2D-3D Registration Between Image and LiDAR Point Cloud for Vehicle LocalizationabstractRobot localization using a built map is essential for a variety of tasks including accurate navigation and mobile manipulation. A popular approach to robot localization is based on image-to-point cloud registration, which combines illumination-invariant LiDAR-based mapping with economical image-based localization. However, the recent works for image-to-point cloud registration either divide the registration into separate modules or project the point cloud to the depth image to register the RGB and depth images. In this paper, we present I2PNet, a novel end-to-end 2D-3D registration network, which directly registers the raw 3D point cloud with the 2D RGB image using differential modules with a united target. The 2D-3D cost volume module for differential 2D-3D association is proposed to bridge feature extraction and pose regression. The soft point-to-pixel correspondence is implicitly constructed on the intrinsic-independent normalized plane in the 2D-3D cost volume module. Moreover, we introduce an outlier mask prediction module to filter the outliers in the 2D-3D association before pose regression. Furthermore, we propose the coarse-to-fine 2D-3D registration architecture to increase localization accuracy. Extensive localization experiments are conducted on the KITTI, nuScenes, M2DGR, Argoverse, Waymo, and Lyft5 datasets. The results demonstrate that I2PNet outperforms the state-of-the-art by a large margin and has a higher efficiency than the previous works. Moreover, we extend the application of I2PNet to the camera-LiDAR online calibration and demonstrate that I2PNet outperforms recent approaches on the online calibration task. Source codes are released athttps://github.com/IRMVLab/I2PNet. Guangming Wang 0001, Yanfeng Guo, Zhe Liu 0022, Yixiang Zhu, Wolfram Burgard, Hesheng Wang 0001 |
IEEE Trans. Robotics | 5 |
| 2024 | DifFlow3D: Toward Robust Uncertainty-Aware Scene Flow Estimation with Iterative Diffusion-Based RefinementabstractScene flow estimation, which aims to predict per-point 3D displacements of dynamic scenes, is a fundamen-tal task in the computer vision field. However, previ-ous works commonly suffer from unreliable correlation caused by locally constrained searching ranges, and struggle with accumulated inaccuracy arising from the coarse-to-fine structure. To alleviate these problems, we propose a novel uncertainty-aware scene flow estimation network(DifFlow3D) with the diffusion probabilistic model. Iter-ative diffusion-based refinement is designed to enhance the correlation robustness and resilience to challenging cases, e.g. dynamics, noisy inputs, repetitive patterns, etc. To re-strain the generation diversity, three key flow-related features are leveraged as conditions in our diffusion model. Furthermore, we also develop an uncertainty estimation module within diffusion to evaluate the reliability of esti-mated scene flow. Our DifFlow3D achieves state-of-the-art performance, with 24.0% and 29.1% EPE3D reduction respectively on FlyingThings3D and KITTI 2015 datasets. Notably, our method achieves an unprecedented millimeter-level accuracy (O.0078m in EPE3D) on the KITTI dataset. Additionally, our diffusion-based refinement paradigm can be readily integrated as a plug-and-play module into ex-isting scene flow networks, significantly increasing their estimation accuracy. Codes are released at https:// github.com/IRMVLab/DifFlow3D. Jiuming Liu, Guangming Wang 0001, Weicai Ye, Chaokang Jiang, Jinru Han, Zhe Liu 0022, Guofeng Zhang 0001, Dalong Du, Hesheng Wang 0001 |
CVPR | 6 |
| 2024 | Frontier-Enhanced Topological Memory with Improved Exploration Awareness for Embodied Visual Navigation
Xinru Cui, Qiming Liu 0001, Zhe Liu 0022, Hesheng Wang 0001 |
ECCV (70) | 3 |
| 2024 | DVLO: Deep Visual-LiDAR Odometry with Local-to-Global Feature Fusion and Bi-directional Structure Alignment
Jiuming Liu, Dong Zhuo, Zhiheng Feng, Siting Zhu 0001, Chensheng Peng, Zhe Liu 0022, Hesheng Wang 0001 |
ECCV (10) | 6 |
| 2024 | EMIE-MAP: Large-Scale Road Surface Reconstruction Based on Explicit Mesh and Implicit Encoding
Guangming Wang 0001, Tiankun Zhao, Dongchao Gao, Zhe Liu 0022, Hesheng Wang 0001 |
ECCV (87) | 8 |
| 2024 | Traffic Flow Learning Enhanced Large-Scale Multi-Robot Cooperative Path Planning Under UncertaintiesabstractRobotic systems with hundreds or even thousands of robots are widely implemented in logistic and industrial applications. In such systems, cooperative path planning is of great importance, as local congestion and motion conflict may greatly degrade system performance, especially in the presence of uncertainties. Our idea is to consider traffic flow equilibrium in path planning to relieve any potential congestion and increase efficiency. In this paper, we propose a hierarchical framework, which includes a traffic flow prediction layer, a sector-level planning layer, and a road-level coordination layer. In traffic flow prediction, we propose a spatio-temporal graph neural network that integrates local information to predict the evolution of future robot density distribution. In sector-level planning, we generate sector-level paths that consider travel distance and traffic flow equilibrium simultaneously. In road-level coordination, we implement the conflict-based search algorithm within each sector to ensure conflict-free local paths. In addition, we also explicitly consider motion/communication uncertainties that are unavoided in practical systems. We validate our effectiveness in simulations with over 1000 robots, what’s more, real experiments are provided. Xingyao Han, Xinye Xiong, Qiming Liu 0001, Shunbo Zhou, Zhe Liu 0022 |
ICRA | 7 |
| 2024 | Toward Universal and Scalable Road Graph Partitioning for Efficient Multi-Robot Path PlanningabstractTo date, multi-robot path planning has primarily been addressed by centralized solvers, typically aiming to maintain optimality. However, given its NP-hard nature, directly applying existing solvers in large and complex scenarios proves inefficient. A promising alternative lies in adopting a divide- and-conquer strategy to break down the problem into manageable sub-problems. In this work, we propose a systematic, universal, and scalable graph partitioning method, aiming to automatically divide any real-world environment into multiple regions. Building upon this, we convert the path planning on the entire graph into distributed sub-region path planning and devise corresponding inter-regional strategies. Our work can be easily implementable in practical systems and effectively enhances the scalability of existing solvers. Experimentally, our approach contributes to a tenfold improvement in computational efficiency while only sacrificing about 10% of optimality. Xingyao Han, Zhe Liu 0022, Shunbo Zhou, Hesheng Wang 0001 |
IROS | 3 |
| 2024 | Cooperative Path Planning for Four-Way Shuttle Vehicles in Storage and Retrieval Systems: A Hierarchically Dynamic Graph-Based ApproachabstractRecently, Shuttle-based Storage and Retrieval Systems (SBS/RSs) have garnered significant attention from both academia and industry, owing to their high spatial utilization and rapid response speed. However, the weak connectivity of roadmaps in densely stored environments increases the likelihood of congestion and deadlocks when multiple four-way shuttles operate simultaneously, thereby imposing greater demands for collaborative path planning. Instead of exhaustively coordinating the shuttle motions during the off-line planning or online local control stages, we solve the cooperative path planning challenge from the perspective of altering the road graph structure dynamically. More specifically, we propose an approach to automatically transfer the typical road graph of SBS/RSs into a hierarchical graph with a reduced size, and then dynamically adjust its edge properties to prohibit any motion conflicts. In this manner, the planning problem of large-scale shuttle groups can be easily resolved and all the potential congestions can be eliminated inherently. Finally, we build a complete multi-shuttle cooperative path planning system adaptable for large-scale problems. Xingyao Han, Yuhong Tan, Zhe Liu 0022, Hesheng Wang 0001 |
IROS | 4 |
| 2024 | Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene RepresentationabstractIn the context of visual navigation in unknown scenes, both “exploration” and “exploitation” are equally crucial. Robots must first establish environmental cognition through exploration and then utilize the cognitive information to accomplish target searches. However, most existing methods for image-goal navigation prioritize target search over the generation of exploratory behavior. To address this, we propose the Navigation with Uncertainty-driven Exploration (NUE) pipeline, which uses an implicit and compact scene representation, NeRF, as a cognitive structure. We estimate the uncertainty of NeRF and augment the exploratory ability by the uncertainty to in turn facilitate the construction of implicit representation. Simultaneously, we extract memory information from NeRF to enhance the robot’s reasoning ability for determining the location of the target. Ultimately, we seamlessly combine the two generated abilities to produce navigational actions. Our pipeline is end-to-end, with the environmental cognitive structure being constructed online. Extensive experimental results on image-goal navigation demonstrate the capability of our pipeline to enhance exploratory behaviors, while also enabling a natural transition from the exploration to exploitation phase. This enables our model to outperform existing memory-based cognitive navigation structures in terms of navigation performance. Project page: https://github.com/IRMVLab/NUE-NeRF-nav Qiming Liu 0001, Zhe Liu 0022, Hesheng Wang 0001 |
IROS | 3 |
| 2024 | Adaptive Optimization Tracking Control for an Unmanned Aerial Vehicle with Disturbance SuppressionabstractAiming at the tracking control for an unmanned aerial vehicle (UAV) with uncertainty and external disturbances, an adaptive optimization control method is proposed based on the backstepping method. The model uncertainty in the UAV is estimated using neural networks. Moreover, a disturbance observer is designed to estimate and counteract the disturbance in real-time. In addition, reinforcement learning (RL) is used to solve optimal problems and overcome the difficult issues of the Hamilton-Jacobi-Bellman (HJB) equation. In the controller design, dynamic surfaces are employed to avoid complex derivation issues in the virtual controller and improve operating efficiency. Furthermore, the stability of the UAV system is proved by the Lyapunov theory. Finally, the effectiveness and superiority of the proposed method are verified through numerical simulations and real-world experiments. Meiying Yang, Xingyu Xia, Zhe Liu 0022 |
SMC | 4 |
| 2024 | Integrating Neural Radiance Fields End-to-End for Cognitive Visuomotor NavigationabstractWe propose an end-to-end visuomotor navigation framework that leverages Neural Radiance Fields (NeRF) for spatial cognition. To the best of our knowledge, this is the first effort to integrate such implicit spatial representation with embodied policy end-to-end for cognitive decision-making. Consequently, our system does not necessitate modularized designs nor transformations into explicit scene representations for downstream control. The NeRF-based memory is constructed online during navigation, without relying on any environmental priors. To enhance the extraction of decision-critical historical insights from the rigid and implicit structure of NeRF, we introduce a spatial information extraction mechanism named Structural Radiance Attention (SRA). SRA empowers the agent to grasp complex scene structures and task objectives, thus paving the way for the development of intelligent behavioral patterns. Our comprehensive testing in image-goal navigation tasks demonstrates that our approach significantly outperforms existing navigation models. We demonstrate that SRA markedly improves the agent's understanding of both the scene and the task by retrieving historical information stored in NeRF memory. The agent also learns exploratory awareness from our pipeline to better adapt to low signal-to-noise memory signals in unknown scenes. We deploy our navigation system on a mobile robot in real-world scenarios, where it exhibits evident cognitive capabilities while ensuring real-time performance. Qiming Liu 0001, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Toward Learning-Based Visuomotor Navigation With Neural Radiance FieldsabstractCreating memory representations is essential for developing viable navigation strategies for intelligent agents. Although neural radiance fields (NeRFs) have shown great promise as a novel method for spatial representation, their potential for integration into learning-based navigation as a memory structure has been largely overlooked in the existing literature. In this article, we introduce a navigation pipeline that incorporates NeRF into visuomotor navigation. Initially, we present a derivative radiance field that facilitates one-shot pose and depth estimation from a single query image. By assuming equivalence between density and space occupancy, we generate a geometric accessibility map based on an offline-constructed NeRF prior. Utilizing the above information, we design a global planner that decomposes long-term tasks by performing waypoint estimation and rendering. Finally, we employ an imitation-learned local controller to achieve a reliable navigation policy. Our pipeline effectively utilizes NeRF's compact spatial representation for task decomposition and action generation, enabling efficient navigation. Experimental results highlight its advantages over recent implicit and explicit memory approaches in image-goal navigation tasks. Moreover, we conduct interpretability studies and apply our algorithm in real-world scenarios to further attest to its practicality and effectiveness. Qiming Liu 0001, Nanxi Chen, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Learning of Long-Horizon Sparse-Reward Robotic Manipulator Tasks With Base ControllersabstractDeep reinforcement learning (DRL) enables robots to perform some intelligent tasks end-to-end. However, there are still many challenges for long-horizon sparse-reward robotic manipulator tasks. On the one hand, a sparse-reward setting causes exploration inefficient. On the other hand, exploration using physical robots is of high cost and unsafe. In this article, we propose a method of learning long-horizon sparse-reward tasks utilizing one or more existing traditional controllers named base controllers in this article. Built upon deep deterministic policy gradients (DDPGs), our algorithm incorporates the existing base controllers into stages of exploration, value learning, and policy update. Furthermore, we present a straightforward way of synthesizing different base controllers to integrate their strengths. Through experiments ranging from stacking blocks to cups, it is demonstrated that the learned state-based or image-based policies steadily outperform base controllers. Compared to previous works of learning from demonstrations, our method improves sample efficiency by orders of magnitude and improves performance. Overall, our method bears the potential of leveraging existing industrial robot manipulation systems to build more flexible and intelligent controllers. Guangming Wang 0001, Minjian Xin, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Biomimetic Morphing Quadrotor Inspired by Eagle Claw for Dynamic GraspingabstractThis paper presents a novel biomimetic morphing quadrotor design inspired by the morphology of an eagle claw during prey capture. The arms of the quadrotor are capable of vertical folding to enable dynamic grasping, mimicking the transition of the eagle claw from an open to a closed state. This transition is achieved through the rotation of a central servomotor and the associated movement of 20 links. Thanks to the closed-loop multi-link structure of the frame, the propellers of the quadrotor remain in a fixed orientation when the arms are folded, allowing for system stabilization at any arm rotation angle. The geometric property of the whole frame is analyzed to determine the relationships and constraints of the links, which is important in experimental vehicle fabrication. To handle possible physical property changes and external disturbances during grasping, the adaptive sliding mode controllers are applied. To deal with objects of unknown size in grasping tasks, an admittance filter is proposed for adaptive morphology. While in flight, our proposed morphing quadrotor is able to rapidly or continuously transition to any configuration within its range smoothly. Experimental results show the ability of the quadrotor to dynamically grasp various unknown objects at 0.4m/s without additional tools, as well as its versatility in traversal of narrow spaces and perching. Mengxin Xu, Qixin De, Dafang Yu, An Hu, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Robotics | 5 |
| 2023 | TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR OdometryabstractRecently, transformer architecture has gained great success in the computer vision community, such as image classification, object detection, etc. Nonetheless, its application for 3D vision remains to be explored, given that point cloud is inherently sparse, irregular, and unordered. Furthermore, existing point transformer frameworks usually feed raw point cloud of N×3 dimension into transformers, which limits the point processing scale because of their quadratic computational costs to the input size N. In this paper, we rethink the structure of point transformer. Instead of directly applying transformer to points, our network (TransLO) can process tens of thousands of points simultaneously by projecting points onto a 2D surface and then feeding them into a local transformer with linear complexity. Specifically, it is mainly composed of two components: Window-based Masked transformer with Self Attention (WMSA) to capture long-range dependencies; Masked Cross-Frame Attention (MCFA) to associate two frames and predict pose estimation. To deal with the sparsity issue of point cloud, we propose a binary mask to remove invalid and dynamic points. To our knowledge, this is the first transformer-based LiDAR odometry network. The experiment results on the KITTI odometry dataset show that our average rotation and translation RMSE achieves 0.500°/100m and 0.993% respectively. The performance of our network surpasses all recent learning-based methods and even outperforms LOAM on most evaluation sequences.Codes will be released on https://github.com/IRMVLab/TransLO. Jiuming Liu, Guangming Wang 0001, Chaokang Jiang, Zhe Liu 0022, Hesheng Wang 0001 |
AAAI | 4 |
| 2023 | RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud RegistrationabstractAlthough point cloud registration has achieved remarkable advances in object-level and indoor scenes, large-scale registration methods are rarely explored. Challenges mainly arise from the huge point number, complex distribution, and outliers of outdoor LiDAR scans. In addition, most existing registration works generally adopt a two-stage paradigm: They first find correspondences by extracting discriminative local features and then leverage estimators (eg. RANSAC) to filter outliers, which are highly dependent on well-designed descriptors and post-processing choices. To address these problems, we propose an end-to-end transformer network (RegFormer) for large-scale point cloud alignment without any further post-processing. Specifically, a projection-aware hierarchical transformer is proposed to capture long-range dependencies and filter outliers by extracting point features globally. Our transformer has linear complexity, which guarantees high efficiency even for large-scale scenes. Furthermore, to effectively reduce mismatches, a bijective association transformer is designed for regressing the initial transformation. Extensive experiments on KITTI and NuScenes datasets demonstrate that our RegFormer achieves competitive performance in terms of both accuracy and efficiency. Codes are available at https://github.com/IRMVLab/RegFormer. Jiuming Liu, Guangming Wang 0001, Zhe Liu 0022, Chaokang Jiang, Marc Pollefeys, Hesheng Wang 0001 |
ICCV | 3 |
| 2023 | RLSAC: Reinforcement Learning enhanced Sample Consensus for End-to-End Robust EstimationabstractRobust estimation is a crucial and still challenging task, which involves estimating model parameters in noisy environments. Although conventional sampling consensus-based algorithms sample several times to achieve robustness, these algorithms cannot use data features and historical information effectively. In this paper, we propose RLSAC, a novel Reinforcement Learning enhanced SAmple Consensus framework for end-to-end robust estimation. RLSAC employs a graph neural network to utilize both data and memory features to guide exploring directions for sampling the next minimum set. The feedback of downstream tasks serves as the reward for unsupervised training. Therefore, RL-SAC can avoid differentiating to learn the features and the feedback of downstream tasks for end-to-end robust estimation. In addition, RLSAC integrates a state transition module that encodes both data and memory features. Our experimental results demonstrate that RLSAC can learn from features to gradually explore a better hypothesis. Through analysis, it is apparent that RLSAC can be easily transferred to other sampling consensus-based robust estimation tasks. To the best of our knowledge, RLSAC is also the first method that uses reinforcement learning to sample consensus for end-to-end robust estimation. We release our codes at https://github.com/IRMVLab/RLSAC. Chang Nie, Guangming Wang 0001, Zhe Liu 0022, Luca Cavalli, Marc Pollefeys, Hesheng Wang 0001 |
ICCV | 3 |
| 2023 | Anomaly Detection For Robust Autonomous NavigationabstractHuman drivers are remarkably robust against various unexpected occurring variations and corruptions by understanding temporal changes and traffic scenes. In contrast, the neural network based autonomous navigation system can be easily affected by sensor data anomaly, like occlusion, sensor noise, challenging weather and illumination conditions. Such external disturbances are inevitable in practical driving applications. In this paper, we develop a semi-supervised anomaly detection module to detect the corrupted data while extracting the traffic scenario features. We further introduce an end-to-end robust autonomous navigation framework based on the idea that the consecutive frames of clean data depict a similar traffic scenario and the differences among the sequential data imply the dynamic state changes. By taking into consideration both spatial traffic scenario and temporal environmental variation, the model is able to achieve robust navigation against sensor data corruptions. We conduct experiments in CARLA platform and the evaluation results show the effectiveness of the proposed method. Kefan Jin, Fan Mu, Xingyao Han, Guangming Wang 0001, Zhe Liu 0022 |
ICRA | 5 |
| 2023 | Self-supervised Multi-frame Monocular Depth Estimation with Pseudo-LiDAR Pose EnhancementabstractDepth estimation is one of the most important tasks in scene understanding. In the existing joint self-supervised learning approaches of depth-pose estimation, depth estimation and pose estimation networks are independent of each other. They only use the adjacent image frames for pose estimation and lack the use of the estimated geometric information. To enhance the depth-pose association, we propose a monocular multi-frame unsupervised depth estimation framework, named PLPE-Depth. There are a depth estimation network and two pose estimation networks with image input and pseudo-LiDAR input. The main idea of our approach is to use the pseudo-LiDAR reconstructed from the depth map to estimate the pose of adjacent frames. We propose depth re-estimation with a better pose between the image pose and the pseudo-LiDAR pose to improve the accuracy of estimation. Besides, we improve the reconstruction loss and design a pseudo-LiDAR pose enhancement loss to facilitate the joint learning. Our approach enhances the use of the estimated depth information and strengthens the coupling between depth estimation and pose estimation. Experiments on the KITTI dataset show that our depth estimation achieves state-of-the-art performance at low resolution. Our source codes will be released on https://github.com/IRMVLabIPLPE-Depth. Guangming Wang 0001, Jiquan Zhong, Hesheng Wang 0001, Zhe Liu 0022 |
ICRA | 5 |
| 2023 | See What the Robot Can't See: Learning Cooperative Perception for Visual NavigationabstractWe consider the problem of navigating a mobile robot towards a target in an unknown environment that is endowed with visual sensors, where neither the robot nor the sensors have access to global positioning information and only use first-person- view images. In order to overcome the need for positioning, we train the sensors to encode and communicate relevant viewpoint information to the mobile robot, whose objective it is to use this information to navigate to the target along the shortest path. We overcome the challenge of enabling all the sensors (even those that cannot directly see the target) to predict the direction along the shortest path to the target by implementing a neighborhood-based feature aggregation module using a Graph Neural Network (GNN) architecture. In our experiments, we first demonstrate generalizability to previously unseen environments with various sensor layouts. Our results show that by using communication between the sensors and the robot, we achieve up to 2.0 × improvement in SPL (Success weighted by Path Length) when compared to a communication-free baseline. This is done without requiring a global map, positioning data, nor pre-calibration of the sensor network. Second, we perform a zero-shot transfer of our model from simulation to the real world. Laboratory experiments demonstrate the feasibility of our approach in various cluttered environments. Finally, we showcase examples of successful navigation to the target while both the sensor network layout as well as obstacles are dynamically reconfigured as the robot navigates. We provide a video demo11https://www.youtube.com/watch?v=kcrnr6RUgucw, the dataset, trained models, and source code22https://github.com/proroklab/sensor-guided-visual-nav. Jan Blumenkamp, Qingbiao Li, Binyu Wang, Zhe Liu 0022, Amanda Prorok |
IROS | 4 |
| 2023 | Lyapunov Constrained Safe Reinforcement Learning for Multicopter Visual ServoingabstractTraditional methods based on Lyapunov analysis and learning-based approaches such as reinforcement learning (RL) are two powerful tools in visual servo tasks. Traditional methods are interpretable and their stability can be guar-anteed by Lyapunov analysis. However, they tend to have a high dependency on an accurate system dynamic model. RL approaches learn to act based on past experiences and thus have higher adaptability on disturbances and errors. However, the training process is long and its safety or stability is generally hard to guarantee, making real-world training risky. In this paper, we propose a residual RL framework for training a multicopter to finish visual servo tasks under disturbances, guided by the system safety in terms of system Lyapunov function. Such an approach compensates for the lack of disturbance-rejection ability of the traditional method, and optimizes stability explicitly so the RL agent makes safer actions both during the training and in the final policy. A comparison between our approach and the baselines is provided in simulation, and real-world experiments on a multicopter are also carried out to show our effectiveness. We believe that this work moves one step toward achieving RL applications on real-world robotic systems. Dafang Yu, Mengxin Xu, Zhe Liu 0022, Hesheng Wang 0001 |
IROS | 3 |
| 2023 | Efficient 3D Deep LiDAR OdometryabstractAn efficient 3D point cloud learning architecture, named EfficientLO-Net, for LiDAR odometry is first proposed in this article. In this architecture, the projection-aware representation of the 3D point cloud is proposed to organize the raw 3D point cloud into an ordered data form to achieve efficiency. The Pyramid, Warping, and Cost volume (PWC) structure for the LiDAR odometry task is built to estimate and refine the pose in a coarse-to-fine approach. A projection-aware attentive cost volume is built to directly associate two discrete point clouds and obtain embedding motion patterns. Then, a trainable embedding mask is proposed to weigh the local motion patterns to regress the overall pose and filter outlier points. The trainable pose warp-refinement module is iteratively used with embedding mask optimized hierarchically to make the pose estimation more robust for outliers. The entire architecture is holistically optimized end-to-end to achieve adaptive learning of cost volume and mask, and all operations involving point cloud sampling and grouping are accelerated by projection-aware 3D feature learning methods. The superior performance and effectiveness of our LiDAR odometry architecture are demonstrated on KITTI, M2DGR, and Argoverse datasets. Our method outperforms all recent learning-based methods and even the geometry-based approach, LOAM with mapping optimization, on most sequences of KITTI odometry dataset. We open sourced our codes at: https://github.com/IRMVLab/EfficientLO-Net. Guangming Wang 0001, Xinrui Wu, Shuyang Jiang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | 3D Hierarchical Refinement and Augmentation for Unsupervised Learning of Depth and Pose From Monocular VideoabstractDepth and ego-motion estimations are essential for the localization and navigation of autonomous robots and autonomous driving. Recent studies make it possible to learn the per-pixel depth and ego-motion from the unlabeled monocular video. In this paper, a novel unsupervised training framework is proposed with 3D hierarchical refinement and augmentation using explicit 3D geometry. In this framework, the depth and pose estimations are hierarchically and mutually coupled to refine the estimated pose layer by layer. The intermediate view image is proposed and synthesized by warping the pixels in an image with the estimated depth and coarse pose. Then, the residual pose transformation can be estimated from the new view image and the image of the adjacent frame to refine the coarse pose. The iterative refinement is implemented in a differentiable manner in this paper, making the whole framework optimized uniformly. Meanwhile, a new image augmentation method is proposed for the pose estimation by synthesizing a new view image, which creatively augments the pose in 3D space but gets a new augmented 2D image. The experiments on KITTI demonstrate that our depth estimation achieves state-of-the-art performance and even surpasses recent approaches that utilize other auxiliary tasks. Our visual odometry outperforms all recent unsupervised monocular learning-based methods and achieves competitive performance to the geometry-based method, ORB-SLAM2 with back-end optimization. The source codes will be released soon at:https://github.com/IRMVLab/HRANet. Guangming Wang 0001, Jiquan Zhong, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Graph Relational Reinforcement Learning for Mobile Robot Navigation in Large-Scale Crowded EnvironmentsabstractMobile robot autonomous navigation in large-scale environments with crowded dynamic objects and static obstacles is still an essential yet challenging task. Recent works have demonstrated the potential of using deep reinforcement learning to enable autonomous navigation in crowds. However, only considering the human-robot interactions results in short-sighted and unsafe behaviors, and they typically use hand-crafted features and assume the global observation range, leading to large performance declines in large-scale crowded environments. Recent advances have shown the power of graph neural networks to learn local interactions among surrounding objects. In this paper, we consider autonomous navigation task in large-scale environments with crowded static and dynamic objects (such as humans). Particularly, local interactions among dynamic objects are learned for better-understanding their moving tendency and relational graph learning is introduced for aggregating both the object-object interactions and object-robot interactions. In addition, local observations are transformed into graphical inputs to achieve the scalability to various number of surrounding dynamic objects and various static obstacle patterns, and the globally guided reinforcement learning strategy is introduced to achieve the fixed-sized learning model even in large-scale complex environments. Simulation results validate our generalizability to various environments and advanced performance compared with existing works in large-scale crowded environments. In particular, our method with only local observations performs better than the benchmarks with global complete observability. Finally, physical robotic experiments demonstrate our effectiveness and practical applicability in real scenarios. Zhe Liu 0022, Jiaming Li 0011, Guangming Wang 0001, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Visual Servoing of Rigid-Link Flexible-Joint Manipulators in the Presence of Unknown Camera Parameters and Boundary OutputabstractExact position control of the flexible-joint manipulator (FJM) is a challenging task due to the manipulator’s nonlinearity and underactuated characteristic. For the three-dimensional (3-D) rigid-link FJM, the image-based visual servoing (IBVS) approach with link-position output constraint and unknown camera parameters is investigated in this article. The controller is designed to deal with three problems. First, the visual servoing control law is divided into two domains based on the singular perturbation method, where the control input for the fast subsystem is developed to damp out the flexible joint’s vibration. Second, the adaptive updating law for estimating the unknown camera parameters is presented in the slow subsystem. Third, to guarantee that the rigid-link position keep in set constraints, the controller for the slow subsystem is designed by introducing a barrier Lyapunov function. According to the Lyapunov theorem of stability, the proposed controller for the FJM is theoretically proved to be asymptotically stable. Several numerical simulation experiments are provided to illustrate that the presented control scheme is effective. Zhe Liu 0022, Tian Hao, Hesheng Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | What Matters for 3D Scene Flow Network
Guangming Wang 0001, Yunzhe Hu, Zhe Liu 0022, Yiyang Zhou, Masayoshi Tomizuka, Hesheng Wang 0001 |
ECCV (33) | 3 |
| 2022 | Visuomotor Reinforcement Learning for Multirobot Cooperative NavigationabstractThis article investigates the multirobot cooperative navigation problem based on raw visual observations. A fully end-to-end learning framework is presented, which leverages graph neural networks to learn local motion coordination and utilizes deep reinforcement learning to generate visuomotor policy that enables each robot to move to its goal without the need of environment map and global positioning information. Experimental results show that, with a few tens of robots, our approach achieves comparable performance with the state-of-the-art imitation learning-based approaches with bird-view state inputs. We also illustrate our generalizability to crowded and large environments and our scalability to ten times number of the training robots. In addition, we demonstrate that our model trained for multirobot case can also improve the success rate in the single-robot navigation task in unseen environments. Note to Practitioners—With the development of intelligent industrial and logistic systems, robotic transportation systems are widely implemented. However, existing multirobot path coordination and navigation approaches are basically under some unreasonable assumptions, which are very hard to be implemented in practical scenarios. This article aims to greatly promote the real application of learning-based multirobot cooperative navigation approach, in order to achieve the following. First, we introduce an end-to-end reinforcement learning framework instead of the commonly used imitation learning strategy, as the latter one needs exhaustive training data to cover all the scenarios and does not have the required generalizability. Second, we directly use the raw sensor data instead of the commonly used bird-eye-view semantic observations, as the latter one is generally not representative of practical application scenario from the robot perspective and cannot solve the occlusion issue. Third, we interpret our learned model to illustrate which parts of the input and shared observations contribute most to the robots’ final actions. The above interpretability ensures predictability (thus safety) of our visuomotor policy in practical applications. Our learned visuomotor policy has the ability to coordinate dozens of robots by only using raw visual observations in unknown environments without map nor global localization information, this is the first time in the literature. Our future work includes solving the sim-to-real issue and conducting physical experiments. Zhe Liu 0022, Qiming Liu 0001, Ling Tang 0002, Kefan Jin, Hongye Wang, Ming Liu 0001, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2022 | Integrated Task Allocation and Path Coordination for Large-Scale Robot Networks With UncertaintiesabstractArtificial intelligence-enhanced autonomous unmanned systems, such as large-scale autonomous robot networks, are widely used in logistic and industrial applications. In this article, we address the integrated task assignment, path planning, and coordination problem applied for large-scale robot networks with the existence of uncertainties. In particular, a novel generalized conflict graph is designed which encodes the traveling time cost of the subsequent path planning result of each task-robot assignment and also includes the predicted path conflicts of each two assignments. An integrated optimization problem which aims to minimize the total traveling cost and potential path conflicts simultaneously is first formulated and then transformed into a linear programming instance to obtain the optimal solution. In particular, to satisfy the real-time requirement in large-scale systems, a greedy solution is presented which has the near-optimal performance but can decrease the computational complexity by orders of magnitude. The optimality, scalability, robustness, and efficiency of our approach are demonstrated by comprehensive comparisons with existing state-of-the-art approaches. Note to Practitioners—With the development of artificial intelligence techniques, large-scale autonomous robot networks are increasingly used in the logistic warehouses, unmanned container terminals, and intelligence transportation systems. This article considers the large-scale networks with hundreds or even thousands of unmanned robots which are implemented in lifelong transportation systems with uncertainties existed in practical execution process. Our main concept is to simultaneously minimize the total time cost of all the tasks and the potential motion conflicts among all the robots in the subsequent execution stage, thus alleviating robot congestions, balancing traffic distributions, increasing system efficiency, and improving the robustness and scalability. Lifelong simulations with thousand robots illustrate that our approach can reduce more than 30% of the time steps consumed for coordinating robot motion conflicts, and in the meantime, the throughput, and overall system efficiency are improved. However, simulation results show that the system improvement decreases in the presence of extreme high uncertainties (such as temporary motion and communication failures of the robot), due to the inaccurate conflict prediction in the integrated optimization stage. Our future work includes the deep-learning-based traffic evolution prediction and the online reallocation and planning in highly dynamic scenarios. Zhe Liu 0022, Huanshu Wei, Hongye Wang, Haoang Li, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2022 | Fully Uncalibrated Image-Based Visual Servoing of 2DOFs Planar Manipulators With a Fixed CameraabstractWe consider the uncalibrated vision-based control problem of robotic manipulators in this work. Though lots of approaches have been proposed to solve this problem, they usually require calibration (offline or online) of the camera parameters in the implementation, and the control performance may be largely affected by parameter estimation errors. In this work, we present new fully uncalibrated visual servoing approaches for position control of the 2DOFs planar manipulator with a fixed camera. In the proposed approaches, no camera calibration is required, and numerical optimization algorithms or adaptive laws for parameter estimation are not needed. One benefit of such features is that exponential convergence of the image position errors can be ensured regardless of the camera parameter uncertainties. Generally, existing uncalibrated approaches only can guarantee asymptotical convergence of the position errors. Moreover, different from most existing approaches which assume that the robot motion plane and the image plane are parallel, one of the proposed approaches allows the camera to be installed at a general pose. This also simplifies the controller implementation and improves the system design flexibility. Finally, simulation and experimental results are provided to illustrate the effectiveness of the presented fully uncalibrated visual servoing approaches. Xinwu Liang, Hesheng Wang 0001, Yun-Hui Liu 0001, Bing You, Zhe Liu 0022, Zhongliang Jing, Weidong Chen 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | Spherical Interpolated Convolutional Network With Distance-Feature Density for 3-D Semantic Segmentation of Point CloudsabstractThe semantic segmentation of point clouds is an important part of the environment perception for robots. However, it is difficult to directly adopt the traditional 3-D convolution kernel to extract features from raw 3-D point clouds because of the unstructured property of point clouds. In this article, a spherical interpolated convolution operator is proposed to replace the traditional grid-shaped 3-D convolution operator. In addition, this article analyzes the defect of point cloud interpolation methods based on the distance as the interpolation weight and proposes the self-learned distance-feature density by combining the distance and the feature correlation. The proposed method makes the feature extraction of the spherical interpolated convolution network more rational and effective. The effectiveness of the proposed network is demonstrated on the 3-D semantic segmentation task of point clouds. Experiments show that the proposed method achieves good performance on the ScanNet dataset and Paris-Lille-3D dataset. The comparison experiments with the traditional grid-shaped 3-D convolution operator demonstrated that the newly proposed feature extraction operator improves the accuracy of the network and reduces the parameters of the network. The source codes will be released on https://github.com/IRMVLab/SIConv. Guangming Wang 0001, Yehui Yang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Viewpoint-Invariant Loop Closure Detection Using Step-Wise Learning With Controlling Embeddings of LandmarksabstractThe research on loop closure detection has been carried out for the last many years; however, loop closure detection efficiency is still not that good and has been affected by many factors, including illumination conditions, weather conditions, seasons, and viewpoint changes. The research on loop closure detection from a different viewpoint is still an open research problem. The paper proposes an efficient solution to loop closure detection from different viewpoints by using landmarks instead of whole frames and taking the deep and robust features of deep learning instead of handcrafted features. A different kind of training approach is used to train the deep CNN to get highly abstract embeddings of input landmarks. The approach has many advantages over the traditional training approach and can solve many complex problems. This paper has used this approach to force similar or closer embeddings for similar landmarks and is forced to have a large gap in embeddings of different landmarks. The proposed method endeavors viewpoint invariant features, and the astonishing power of deep learning makes the features robust to viewpoint changes, occlusions, and illumination variations. The proposed visual SLAM system is tested on six publicly available datasets, and the results are compared with the most popular Bag of Words methods like DBoW2, DBoW3, and state-of-the-art deep learning methods AlexNet, ResNeXt, FlyNet, AFDPR and Siamese network. The results show that our method is efficient in finding loop closures candidates from different viewpoints. Code is available athttps://github.com/IRMVLab/Step-wise-Learning. Azam Rafique Memon, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Learning To Identify Correct 2D-2D Line Correspondences on SphereabstractGiven a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both structured and unstructured scenes. Instead of geometric constraint, we leverage the spatial regularity on sphere. Specifically, we propose to map line correspondences into vectors tangent to sphere. We use these vectors to encode both angular and positional variations of image lines, which is more reliable and concise than directly using inclinations, midpoints or endpoints of image lines. Neighboring vectors mapped from correct matches exhibit a spatial regularity called local trend consistency, regardless of the type of scenes. To encode this regularity, we design a neural network and also propose a novel loss function that enforces the smoothness constraint of vector field. In addition, we establish a large real-world dataset for image line matching. Experiments showed that our approach outperforms state-of-the-art ones in terms of accuracy, efficiency and robustness, and also leads to high generalization. Haoang Li, Kai Chen 0028, Ji Zhao 0001, Jiangliu Wang, Pyojin Kim, Zhe Liu 0022, Yun-Hui Liu 0001 |
CVPR | 6 |
| 2021 | PWCLO-Net: Deep LiDAR Odometry in 3D Point Clouds Using Hierarchical Embedding Mask OptimizationabstractA novel 3D point cloud learning model for deep LiDAR odometry, named PWCLO-Net, using hierarchical embedding mask optimization is proposed in this paper. In this model, the Pyramid, Warping, and Cost volume (PWC) structure for the LiDAR odometry task is built to refine the estimated pose in a coarse-to-fine approach hierarchically. An attentive cost volume is built to associate two point clouds and obtain embedding motion patterns. Then, a novel trainable embedding mask is proposed to weigh the local motion patterns of all points to regress the overall pose and filter outlier points. The estimated current pose is used to warp the first point cloud to bridge the distance to the second point cloud, and then the cost volume of the residual motion is built. At the same time, the embedding mask is optimized hierarchically from coarse to fine to obtain more accurate filtering information for pose refinement. The trainable pose warp-refinement process is iteratively used to make the pose estimation more robust for outliers. The superior performance and effectiveness of our LiDAR odometry model are demonstrated on KITTI odometry dataset. Our method outperforms all recent learning-based methods and outperforms the geometry-based approach, LOAM with mapping optimization, on most sequences of KITTI odometry dataset. Our source codes will be released on https://github.com/IRMVLab/PWCLONet. Guangming Wang 0001, Xinrui Wu, Zhe Liu 0022, Hesheng Wang 0001 |
CVPR | 3 |
| 2021 | Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point EstimationabstractExisting vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts a spherical probability map of VP. Based on this map, we can detect all the VPs. Our method is reliable thanks to four technical novelties. First, we leverage the icosahedral spherical representation to express our probability map. This representation provides uniform pixel distribution, and thus facilitates estimating arbitrary positions of VPs. Second, we design a loss function that enforces the antipodal symmetry and sparsity of our spherical probability map to prevent over-fitting. Third, we generate the ground truth probability map that reasonably expresses the locations and uncertainties of VPs. This map unnecessarily peaks at noisy annotated VPs, and also exhibits various anisotropic dispersions. Fourth, given a predicted probability map, we detect VPs by fitting a Bingham mixture model. This strategy can robustly handle close VPs and provide the confidence level of VP useful for practical applications. Experiments showed that our method achieves the best compromise between generality, accuracy, and efficiency, compared with state-of-the-art approaches. Haoang Li, Kai Chen 0028, Pyojin Kim, Kuk-Jin Yoon, Zhe Liu 0022, Kyungdon Joo, Yun-Hui Liu 0001 |
ICCV | 5 |
| 2021 | A Registration-aided Domain Adaptation Network for 3D Point Cloud Based Place RecognitionabstractIn the field of large-scale SLAM for autonomous driving and mobile robotics, 3D point cloud based place recognition has aroused significant research interest due to its robustness to changing environments with drastic daytime and weather variance. However, it is time-consuming and effort-costly to obtain high-quality point cloud data for place recognition model training and ground truth for registration in the real world. To this end, a novel registration-aided 3D domain adaptation network for point cloud based place recognition is proposed. A structure-aware registration network is introduced to help to learn features with geometric information and a 6-DoFs pose between two point clouds with partial overlap can be estimated. The model is trained through a synthetic virtual LiDAR dataset through GTA-V with diverse weather and daytime conditions and domain adaptation is implemented to the real-world domain by aligning the global features. Our results outperform state-of-the-art 3D place recognition baselines or achieve comparable on the real-world Oxford RobotCar dataset with the visualization of registration on the virtual dataset. Zhijian Qiao, Hanjiang Hu, Weiang Shi, Zhe Liu 0022, Hesheng Wang 0001 |
IROS | 5 |
| 2021 | Prediction, Planning, and Coordination of Thousand-Warehousing-Robot Networks With Motion and Communication UncertaintiesabstractIn this article, we focus on resolving the traffic flow prediction, robot path planning, and motion coordination problems in large-scale warehousing robotics systems with thousand-robot networks. The warehousing environment is partitioned into several sectors, and a hierarchical framework is developed, which includes a centralized prediction and planning level and a decentralized local coordination level. In the centralized level, a traffic flow prediction algorithm is first proposed to predict the evolution of the robot density distribution in a future horizon and estimate the future traffic heat value of each sector. Based on this, the sector-level robot path can be generated in the time-expended sector graph by comprehensively considering the traveling distance and the predicted traffic heat value and will be dynamically updated by considering the most recent traffic information. In the coordination level, local cooperative A* algorithm, incorporated with the conflict-based searching strategy, is implemented within each sector to generate conflict-free road-level paths for all the robots in the sector simultaneously, and the rolling planning scheme is utilized in order to immediately react to robot motion uncertainties and communication disconnections. The effectiveness and practical applicability of the proposed approach are validated by large-scale simulations with more than one 1000 robots and real laboratory experiments.Note to Practitioners—Considering practical situations and requirements in industrial warehouses and automated logistics systems, this article resolves the life-long planning and coordination problems of large-scale robot networks and ensures the practical execution performance in the presence of robot motion uncertainties and temporary communication disconnections. Our main idea is to reduce robot congestions and improve warehouse working efficiency by balancing the traffic flow in the whole environment. To achieve this, we present a traffic flow prediction algorithm to estimate the robot density distribution in a future horizon and take this information into consideration in sector-level path planning. The reliability, scalability, and the real-time performance of the proposed solution are achieved by the presented hierarchical system framework and the dynamic planning scheme. The proposed concept and approach can also be used to coordinate other large-scale systems with multirobot or multi-AGV networks. Simulation and experimental results suggest that the proposed solution is effective and practically applicable, but a saturation phenomenon of the system capacity can be observed under a very heavy workload. In the future, we will investigate the relation between the maximum system capacity and the environment structure and make further efforts to optimize the environment structure and road layout in order to improve the warehouse working efficiency. Zhe Liu 0022, Hesheng Wang 0001, Huanshu Wei, Ming Liu 0001, Yun-Hui Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2021 | DASGIL: Domain Adaptation for Semantic and Geometric-Aware Image-Based LocalizationabstractLong-Term visual localization under changing environments is a challenging problem in autonomous driving and mobile robotics due to season, illumination variance, etc. Image retrieval for localization is an efficient and effective solution to the problem. In this paper, we propose a novel multi-task architecture to fuse the geometric and semantic information into the multi-scale latent embedding representation for visual place recognition. To use the high-quality ground truths without any human effort, the effective multi-scale feature discriminator is proposed for adversarial training to achieve the domain adaptation from synthetic virtual KITTI dataset to real-world KITTI dataset. The proposed approach is validated on the Extended CMU-Seasons dataset and Oxford RobotCar dataset through a series of crucial comparison experiments, where our performance outperforms state-of-the-art baselines for retrieval-based localization and large-scale place recognition under the challenging environment. Hanjiang Hu, Zhijian Qiao, Ming Cheng 0004, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Hierarchical Attention Learning of Scene Flow in 3D Point CloudsabstractScene flow represents the 3D motion of every point in the dynamic environments. Like the optical flow that represents the motion of pixels in 2D images, 3D motion representation of scene flow benefits many applications, such as autonomous driving and service robot. This paper studies the problem of scene flow estimation from two consecutive 3D point clouds. In this paper, a novel hierarchical neural network with double attention is proposed for learning the correlation of point features in adjacent frames and refining scene flow from coarse to fine layer by layer. The proposed network has a new more-for-less hierarchical architecture. The more-for-less means that the number of input points is greater than the number of output points for scene flow estimation, which brings more input information and balances the precision and resource consumption. In this hierarchical architecture, scene flow of different levels is generated and supervised respectively. A novel attentive embedding module is introduced to aggregate the features of adjacent points using a double attention method in a patch-to-patch manner. The proper layers for flow embedding and flow supervision are carefully considered in our network designment. Experiments show that the proposed network outperforms the state-of-the-art performance of 3D scene flow estimation on the FlyingThings3D and KITTI Scene Flow 2015 datasets. We also apply the proposed network to the realistic LiDAR odometry task, which is a key problem in autonomous driving. The experiment results demonstrate that our proposed network can outperform the ICP-based method and shows good practical application ability. The source codes will be released on https://github.com/IRMVLab/HALFlow. Guangming Wang 0001, Xinrui Wu, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World
Haoang Li, Pyojin Kim, Ji Zhao 0001, Kyungdon Joo, Zhe Liu 0022, Yun-Hui Liu 0001 |
ECCV (22) | 6 |
| 2020 | Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual OdometryabstractGiven a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In contrast, we propose a novel approach based on the robust "L2-minimizing estimate" (L2E) loss. We first define a novel cost function by integrating the projection constraint into the L2E loss. Then to efficiently obtain the global minimum of this function, we propose a hybrid strategy of a local optimizer and branch-and-bound. For branch-and-bound, we derive effective function bounds. Our approach can handle high outlier ratios, leading to high robustness. It can run reliably regardless of whether the initial pose is appropriate, providing high generality. Moreover, given a decent initial pose, it is suitable for real-time applications. Experiments on synthetic and real-world datasets showed that our approach outperforms state-of-the-art methods in terms of robustness and/or efficiency. Haoang Li, Wen Chen 0021, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 6 |
| 2020 | A Synchronization Approach for Achieving Cooperative Adaptive Cruise Control Based Non-Stop Intersection PassingabstractCooperative adaptive cruise control (CACC) of intelligent vehicles contributes to improving cruise control performance, reducing traffic congestion, saving energy and increasing traffic flow capacity. In this paper, we resolve the CACC problem from the viewpoint of synchronization control, our main idea is to introduce the spatial-temporal synchronization mechanism into vehicle platoon control to achieve the robust CACC and to further realize the non-stop intersection control. Firstly, by introducing the cross-coupling based space synchronization mechanism, a distributed control algorithm is presented to achieve the single-lane CACC in the presence of vehicle-to-vehicle (V2V) communications, which enables autonomous vehicles to track the desired platoon trajectory while synchronizing their longitudinal velocities to keeping the expected inter-vehicle distance. Secondly, by designing the enter-time scheduling mechanism (temporal synchronization), a high-level intersection control strategy is proposed to command vehicles to form a virtual platoon to pass through the intersection without stopping. Thirdly, a Lyapunov-based time-domain stability analysis approach is presented. Compared with the traditional string stability based approach, the proposed approach guarantees the global asymptotical convergence of the proposed CACC system. Experiments in the small-scale simulated system demonstrate the effectiveness of the proposed approach. Zhe Liu 0022, Huanshu Wei, Hanjiang Hu, Chuanzhe Suo, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICRA | 1 |
| 2020 | Online Trajectory Planning for an Industrial Tractor Towing Multiple Full TrailersabstractThis paper presents a novel solution for online trajectory planning of a full-size tractor-trailers vehicle composed of a car-like tractor and arbitrary number of passive full trailers. The motion planning problem for such systems was rarely addressed due to the complex nonlinear dynamics. A simulation-based prediction method is proposed to easily handle the complicated nonlinear dynamics and efficiently generate the obstacle-free and dynamically feasible trajectories. The vehicle dynamics model and a two-layer controller are used in the prediction. Implementation results on the real-world full-size industrial tractor-trailers vehicle are presented to validate the performance of the proposed methods. Wen Chen 0021, Shunbo Zhou, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 4 |
| 2020 | CUHK-AHU Dataset: Promoting Practical Self-Driving Applications in the Complex Airport Logistics, Hill and Urban EnvironmentsabstractThis paper presents a novel dataset targeting three types of challenging environments for autonomous driving, i.e., the industrial logistics environment, the undulating hill environment and the mixed complex urban environment. To the best of the author's knowledge, similar dataset has not been published in the existing public datasets, especially for the logistics environment collected in the functioning Hong Kong Air Cargo Terminal (HACT). Structural changes always suddenly appeared in the airport logistics environment due to the frequent movement of goods in and out. In the structureless and noisy hill environment, the non-flat plane movement is usual. In the mixed complex urban environment, the highly dynamic residence blocks, sloped roads and highways are included in a single collection. The presented dataset includes LiDAR, image, IMU and GPS data by repeatedly driving along several paths to capture the structural changes, the illumination changes and the different degrees of undulation of the roads. The baseline trajectories are provided which are estimated by Simultaneous Localization and Mapping (SLAM). Wen Chen 0021, Zhe Liu 0022, Shunbo Zhou, Haoang Li, Yun-Hui Liu 0001 |
IROS | 2 |
| 2020 | End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondencesabstract3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learning based approach to resolve the point cloud registration problem. Firstly, the revised LPD-Net is introduced to extract features and aggregate them with the graph network. Secondly, the self-attention mechanism is utilized to enhance the structure information in the point cloud and the cross-attention mechanism is designed to enhance the corresponding information between the two input point clouds. Based on which, the virtual corresponding points can be generated by a soft pointer based method, and finally, the point cloud registration problem can be solved by implementing the SVD method. Comparison results in ModelNet40 dataset validate that the proposed approach reaches the state-of-the-art in point cloud registration tasks and experiment resutls in KITTI dataset validate the effectiveness of the proposed approach in real applications. Huanshu Wei, Zhijian Qiao, Zhe Liu 0022, Chuanzhe Suo, Peng Yin 0001, Yueling Shen, Haoang Li, Hesheng Wang 0001 |
IROS | 3 |
| 2020 | Robust Dynamic State Estimation for Lateral Control of an Industrial Tractor Towing Multiple Passive TrailersabstractIn this paper, we propose a dynamic state estimation framework for lateral control of a heavy tractor-trailers system using only mass-produced low-cost sensors. This issue is challenging since the lateral velocity of the lead tractor is difficult to measure directly. The performance of existing dynamic model-based estimation methods will also be degraded, as different trailers and payloads cause the tractor model parameters to change. We address this issue by incorporating a kinematic estimator into a dynamic model-based estimation scheme. Accurate and reliable tire cornering stiffness and dynamics-informed lateral velocity of the lead tractor can be output in real-time by using our method. The stability and robustness of the proposed method are theoretically proved. The feasibility of our method is verified by full-scale experiments. It is also verified that the estimated model parameters and lateral states do improve the control performance by integrating the estimator into a lateral control system. Shunbo Zhou, Wen Chen 0021, Zhe Liu 0022, Hesheng Wang 0001, Yun-Hui Liu 0001 |
IROS | 4 |
| 2020 | Calibration-Free Image-Based Trajectory Tracking Control of Mobile Robots With an Overhead CameraabstractTo make the controller implementation easier and to enhance the system robustness and control performance in the presence of the camera parameter uncertainties, it is very desired to develop vision-based control approaches without any offline or online camera calibration. In this article, we propose a new calibration-free image-based trajectory tracking control scheme for nonholonomic mobile robots with a truly uncalibrated fixed camera. By developing a novel camera-parameter-independent kinematic model, both offline and online camera calibration can be avoided in the proposed scheme, and any knowledge of the camera is not needed in the controller design. The proposed trajectory tracking control scheme can guarantee exponential convergence of the image position and velocity tracking errors. To illustrate the performance of the proposed scheme, experimental results are provided in this article. Note to Practitioners-This article was motivated by the vision-based motion control problem of mobile robots in uncalibrated environments. Existing vision-based motion control approaches for nonholonomic mobile robots generally depend on offline precise/coarse or online numerical/adaptive calibration of the camera intrinsic and extrinsic parameters and require precise or coarse knowledge of the camera in their implementation. This article presents a novel calibration-free image-based trajectory tracking control scheme, which can be implemented easily in real environments without any offline or online calibration of the camera parameters and can be used to efficiently control the motion of nonholonomic mobile robots with an arbitrarily placed and truly unknown overhead camera. Experimental results show that the proposed scheme can achieve satisfactory trajectory tracking control performance despite the lack of any knowledge about the camera intrinsic and extrinsic parameters and the presence of unknown camera lens distortions, and hence, can provide a simple but efficient solution to the vision-based motion control problem of nonholonomic mobile robots. Xinwu Liang, Hesheng Wang 0001, Yun-Hui Liu 0001, Bing You, Zhe Liu 0022, Weidong Chen 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2020 | Purely Image-Based Pose Stabilization of Nonholonomic Mobile Robots With a Truly Uncalibrated Overhead CameraabstractAlthough many vision-based control methods have been proposed for nonholonomic mobile robots, in their implementation, it is usually necessary to calibrate the camera intrinsic and/or extrinsic parameters using offline/online parameter estimation algorithms or online adaptation laws. To avoid the tediousness of camera calibration and to make the system performance highly robust to camera parameter uncertainties, in this article, we propose novel image-based pose stabilization control approaches for nonholonomic mobile robots with a truly uncalibrated overhead fixed camera. In the proposed approaches, only image position information of three feature points from an overhead camera is used for controller design, while information from other sensors (such as wheel encoders) is not required. Furthermore, either offline or online camera calibration is not necessary, and no knowledge about the camera intrinsic and extrinsic parameters is needed, which also can greatly simplify the controller implementation. Simulation and experimental results are given to demonstrate the feasibility and effectiveness of the proposed purely image-based pose stabilization approaches. Xinwu Liang, Hesheng Wang 0001, Yun-Hui Liu 0001, Zhe Liu 0022, Bing You, Zhongliang Jing, Weidong Chen 0001 |
IEEE Trans. Robotics | 4 |
| 2020 | A Self-Repairing Algorithm With Optimal Repair Path for Maintaining Motion Synchronization of Mobile Robot NetworkabstractIn this paper, we consider the self-repairing problem from the viewpoint of robotics and our objective is not only to restore the logical network topology but also to maintain the motion synchronization of the physical mobile robot formation. A gradient-based self-repairing algorithm which only relies on the local interactions among coupling robots is presented. More specifically, aiming to optimize the repair path in a distributed manner, a gradient generation and diffusion mechanism is presented first, which can generate a stable gradient distribution in the robot formation. Then, based on the recursive self-repairing technique and the proposed gradient distribution, several self-repairing rules as well as the corresponding individual control method are presented to solve the self-repairing problem. The improvement of the proposed algorithm on the motion synchronism of the robot formation and the optimality of the selected repair path are proved by theoretical analyses. Finally, the effectiveness and the practical applicability of the proposed algorithm are validated by simulations and real experiments. Zhe Liu 0022, Weidong Chen 0001, Hesheng Wang 0001, Yun-Hui Liu 0001, Xiangyu Fu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan WorldabstractThe image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and the camera frame. To compute this rotation, state-of-the-art methods are based on either data sampling or parameter search, and they fail to guarantee the accuracy and efficiency simultaneously. In contrast, we propose to hybridize these two strategies. We first compute two degrees of freedom (DOF) of the above rotation by two sampled image lines, and then search for the optimal third DOF based on the branch-and-bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search is not sensitive to noise and achieves quasi-global optimality in terms of maximizing the number of inliers. Experiments on synthetic and real-world images showed that our method outperforms state-of-the-art approaches in terms of accuracy and/or efficiency. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICCV | 5 |
| 2019 | LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment AnalysisabstractPoint cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named LPD-Net (Large-scale Place Description Network), which can extract discriminative and generalizable global descriptors from the raw 3D point cloud. Two modules, the adaptive local feature extraction module and the graph-based neighborhood aggregation module, are proposed, which contribute to extract the local structures and reveal the spatial distribution of local features in the large-scale point cloud, with an end-to-end manner. We implement the proposed global descriptor in solving point cloud based retrieval tasks to achieve the large-scale place recognition. Comparison results show that our LPD-Net is much better than PointNetVLAD and reaches the state-of-the-art. We also compare our LPD-Net with the vision-based solutions to show the robustness of our approach to different weather and light conditions. Zhe Liu 0022, Shunbo Zhou, Chuanzhe Suo, Peng Yin 0001, Wen Chen 0021, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICCV | 1 |
| 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAMabstractA wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocular SLAM. First, we robustly cluster image lines. Based on these clusters, we compute the local Atlanta frames in the camera frame by solving polynomial equations. Our method provides the global optimum and satisfies inherent geometric constraints. Second, we define the posterior probabilities to refine the initial clusters and Atlanta frames alternately by the maximum a posteriori estimation. Third, based on multiple local Atlanta frames, we compute the global Atlanta frames in the world frame using Kalman filtering. We optimize rotations by the global alignment and then refine translations and 3D line-based map under the directional constraints. Experiments on both synthesized and real data have demonstrated that our approach outperforms state-of-the-art methods. Haoang Li, Yazhou Xing, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 5 |
| 2019 | A Hierarchical Framework for Coordinating Large-Scale Robot NetworksabstractIn this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guarantee the coordination performance in the presence of robot motion uncertainties. A hierarchical path planning and motion coordination structure is presented. The environment is divided into several sectors and a traffic heat-map is presented to describe the current sector-level traffic condition. In path planning level, the sector-level path is calculated by considering the path distance, the current traffic condition and the current robot uncertainty. In motion coordination level, local cooperative A* algorithm and conflict-based searching strategy are utilized within each sector to generate the collision-free local path of each robot in a rolling planning manner. The effectiveness and practical applicability of the proposed approach are validated by simulations with more than one thousand robots and real experiments. Zhe Liu 0022, Shunbo Zhou, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICRA | 1 |
| 2019 | Vision-Based Dynamic Control of Car-Like Mobile RobotsabstractMost existing controllers for Car-Like Mobile Robots (CLMR) are designed to handle dynamic effects by decoupling speed and steering controls, also assume that full states are accessible, which are unrealistic for real-world applications. This paper presents a combined speed and steering control system for CLMR. To provide the essential state for the controller, a newly developed visual algorithm is adopted for estimating the high-update rate longitudinal and lateral velocities of the robot which cannot be accurately measured by wheel encoders due to the skidding and slipping effects. The stability of the proposed system can be guaranteed by Lyapunov method since the velocity estimation error, the speed tracking error and the lateral deviation converging to zero simultaneously. Real-world experiments are conducted on an electric autonomous tractor with online estimation to demonstrate the feasibility of the approach. Shunbo Zhou, Zhe Liu 0022, Chuanzhe Suo, Hesheng Wang 0001, Yun-Hui Liu 0001 |
ICRA | 2 |
| 2019 | Retrieval-based Localization Based on Domain-invariant Feature Learning under Changing EnvironmentsabstractVisual localization is a crucial problem in mobile robotics and autonomous driving. One solution is to retrieve images with known pose from a database for the localization of query images. However, in environments with drastically varying conditions (e.g. illumination changes, seasons, occlusion, dynamic objects), retrieval-based localization is severely hampered and becomes a challenging problem. In this paper, a novel domain-invariant feature learning method (DIFL) is proposed based on ComboGAN, a multi-domain image translation network architecture. By introducing a feature consistency loss (FCL) between the encoded features of the original image and translated image in another domain, we are able to train the encoders to generate domain-invariant features in a self-supervised manner. To retrieve a target image from the database, the query image is first encoded using the encoder belonging to the query domain to obtain a domain-invariant feature vector. We then preform retrieval by selecting the database image with the most similar domain-invariant feature vector. We validate the proposed approach on the CMU-Seasons dataset, where we outperform state-of-the-art learning-based descriptors in retrieval-based localization for high and medium precision scenarios. Hanjiang Hu, Hesheng Wang 0001, Zhe Liu 0022, Chenguang Yang 0004, Weidong Chen 0001, Le Xie 0002 |
IROS | 3 |
| 2019 | SeqLPD: Sequence Matching Enhanced Loop-Closure Detection Based on Large-Scale Point Cloud Description for Self-Driving VehiclesabstractPlace recognition and loop-closure detection are main challenges in the localization, mapping and navigation tasks of self-driving vehicles. In this paper, we solve the loop-closure detection problem by incorporating the deep-learning based point cloud description method and the coarse-to-fine sequence matching strategy. More specifically, we propose a deep neural network to extract a global descriptor from the original large-scale 3D point cloud, then based on which, a typical place analysis approach is presented to investigate the feature space distribution of the global descriptors and select several super keyframes. Finally, a coarse-to-fine strategy, which includes a super keyframe based coarse matching stage and a local sequence matching stage, is presented to ensure the loop-closure detection accuracy and real-time performance simultaneously. Thanks to the sequence matching operation, the proposed approach obtains an improvement against the existing deep-learning based methods. Experiment results on a self-driving vehicle validate the effectiveness of the proposed loop-closure detection algorithm. Zhe Liu 0022, Chuanzhe Suo, Shunbo Zhou, Fan Xu 0004, Huanshu Wei, Wen Chen 0021, Hesheng Wang 0001, Xinwu Liang, Yun-Hui Liu 0001 |
IROS | 1 |
| 2019 | Modelling and Dynamic Tracking Control of Industrial Vehicles with Tractor-trailer StructureabstractExisting works on control of tractor-trailers systems only consider the kinematics model without taking dynamics into account. Also, most of them treat the issue as a pure control theory problem whose solutions are difficult to implement. This paper presents a trajectory tracking control approach for a full-scale industrial tractor-trailers vehicle composed of a carlike tractor and arbitrary number of passive full trailers. To deal with dynamic effects of trailing units, a force sensor is innovatively installed at the connection between the tractor and the first trailer to measure the forces acting on the tractor. The tractor's dynamic model that explicitly accounts for the measured forces is derived. A tracking controller that compensates the pulling/pushing forces in real time and simultaneously drives the system onto desired trajectories is proposed. The propulsion map between throttle opening and the propulsion force is proposed to be modeled with a fifth-order polynomial. The parameters are estimated by fitting experimental data, in order to provide accurate driving force. Stability of the control algorithm is rigorously proved by Lyapunov methods. Experiments of full-size vehicles are conducted to validate the performance of the control approach. Zhe Liu 0022, Shunbo Zhou, Wen Chen 0021, Chuanzhe Suo, Yun-Hui Liu 0001 |
IROS | 2 |
| 2019 | DSESP: Dual sparsity estimation subspace pursuit for the compressive sensing based close-loop ecg monitoring structure
Wenbin Yu 0001, Cailian Chen, Zhe Liu 0022, Bo Yang 0006, Xin-Ping Guan |
Peer-to-Peer Netw. Appl. | 3 |
| 2018 | A Failure-Tolerant Approach to Synchronous Formation Control of Mobile Robots Under Communication DelaysabstractRobot malfunction is inevitable in practical applications of the robot formation control due to uncontrolled crashing, system malfunction or communication loss. In this paper, we study the synchronous formation control problem in the presence of robot malfunctions. Our main idea is to improve the network connectivity and motion synchronism of the robot formation through a series of topology switchings and robot replacements. Firstly, the synchronous formation control method is introduced which enables the robots to tracking their desired trajectories while keeping predefined formation shapes. Secondly, a recursive switched topology control strategy is proposed to restore the formation shape as well as to improve the network connectivity and motion synchronism in the presence of robot malfunctions. Thirdly, the convergence analysis of the proposed control system is presented and a sufficient condition is obtained under an average dwell time scheme. What's more, the proposed approach is fully distributed and the communication delays between neighboring robots also have been taken into consideration. Simulation results demonstrate the effectiveness of the proposed approach. Zhe Liu 0022, Hesheng Wang 0001, Yun-Hui Liu 0001, Weidong Chen 0001 |
ICRA | 1 |
| 2018 | Stabilize an Unsupervised Feature Learning for LiDAR-based Place RecognitionabstractPlace recognition is one of the major challenges for the LiDAR-based effective localization and mapping task. Traditional methods are usually relying on geometry matching to achieve place recognition, where a global geometry map need to be restored. In this paper, we accomplish the place recognition task based on an end-to-end feature learning framework with the LiDAR inputs. This method consists of two core modules, a dynamic octree mapping module that generates local 2D maps with the consideration of the robot's motion; and an unsupervised place feature learning module which is an improved adversarial feature learning network with additional assistance for the long-term place recognition requirement. More specially, in place feature learning, we present an additional Generative Adversarial Network with a designed Conditional Entropy Reduction module to stabilize the feature learning process in an unsupervised manner. We evaluate the proposed method on the Kitti dataset and North Campus Long-Term LiDAR dataset. Experimental results show that the proposed method outperforms state-of-the-art in place recognition tasks under long-term applications. What's more, the feature size and inference efficiency in the proposed method are applicable in real-time performance on practical robotic platforms. Peng Yin 0001, Zhe Liu 0022, Lu Li 0018, Hadi Salman, Weiliang Xu 0001, Hesheng Wang 0001, Howie Choset |
IROS | 3 |
| 2018 | Vision-Based State Estimation and Trajectory Tracking Control of Car-Like Mobile Robots with Wheel Skidding and SlippingabstractMost existing trajectory tracking controllers are based on non-skidding and non-slipping assumptions, also assume that full states are accessible, which is unrealistic for real-world applications due to tire-road interaction. This paper presents a novel vision-based approach to achieve high performance tracking control of a Car-Like Mobile Robot (CLMR) with wheel skidding and slippage. A visual estimation algorithm is proposed to provide reliable position, velocity, skidding and slipping information to close the control loop. The stability of the proposed system can be guaranteed by Lyapunov method since the position tracking error and the estimation error converge to zero simultaneously. Simulation is made to validate the effectiveness of the developed controller in the presence of skidding and slipping with online visual estimator. Shunbo Zhou, Zhiqiang Miao, Zhe Liu 0022, Hesheng Wang 0001, Haoyao Chen, Yun-Hui Liu 0001 |
IROS | 3 |
| 2017 | Privacy-preserving design for emergency response scheduling system in medical social networks
Wenbin Yu 0001, Zhe Liu 0022, Cailian Chen, Bo Yang 0006, Xin-Ping Guan |
Peer-to-Peer Netw. Appl. | 2 |
| 2016 | An Incidental Delivery Based Method for Resolving Multirobot Pairwised Transportation ProblemsabstractThis paper presents a multirobot pairwised transportation (MRPWT) approach for factory automated material and product deliveries. We consider MRPWT from the viewpoint of robotics and incorporate practical factory application constraints in the transportation method design. The proposed MRPWT approach is a two-level hybrid planning method, consisting of an incidental delivery based single robot level planner and a simulated annealing based robot group level planner. Each robot resolves its individual transportation plan incidentally to reduce the transportation cost, whereas the group level planner utilizes predefined random actions to search the task assignment solution space and then incorporates the simulated annealing algorithm to resolve the MRPWT problem as a combinatorial optimization problem. By implementing a distributed auction mechanism, the proposed MRPWT approach can be further extended to resolve the online task allocation or reallocation problem in dynamic environments. Experiments performed on a group of mobile robots successfully demonstrate the effectiveness and the practical applicability of the proposed MRPWT approach for factory automated material and product deliveries. Zhe Liu 0022, Hesheng Wang 0001, Weidong Chen 0001, Junzhi Yu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | A gradient-based self-healing algorithm for mobile robot formationabstractIn this paper, we investigate the self-healing problem of mobile robot formation after some robots have been damaged, and present a gradient-based algorithm which enables mobile robots to restore the topology of the formation through local interactions among neighboring robots. Firstly, in order to optimize the repair path in a distributed manner, a gradient generation and diffusion mechanism is proposed to generate a specific gradient distribution in the formation. Then, utilizing several predefined path selection rules, a path selection algorithm is presented to guarantee the optimality of the selected repair path. Furthermore, several optimization indices are presented to quantitatively characterize the performance of self-healing algorithms. Finally, the effectiveness of the proposed algorithm is validated by numerical simulations and the simulation results show that the proposed algorithm can restore the topology of the formation with the fewer repair robots and lower energy consumptions. Zhe Liu 0022, Jianjun Ju, Weidong Chen 0001, Xiangyu Fu, Hesheng Wang 0001 |
IROS | 1 |