VLDB 2026 Research / reviewers in the wild / expert
Haoang Li
dblp:199/8347
· DBLP profile ↗
41ranked-venue papers
14as first author
25since 2021 · last 2026
0000-0002-1576-9408ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 13 first-author · 20 since 2021Systems, architecture and hardware · 18 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot PerceiverabstractRecent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispersed. To guide the visual attention grounding on the correct target, we propose ReconVLA, a reconstructive VLA model with an implicit grounding paradigm. Conditioned on the model's visual outputs, a diffusion transformer aims to reconstruct the gaze region of the image, which corresponds to the target manipulated objects. This process prompts the VLA model to learn fine-grained representations and accurately allocate visual attention, thus effectively leveraging task-specific visual information and conducting precise manipulation. Moreover, we curate a large-scale pretraining dataset comprising over 100k trajectories and 2 million data samples from open-source robotic datasets, further boosting the model’s generalization in visual reconstruction. Extensive experiments in simulation and the real world demonstrate the superiority of our implicit grounding method, showcasing its capabilities of precise manipulation and generalization. Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Haodong Yan, Haoang Li |
AAAI | 10 |
| 2026 | Spatial-Aware and Viewpoint-Robust Vision-Language NavigationabstractVision-language navigation (VLN) requires an agent to follow visual observations and language instructions to navigate within a 3D environment. Most prior VLN approaches utilize recurrent units, topological graphs, or grid maps to represent the agent’s previously explored environment. However, these methods face two key limitations: (1) difficulty in accurately and adequately representing multi-level spatial information, and (2) a lack of robustness in semantic feature extraction against viewpoint variations. To overcome these challenges, we design a novel spatial-aware scene representation (SSR) as well as a viewpoint-robust feature extraction method. Specifically, for SSR, we first construct a multi-layer map that models the scene at different granularities, incorporating waypoints, objects, rooms, and floors. We then propose to extract spatial features from this multi-layer map based on a heterogeneous graph transformer and align them with the instructions, effectively guiding the agent’s planning process. As to viewpoint-robust feature extraction, we integrate 3D Gaussian splatting to capture the 3D geometry and visual texture of the environment, enabling novel view object observation and feature extraction. This improves observation coverage and enhances the viewpoint robustness of feature extraction. Extensive experiments demonstrated that our approach achieves state-of-the-art performance on two VLN benchmarks and shows strong performance in real-world scenarios. Source codes will be publicly available upon paper acceptance. Zhide Zhong, Xiangchen Liu, Xinhu Zheng, Zhe Liu 0022, Hesheng Wang 0001, Haoang Li |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | DynSUP: Dynamic Gaussian Splatting From an Unposed Image PairabstractRecent advances in 3D Gaussian Splatting have shown promising results. Existing methods typically assume static scenes and/or multiple images with prior poses. Dynamics, sparse views, and unknown poses significantly increase the problem complexity due to insufficient geometric constraints. To overcome this challenge, we propose a method that can use only two images without prior poses to fit Gaussians in dynamic environments. To achieve this, we introduce two technical contributions. First, we propose an object-level two-view bundle adjustment. This strategy decomposes dynamic scenes into piece-wise rigid components, and jointly estimates the relative camera motion and dynamic object motions for dynamic Gaussian initialization. Second, we design an SE(3) field-driven Gaussian training method. It enables fine-grained motion modeling through learnable per-Gaussian transformations. Our method leads to high-fidelity novel view synthesis of dynamic scenes while accurately preserving temporal consistency and object motion. Experiments on both synthetic and real-world datasets demonstrate that our method significantly outperforms state-of-the-art approaches designed for the cases of static environments, multiple images, and/or known poses. Our project page is available at https://colin-de.github.io/DynSUP/. Weihang Li, Shenhan Qian, Benjamin Busam, Daniel Cremers, Haoang Li |
IEEE Trans. Image Process. | 6 |
| 2026 | SCSV: Spatial-Temporal Consistent Dynamic 3D Scene Generation From Sparse ViewsabstractGenerating dynamic scenes from images has gained increasing attention. Existing methods have two major limitations: 1) they can hardly handle sparse images which exhibit limited geometry constraints and insufficient motion; 2) they struggle to maintain spatial-temporal consistency when rendering multi-view videos. To address these limitations, we propose SCSV, a spatial-temporal consistent dynamic scene generation method from sparse views. Our method consists of two stages: scene reconstruction and scene expansion, both of which decouple background and foreground. In the scene reconstruction stage, we first interpolate a set of images between the input images based on a video generation model, followed by the optimization of the scene Gaussian from the interpolated and input images. To improve the spatial-temporal consistency of the reconstructed scene, we propose an uncertainty-aware Gaussian training approach, which introduces adaptive weights of images and pixels. In the scene expansion stage, for background, we render novel views and refine them with a geometry-aware diffusion process. These refined images are then used to incrementally add the Gaussians. As to foreground, we generate human motion according to previous motion, enabling temporal coherent generation of motion. To further enhance the physical plausibility, we integrate the expanded foreground into the background using a gravity-aware alignment. Experiments on NeuMan, Bonn, and EMDB datasets demonstrate that our SCSV achieves superior performance compared to state-of-the-art methods. The code will be released upon acceptance. Shunbo Zhou, Jun Ma 0008, Hesheng Wang 0001, Haoang Li |
IEEE Trans. Image Process. | 8 |
| 2025 | Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldabstractDetermining the vanishing points (VPs) in a Manhattan world, as a fundamental task in many 3D vision applications, consists of jointly inferring the line-VP association and locating each VP. Existing methods are, however, either sub-optimal solvers or pursuing global optimality at a significant cost of computing time. In contrast to prior works, we introduce convex relaxation techniques to solve this task for the first time. Specifically, we employ a "soft" association scheme, realized via a truncated multi-selection error, that allows for joint estimation of VPs’ locations and line-VP associations. This approach leads to a primal problem that can be reformulated into a quadratically constrained quadratic programming (QCQP) problem, which is then relaxed into a convex semidefinite programming (SDP) problem. To solve this SDP problem efficiently, we present a globally optimal outlier-robust iterative solver (called GlobustVP), which independently searches for one VP and its associated lines in each iteration, treating other lines as outliers. After each independent update of all VPs, the mutual orthogonality between the three VPs in a Manhattan world is reinforced via local refinement. Extensive experiments on both synthetic and real-world data demonstrate that GlobustVP achieves a favorable balance between efficiency, robustness, and global optimality compared to previous works. The code is publicly available at github.com/wu-cvgl/GlobustVP. Bangyan Liao, Zhenjun Zhao, Haoang Li, Yi Zhou 0010, Yingping Zeng, Peidong Liu 0001 |
CVPR | 3 |
| 2025 | Interactive Navigation for Legged Manipulators with Learned Arm-Pushing ControllerabstractInteractive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in narrow or constrained spaces where the robot’s dimensions restrict its manipulation capabilities. This paper introduces a novel interactive navigation framework for legged manipulators, featuring an active arm-pushing mechanism that enables the robot to reposition movable obstacles in space-constrained environments. To this end, we develop a reinforcement learning-based arm-pushing controller with a two-stage reward strategy for object manipulation. Specifically, this strategy first directs the manipulator to a designated pushing zone to achieve a kinematically feasible contact configuration. Then, the end effector is guided to maintain its position at appropriate contact points for stable object displacement while preventing toppling. The simulations validate the robustness of the arm-pushing controller, showing that the two-stage reward strategy improves policy convergence and long-term performance. Real-World experiments further demonstrate the effectiveness of the proposed navigation framework, which achieves shorter paths and reduced traversal time. The open-source project can be found at https://zhihaibi.github.io/interactive-push.github.io/. Zhihai Bi, Kai Chen 0006, Chunxin Zheng, Yulin Li 0001, Haoang Li, Jun Ma 0008 |
IROS | 5 |
| 2025 | STG-Avatar: Animatable Human Avatars via Spacetime GaussianabstractRealistic animatable human avatars from monocular videos are crucial for advancing human-robot interaction and enhancing immersive virtual experiences. While recent research on 3DGS-based human avatars has made progress, it still struggles with accurately representing detailed features of non-rigid objects (e.g., clothing deformations) and dynamic regions (e.g., rapidly moving limbs). To address these challenges, we present STG-Avatar, a 3DGS-based framework for high-fidelity animatable human avatar reconstruction. Specifically, our framework introduces a rigid-nonrigid coupled deformation framework that synergistically integrates Spacetime Gaussians (STG) with linear blend skinning (LBS). In this hybrid design, LBS enables real-time skeletal control by driving global pose transformations, while STG complements it through spacetime-adaptive optimization of 3D Gaussians. Furthermore, we employ optical flow to identify high-dynamic regions and guide the adaptive densification of 3D Gaussians in these regions. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines in both reconstruction quality and operational efficiency, achieving superior quantitative metrics while retaining real-time rendering capabilities. Our code is available at https://github.com/jiangguangan/STG-Avatar Guangan Jiang, Tianzi Zhang, Zhenjun Zhao, Haoang Li, Hongyu Wang 0001 |
IROS | 5 |
| 2025 | SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban EnvironmentsabstractUnmanned Aerial Vehicles (UAVs) have emerged as versatile tools across various sectors, driven by their mobility and adaptability. This paper introduces SkyVLN, a novel framework integrating vision-and-language navigation (VLN) with Nonlinear Model Predictive Control (NMPC) to enhance UAV autonomy in complex urban environments. Unlike traditional navigation methods, SkyVLN leverages Large Language Models (LLMs) to interpret natural language instructions and visual observations, enabling UAVs to navigate through dynamic 3D spaces with improved accuracy and robustness. We present a multimodal navigation agent equipped with a fine-grained spatial verbalizer and a history path memory mechanism. These components allow the UAV to disambiguate spatial contexts, handle ambiguous instructions, and backtrack when necessary. The framework also incorporates an NMPC module for dynamic obstacle avoidance, ensuring precise trajectory tracking and collision prevention. To validate our approach, we developed a high-fidelity 3D urban simulation environment using AirSim, featuring realistic imagery and dynamic urban elements. Extensive experiments demonstrate that SkyVLN significantly improves navigation success rates and efficiency, particularly in new and unseen environments. Tianshun Li, Tianyi Huai, Yichun Gao, Haoang Li, Xinhu Zheng |
IROS | 5 |
| 2025 | RMG: Real-Time Expressive Motion Generation with Self-collision Avoidance for 6-DOF Companion Robotic ArmsabstractThe six-degree-of-freedom (6-DOF) robotic arm has gained widespread application in human-coexisting environments. While previous research has predominantly focused on functional motion generation, the critical aspect of expressive motion in human-robot interaction remains largely unexplored. This paper presents a novel real-time motion generation planner that enhances interactivity by creating expressive robotic motions between arbitrary start and end states within predefined time constraints. Our approach involves three key contributions: first, we develop a mapping algorithm to construct an expressive motion dataset derived from human dance movements; second, we train motion generation models in both Cartesian and joint spaces using this dataset; third, we introduce an optimization algorithm that guarantees smooth, collision-free motion while maintaining the intended expressive style. Experimental results demonstrate the effectiveness of our method, which can generate expressive and generalized motions in under 0.5 seconds while satisfying all specified constraints. Jiansheng Li, Haotian Song, Haoang Li, Jinni Zhou, Qiang Nie |
IROS | 3 |
| 2025 | RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot ManipulationabstractThis paper introduces RoboDexVLM, an innovative framework for robot task planning and grasp detection tailored for a collaborative manipulator equipped with a dexterous hand. Previous methods focus on simplified and limited manipulation tasks, which often neglect the complexities associated with grasping a diverse array of objects in a long-horizon manner. In contrast, our proposed framework utilizes a dexterous hand capable of grasping objects of varying shapes and sizes while executing tasks based on natural language commands. The proposed approach has the following core components: First, a robust task planner with a task-level recovery mechanism that leverages vision-language models (VLMs) is designed, which enables the system to interpret and execute open-vocabulary commands for long sequence tasks. Second, a language-guided dexterous grasp perception algorithm is presented based on robot kinematics and formal methods, tailored for zero-shot dexterous manipulation with diverse objects and commands. Comprehensive experimental results validate the effectiveness, adaptability, and robustness of RoboDexVLM in handling long-horizon scenarios and performing dexterous grasping. These results highlight the framework’s ability to operate in complex environments, showcasing its potential for open-vocabulary dexterous manipulation. Our open-source project page can be found at https://henryhcliu.github.io/robodexvlm. Haichao Liu 0003, Sikai Guo, Pengfei Mai, Jiahang Cao, Haoang Li, Jun Ma 0008 |
IROS | 5 |
| 2025 | PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel DecodingabstractVision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. Therefore, accelerating VLA integrated with action chunking is an urgent need. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52× execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks. Wenxuan Song, Pengxiang Ding, Han Zhao 0008, Zhide Zhong, ZongYuan Ge, Jun Ma 0008, Haoang Li |
IROS | 12 |
| 2025 | L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR ModelabstractSemantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operations, imposing significant computational requirements on the platform during training and testing. This paper proposes L2COcc, a lightweight camera-centric SSC framework that also accommodates LiDAR inputs. With our proposed efficient voxel transformer (EVT) and cross-modal knowledge modules, including feature similarity distillation (FSD), TPV distillation (TPVD) and prediction alignment distillation (PAD), our method substantially reduce computational burden while maintaining high accuracy. The experimental evaluations demonstrate that our proposed method surpasses the current state-of-the-art vision-based SSC methods regarding accuracy on both the SemanticKITTI and SSCBench-KITTI-360 benchmarks, respectively. Additionally, our method is more lightweight, exhibiting a reduction in both memory consumption and inference time by over 23% compared to the current state-of-the-arts method. Code is available at our project page: https://studyingfufu.github.io/L2COcc/. Yukai Ma, Sheng Tao, Haoang Li, Zongzhi Zhu, Yong Liu 0007, Xingxing Zuo 0001 |
IROS | 5 |
| 2025 | EndoFlow-SLAM: Real-Time Endoscopic SLAM with Flow-Constrained Gaussian Splatting
Taoyu Wu, Yiyi Miao, Zhuoxiao Li, Haocheng Zhao, Kang Dang, Jionglong Su, Limin Yu, Haoang Li |
MICCAI (9) | 8 |
| 2024 | GlobalPointer: Large-Scale Plane Adjustment with Bi-Convex Relaxation
Bangyan Liao, Zhenjun Zhao, Haoang Li, Daniel Cremers, Peidong Liu 0001 |
ECCV (59) | 4 |
| 2024 | Structured-NeRF: Hierarchical Scene Graph with Neural Representation
Zhide Zhong, Jiakai Cao, Songen Gu, Sirui Xie, Liyi Luo, Hao Zhao 0002, Guyue Zhou, Haoang Li, Zike Yan |
ECCV (35) | 8 |
| 2024 | Physically-Based Photometric Bundle Adjustment in Non-Lambertian EnvironmentsabstractPhotometric bundle adjustment (PBA) is widely used in estimating the camera pose and 3D geometry by assuming a Lambertian world. However, the assumption of photometric consistency is often violated since the non-diffuse reflection is common in real-world environments. The photometric inconsistency significantly affects the reliability of existing PBA methods. To solve this problem, we propose a novel physically-based PBA method. Specifically, we introduce the physically-based weights regarding material, illumination, and light path. These weights distinguish the pixel pairs with different levels of photometric inconsistency. We also design corresponding models for material estimation based on sequential images and illumination estimation based on point clouds. In addition, we establish the first SLAM-related dataset of non-Lambertian scenes with complete ground truth of illumination and material. Extensive experiments demonstrated that our PBA method outperforms existing approaches in accuracy. Junpeng Hu, Haodong Yan, Mariia Gladkova, Yun-Hui Liu 0001, Daniel Cremers, Haoang Li |
IROS | 8 |
| 2024 | Efficient and Robust Point Cloud Registration via Heuristics-Guided Parameter SearchabstractEstimating the rigid transformation with 6 degrees of freedom based on a putative 3D correspondence set is a crucial procedure in point cloud registration. Existing correspondence identification methods usually lead to large outlier ratios (>95% is common), underscoring the significance of robust registration methods. Many researchers turn to parameter search-based strategies (e.g., Branch-and-Bround) for robust registration. Although related methods show high robustness, their efficiency is limited to the high-dimensional search space. This paper proposes a heuristics-guided parameter search strategy to accelerate the search while maintaining high robustness. We first sample some correspondences (i.e., heuristics) and then just need to sequentially search the feasible regions that make each sample an inlier. Our strategy largely reduces the search space and can guarantee accuracy with only a few inlier samples, therefore enjoying an excellent trade-off between efficiency and robustness. Since directly parameterizing the 6-dimensional nonlinear feasible region for efficient search is intractable, we construct a three-stage decomposition pipeline to reparameterize the feasible region, resulting in three lower-dimensional sub-problems that are easily solvable via our strategy. Besides reducing the searching dimension, our decomposition enables the leverage of 1-dimensional interval stabbing at all three stages for searching acceleration. Moreover, we propose a valid sampling strategy to guarantee our sampling effectiveness, and a compatibility verification setup to further accelerate our search. Extensive experiments on both simulated and real-world datasets demonstrate that our approach exhibits comparable robustness with state-of-the-art methods while achieving a significant efficiency boost. Haoang Li, Liangzu Peng, Yinlong Liu, Yun-Hui Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning Accurate 3D Shape Based on Stereo Polarimetric ImagingabstractShape from Polarization (SfP) aims to recover surface normal using the polarization cues of light. The accuracy of existing SfP methods is affected by two main problems. First, the ambiguity of polarization cues partially results in false normal estimation. Second, the widely-used assumption about orthographic projection is too ideal. To solve these problems, we propose the first approach that com-bines deep learning and stereo polarization information to recover not only normal but also disparity. Specifically, for the ambiguity problem, we design a Shape Consistency-based Mask Prediction (SCMP) module. It exploits the inherent consistency between normal and disparity to identify the areas with false normal estimation. We replace the unreliable features enclosed by these areas with new features extracted by global attention mechanism. As to the orthographic projection problem, we propose a novel Viewing Direction-aided Positional Encoding (VDPE) strategy. This strategy is based on the unique pixel-viewing direction encoding, and thus enables our neural network to handle the non-orthographic projection. In addition, we establish a real-world stereo SfP dataset that contains various object categories and illumination conditions. Experiments showed that compared with existing SfP methods, our approach is more accurate. Moreover, our approach shows higher robustness to light variation. Haoang Li, Kejing He 0002, Congying Sui, Bin Li 0082, Yun-Hui Liu 0001 |
CVPR | 2 |
| 2023 | DDIT: Semantic Scene Completion via Deformable Deep Implicit TemplatesabstractScene reconstructions are often incomplete due to occlusions and limited viewpoints. There have been efforts to use semantic information for scene completion. However, the completed shapes may be rough and imprecise since respective methods rely on 3D convolution and/or lack effective shape constraints. To overcome these limitations, we propose a semantic scene completion method based on deformable deep implicit templates (DDIT). Specifically, we complete each segmented instance in a scene by deforming a template with a latent code. Such a template is expressed by a deep implicit function in the canonical frame. It abstracts the shape prior of a category, and thus can provide constraints on the overall shape of an instance. Latent code controls the deformation of template to guarantee fine details of an instance. For code prediction, we design a neural network that leverages both intra-and inter-instance information. We also introduce an algorithm to transform instances between the world and canonical frames based on geometric constraints and a hierarchical tree. To further improve accuracy, we jointly optimize the latent code and transformation by enforcing the zero-valued isosurface constraint. In addition, we establish a new dataset to solve different problems of existing datasets. Experiments showed that our DDIT outperforms state-of-the-art approaches. Haoang Li, Jinhu Dong, Binghui Wen, Yun-Hui Liu 0001, Daniel Cremers |
ICCV | 1 |
| 2023 | Hong Kong World: Leveraging Structural Regularity for Line-Based SLAMabstractManhattan and Atlanta worlds hold for the structured scenes with only vertical and horizontal dominant directions (DDs). To describe the scenes with additional sloping DDs, a mixture of independent Manhattan worlds seems plausible, but may lead to unaligned and unrelated DDs. By contrast, we propose a novel structural model called Hong Kong world. It is more general than Manhattan and Atlanta worlds since it can represent the environments with slopes, e.g., a city with hilly terrain, a house with sloping roof, and a loft apartment with staircase. Moreover, it is more compact and accurate than a mixture of independent Manhattan worlds by enforcing the orthogonality constraints between not only vertical and horizontal DDs, but also horizontal and sloping DDs. We further leverage the structural regularity of Hong Kong world for the line-based SLAM. Our SLAM method is reliable thanks to three technical novelties. First, we estimate DDs/vanishing points in Hong Kong world in a semi-searching way. We use a new consensus voting strategy for search, instead of traditional branch and bound. This method is the first one that can simultaneously determine the number of DDs, and achieve quasi-global optimality in terms of the number of inliers. Second, we compute the camera pose by exploiting the spatial relations between DDs in Hong Kong world. This method generates concise polynomials, and thus is more accurate and efficient than existing approaches designed for unstructured scenes. Third, we refine the estimated DDs in Hong Kong world by a novel filter-based method. Then we use these refined DDs to optimize the camera poses and 3D lines, leading to higher accuracy and robustness than existing optimization algorithms. In addition, we establish the first dataset of sequential images in Hong Kong world. Experiments showed that our approach outperforms state-of-the-art methods in terms of accuracy and/or efficiency. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Pyojin Kim, Kyungdon Joo, Zhenjun Zhao, Yun-Hui Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Deterministic Point Cloud Registration via Novel Transformation DecompositionabstractGiven a set of putative 3D-3D point correspondences, we aim to remove outliers and estimate rigid transformation with 6 degrees of freedom (DOF). Simultaneously estimating these 6 DOF is time-consuming due to high-dimensional parameter space. To solve this problem, it is common to decompose 6 DOF, i.e. independently compute 3-DOF rotation and 3-DOF translation. However, high non-linearity of 3-DOF rotation still limits the algorithm efficiency, especially when the number of correspondences is large. In contrast, we propose to decompose 6 DOF into$(2+1)$and$(1+2)\ DOF$. Specifically,$(2+1)DOF$represent 2-DOF rotation axis and 1-DOF displacement along this rotation axis.$(1+2)\ DOF$indicate 1-DOF rotation angle and 2-DOF displacement orthogonal to the above rotation axis. To compute these DOF, we design a novel two-stage strategy based on inlier set maximization. By leveraging branch and bound, we first search for$(2+1)\ DOF$, and then the remaining$(1+2)\ DOF$. Thanks to the proposed transformation decomposition and two-stage search strategy, our method is deterministic and leads to low computational complexity. We extensively compare our method with state-of-the-art approaches. Our method is more accurate and robust than the approaches that provide similar efficiency to ours. Our method is more efficient than the approaches whose accuracy and robustness are comparable to ours. Wen Chen 0021, Haoang Li, Qiang Nie, Yun-Hui Liu 0001 |
CVPR | 2 |
| 2022 | Quasi-Globally Optimal and Near/True Real-Time Vanishing Point Estimation in Manhattan WorldabstractImage lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim to cluster them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and camera frame. To estimate three degrees of freedom (DOF) of this rotation, state-of-the-art methods are based on either data sampling or parameter search. However, they fail to guarantee high accuracy and efficiency simultaneously. In contrast, we propose a set of approaches that hybridize these two strategies. We first constrain two or one DOF of the rotation by two or one sampled image line. Then we search for the remaining one or two DOF based on branch and bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search achieves quasi-global optimality. Specifically, it guarantees to retrieve the maximum number of inliers on the condition that two or one DOF is constrained. Our hybridization of two-line sampling and one-DOF search can estimate VPs in real time. Our hybridization of one-line sampling and two-DOF search can estimate VPs in near real time. Experiments on both synthetic and real-world datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and/or efficiency. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Integrated Task Allocation and Path Coordination for Large-Scale Robot Networks With UncertaintiesabstractArtificial intelligence-enhanced autonomous unmanned systems, such as large-scale autonomous robot networks, are widely used in logistic and industrial applications. In this article, we address the integrated task assignment, path planning, and coordination problem applied for large-scale robot networks with the existence of uncertainties. In particular, a novel generalized conflict graph is designed which encodes the traveling time cost of the subsequent path planning result of each task-robot assignment and also includes the predicted path conflicts of each two assignments. An integrated optimization problem which aims to minimize the total traveling cost and potential path conflicts simultaneously is first formulated and then transformed into a linear programming instance to obtain the optimal solution. In particular, to satisfy the real-time requirement in large-scale systems, a greedy solution is presented which has the near-optimal performance but can decrease the computational complexity by orders of magnitude. The optimality, scalability, robustness, and efficiency of our approach are demonstrated by comprehensive comparisons with existing state-of-the-art approaches. Note to Practitioners—With the development of artificial intelligence techniques, large-scale autonomous robot networks are increasingly used in the logistic warehouses, unmanned container terminals, and intelligence transportation systems. This article considers the large-scale networks with hundreds or even thousands of unmanned robots which are implemented in lifelong transportation systems with uncertainties existed in practical execution process. Our main concept is to simultaneously minimize the total time cost of all the tasks and the potential motion conflicts among all the robots in the subsequent execution stage, thus alleviating robot congestions, balancing traffic distributions, increasing system efficiency, and improving the robustness and scalability. Lifelong simulations with thousand robots illustrate that our approach can reduce more than 30% of the time steps consumed for coordinating robot motion conflicts, and in the meantime, the throughput, and overall system efficiency are improved. However, simulation results show that the system improvement decreases in the presence of extreme high uncertainties (such as temporary motion and communication failures of the robot), due to the inaccurate conflict prediction in the integrated optimization stage. Our future work includes the deep-learning-based traffic evolution prediction and the online reallocation and planning in highly dynamic scenarios. Zhe Liu 0022, Huanshu Wei, Hongye Wang, Haoang Li, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Learning To Identify Correct 2D-2D Line Correspondences on SphereabstractGiven a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both structured and unstructured scenes. Instead of geometric constraint, we leverage the spatial regularity on sphere. Specifically, we propose to map line correspondences into vectors tangent to sphere. We use these vectors to encode both angular and positional variations of image lines, which is more reliable and concise than directly using inclinations, midpoints or endpoints of image lines. Neighboring vectors mapped from correct matches exhibit a spatial regularity called local trend consistency, regardless of the type of scenes. To encode this regularity, we design a neural network and also propose a novel loss function that enforces the smoothness constraint of vector field. In addition, we establish a large real-world dataset for image line matching. Experiments showed that our approach outperforms state-of-the-art ones in terms of accuracy, efficiency and robustness, and also leads to high generalization. Haoang Li, Kai Chen 0028, Ji Zhao 0001, Jiangliu Wang, Pyojin Kim, Zhe Liu 0022, Yun-Hui Liu 0001 |
CVPR | 1 |
| 2021 | Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point EstimationabstractExisting vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts a spherical probability map of VP. Based on this map, we can detect all the VPs. Our method is reliable thanks to four technical novelties. First, we leverage the icosahedral spherical representation to express our probability map. This representation provides uniform pixel distribution, and thus facilitates estimating arbitrary positions of VPs. Second, we design a loss function that enforces the antipodal symmetry and sparsity of our spherical probability map to prevent over-fitting. Third, we generate the ground truth probability map that reasonably expresses the locations and uncertainties of VPs. This map unnecessarily peaks at noisy annotated VPs, and also exhibits various anisotropic dispersions. Fourth, given a predicted probability map, we detect VPs by fitting a Bingham mixture model. This strategy can robustly handle close VPs and provide the confidence level of VP useful for practical applications. Experiments showed that our method achieves the best compromise between generality, accuracy, and efficiency, compared with state-of-the-art approaches. Haoang Li, Kai Chen 0028, Pyojin Kim, Kuk-Jin Yoon, Zhe Liu 0022, Kyungdon Joo, Yun-Hui Liu 0001 |
ICCV | 1 |
| 2020 | Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World
Haoang Li, Pyojin Kim, Ji Zhao 0001, Kyungdon Joo, Zhe Liu 0022, Yun-Hui Liu 0001 |
ECCV (22) | 1 |
| 2020 | Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual OdometryabstractGiven a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In contrast, we propose a novel approach based on the robust "L2-minimizing estimate" (L2E) loss. We first define a novel cost function by integrating the projection constraint into the L2E loss. Then to efficiently obtain the global minimum of this function, we propose a hybrid strategy of a local optimizer and branch-and-bound. For branch-and-bound, we derive effective function bounds. Our approach can handle high outlier ratios, leading to high robustness. It can run reliably regardless of whether the initial pose is appropriate, providing high generality. Moreover, given a decent initial pose, it is suitable for real-time applications. Experiments on synthetic and real-world datasets showed that our approach outperforms state-of-the-art methods in terms of robustness and/or efficiency. Haoang Li, Wen Chen 0021, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 1 |
| 2020 | A Synchronization Approach for Achieving Cooperative Adaptive Cruise Control Based Non-Stop Intersection PassingabstractCooperative adaptive cruise control (CACC) of intelligent vehicles contributes to improving cruise control performance, reducing traffic congestion, saving energy and increasing traffic flow capacity. In this paper, we resolve the CACC problem from the viewpoint of synchronization control, our main idea is to introduce the spatial-temporal synchronization mechanism into vehicle platoon control to achieve the robust CACC and to further realize the non-stop intersection control. Firstly, by introducing the cross-coupling based space synchronization mechanism, a distributed control algorithm is presented to achieve the single-lane CACC in the presence of vehicle-to-vehicle (V2V) communications, which enables autonomous vehicles to track the desired platoon trajectory while synchronizing their longitudinal velocities to keeping the expected inter-vehicle distance. Secondly, by designing the enter-time scheduling mechanism (temporal synchronization), a high-level intersection control strategy is proposed to command vehicles to form a virtual platoon to pass through the intersection without stopping. Thirdly, a Lyapunov-based time-domain stability analysis approach is presented. Compared with the traditional string stability based approach, the proposed approach guarantees the global asymptotical convergence of the proposed CACC system. Experiments in the small-scale simulated system demonstrate the effectiveness of the proposed approach. Zhe Liu 0022, Huanshu Wei, Hanjiang Hu, Chuanzhe Suo, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICRA | 6 |
| 2020 | CUHK-AHU Dataset: Promoting Practical Self-Driving Applications in the Complex Airport Logistics, Hill and Urban EnvironmentsabstractThis paper presents a novel dataset targeting three types of challenging environments for autonomous driving, i.e., the industrial logistics environment, the undulating hill environment and the mixed complex urban environment. To the best of the author's knowledge, similar dataset has not been published in the existing public datasets, especially for the logistics environment collected in the functioning Hong Kong Air Cargo Terminal (HACT). Structural changes always suddenly appeared in the airport logistics environment due to the frequent movement of goods in and out. In the structureless and noisy hill environment, the non-flat plane movement is usual. In the mixed complex urban environment, the highly dynamic residence blocks, sloped roads and highways are included in a single collection. The presented dataset includes LiDAR, image, IMU and GPS data by repeatedly driving along several paths to capture the structural changes, the illumination changes and the different degrees of undulation of the roads. The baseline trajectories are provided which are estimated by Simultaneous Localization and Mapping (SLAM). Wen Chen 0021, Zhe Liu 0022, Shunbo Zhou, Haoang Li, Yun-Hui Liu 0001 |
IROS | 5 |
| 2020 | End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondencesabstract3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learning based approach to resolve the point cloud registration problem. Firstly, the revised LPD-Net is introduced to extract features and aggregate them with the graph network. Secondly, the self-attention mechanism is utilized to enhance the structure information in the point cloud and the cross-attention mechanism is designed to enhance the corresponding information between the two input point clouds. Based on which, the virtual corresponding points can be generated by a soft pointer based method, and finally, the point cloud registration problem can be solved by implementing the SVD method. Comparison results in ModelNet40 dataset validate that the proposed approach reaches the state-of-the-art in point cloud registration tasks and experiment resutls in KITTI dataset validate the effectiveness of the proposed approach in real applications. Huanshu Wei, Zhijian Qiao, Zhe Liu 0022, Chuanzhe Suo, Peng Yin 0001, Yueling Shen, Haoang Li, Hesheng Wang 0001 |
IROS | 7 |
| 2020 | Robust Estimation of Absolute Camera Pose via Intersection Constraint and Flow ConsensusabstractEstimating the absolute camera pose requires 3D-to-2D correspondences of points and/or lines. However, in practice, these correspondences are inevitably corrupted by outliers, which affects the pose estimation. Existing outlier removal strategies for robust pose estimation have some limitations. They are only applicable to points, rely on prior pose information, or fail to handle high outlier ratios. By contrast, we propose a general and accurate outlier removal strategy. It can be integrated with various existing pose estimation methods originally vulnerable to outliers, and is applicable to points, lines, and the combination of both. Moreover, it does not rely on any prior pose information. Our strategy has a nested structure composed of the outer and inner modules. First, our outer module leverages our intersection constraint, i.e., the projection rays or planes defined by inliers intersect at the camera center. Our outer module alternately computes the inlier probabilities of correspondences and estimates the camera pose. It can run reliably and efficiently under high outlier ratios. Second, our inner module exploits our flow consensus. The 2D displacement vectors or 3D directed arcs generated by inliers exhibit a common directional regularity, i.e., follow a dominant trend of flow. Our inner module refines the inlier probabilities obtained at each iteration of our outer module. This refinement improves the accuracy and facilitates the convergence of our outer module. Experiments on both synthetic data and real-world images have shown that our method outperforms state-of-the-art approaches in terms of accuracy and robustness. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan WorldabstractThe image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and the camera frame. To compute this rotation, state-of-the-art methods are based on either data sampling or parameter search, and they fail to guarantee the accuracy and efficiency simultaneously. In contrast, we propose to hybridize these two strategies. We first compute two degrees of freedom (DOF) of the above rotation by two sampled image lines, and then search for the optimal third DOF based on the branch-and-bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search is not sensitive to noise and achieves quasi-global optimality in terms of maximizing the number of inliers. Experiments on synthetic and real-world images showed that our method outperforms state-of-the-art approaches in terms of accuracy and/or efficiency. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICCV | 1 |
| 2019 | LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment AnalysisabstractPoint cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named LPD-Net (Large-scale Place Description Network), which can extract discriminative and generalizable global descriptors from the raw 3D point cloud. Two modules, the adaptive local feature extraction module and the graph-based neighborhood aggregation module, are proposed, which contribute to extract the local structures and reveal the spatial distribution of local features in the large-scale point cloud, with an end-to-end manner. We implement the proposed global descriptor in solving point cloud based retrieval tasks to achieve the large-scale place recognition. Comparison results show that our LPD-Net is much better than PointNetVLAD and reaches the state-of-the-art. We also compare our LPD-Net with the vision-based solutions to show the robustness of our approach to different weather and light conditions. Zhe Liu 0022, Shunbo Zhou, Chuanzhe Suo, Peng Yin 0001, Wen Chen 0021, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICCV | 7 |
| 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAMabstractA wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocular SLAM. First, we robustly cluster image lines. Based on these clusters, we compute the local Atlanta frames in the camera frame by solving polynomial equations. Our method provides the global optimum and satisfies inherent geometric constraints. Second, we define the posterior probabilities to refine the initial clusters and Atlanta frames alternately by the maximum a posteriori estimation. Third, based on multiple local Atlanta frames, we compute the global Atlanta frames in the world frame using Kalman filtering. We optimize rotations by the global alignment and then refine translations and 3D line-based map under the directional constraints. Experiments on both synthesized and real data have demonstrated that our approach outperforms state-of-the-art methods. Haoang Li, Yazhou Xing, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 1 |
| 2019 | A Hierarchical Framework for Coordinating Large-Scale Robot NetworksabstractIn this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guarantee the coordination performance in the presence of robot motion uncertainties. A hierarchical path planning and motion coordination structure is presented. The environment is divided into several sectors and a traffic heat-map is presented to describe the current sector-level traffic condition. In path planning level, the sector-level path is calculated by considering the path distance, the current traffic condition and the current robot uncertainty. In motion coordination level, local cooperative A* algorithm and conflict-based searching strategy are utilized within each sector to generate the collision-free local path of each robot in a rolling planning manner. The effectiveness and practical applicability of the proposed approach are validated by simulations with more than one thousand robots and real experiments. Zhe Liu 0022, Shunbo Zhou, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001 |
ICRA | 5 |
| 2019 | Line-based Absolute and Relative Camera Pose Estimation in Structured Environmentsabstract3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the absolute and relative rotations and reformulate the problem as computing the rotation from the Manhattan frame to the camera frame. To compute this rotation, we propose an accurate and efficient two-step method. We first estimate its two degrees of freedom (DOF) by two image lines, and then estimate its third DOF by another image line. For these lines, we assume their associated 3D lines are mutually orthogonal, or two 3D lines are parallel to each other and orthogonal to the third. Thanks to our two-step DOF estimation, our absolute and relative pose estimation methods are accurate and efficient. Moreover, our relative pose estimation method relies on weaker assumptions or less correspondences than existing approaches. We also propose a novel strategy to reject outliers and identify dominant directions of the scene. We integrate it into our pose estimation methods, and show that it is more robust than RANSAC. Experiments on synthetic and real-world datasets demonstrated that our methods outperform state-of-the-art approaches. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Kai Chen 0028, Yun-Hui Liu 0001 |
IROS | 1 |
| 2018 | A Monocular SLAM System Leveraging Structural Regularity in Manhattan WorldabstractThe structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The foremost contribution of the proposed system is a structural feature-based optimization module which contains three novel optimization strategies. First, a rotation optimization strategy using the parallelism and orthogonality of 3D lines is presented. We propose a global binding method to compute an accurate estimation of the absolute rotation of the camera. Then we propose an approach for calculating the relative rotation to further refine the absolute rotation. Second, a translation optimization strategy leveraging coplanarity is proposed. Coplanar features are effectively identified, and we leverage them by a unified model handling both points and lines to calculate the relative translation, and then the optimal absolute translation. Third, a 3D line optimization strategy utilizing parallelism, orthogonality and coplanarity simultaneously is proposed to obtain an accurate 3D map consisting of structural line segments with low computational complexity. Experiments in man-made environments have demonstrated that the proposed system outperforms existing state-of-the-art monocular SLAM systems in terms of accuracy and robustness. Haoang Li, Jian Yao 0002, Jean-Charles Bazin, Xiaohu Lu, Yazhou Xing, Kang Liu 0003 |
ICRA | 1 |
| 2018 | Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector FieldabstractEstimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect outliers by leveraging the fact that only inliers comply with two effective consensuses, i.e., 3D ray bundle consensus and 2D vector field consensus. Our strategy has a nested structure. First, the outer module utilizes the 3D ray bundle consensus. We define the likelihood based on the probabilistic mixture model and maximize it by the expectation-maximization (EM) algorithm. The inlier probability of each correspondence and the camera pose are determined alternately. Second, the inner module exploits the 2D vector field consensus to refine the probabilities obtained by the outer module. The refinement based on the Bayesian rule facilitates the convergence of the outer module and improves the accuracy of the entire framework. Our strategy can be integrated into various existing camera pose estimation methods which are originally vulnerable to outliers. Experiments on both synthesized data and real images have shown that our approach outperforms state-of-the-art outlier rejection methods in terms of accuracy and robustness. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Jian Yao 0002 |
IROS | 1 |
| 2017 | Combining points and lines for camera pose estimation and optimization in monocular visual odometryabstractIn this paper, we propose a unified model for camera pose estimation and a novel strategy for pose optimization by combining points and lines in monocular visual odometry. Our proposed unified model treats point and line features equivalently, which is applicable for all the minimal cases requiring the minimum number 3 of point or/and line features and can be easily extended for various circumstances with more additional observations. The core idea is to directly retrieve all stationary points of a cost function which is minimized by the first-order optimality condition without initialization or iteration. The estimated pose is reliable due to robust geometric constraints and the reliable algebraic solver. To refine the camera pose, we propose a novel optimization strategy to minimize the unconstrained Sampson error by taking specific uncertainty for each feature into account to penalize noise more reasonably. Moreover, it is simpler than the conventional bundle adjustment by avoiding the high-dimensional parameter searching. Experimental results on simulated data and real images have sufficiently demonstrated the superiority of our proposed camera pose estimation and optimization method by comparing with state-of-the-art monocular algorithms. Haoang Li, Jian Yao 0002, Xiaohu Lu |
IROS | 1 |
| 2017 | 2-Line Exhaustive Searching for Real-Time Vanishing Point Estimation in Manhattan WorldabstractThis paper presents a very simple and efficient algorithm to estimate 1, 2 or 3 orthogonal vanishing point(s) on a calibrated image in Manhattan world. Unlike the traditional methods which apply 1, 3, 4, or 6 line(s) to generate vanishing point hypotheses, we propose to use 2 lines to get the first vanishing point v1, then uniformly take sample of the second vanishing point v2 on the great circle of v1 on the equivalent sphere, and finally calculate the third vanishing point v3 by the cross-product of v1 and v2. There are three advantages of the proposed method over traditional multi-line method. First, the 2-line model is much more robust and reliable than the multi-line method, which can be applied in the scene with 1, 2 or 3 orthogonal vanishing point(s). Second, the probability of the 2-line model being formed of inner line segments can be calculated given the outlier ratio, which means that the number of iterations can be determined, and thus the estimation of vanishing points can be performed in a very simple exhaustive way instead of the traditional RANSAC method. Third, the real-time performance is achieved by building a polar grid for the line intersection points, which functions as a lookup table for the validation of vanishing point hypotheses. Our algorithm has been validated successfully in the YUD dataset and sets of challenging real images. Xiaohu Lu, Jian Yao 0002, Haoang Li |
WACV | 3 |
| 2017 | Optimal seamline detection in dynamic scenes via graph cuts for image mosaicking
Li Li 0047, Jian Yao 0002, Haoang Li, Menghan Xia, Wei Zhang 0021 |
Mach. Vis. Appl. | 3 |