VLDB 2026 Research / reviewers in the wild / expert
Liangjun Zhang
dblp:25/2691
· DBLP profile ↗
67ranked-venue papers
13as first author
43since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 7 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 19 since 2021Systems, architecture and hardware · 23 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EICSeg: Universal Medical Image Segmentation via Explicit In-Context LearningabstractDeep learning models for medical image segmentation often struggle with task-specific characteristics, limiting their generalization to unseen tasks with new anatomies, labels, or modalities. Retraining or fine-tuning these models requires substantial human effort and computational resources. To address this, in-context learning (ICL) has emerged as a promising paradigm, enabling query image segmentation by conditioning on example image-mask pairs provided as prompts. Unlike previous approaches that rely on implicit modeling or non-end-to-end pipelines, we redefine the core interaction mechanism in ICL as an explicit retrieval process, termed E-ICL, benefiting from the emergence of vision foundation models (VFMs). E-ICL captures dense correspondences between queries and prompts at minimal learning cost and leverages them to dynamically weight multi-class prompt masks. Built upon E-ICL, we propose EICSeg, the first end-to-end ICL framework that integrates complementary VFMs for universal medical image segmentation. Specifically, we introduce a lightweight SD-Adapter to bridge the distinct functionalities of the VFMs, enabling more accurate segmentation predictions. To fully exploit the potential of EICSeg, we further design a scalable self-prompt training strategy and an adaptive token-to-image prompt selection mechanism, facilitating both efficient training and inference. EICSeg is trained on 47 datasets covering diverse modalities and segmentation targets. Experiments on nine unseen datasets demonstrate its strong few-shot generalization ability, achieving an average Dice score of 74.0%, outperforming existing in-context and few-shot methods by 4.5%, and reducing the gap to task-specific models to 10.8%. Even with a single prompt, EICSeg achieves a competitive average Dice score of 60.1%. Notably, it performs automatic segmentation without manual prompt engineering, delivering results comparable to interactive models while requiring minimal labeled data. Source code will be available at https://github.com/zerone-fg/EICSeg. Shiao Xie, Liangjun Zhang, Ziwei Niu, Fanfan Ye, Qiaoyong Zhong, Di Xie, Yen-Wei Chen 0001, Lanfen Lin |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Triple-Prompt Controllable Diffusion for Universal Data Augmentation in Medical Image SegmentationabstractMedical image segmentation is a crucial yet challenging task in image analysis across diverse anatomical structures. Current segmentation models heavily depend on large-scale datasets, which are laborious to collect and annotate. While generative models offer a promising alternative for data augmentation, most existing approaches are limited to single-modality outputs, either synthetic images or segmentation masks. Moreover, these methods often lack flexible conditioning mechanisms and struggle to capture the rich contextual dependencies inherent in anatomical structures. To address these challenges, in this paper, we propose TPCDM, a novel framework that co-synthesizes high-fidelity paired medical images and segmentation masks through a unified Triple-Prompt Conditional Diffusion Model. At the heart of TPCDM lies a newly defined joint image-label generation paradigm, termed Coordinated Distribution Learning, governed by three synergistic prompts: (1) a text prompt encoding global anatomical semantics; (2) a spatial prompt enforcing pixel-wise spatial coherence; (3) a task prompt dynamically adapting to diverse distributions. Furthermore, TPCDM disentangles instance-wise annotations into semantic masks and distance maps, enabling seamless extension to instance segmentation tasks. Extensive experiments on four benchmarks demonstrate that TPCDM achieves superior synthesis quality. Besides, incorporating the synthesized samples leads to state-of-the-art performance in both downstream semantic and instance segmentation tasks, while also delivering significant improvements under limited labeled data. Shiao Xie, Hongyi Wang 0002, Liangjun Zhang, Ziwei Niu, Yen-Wei Chen 0001, Lanfen Lin |
ECAI | 4 |
| 2025 | NeuS-PIR: Learning Relightable Neural Surface Using Pre-Integrated RenderingabstractIn this paper, we propose NeuS-PIR, a novel approach for learning relightable neural surfaces using pre-integrated rendering from multi-view image observations. Unlike traditional methods based on NeRFs or discrete mesh representations, our approach employs an implicit neural surface representation to reconstruct high-quality geometry. This representation enables the factorization of the radiance field into two components: a spatially varying material field and an all-frequency lighting model. By jointly optimizing this factorization with a differentiable pre-integrated rendering framework, and material encoding regularization, our method effectively addresses the ambiguity in geometry reconstruction, leading to improved disentanglement and refinement of scene properties. Furthermore, we introduce a technique to distill indirect illumination fields, capturing complex lighting effects such as inter-reflections. As a result, NeuS-PIR enables advanced applications like relighting, which can be seamlessly integrated into modern graphics engines. Extensive qualitative and quantitative experiments on both synthetic and real datasets demonstrate that NeuS-PIR outperforms existing methods across various tasks. Source code is available at https://github.com/Sheldonmao/NeuSPIR. Shi Mao, Chenming Wu, Zhelun Shen, Dayan Wu, Liangjun Zhang |
Comput. Vis. Media | 6 |
| 2025 | Constraining multimodal distribution for domain adaptation in stereo matching
Zhelun Shen, Chenming Wu, Zhibo Rao, Lina Liu 0010, Yuchao Dai, Liangjun Zhang |
Pattern Recognit. | 7 |
| 2025 | DGNR: Density-Guided Neural Point Rendering of Large Driving ScenesabstractDespite the recent success of Neural Radiance Field (NeRF), it is still challenging to render large-scale driving scenes with long trajectories, particularly when the rendering quality and efficiency are in high demand. Existing methods for such scenes usually involve with spatial warping, geometric supervision from zero-shot normal or depth estimation, or scene division strategies, where the synthesized views are often blurry or fail to meet the requirement of efficient rendering. To address the above challenges, this paper presents a novel framework that learns a density space from the scenes to guide the construction of a point-based renderer, dubbed as DGNR (Density-Guided Neural Rendering). In DGNR, geometric priors are no longer needed, which can be intrinsically learned from the density space through volumetric rendering. Specifically, we make use of a differentiable renderer to synthesize images from the neural density features obtained from the learned density space. A density-based fusion module and geometric regularization are proposed to optimize the density space. By conducting experiments on a widely used autonomous driving dataset, we have validated the effectiveness of DGNR in synthesizing photorealistic driving scenes and achieving real-time capable rendering. Our project page is available athttps://github.com/JOP-Lee/DGNR-Rendering. Note to Practitioners—While Neural Radiance Field (NeRF) has been gaining attraction, it is still challenging to create highly detailed, efficient renderings of large driving scenes. Current methods often resort to spatial warping, geometric guidance from tools like zero-shot normal or depth estimates, or dividing the scene into smaller parts. Unfortunately, these techniques can result in blurred images or fail to meet efficiency needs. To solve these challenges, we introduce a learned density space to build a point-based renderer, termed Density-Guided Neural Rendering (DGNR). With DGNR, we no longer need geometric priors because the density space can inherently learn them through volume rendering. Specifically, we use a flexible renderer to create images from the neural density features derived from the learned density space. We have also proposed a density-based fusion module and geometric regularization to optimize the density space. We evaluated DGNR on a popular autonomous driving dataset and found it to be effective in creating realistic driving scenes and capable of real-time rendering. Project page:https://github.com/JOP-Lee/DGNR-Rendering. Zhuopeng Li, Chenming Wu, Liangjun Zhang, Jianke Zhu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | TransBridge: Boost 3D Object Detection by Scene-Level Completion With Transformer Decoderabstract3D object detection is essential in autonomous driving, providing vital information about moving objects and obstacles. Detecting objects in distant regions with only a few LiDAR points is still a challenge, and numerous strategies have been developed to address point cloud sparsity through densification. This paper presents a joint completion and detection framework that improves the detection feature in sparse areas while maintaining costs unchanged. Specifically, we proposeTransBridge, a novel transformer-based up-sampling block that fuses the features from the detection and completion networks. The detection network can benefit from acquiring implicit completion features derived from the completion network. Additionally, we design theDynamic-Static Reconstruction(DSRecon) module to produce dense LiDAR data for the completion network, meeting the requirement for dense point cloud ground truth. Furthermore, we employ the transformer mechanism to establish connections between channels and spatial relations, resulting in a high-resolution feature map used for completion purposes. Extensive experiments on the nuScenes and Waymo datasets demonstrate the effectiveness of the proposed framework. The results show that our framework consistently improves end-to-end 3D object detection, with the mean average precision (mAP) ranging from 0.7 to 1.5 across multiple methods, indicating its generalization ability. For the two-stage detection framework, it also boosts the mAP up to 5.78 points. Qinghao Meng, Chenming Wu, Liangjun Zhang, Jianbing Shen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | HO-Gaussian: Hybrid Optimization of 3D Gaussian Splatting for Urban Scenes
Zhuopeng Li, Chenming Wu, Jianke Zhu, Liangjun Zhang |
ECCV (60) | 5 |
| 2024 | 3D Human Pose Estimation via Non-causal Retentive Networks
Kaili Zheng, Feixiang Lu, Yihao Lv, Liangjun Zhang, Chenyi Guo, Ji Wu 0002 |
ECCV (33) | 4 |
| 2024 | LiDAR-CS Dataset: LiDAR Point Cloud Dataset with Cross-Sensors for 3D Object DetectionabstractOver the past few years, there has been remarkable progress in research on 3D point clouds and their use in autonomous driving scenarios has become widespread. However, deep learning methods heavily rely on annotated data and often face domain generalization issues. Unlike 2D images whose domains usually pertain to the texture information present in them, the features derived from a 3D point cloud are affected by the distribution of the points. The lack of a 3D domain adaptation benchmark leads to the common practice of training a model on one benchmark (e.g. Waymo) and then assessing it on another dataset (e.g. KITTI). This setting results in two distinct domain gaps: scenarios and sensors, making it difficult to analyze and evaluate the method accurately. To tackle this problem, this paper presents ${\color{Red}\text{LiDAR}}$ Dataset with ${\color{Red}\text{C}}{\text{ross}} - {\color{Red}\text{S}}{\text{ensors}}$ (LiDAR-CS Dataset), which contains large-scale annotated LiDAR point cloud under six groups of different sensors but with the same corresponding scenarios, captured from hybrid realistic LiDAR simulator. To our knowledge, LiDAR-CS Dataset is the first dataset that addresses the sensor-related gaps in the domain of 3D object detection in real traffic. Furthermore, we evaluate and analyze the performance using various baseline detectors and demonstrated its potential applications. Project page: https://opendriving.github.io/lidar-cs. Dingfu Zhou, Chenming Wu, Chulin Tang, Cheng-Zhong Xu 0001, Liangjun Zhang |
ICRA | 7 |
| 2024 | Safety-Critical Scenario Generation Via Reinforcement Learning Based EditingabstractGenerating safety-critical scenarios is essential for testing and verifying the safety of autonomous vehicles. Traditional optimization techniques suffer from the curse of dimensionality and limit the search space to fixed parameter spaces. To address these challenges, we propose a deep reinforcement learning approach that generates scenarios by sequential editing, such as adding new agents or modifying the trajectories of the existing agents. Our framework employs a reward function consisting of both risk and plausibility objectives. The plausibility objective leverages generative models, such as a variational autoencoder, to learn the likelihood of the generated parameters from the training datasets; It penalizes the generation of unlikely scenarios. Our approach overcomes the dimensionality challenge and explores a wide range of safety-critical scenarios. Our evaluation demonstrates that the proposed method generates safety-critical scenarios of higher quality compared with previous approaches. Haolan Liu, Liangjun Zhang, Siva Kumar Sastry Hari, Jishen Zhao |
ICRA | 2 |
| 2024 | HERO-SLAM: Hybrid Enhanced Robust Optimization of Neural SLAMabstractSimultaneous Localization and Mapping (SLAM) is a fundamental task in robotics, driving numerous applications such as autonomous driving and virtual reality. Recent progress on neural implicit SLAM has shown encouraging and impressive results. However, the robustness of neural SLAM, particularly in challenging or data-limited situations, remains an unresolved issue. This paper presents HERO-SLAM, a Hybrid Enhanced Robust Optimization method for neural SLAM, which combines the benefits of neural implicit field and feature-metric optimization. This hybrid method optimizes a multi-resolution implicit field and enhances robustness in challenging environments with sudden viewpoint changes or sparse data collection. Our comprehensive experimental results on benchmarking datasets validate the effectiveness of our hybrid approach, demonstrating its superior performance over existing implicit field-based methods in challenging scenarios. HERO-SLAM provides a new pathway to enhance the stability, performance, and applicability of neural SLAM in real-world scenarios. Project page: https://hero-slam.github.io. Zhe Xin, Yufeng Yue, Liangjun Zhang, Chenming Wu |
ICRA | 3 |
| 2024 | VIHE: Virtual In-Hand Eye Transformer for 3D Robotic ManipulationabstractIn this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io. Weiyao Wang 0002, Shiyu Jin, Gregory D. Hager, Liangjun Zhang |
IROS | 5 |
| 2024 | RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language ModelabstractRecent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses this limitation by investigating the potential of adopting the reasoning ability of LLMs for generating numerical predictions in robotics tasks, specifically for robotic grasping. We propose Reasoning Tuning, a novel method that integrates a reasoning phase before prediction during training, leveraging the extensive prior knowledge and advanced reasoning abilities of LLMs. This approach enables LLMs, notably with multi-modal capabilities, to generate accurate numerical outputs like grasp poses that are context-aware and adaptable through conversations. Additionally, we present the Reasoning Tuning VLM Grasp dataset, carefully curated to facilitate the adaptation of LLMs to robotic grasping. Extensive validation on both grasping datasets and real-world experiments underscores the adaptability of multi-modal LLMs for numerical prediction tasks in robotics. This not only expands their applicability but also bridges the gap between text-based planning and direct robot control, thereby maximizing the potential of LLMs in robotics. Our dataset will be released. More details and videos of this work are available on our project page: https://sites.google.com/view/rt-grasp. Jinxuan Xu, Shiyu Jin, Liangjun Zhang |
IROS | 5 |
| 2024 | AGDF-Net: Learning Domain Generalizable Depth Features With Adaptive Guidance FusionabstractCross-domain generalizable depth estimation aims to estimate the depth of target domains (i.e., real-world) using models trained on the source domains (i.e., synthetic). Previous methods mainly use additional real-world domain datasets to extract depth specific information for cross-domain generalizable depth estimation. Unfortunately, due to the large domain gap, adequate depth specific information is hard to obtain and interference is difficult to remove, which limits the performance. To relieve these problems, we propose a domain generalizable feature extraction network with adaptive guidance fusion (AGDF-Net) to fully acquire essential features for depth estimation at multi-scale feature levels. Specifically, our AGDF-Net first separates the image into initial depth and weak-related depth components with reconstruction and contrary losses. Subsequently, an adaptive guidance fusion module is designed to sufficiently intensify the initial depth features for domain generalizable intensified depth features acquisition. Finally, taking intensified depth features as input, an arbitrary depth estimation network can be used for real-world depth estimation. Using only synthetic datasets, our AGDF-Net can be applied to various real-world datasets (i.e., KITTI, NYUDv2, NuScenes, DrivingStereo and CityScapes) with state-of-the-art performances. Furthermore, experiments with a small amount of real-world data in a semi-supervised setting also demonstrate the superiority of AGDF-Net over state-of-the-art approaches. Lina Liu 0010, Xibin Song, Mengmeng Wang 0005, Yuchao Dai, Yong Liu 0007, Liangjun Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | FloorplanNet: Learning Topometric Floorplan Matching for Robot LocalizationabstractGiven a building floorplan, humans can localize themselves by matching the observation of the environment with the floorplan using geometric, semantic, and topological clues. Inspired by this insight, this paper proposes a learning- based topometric robot localization method FloorplanNet, which implements a match between a metric robot map and the potentially inaccurate building floorplan in nonuniform scales and different shapes by semantic information. The method uses a novel Graph Neural Network to learn descriptors of nodes from topometric graphs generated from the input maps. We demonstrate that our method can match the 3D point cloud sub-map generated by the robot during the SLAM process with the 2D map. Furthermore, we apply our map-matching algorithm for real-world robot localization. We evaluate our method on several publicly available real-world datasets. Even though our network is solely trained using simulation data, our method demonstrates high robustness and effectiveness in real- world indoor environments and outperforms the existing SOTA map-matching algorithms. We further develop a simulator that automatically creates and annotates the required training data to train our neural networks. The method and simulator are released at: https://github.com/fengdelin/FloorplanNet.git Delin Feng, Zhenpeng He, Sören Schwertfeger, Liangjun Zhang |
ICRA | 5 |
| 2023 | VINet: Visual and Inertial-based Terrain Classification and Adaptive Navigation over Unknown TerrainabstractWe present a visual and inertial-based terrain classification network (VINet) for robotic navigation over different traversable surfaces. We use a novel navigation-based labeling scheme for terrain classification and generalization on unknown surfaces. Our proposed perception method and adaptive scheduling control framework can make predictions according to terrain navigation properties and lead to better performance on both terrain classification and navigation control on known and unknown surfaces. Our VINet can achieve 98.37% in terms of accuracy under supervised setting on known terrains and improve the accuracy by 8.51% on unknown terrains compared to previous methods. We deploy VINet on a mobile tracked robot for trajectory following and navigation on different terrains, and we demonstrate an improvement of 10.3% compared to a baseline controller in terms of RMSE. Tianrui Guan, Ruitao Song, Zhixian Ye, Liangjun Zhang |
ICRA | 4 |
| 2023 | Interpretable and Flexible Target-Conditioned Neural Planners For Autonomous VehiclesabstractLearning-based approaches to autonomous vehicle planners have the potential to scale to many complicated real-world driving scenarios by leveraging huge amounts of driver demonstrations. However, prior work only learns to estimate a single planning trajectory, while there may be multiple acceptable plans in real-world scenarios. To solve the problem, we propose an interpretable neural planner to regress a heatmap, which effectively represents multiple potential goals in the bird's-eye view for an autonomous vehicle. The planner employs an adaptive Gaussian kernel and relaxed hourglass loss to better capture the uncertainty of planning problems. We also use a negative Gaussian kernel to add supervision to the heatmap regression, enabling the model to learn collision avoidance effectively. Our systematic evaluation on the Lyft Open Dataset across a diverse range of real-world driving scenarios shows that our model achieves a safer and more flexible driving performance than prior works. Haolan Liu, Jishen Zhao, Liangjun Zhang |
ICRA | 3 |
| 2023 | Boosting Feedback Efficiency of Interactive Reinforcement Learning by Adaptive Learning from ScoresabstractInteractive reinforcement learning has shown promise in learning complex robotic tasks. However, the process can be human-intensive due to the requirement of a large amount of interactive feedback. This paper presents a new method that uses scores provided by humans instead of pairwise preferences to improve the feedback efficiency of interactive reinforcement learning. Our key insight is that scores can yield significantly more data than pairwise preferences. Specifically, we require a teacher to interactively score the full trajectories of an agent to train a behavioral policy in a sparse reward environment. To avoid unstable scores given by humans negatively impacting the training process, we propose an adaptive learning scheme. This enables the learning paradigm to be insensitive to imperfect or unreliable scores. We extensively evaluate our method for robotic locomotion and manipulation tasks. The results show that the proposed method can efficiently learn near-optimal policies by adaptive learning from scores while requiring less feedback compared to pairwise preference learning methods. The source codes are publicly available at https://github.com/SSKKai/Interactive-Scoring-IRL. Chenming Wu, Ying Li 0036, Liangjun Zhang |
IROS | 4 |
| 2023 | GOATS: Goal Sampling Adaptation for Scooping with Curriculum Reinforcement LearningabstractIn this work, we first formulate the problem of robotic water scooping using goal-conditioned reinforcement learning. This task is particularly challenging due to the complex dynamics of fluid and the need to achieve multi-modal goals. The policy is required to successfully reach both position goals and water amount goals, which leads to a large convoluted goal state space. To overcome these challenges, we introduce Goal Sampling Adaptation for Scooping (GOATS), a curriculum reinforcement learning method that can learn an effective and generalizable policy for robot scooping tasks. Specifically, we use a goal-factorized reward formulation and interpolate position goal distributions and amount goal distributions to create curriculum throughout the learning process. As a result, our proposed method can outperform the baselines in simulation and achieves 5.46% and 8.71% amount errors on bowl scooping and bucket scooping tasks, respectively, under 1000 variations of initial water states in the tank and a large goal state space. Besides being effective in simulation environments, our method can efficiently adapt to noisy real-robot water-scooping scenarios with diverse physical configurations and unseen settings, demonstrating superior efficacy and generalizability. The videos of this work are available on our project page: https://sites.google.com/view/goatscooping. Yaru Niu, Shiyu Jin, Zeqing Zhang, Ding Zhao, Liangjun Zhang |
IROS | 6 |
| 2023 | MapNeRF: Incorporating Map Priors into Neural Radiance Fields for Driving View SimulationabstractSimulating camera sensors is a crucial task in autonomous driving. Although neural radiance fields are exceptional at synthesizing photorealistic views in driving simulations, they still fail to generate extrapolated views. This paper proposes to incorporate map priors into neural radiance fields to synthesize out-of-trajectory driving views with semantic road consistency. The key insight is that map information can be utilized as a prior to guiding the training of the radiance fields with uncertainty. Specifically, we utilize the coarse ground surface as uncertain information to supervise the density field and warp depth with uncertainty from unknown camera poses to ensure multi-view consistency. Experimental results demonstrate that our approach can produce semantic consistency in deviated views for vehicle camera simulation. The supplementary video can be viewed at https://youtu.be/jEQWr-Rfh3A. Chenming Wu, Jiadai Sun, Zhelun Shen, Liangjun Zhang |
IROS | 4 |
| 2023 | Digging into Depth Priors for Outdoor Neural Radiance FieldsabstractNeural Radiance Fields (NeRFs) have demonstrated impressive performance in vision and graphics tasks, such as novel view synthesis and immersive reality. However, the shape-radiance ambiguity of radiance fields remains a challenge, especially in the sparse viewpoints setting. Recent work resorts to integrating depth priors into outdoor NeRF training to alleviate the issue. However, the criteria for selecting depth priors and the relative merits of different priors have not been thoroughly investigated. Moreover, the relative merits of selecting different approaches to use the depth priors is also an unexplored problem. In this paper, we provide a comprehensive study and evaluation of employing depth priors to outdoor neural radiance fields, covering common depth sensing technologies and most application ways. Specifically, we conduct extensive experiments with two representative NeRF methods equipped with four commonly-used depth priors and different depth usages on two widely used outdoor datasets. Our experimental results reveal several interesting findings that can potentially benefit practitioners and researchers in training their NeRF models with depth priors. Project page: https://cwchenwang.github.io/outdoor-nerf-depth Chen Wang 0049, Jiadai Sun, Lina Liu 0010, Chenming Wu, Zhelun Shen, Dayan Wu, Yuchao Dai, Liangjun Zhang |
ACM Multimedia | 8 |
| 2023 | Digging Into Uncertainty-Based Pseudo-Label for Robust Stereo MatchingabstractDue to the domain differences and unbalanced disparity distribution across multiple datasets, current stereo matching approaches are commonly limited to a specific dataset and generalize poorly to others. Such domain shift issue is usually addressed by substantial adaptation on costly target-domain ground-truth data, which cannot be easily obtained in practical settings. In this paper, we propose to dig into uncertainty estimation for robust stereo matching. Specifically, to balance the disparity distribution, we employ a pixel-level uncertainty estimation to adaptively adjust the next stage disparity searching space, in this way driving the network progressively prune out the space of unlikely correspondences. Then, to solve the limited ground truth data, an uncertainty-based pseudo-label is proposed to adapt the pre-trained model to the new domain, where pixel-level and area-level uncertainty estimation are proposed to filter out the high-uncertainty pixels of predicted disparity maps and generate sparse while reliable pseudo-labels to align the domain gap. Experimentally, our method shows strong cross-domain, adapt, and joint generalization and obtains 1st place on the stereo task of Robust Vision Challenge 2020. Additionally, our uncertainty-based pseudo-labels can be extended to train monocular depth estimation networks in an unsupervised way and even achieves comparable performance with the supervised methods. Zhelun Shen, Xibin Song, Yuchao Dai, Dingfu Zhou, Zhibo Rao, Liangjun Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | WSAMF-Net: Wavelet Spatial Attention-Based MultiStream Feedback Network for Single Image DehazingabstractSingle image-based dehazing has achieved remarkable progress with the development of deep learning technologies. End-to-end neural networks have been proposed to learn a direct hazy-to-clear image translation to recover the clear structures and edges cues from the hazy inputs. However, the frequency domain information is explored insufficiently and lots of intermediate structure and texture related cues of current dehazing networks are ignored, which limits the performances of current approaches. To handle these limitations mentioned above, a wavelet spatial attention based multi-stream feedback network (WSAMF-Net) is proposed for effective single image dehazing. Specifically, the proposed wavelet spatial attention utilizes both frequency-domain and spatial-domain information to enhance the extracted features for better structures and edges. Meanwhile, an enhanced multi-stream based cross feature fusion strategy, including vertical and horizontal attentions, is proposed to reweight and fuse the intermediate features of each stream to acquire more meaningful aggregated features, while the weight sharing strategy is used to achieve a good trade-off between performance and parameters. Besides, feedback mechanism is also designed to provide strong reconstruction ability. Furthermore, we propose a critical real-world industrial dataset (IDS) with images captured in real-world industrial quarry scenarios for research uses. Extensive experiments on various benchmarking datasets, including both synthetic and real-world datasets, demonstrate the superiority of our WSAMF-Net over state-of-the-art single image dehazing methods. The IDS dataset will be available athttps://github.com/XBSong/IDS-Datasethttps://github.com/XBSong/IDS-Dataset. Xibin Song, Dingfu Zhou, Wei Li 0143, Haodong Ding, Yuchao Dai, Liangjun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | TUSR-Net: Triple Unfolding Single Image Dehazing With Self-Regularization and Dual Feature to Pixel AttentionabstractSingle image dehazing is a challenging and ill-posed problem due to severe information degeneration of images captured in hazy conditions. Remarkable progresses have been achieved by deep-learning based image dehazing methods, where residual learning is commonly used to separate the hazy image into clear and haze components. However, the nature of low similarity between haze and clear components is commonly neglected, while the lack of constraint of contrastive peculiarity between the two components always restricts the performance of these approaches. To deal with these problems, we propose an end-to-end self-regularized network (TUSR-Net) which exploits the contrastive peculiarity of different components of the hazy image, i.e, self-regularization (SR). In specific, the hazy image is separated into clear and hazy components and constraint between different image components, i.e., self-regularization, is leveraged to pull the recovered clear image closer to groundtruth, which largely promotes the performance of image dehazing. Meanwhile, an effective triple unfolding framework combined with dual feature to pixel attention is proposed to intensify and fuse the intermediate information in feature, channel and pixel levels, respectively, thus features with better representational ability can be obtained. Our TUSR-Net achieves better trade-off between performance and parameter size with weight-sharing strategy and is much more flexible. Experiments on various benchmarking datasets demonstrate the superiority of our TUSR-Net over state-of-the-art single image dehazing methods. Xibin Song, Dingfu Zhou, Wei Li 0143, Yuchao Dai, Zhelun Shen, Liangjun Zhang, Hongdong Li |
IEEE Trans. Image Process. | 6 |
| 2022 | Rec2Real: Semantics-Guided Photo-Realistic Image Synthesis Using Rough Urban Reconstruction Models
Feixiang Lu, Tiancheng Xu, Liangjun Zhang |
CGI | 4 |
| 2022 | PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching
Zhelun Shen, Yuchao Dai, Xibin Song, Zhibo Rao, Dingfu Zhou, Liangjun Zhang |
ECCV (32) | 6 |
| 2022 | Semi-supervised 3D Object Detection with Proficient Teachers
Junbo Yin, Dingfu Zhou, Liangjun Zhang, Cheng-Zhong Xu 0001, Jianbing Shen, Wenguan Wang |
ECCV (38) | 4 |
| 2022 | ProposalContrast: Unsupervised Pre-training for LiDAR-Based 3D Object Detection
Junbo Yin, Dingfu Zhou, Liangjun Zhang, Cheng-Zhong Xu 0001, Jianbing Shen, Wenguan Wang |
ECCV (39) | 3 |
| 2022 | Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose DictionaryabstractWith the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds a phoneme-pose dictionary and trains a generative adversarial network (GAN) to generate video from interpolated phoneme poses. Compared to audio-driven video generation algorithms, our approach has a number of advantages: 1) It only needs about 1 min of the training data, which is significantly less than audio-driven approaches; 2) It is more flexible and not subject to vulnerability due to speaker variation; 3) It significantly reduces the preprocessing and training time from several days for audio-based methods to 4 hours, which is 10 times faster. We perform extensive experiments to compare the proposed method with state-of-the-art talking face generation methods on a benchmark dataset and datasets of our own. The results demonstrate the effectiveness and superiority of our approach. Jiahong Yuan, Miao Liao, Liangjun Zhang |
ICASSP | 4 |
| 2022 | Imitation Learning and Model Integrated Excavator Trajectory PlanningabstractAutomated excavation is promising to improve the safety and efficiency of excavators, and trajectory planning is one of the most important techniques. In this paper, we propose a two-stage method that integrates data-driven imitation learning and model-based trajectory optimization to generate optimal trajectories for autonomous excavators. We firstly train a deep neural network using demonstration data to mimic the operation patterns of human experts under various terrain states including their geometry shape and material type. Then, we use a stochastic trajectory optimization method to improve the trajectory generated by the neural network to guarantee kinematics feasibility, improve smoothness, satisfy hard constraints, and achieve desired excavation volumes. We test the proposed algorithm on a Franka robot arm equipped with a bucket end-effector. We further evaluate our method on different material types, such as sand and rigid blocks. The ex-perimental results show that the proposed two-stage algorithm by combining expert knowledge and model optimization can increase the excavation weights by up to 24.77% meanwhile with low variance. Qiangqiang Guo, Zhixian Ye, Liangjun Zhang |
IROS | 4 |
| 2022 | Excavation of Fragmented Rocks with Multi-modal Model-based Reinforcement LearningabstractThis paper presents a multi-modal model-based reinforcement learning (MBRL) approach to the excavation of fragmented rocks, which are very challenging to model due to their highly variable sizes and geometries, and visual occlusions. A multi-modal recurrent neural network (RNN) learns the dynamics of bucket-terrain interaction from a small physical dataset, with a discrete set of motion primitives encoded with domain knowledge as the action space. Then a model predictive controller (MPC) tracks a global reference path using multi-modal feedback. We show that our RNN-based dynamics function achieves lower prediction errors compared to a feed-forward neural network baseline, and the MPC is able to significantly outperform manually designed strategies on such a challenging task. Yifan Zhu 0020, Liangjun Zhang |
IROS | 3 |
| 2022 | Context-Aware 3D Object Detection From a Single Image in Autonomous DrivingabstractCamera sensors have been widely used in Driver-Assistance and Autonomous Driving Systems due to their rich texture information. Recently, with the development of deep learning techniques, many approaches have been proposed to detect objects in 3D from a single frame, however, there is still much room for improvement. In this paper, we generally review the recently proposed state-of-the-art monocular-based 3D object detection approaches first. Based on the analysis of the disadvantage of previous center-based frameworks, a novel feature aggregation strategy has been proposed to boost the 3D object detection by exploring the context information. Specifically, an Instance-Guided Spatial Attention (IGSA) module is proposed to collect the local instance information and the Channel-Wise Feature Attention (CWFA) module is employed for aggregating the global context information. In addition, an instance-guided object regression strategy is also proposed to alleviate the influence of center location prediction uncertainty in the inference process. Finally, the proposed approach has been verified on the public 3D object detection benchmark. The experimental results show that the proposed approach can significantly boost the performance of the baseline method on both 3D detection and 2D Bird’s-Eye View among all three categories. Furthermore, our method outperforms all the monocular-based methods (even these trained with depth as auxiliary inputs) and achieves state-of-the-art performance on the KITTI benchmark. Dingfu Zhou, Xibin Song, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches. Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang |
IEEE Trans. Multim. | 8 |
| 2021 | FCFR-Net: Feature Fusion based Coarse-to-Fine Residual Learning for Depth CompletionabstractDepth completion aims to recover a dense depth map from a sparse depth map with the corresponding color image as input. Recent approaches mainly formulate the depth completion as a one-stage end-to-end learning task, which outputs dense depth maps directly. However, the feature extraction and supervision in one-stage frameworks are insufficient, limiting the performance of these approaches. To address this problem, we propose a novel end-to-end residual learning framework, which formulates the depth completion as a two-stage learning task, i.e., a sparse-to-coarse stage and a coarse-to-fine stage. First, a coarse dense depth map is obtained by a simple CNN framework. Then, a refined depth map is further obtained using a residual learning strategy in the coarse-to-fine stage with coarse depth map and color image as input. Specially, in the coarse-to-fine stage, a channel shuffle extraction operation is utilized to extract more representative features from color image and coarse depth map, and an energy based fusion operation is exploited to effectively fuse these features obtained by channel shuffle operation, thus leading to more accurate and refined depth maps. We achieve SoTA performance in RMSE on KITTI benchmark. Extensive experiments on other datasets future demonstrate the superiority of our approach over current state-of-the-art depth completion approaches. Lina Liu 0010, Xibin Song, Xiaoyang Lyu, Junwei Diao, Mengmeng Wang 0005, Yong Liu 0007, Liangjun Zhang |
AAAI | 7 |
| 2021 | LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object DetectionabstractAnnotating the LiDAR point cloud is crucial for deep learning-based 3D object detection tasks. Due to expensive labeling costs, data augmentation has been taken as a necessary module and plays an important role in training the neural network. "Copy" and "paste" (i.e., GT-Aug) is the most commonly used data augmentation strategy, however, the occlusion between objects has not been taken into consideration. To handle the above limitation, we propose a rendering-based LiDAR augmentation frame-work (i.e., LiDAR-Aug) to enrich the training data and boost the performance of LiDAR-based 3D object detectors. The proposed LiDAR-Aug is a plug-and-play module that can be easily integrated into different types of 3D object detection frameworks. Compared to the traditional object augmentation methods, LiDAR-Aug is more realistic and effective. Finally, we verify the proposed framework on the public KITTI dataset with different 3D object detectors. The experimental results show the superiority of our method compared to other data augmentation strategies. We plan to make our data and code public to help other researchers reproduce our results. Xinxin Zuo, Dingfu Zhou, Shengze Jin, Sen Wang 0003, Liangjun Zhang |
CVPR | 6 |
| 2021 | Self-supervised Monocular Depth Estimation for All Day Images using Domain SeparationabstractRemarkable results have been achieved by DCNN based self-supervised depth estimation approaches. However, most of these approaches can only handle either day-time or night-time images, while their performance degrades for all-day images due to large domain shift and the variation of illumination between day and night images. To relieve these limitations, we propose a domain-separated network for self-supervised depth estimation of all-day images. Specifically, to relieve the negative influence of disturbing terms (illumination, etc.), we partition the information of day and night image pairs into two complementary sub-spaces: private and invariant domains, where the former contains the unique information (illumination, etc.) of day and night images and the latter contains essential shared information (texture, etc.). Meanwhile, to guarantee that the day and night images contain the same information, the domain-separated network takes the day-time images and corresponding night-time images (generated by GAN) as input, and the private and invariant feature extractors are learned by orthogonality and similarity loss, where the domain gap can be alleviated, thus better depth maps can be expected. Meanwhile, the reconstruction and photometric losses are utilized to estimate complementary information and depth maps effectively. Experimental results demonstrate that our approach achieves state-of-the-art depth estimation results for all-day images on the challenging Oxford RobotCar dataset, proving the superiority of our proposed approach. Code and data split are available at https://github.com/LINA-lln/ADDS-DepthNet. Lina Liu 0010, Xibin Song, Mengmeng Wang 0005, Yong Liu 0007, Liangjun Zhang |
ICCV | 5 |
| 2021 | AutoShape: Real-Time Shape-Aware Monocular 3D Object DetectionabstractExisting deep learning-based approaches for monocular 3D object detection in autonomous driving often model the object as a rotated 3D cuboid while the object’s geometric shape has been ignored. In this work, we propose an approach for incorporating the shape-aware 2D/3D constraints into the 3D detection framework. Specifically, we employ the deep neural network to learn distinguished 2D keypoints in the 2D image domain and regress their corresponding 3D coordinates in the local 3D object coordinate first. Then the 2D/3D geometric constraints are built by these correspondences for each object to boost the detection performance. For generating the ground truth of 2D/3D keypoints, an automatic model-fitting approach has been proposed by fitting the deformed 3D object model and the object mask in the 2D image. The proposed framework has been verified on the public KITTI dataset and the experimental results demonstrate that by using additional geometrical constraints the detection performance has been significantly improved as compared to the baseline method. More importantly, the proposed framework achieves state-of-the-art performance with real time. Data and code will be available at https://github.com/zongdai/AutoShape Zongdai Liu, Dingfu Zhou, Feixiang Lu, Liangjun Zhang |
ICCV | 5 |
| 2021 | Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVISabstractWe present a novel approach to robustly detect and perceive vehicles in different camera views as part of a cooperative vehicle-infrastructure system (CVIS). Our formulation is designed for arbitrary camera views and makes no assumptions about intrinsic or extrinsic parameters. First, to deal with multi-view data scarcity, we propose a part-assisted novel view synthesis algorithm for data augmentation. We train a part-based texture inpainting network in a self-supervised manner. Then we render the textured model into the background image with the target 6-DoF pose. Second, to handle various camera parameters, we present a new method that produces dense mappings between image pixels and 3D points to perform robust 2D/3D vehicle parsing. Third, we build the first CVIS dataset for bench-marking, which annotates more than 1540 images (14017 instances) from real-world traffic scenarios. We combine these novel algorithms and datasets to develop a robust approach for 2D/3D vehicle parsing for CVIS. In practice, our approach outperforms SOTA methods on 2D detection, in-stance segmentation, and 6-DoF pose estimation by 3.8%, 4.3%, and 2.9%, respectively. Feixiang Lu, Zongdai Liu, Liangjun Zhang, Dinesh Manocha |
ICCV | 4 |
| 2021 | TaskNet: A Neural Task Planner for Autonomous ExcavatorabstractWe present a novel task planner - TaskNet for an autonomous excavator based on a data-driven method, which plans feasible task-level sequence by learning from demonstration data. Given a high-level excavation objective, our TaskNet planner can decompose it into sub-tasks, each of which can be further decomposed into task primitives with specifications. We train our TaskNet using an excavation trace generator and evaluate its performance using a 3D physically-based terrain and excavator simulator. As compared to imitation learning-based methods, the experimental results show that TaskNet can effectively learn task decomposition strategies. The resulting sequences of task primitives can be used as inputs by any excavator motion planner for generating feasible joint-level trajectories. We further validate TaskNet on a state-of-the-art autonomous excavator hardware and software system. The 49-ton autonomous excavator can successfully perform material loading tasks. Jinxin Zhao, Liangjun Zhang |
ICRA | 2 |
| 2021 | MapFusion: A General Framework for 3D Object Detection with HDMapsabstract3D object detection is a key perception component in autonomous driving. Most recent approaches are based on LiDAR sensors only or fused with cameras. Maps (e.g., High Definition Maps), a basic infrastructure for intelligent vehicles, however, have not been well exploited for boosting object detection tasks. In this paper, we propose a simple but effective framework - MapFusion to integrate the map information into modern 3D object detector pipelines. In particular, we design a FeatureAgg module for HD Map feature extraction and fusion, and a MapSeg module as an auxiliary segmentation head for the detection backbone. Our proposed MapFusion is detector independent and can be easily integrated into different detectors. The experimental results of three different baselines on large public autonomous driving dataset demonstrate the superiority of the proposed framework. By fusing the map information, we can achieve 1.27 to 2.79 points improvements for mean Average Precision (mAP) on three strong 3D object detection baselines. Dingfu Zhou, Xibin Song, Liangjun Zhang |
IROS | 4 |
| 2021 | Look Before You Act: Boosting Pseudo-LiDAR with Online Semantic EmbeddingabstractVision-based 3D object detection is a research focus in the field of autonomous driving system. While recently proposed pseudo-LiDAR is a promising solution, its performance is severely restricted by the image-based depth estimator, leading to a considerable performance gap against the LiDAR-based counterparts. In this paper, substantial advances are developed along an orthogonal direction to the previous efforts in the pseudo-LiDAR pipeline. Concretely, we propose a plug- and-play module, called Online Semantic Embedding (OSE), aligning image semantics with the pseudo-LiDAR detection in an end-to-end manner. On the KITTI object detection benchmark, existing stereo-based baselines integrated with our approach show impressive improvements without bells and whistles. Furthermore, we emphasize that OSE works in retrieving the performance under geometric imperfection conditions. Liangjun Zhang, Di Xie, Shiliang Pu |
IROS | 1 |
| 2021 | Large Scale Autonomous Driving Scenarios Clustering with Self-supervised Feature ExtractionabstractThe clustering of autonomous driving scenario data can substantially benefit the autonomous driving validation and simulation systems by improving the simulation tests' completeness and fidelity. This article proposes a comprehensive data clustering framework for a large set of vehicle driving data. Existing algorithms utilize handcrafted features whose quality relies on the judgments of human experts. Additionally, the related feature compression methods are not scalable for a large dataset. Our approach thoroughly considers the traffic elements, including both in-traffic agent objects and map information. Meanwhile, we proposed a self-supervised deep learning approach for spatial and temporal feature extraction to avoid biased data representation. With the newly designed driving data clustering evaluation metrics based on data-augmentation, the accuracy assessment does not require a human-labeled dataset, which is subject to human bias. Via such unprejudiced evaluation metrics, we have shown our approach surpasses the existing methods that rely on handcrafted feature extractions. Jinxin Zhao, Zhixian Ye, Liangjun Zhang |
IV | 4 |
| 2021 | MLDA-Net: Multi-Level Dual Attention-Based Network for Self-Supervised Monocular Depth EstimationabstractThe success of supervised learning-based single image depth estimation methods critically depends on the availability of large-scale dense per-pixel depth annotations, which requires both laborious and expensive annotation process. Therefore, the self-supervised methods are much desirable, which attract significant attention recently. However, depth maps predicted by existing self-supervised methods tend to be blurry with many depth details lost. To overcome these limitations, we propose a novel framework, named MLDA-Net, to obtain per-pixel depth maps with shaper boundaries and richer depth details. Our first innovation is a multi-level feature extraction (MLFE) strategy which can learn rich hierarchical representation. Then, a dual-attention strategy, combining global attention and structure attention, is proposed to intensify the obtained features both globally and locally, resulting in improved depth maps with sharper boundaries. Finally, a reweighted loss strategy based on multi-level outputs is proposed to conduct effective supervision for self-supervised depth estimation. Experimental results demonstrate that our MLDA-Net framework achieves state-of-the-art depth prediction results on the KITTI benchmark for self-supervised monocular depth estimation with different input modes and training modes. Extensive experiments on other benchmark datasets further confirm the superiority of our proposed approach. Xibin Song, Wei Li 0143, Dingfu Zhou, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Image Process. | 7 |
| 2020 | RotPredictor: Unsupervised Canonical Viewpoint Learning for Point Cloud ClassificationabstractRecently, significant progress has been achieved in analyzing the 3D point cloud with deep learning techniques. However, existing networks suffer from poor generalization and robustness to arbitrary rotations applied to the input point cloud. Different from traditional strategies that improve the rotation robustness with data augmentation or specifically designed spherical representation or harmonics-based kernels, we propose to rotate the point cloud into a canonical viewpoint for boosting the following downstream target task, e.g., object classification and part segmentation. Specifically, the canonical viewpoint is predicted by the network RotPredictor in an unsupervised way and the loss function is only built on the target task. Our RotPredictor satisfies the rotation equivariance property in (3) approximately and the predication output has the linear relationship with the applied rotation transformation. In addition, the RotPredictor is an independent plug and play module, which can be employed by any point-based deep learning framework without extra burden. Experimental results on the public model classification dataset ModelNet40 show the performance for all baselines can be boosted by integrating the proposed module. In addition, by adding our proposed module, we can achieve the state-of-the-art classification accuracy with 90.2% on the rotation-augmented ModelNet40 benchmark. Dingfu Zhou, Xibin Song, Shengze Jin, Ruigang Yang, Liangjun Zhang |
3DV | 6 |
| 2020 | IAFA: Instance-Aware Feature Aggregation for 3D Object Detection from a Single Image
Dingfu Zhou, Xibin Song, Yuchao Dai, Junbo Yin, Feixiang Lu, Miao Liao, Liangjun Zhang |
ACCV (1) | 8 |
| 2020 | 3D Part Guided Image Editing for Fine-Grained Object UnderstandingabstractHolistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy. Zongdai Liu, Feixiang Lu, Peng Wang 0001, Liangjun Zhang, Ruigang Yang |
CVPR | 5 |
| 2020 | InstanceFusion: Real-time Instance-level 3D Reconstruction Using a Single RGBD CameraabstractAbstract We present InstanceFusion, a robust real‐time system to detect, segment, and reconstruct instance‐level 3D objects of indoor scenes with a hand‐held RGBD camera. It combines the strengths of deep learning and traditional SLAM techniques to produce visually compelling 3D semantic models. The key success comes from our novel segmentation scheme and the efficient instance‐level data fusion, which are both implemented on GPU. Specifically, for each incoming RGBD frame, we take the advantages of the RGBD features, the 3D point cloud, and the reconstructed model to perform instance‐level segmentation. The corresponding RGBD data along with the instance ID are then fused to the surfel‐based models. In order to sufficiently store and update these data, we design and implement a new data structure using the OpenGL Shading Language. Experimental results show that our method advances the state‐of‐the‐art (SOTA) methods in instance segmentation and data fusion by a big margin. In addition, our instance segmentation improves the precision of 3D reconstruction, especially in the loop closure. InstanceFusion system runs 20.5Hz on a consumer‐level GPU, which supports a number of augmented reality (AR) applications (e.g., 3D model registration, virtual interaction, AR map) and robot applications (e.g., navigation, manipulation, grasping). To facilitate future research and reproduce our system more easily, the source code, data, and the trained model are released on Github: https://github.com/Fancomi2017/InstanceFusion . Feixiang Lu, Haotian Peng, Xinhang Yang, Ruizhi Cao, Liangjun Zhang, Ruigang Yang |
Comput. Graph. Forum | 7 |
| 2019 | Compact Reachability Map for Excavator Motion PlanningabstractIn this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property that the basic excavation motion lies on the excavation plane determined by excavator links, we further reduce the construction of the map from 3D Euclidean space to 2D excavation plane. We show the pre-computed reachability map can be used to develop new excavator motion planning approach. By indexing on the pre-computed reachability map, we can efficiently compute the feasible full-bucket trajectory for single step excavation operation. We highlight the results of the reachability map construction and demonstrate the simulation results of motion planning using a commercial dynamic simulator. Yajue Yang, Liangjun Zhang, Xinjing Cheng, Jia Pan 0001, Ruigang Yang |
IROS | 2 |
| 2014 | A Selective Retraction-Based RRT Planner for Various EnvironmentsabstractWe present a novel randomized path planner for rigid robots to efficiently handle various environments that have different characteristics. We first present a bridge line test that can identify narrow passage regions and then selectively performs an optimization-based retraction only at those regions. We also propose a noncolliding line test, which is a dual operator to the bridge line test, as a culling method to avoid generating samples near wide-open free spaces. These two line tests are performed with a small computational overhead. We have tested our method with different benchmarks that have varying amounts of narrow passages. Our method achieves up to several times improvements over prior RRT-based planners and consistently shows the best performance across all the tested benchmarks. OSung Kwon, Liangjun Zhang, Sung-Eui Yoon |
IEEE Trans. Robotics | 3 |
| 2012 | SR-RRT: Selective retraction-based RRT plannerabstractWe present a novel retraction-based planner, selective retraction-based RRT, for efficiently handling a wide variety of environments that have different characteristics. We first present a bridge line-test that can identify regions around narrow passages, and then perform an optimization-based retraction operation selectively only at those regions. We also propose a non-colliding line-test, a dual operator to the bridge line-test, as a culling method to avoid generating samples near wide-open free spaces and thus to generate more samples around narrow passages. These two tests are performed with a small computational overhead and are integrated with a retraction-based RRT. In order to demonstrate benefits of our method, we have tested our method with different benchmarks that have varying amounts of narrow passages. Our method achieves up to 21 times and 3.5 times performance improvements over a basic RRT and an optimization-based retraction RRT, respectively. Furthermore, our method consistently improves the performances of other tested methods across all the tested benchmarks that have or do not have narrow passages. OSung Kwon, Liangjun Zhang, Sung-Eui Yoon |
ICRA | 3 |
| 2011 | Spoken arabic digits recognition based on wavelet neural networksabstractThe paper describes a novel method for discrete speech recognition based on spoken Arabic digit recognition by means of wavelet neural network in which Morlet wavelet is introduced to the hidden layer. The speech signal is extracted by means of Mel Frequency Cepstral Coefficients (MFCCs) and followed by vector quantization (VQ). The experimental results obtained on a spoken Arabic digit dataset proved that it could achieve better accuracy and need less learning time than the proposed method. Lvjun Zhan, Yun Xue 0002, Weixing Zhou, Liangjun Zhang |
SMC | 5 |
| 2010 | Retraction-based RRT planner for articulated modelsabstractWe present a new retraction algorithm for high DOF articulated models and use our algorithm to improve the performance of RRT planners in narrow passages. The retraction step is formulated as a constrained optimization problem and performs iterative refinement on the boundary of C-Obstacle space. We also combine the retraction algorithm with decomposition planners to handle very high DOF articulated models. The performance of our approach is analyzed using Voronoi diagrams and we show that our retraction algorithm provides a good approximation to the ideal RRT-extension in constrained environments. We have implemented our algorithm and tested its performance on robots with more than 40 DOFs in complex environments. In practice, we observe significant performance (2-80X) improvement over prior RRT planners on challenging scenarios with narrow passages. Jia Pan 0001, Liangjun Zhang, Dinesh Manocha |
ICRA | 2 |
| 2010 | A hybrid approach for simulating human motion in constrained environmentsabstractAbstract We present a new algorithm to generate plausible motions for high‐DOF human‐like articulated figures in constrained environments with multiple obstacles. Our approach is general and makes no assumptions about the articulated model or the environment. The algorithm combines hierarchical model decomposition with sample‐based planning to efficiently compute a collision‐free path in tight spaces. Furthermore, we use path perturbation and replanning techniques to satisfy the kinematic and dynamic constraints on the motion. In order to generate realistic human‐like motion, we present a new motion blending algorithm that refines the path computed by the planner with motion capture data to compute a smooth and plausible trajectory. We demonstrate the results of generating motion corresponding to placing or lifting object, walking, and bending for a 38‐DOF articulated model. Copyright © 2010 John Wiley & Sons, Ltd. Jia Pan 0001, Liangjun Zhang, Ming C. Lin, Dinesh Manocha |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Global vector field computation for feedback motion planningabstractWe present a global vector field computation algorithm in configuration spaces for smooth feedback motion planning. Our algorithm performs approximate cell decomposition in the configuration space and approximates the free space using rectanguloid cells. We compute a smooth local vector field for each cell in the free space and address the issue of the smooth composition of the local vector fields between the non-uniform adjacent cells. We show that the integral curve over the computed vector field is guaranteed to converge to the goal configuration, be collision-free, and maintain Cinfinsmoothness. As compared to prior approaches, our algorithm works well on non-convex robots and obstacles.We demonstrate its performance on planar robots with 2 or 3 DOFs, articulated robots composed of 3 serial links and multi-robot systems with 6 DOFs. Liangjun Zhang, Steven M. LaValle, Dinesh Manocha |
ICRA | 1 |
| 2008 | An efficient retraction-based RRT plannerabstractWe present a novel optimization-based retraction algorithm to improve the performance of sample-based planners in narrow passages for 3D rigid robots. The retraction step is formulated as an optimization problem using an appropriate distance metric in the configuration space. Our algorithm computes samples near the boundary of C-obstacle using local contact analysis and uses those samples to improve the performance of RRT planners in narrow passages. We analyze the performance of our planner using Voronoi diagrams and show that the tree can grow closely towards any randomly generated sample. Our algorithm is general and applicable to all polygonal models. In practice, we observe significant speedups over prior RRT planners on challenging scenarios with narrow passages. Liangjun Zhang, Dinesh Manocha |
ICRA | 1 |
| 2008 | Constrained Motion Interpolation with Distance Constraints
Liangjun Zhang, Dinesh Manocha |
WAFR | 1 |
| 2008 | Efficient distance computation in configuration space
Liangjun Zhang, Young J. Kim, Dinesh Manocha |
Comput. Aided Geom. Des. | 1 |
| 2007 | A hybrid approach for complete motion planningabstractWe present an efficient algorithm for complete motion planning that combines approximate cell decomposition (ACD) with probabilistic roadmaps (PRM). Our approach uses ACD to subdivide the configuration space into cells and computes localized roadmaps by generating samples within these cells. We augment the connectivity graph for adjacent cells in ACD with pseudo-free edges that are computed based on localized roadmaps. These roadmaps are used to capture the connectivity of free space and guide the adaptive subdivision algorithm. At the same time, we use cell decomposition to check for path non-existence and generate samples in narrow passages. Overall, our hybrid algorithm combines the efficiency of PRM methods with the completeness of ACD-based algorithms. We have implemented our algorithm on 3-DOF and 4-DOF robots. We demonstrate its performance on planning scenarios with narrow passages or no collision-free paths. In practice, we observe up to 10 times improvement in performance over prior complete motion planning algorithms. Liangjun Zhang, Young J. Kim, Dinesh Manocha |
IROS | 1 |
| 2007 | C-DIST: efficient distance computation for rigid and articulated models in configuration spaceabstractThe problem of distance computation arises in many applications including motion planning, CAD/CAM, dynamic simulation and virtual environments. Most prior work in this area has been restricted to separation or penetration distance computation between two objects. In this paper, we address the problem of computing a measure of distance between two configurations of a rigid or articulated model. The underlying distance metric is defined as the length of the longest displacement vector over the corresponding vertices of the model between two configurations. Our algorithm is based on Chasles theorem in Screw theory, and we show that the maximum distance can be realized only by a vertex of the convex hull of a rigid object. We use this formulation to compute the distance, and present two acceleration techniques to speed up the computation: incremental walking on the dual space of the convex hull and culling vertices on the convex hull using a bounding volume hierarchy (BVH). Our algorithm can be easily extended to articulated models by maximizing the distance over its each link and we also present culling techniques to accelerate the computation. We highlight the performance of our algorithm on many complex models and describe its application to proximity queries and motion planning. Liangjun Zhang, Young J. Kim, Dinesh Manocha |
Symposium on Solid and Physical Modeling | 1 |
| 2007 | Generalized penetration depth computation
Liangjun Zhang, Young J. Kim, Gokul Varadhan, Dinesh Manocha |
Comput. Aided Des. | 1 |
| 2006 | Fast C-obstacle Query Computation for Motion PlanningabstractThe configuration space of a robot is partitioned into free space and C-obstacle space. Most of the prior work in collision detection and motion planning algorithms is targeted towards checking whether a configuration or a 1D path lies in the free space. In this paper, we address the problem of checking whether a C-space primitive or a spatial cell lies completely inside C-obstacle space, without explicitly computing the boundary of C-obstacle. We refer to the problem as the C-obstacle query. We present a fast and conservative algorithm to perform this C-obstacle query. Our algorithm uses the notion of generalized penetration depth that takes into account both translational and rotational motion. We compute the generalized penetration depth for polyhedral objects and compare it with the extent of the motion that the polyhedral robot can undergo. Our approach is general and useful for designing practical algorithms for complete motion planning of rigid robots. We have integrated our query computation algorithm with star-shaped roadmaps (G. Varadhan and D. Manocha, 2005) - a deterministic sampling approach for complete motion planning. We have applied our modified planning algorithm to planar robots undergoing translational and rotational motion in complex 2D environments. Our algorithm is able to perform the C-obstacle query in milliseconds and improves the performance of the complete motion planning algorithm Liangjun Zhang, Young J. Kim, Gokul Varadhan, Dinesh Manocha |
ICRA | 1 |
| 2006 | Reliable implicit surface polygonization using visibility mapping
Gokul Varadhan, Shankar Krishnan, Liangjun Zhang, Dinesh Manocha |
Symposium on Geometry Processing | 3 |
| 2006 | Holoimages
Xianfeng Gu, Song Zhang 0002, Peisen Huang, Liangjun Zhang, Shing-Tung Yau, Ralph R. Martin |
Symposium on Solid and Physical Modeling | 4 |
| 2006 | Generalized penetration depth computationabstractPenetration depth (PD) is a distance metric that is used to describe the extent of overlap between two intersecting objects. Most of the prior work in PD computation has been restricted to translational PD, which is defined as the minimal translational motion that one of the overlapping objects must undergo in order to make the two objects disjoint. In this paper, we extend the notion of PD to take into account both translational and rotational motion to separate the intersecting objects, namely generalized PD. When an object undergoes rigid transformation, some point on the object traces the longest trajectory. The generalized PD between two overlapping objects is defined as the minimum of the longest trajectories of one object under all possible rigid transformations to separate the overlapping objects.We present three new results to compute generalized PD between polyhedral models. First, we show that for two overlapping convex polytopes, the generalized PD is same as the translational PD. Second, when the complement of one of the objects is convex, we pose the generalized PD computation as a variant of the convex containment problem and compute an upper bound using optimization techniques. Finally, when both the objects are non-convex, we treat them as a combination of the above two cases, and present an algorithm that computes a lower and an upper bound on generalized PD. We highlight the performance of our algorithms on different models that undergo rigid motion in the 6-dimensional configuration space. Moreover, we utilize our algorithm for complete motion planning of polygonal robots undergoing translational and rotational motion in a plane. In particular, we use generalized PD computation for checking path non-existence. Liangjun Zhang, Young J. Kim, Gokul Varadhan, Dinesh Manocha |
Symposium on Solid and Physical Modeling | 1 |
| 2006 | A Simple Path Non-existence Algorithm Using C-Obstacle Query
Liangjun Zhang, Young J. Kim, Dinesh Manocha |
WAFR | 1 |
| 2002 | A Feature-Based Collaborative CAD SystemabstractWith the intensification of the competition in manufacture, the distributed technology, whose aim is to promote product design process, has changed the traditional CAD serial design approach. But the distributed design systems also bring some new problems such as design conflict. To avoid this inconsistent situation, there must be some coordination mechanisms. At the same time, these mechanisms must not constrain the freedom of the designers too much to take their creativity away. This paper introduces a feature-based distributed CAD system to support this collaborative work. We analyze the reason of the design conflict. For addressing this conflict, we present feature-based concurrency operation model. In such model, the feature is the basic atom that can be locked and excluded from other designers using. Comparing with part level concurrency system, this mechanism doesn't limit the design flexibility too much. Liangjun Zhang, Min Tang 0001, Ruofeng Tong 0001, Jinxiang Dong |
CSCWD | 1 |
| 2002 | A Mesh Watermarking Approach for Appearance AttributesabstractWe describe an algorithm to watermark appearance attributes, as well as the shape of the mesh. Appearance attributes are potential watermarking primitives and the watermarking approach for them can be generalized from that for the shape. The major challenge of generalization is that the watermarking for appearance attributes has more constraints. We focus on this challenge. Especially for the normal vector, we embed the watermark by modifying its orientation, not magnitude. Results show our scheme effectively improves the capacity and enhances the robustness of mesh watermarking. Liangjun Zhang, Ruofeng Tong 0001, Feiqi Su, Jinxiang Dong |
PG | 1 |