Kangcheng Liu

dblp:257/4913 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0002-8387-3565ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 10 first-author · 19 since 2021Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding boxes to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap raises uncertainty about whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and utilize them to solve problems. To address this limitation, we introduce VP-Bench, aiming to assess MLLMs’ capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models’ ability to perceive VPs in natural scenes, utilizing 100K visualized prompts spanning 8 shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 21 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL-2.5 and Qwen2.5-VL). In addition, we conduct a comprehensive analysis of the factors influencing VP understanding, such as attribute variations and model scale. VP-Bench establishes a new reference framework for studying MLLMs’ ability to comprehend and resolve grounded referring questions.
Mingjie Xu, Jinpeng Chen 0003, Yuzhi Zhao, Jason Chun Lok Li, Zekang Du, Mengyang Wu, Kun Li 0015, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei
AAAI14
2025 AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
abstract
Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations.However, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education.In this work, we propose AesBiasBench, a benchmark designed to evaluate MLLMs along two complementary dimensions: (1) stereotype bias, quantified by measuring variations in aesthetic evaluations across demographic groups; and (2) alignment between model outputs and genuine human aesthetic preferences.Our benchmark covers three subtasks (Aesthetic Perception, Assessment, Empathy) and introduces structured metrics (IFD, NRD, AAS) to assess both bias and alignment.We evaluate 19 MLLMs, including proprietary models (e.g., GPT-4o, Claude-3.5-Sonnet)and open-source models (e.g., InternVL-2.5, Qwen2.5-VL).Results indicate that smaller models exhibit stronger stereotype biases, whereas larger models align more closely with human preferences.Incorporating identity information often exacerbates bias, particularly in emotional judgments.These findings underscore the importance of identity-aware evaluation frameworks in subjective vision-language tasks.
Kun Li 0015, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao
EMNLP5
2025 Dual-Mode Passive Fault-Tolerant Control for Underwater Vehicles with Actuator Faults and Time-Varying Disturbances
abstract
This paper investigates the control problem of underwater vehicles subject to time-varying external disturbances and actuator faults. A novel passive fault-tolerant control (PFTC) scheme is developed to address the coupled disturbance-fault dynamics inherent in underwater vehicle systems. The proposed dual-mode architecture comprises: 1) a robust fault-tolerant control scheme based on high-order sliding mode observers (HOSMOs) for minor fault scenarios, which effectively compensates for bounded disturbances and partial actuator degradation; 2) a conditionally triggered estimation mechanism integrated with fault-tolerant control allocation (FTCA) and HOSMOs for severe fault conditions, enabling fault estimation and model compensation via event-triggered parameter updating. The hybrid architecture ensures computational efficiency by activating the estimation module only when predefined triggering conditions are violated. Comprehensive experimental results validate the superiority of the proposed method in maintaining stability and performance under various fault conditions. This work provides a systematic solution for underwater vehicle control under coupled disturbance-fault conditions, with verified real-time performance and implementation feasibility.
Yizong Chen, Zhiqiang Miao, Kangcheng Liu, Yaonan Wang 0001
IROS4
2025 WLuav: An Air-Ground Robot with High Ground Adaptability and Trajectory Tracking Performance
abstract
Air-ground robots have received more and more attention and applications due to their air-to-ground motion performance and excellent energy efficiency. However, airground robots have many gaps including complex structure mechanisms, low terrain adaptability and low-precision controllers to significantly limit practical application. In this paper, an air-ground robot, wheel-leg unmanned aerial vehicle (WLuav), is proposed based on a five-link wheel leg structure to obtain excellent ground adaptive and air maneuvering capabilities. Based on the improved structure mechanism, a hierarchical adaptive agile controller is proposed to improve its trajectory tracking accuracy and ground adaptability. Besides, a mode switching strategy based on the support force solver is proposed to provide smooth and rapid mode switching. Finally, comprehensive experiments and a benchmark comparison are carried out to validate the performance of the proposed system, where the WLuav system shows excellent ground adaptive performance and trajectory tracking performance, and the energy efficiency can reach 79.46 %.
Zhiqiang Miao, Chuanpeng Niu, Kangcheng Liu, Yaonan Wang 0001
IROS4
2025 GeoScene: Temporal 3D Semantic Scene Completion with Geometric Correlation between Images
abstract
Semantic Scene Completion (SSC) aims to reconstruct the entire 3D scene in terms of both occupancy and semantics, serving as a fundamental task for autonomous driving and robotic systems. Camera-based methods have seen significant advancements due to their low cost and rich visual cues. However, previous approaches have predominantly focused on semantic recovery. This can lead to inaccurate occupancy predictions and, consequently, the failure of downstream tasks such as trajectory planning. To address this limitation, we propose a novel multi-frame matching framework, GeoScene, which reconstructs spatial structures through inter-frame geometric correlations of temporal images and subsequently infers scene semantic information. Specifically, we extract features from distinct frames in the depth dimension and derive depth features by constructing a cost volume. Following this, dot product and voxelization operations are applied between the extracted features and depth features to correct assignment errors. Furthermore, we introduce a surface normal-based regression loss to preserve fine-grained surface structures. Extensive experiments on the SemanticKITTI dataset demonstrate that GeoScene outperforms existing state-of-the-art methods.
Xiaogang Zhang 0002, Hua Chen 0008, Zhiqiang Miao, Yaonan Wang 0001, Kangcheng Liu
IROS6
2025 Bottom-Up and Top-Down Thoughts for Visual Intention Grounding
abstract
RIO (Reasoning Intention-oriented Object) is a visual grounding task aimed at locating object within an image that best matches a given intention. Due to the implicit referential nature of the intention descriptions and the inherent complexity of the scenes depicted in images, existing methods struggle to effectively accomplish this task.In this paper, we propose a novel training-free approach that decouples visual processing and reasoning to effectively solve this task. Our approach comprises two complementary pipelines: a top-down pipeline that first performs textual reasoning before proceeding to visual processing, and a bottom-up pipeline that initially conducts visual processing followed by textual reasoning. We then employ a meticulously designed strategy to merge the outputs of these parallel pipelines, thereby achieving complementarity. Experimental results demonstrate that our approach attains performance levels comparable to those of fine-tuned visual grounding models on the RIO dataset. Additional experiments conducted on the SKVG dataset further attest to the generalizability of our method.
Kangcheng Liu, Junbin Xiao, Rui Zhang 0040, Hanqi Lv, Zidong Du
ICMR1
2025 Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen
Int. J. Comput. Vis.1
2025 Correction: Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen
Int. J. Comput. Vis.1
2025 General 3D Vision-Language Model With Fast Rendering and Pre-Training Vision-Language Alignment
abstract
Current prevailing vision-language models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. The major bottleneck for the current robot 3D scene recognition approach for robotic applications is that these models do not have the capacity to recognize any unseen novel classes beyond the training categories in diverse real-world robot applications such as robot manipulation as well as robot navigation. In the meantime, current state-of-the-art 3D scene understanding approaches primarily require a large number of high-quality labels to train neural networks, which merely perform well in a fully supervised manner. Therefore, we are in urgent need of a framework that can simultaneously be applicable to both 3D point cloud segmentation and detection, particularly in the circumstances where the labels are rather scarce. This work presents a generalized and straightforward framework for dealing with 3D scene understanding when the labeled scenes are quite limited. To extract knowledge for novel categories from the pre-trained vision-language models, we propose a hierarchical feature-aligned pre-training and knowledge distillation strategy to extract and distill meaningful information from large-scale vision-language models, which helps benefit the open-vocabulary scene understanding tasks. To leverage the boundary information, we propose a novel energy-based loss with boundary awareness benefiting from the region-level boundary predictions. To encourage latent instance discrimination and to guarantee efficiency, we propose the unsupervised region-level semantic contrastive learning scheme for point clouds, using confident predictions of the neural network to discriminate the intermediate feature embeddings at multiple stages. In the limited reconstruction case, our proposed approach, termed WS3D++, ranks 1st on the large-scale ScanNet benchmark on both the task of semantic segmentation and instance segmentation. Also, our proposed WS3D++ achieves state-of-the-art data-efficient learning performance on the other large-scale real-scene indoor and outdoor datasets S3DIS and SemanticKITTI. Extensive experiments with both indoor and outdoor scenes demonstrated the effectiveness of our approach in both data-efficient learning and open-world few-shot learning.
Kangcheng Liu, Yong-Jin Liu 0001, Baoquan Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both Students
abstract
The popular methods for semi-supervised semantic segmentation mostly adopt a unitary network model using convolutional neural networks (CNNs) and enforce consistency of the model’s predictions over perturbations applied to the inputs or model. However, such a learning paradigm suffers from two critical limitations: a) learning the discriminative features for the unlabeled data; b) learning both global and local information from the whole image. In this paper, we propose a novel Semi-supervised Learning (SSL) approach, called Transformer-CNN Cohort (TCC), that consists of two students with one based on the vision transformer (ViT) and the other based on the CNN. Our method subtly incorporates the multi-level consistency regularization on the predictions and the heterogeneous feature spaces via pseudo-labeling for the unlabeled data. First, as the inputs of the ViT student are image patches, the feature maps extracted encode crucial class-wise statistics. To this end, we propose class-aware feature consistency distillation (CFCD) that first leverages the outputs of each student as the pseudo labels and generates class-aware feature (CF) maps for knowledge transfer between the two students. Second, as the ViT student has more uniform representations for all layers, we propose consistency-aware cross distillation (CCD) to transfer knowledge between the pixel-wise predictions from the cohort. We validate the TCC framework on Cityscapes and Pascal VOC 2012 datasets, which outperforms existing SSL methods by a large margin. Project page: https://vlislab22.github.io/TCC/.
Xu Zheng 0002, Yunhao Luo 0001, Chong Fu 0001, Kangcheng Liu, Lin Wang 0025
ICRA4
2024 Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning
abstract
One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop closure detection) and pose refinement (i.e., point cloud registration) problem both in isolation or combined. However, few have explicitly considered the relationship between candidate retrieval and correspondence generation in pose estimation, leaving them brittle to substructure ambiguities. To this end, we propose a hierarchical one-shot localization algorithm called Outram that leverages substructures of 3D scene graphs for locally consistent correspondence searching and global substructure-wise outlier pruning. Such a hierarchical process couples the feature retrieval and the correspondence extraction to resolve the substructure ambiguities by conducting a local-to-global consistency refinement. We demonstrate the capability of Outram in a variety of scenarios in multiple large-scale outdoor datasets. Our implementation is open-sourced: https://github.com/Pamphlett/Outram.
Pengyu Yin, Haozhi Cao, Thien-Minh Nguyen, Shenghai Yuan 0001, Kangcheng Liu, Lihua Xie 0001
ICRA6
2024 ECFO: An Efficient Edge Classification-Based Fusion Optimizer for Deep Learning Compilers
abstract
Operation fusion is a critical technique in optimizing deep learning compilers as it enhances computational efficiency by integrating multiple operations into a single computational graph. However, finding an effective fusion strategy is challenging, requiring the definition of an optimization search space and identification of the best strategy within this space. Existing methods, such as heuristic searches and learning-based searches, have significant limitations. Heuristic searches are complex, labor-intensive, and often lack generalizability across different network architectures. On the other hand, learning-based methods demand extensive training and pro-longed search time. To address these challenges, we introduce the Edge Classification-Based Fusion Optimizer (ECFO), a novel approach that reconceptualizes operation fusion as an edge classification problem. By leveraging Graph Neural Networks (GNNs) for efficient graph feature encoding, ECFO streamline the optimization process and significantly reduces computational overhead. Comprehensive evaluations across diverse neural networks demonstrate that ECFO decrease search time by up to 23x and improves inference performance by 3.2%, representing a substantial advancement over existing strategies.
Wei Li 0008, Kangcheng Liu, Lide Xue, Zidong Du, Xishan Zhang, Xuehai Zhou
SMC4
2024 Domain Adaptive LiDAR Point Cloud Segmentation via Density-Aware Self-Training
abstract
Domain adaptive LiDAR point cloud segmentation aims to learn a target segmentation model from labeled source point clouds and unlabelled target point clouds, which has recently attracted increasing attention due to various challenges in point cloud annotation. However, its performance is still very constrained as most existing studies did not well capture data-specific characteristics of LiDAR point clouds. Inspired by the observation that the domain discrepancy of LiDAR point clouds is highly correlated with point density, we design a density-aware self-training (DAST) technique that introduces point density into the self-training framework for domain adaptive point cloud segmentation. DAST consists of two novel and complementary designs. The first is density-aware pseudo labelling that introduces point density for accurate pseudo labelling of target data and effective self-supervised network retraining. The second is density-aware consistency regularization that encourages to learn density-invariant representations by enforcing target predictions to be consistent across points of different densities. Extensive experiments over multiple large-scale public datasets show that DAST achieves superior domain adaptation performance as compared with the state-of-the-art.
Aoran Xiao, Jiaxing Huang 0001, Kangcheng Liu, Dayan Guan, Xiaoqin Zhang 0002, Shijian Lu
IEEE Trans. Intell. Transp. Syst.3
2023 FAC: 3D Representation Learning via Foreground Aware Feature Contrast
abstract
Contrastive learning has recently demonstrated great potential for unsupervised pre-training in 3D scene understanding tasks. However, most existing work randomly selects point features as anchors while building contrast, leading to a clear bias toward background points that often dominate in 3D scenes. Also, object awareness and foreground-to-background discrimination are neglected, making contrastive learning less effective. To tackle these issues, we propose a general foreground-aware feature contrast (FAC) framework to learn more effective point cloud representations in pre-training. FAC consists of two novel contrast designs to construct more effective and informative contrast pairs. The first is building positive pairs within the same foreground segment where points tend to have the same semantics. The second is that we prevent over-discrimination between 3D segments/objects and encourage foreground-to-background distinctions at the segment level with adaptive feature learning in a Siamese correspondence network, which adaptively learns feature correlations within and across point cloud views effectively. Visualization with point activation maps shows that our contrast pairs capture clear correspondences among fore-ground regions during pre-training. Quantitative experiments also show that FAC achieves superior knowledge transfer and data efficiency in various downstream 3D semantic segmentation and object detection tasks. All codes, data, and models are available.
Kangcheng Liu, Aoran Xiao, Xiaoqin Zhang 0002, Shijian Lu, Ling Shao 0001
CVPR1
2023 3D Semantic Segmentation in the Wild: Learning Generalized Models for Adverse-Condition Point Clouds
abstract
Robust point cloud parsing under all-weather conditions is crucial to level-5 autonomy in autonomous driving. However, how to learn a universal 3D semantic segmentation (3DSS) model is largely neglected as most existing benchmarks are dominated by point clouds captured under normal weather. We introduce SemanticSTF, an adverse-weather point cloud dataset that provides dense point-level annotations and allows to study 3DSS under various adverse weather conditions. We study all-weather 3DSS modeling under two setups: 1) domain adaptive 3DSS that adapts from normal-weather data to adverse-weather data; 2) domain generalizable 3DSS that learns all-weather 3DSS models from normal-weather data. Our studies reveal the challenge while existing 3DSS methods encounter adverse-weather data, showing the great value of SemanticSTF in steering the future endeavor along this very meaningful research direction. In addition, we design a domain randomization technique that alternatively randomizes the geometry styles of point clouds and aggregates their embeddings, ultimately leading to a generalizable model that can improve 3DSS under various adverse weather effectively. The SemanticSTF and related codes are available at https://github.com/xiaoaoran/SemanticSTF.
Aoran Xiao, Jiaxing Huang 0001, Weihao Xuan, Ruijie Ren, Kangcheng Liu, Dayan Guan, Abdulmotaleb El Saddik, Shijian Lu, Eric P. Xing
CVPR5
2023 DoubleBee: A Hybrid Aerial-Ground Robot with Two Active Wheels
abstract
In this paper, we present the dynamic model and control of DoubleBee, a novel hybrid aerial-ground vehicle consisting of two propellers mounted on tilting servo motors and two motor-driven wheels. DoubleBee exploits the high energy efficiency of a bicopter configuration in aerial mode, and enjoys the low power consumption of a two-wheel self-balancing robot on the ground. Furthermore, the propeller thrusts act as additional control inputs on the ground, enabling a novel decoupled control scheme where the attitude of the robot is controlled using thrusts and the translational motion is realized using wheels. A prototype of DoubleBee is constructed using commercially available components. The power efficiency and the control performance of the robot are verified through comprehensive experiments. Challenging tasks in indoor and outdoor environments demonstrate the capability of DoubleBee to traverse unstructured environments, fly over and move under barriers, and climb steep and rough terrains.
Muqing Cao, Xinhang Xu, Shenghai Yuan 0001, Kun Cao 0002, Kangcheng Liu, Lihua Xie 0001
IROS5
2023 RM3D: Robust Data-Efficient 3D Scene Parsing via Traditional and Learnt 3D Descriptors-Based Semantic Region Merging
Kangcheng Liu
Int. J. Comput. Vis.1
2023 FG-Net: A Fast and Accurate Framework for Large-Scale LiDAR Point Cloud Understanding
abstract
This work presents FG-Net, a general deep learning framework for large-scale point cloud understanding without voxelizations, which achieves accurate and real-time performance with a single NVIDIA GTX 1080 8G GPU and an i7 CPU. First, a novel noise and outlier filtering method is designed to facilitate the subsequent high-level understanding tasks. For effective understanding purpose, we propose a novel plug-and-play module consisting of correlated feature mining and deformable convolution-based geometric-aware modeling, in which the local feature relationships and point cloud geometric structures can be fully extracted and exploited. For the efficiency issue, we put forward a new composite inverse density sampling (IDS)-based and learning-based operation and a feature pyramid-based residual learning strategy to save the computational cost and memory consumption, respectively. Compared with current methods which are only validated on limited datasets, we have done extensive experiments on eight real-world challenging benchmarks, which demonstrates that our approaches outperform state-of-the-art (SOTA) approaches in terms of accuracy, speed, and memory efficiency. Moreover, weakly supervised transfer learning is also conducted to demonstrate the generalization capacity of our method.
Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen
IEEE Trans. Cybern.1
2023 SVCNet: Scribble-Based Video Colorization Network With Temporal Aggregation
abstract
In this paper, we propose a scribble-based video colorization network with temporal aggregation called SVCNet. It can colorize monochrome videos based on different user-given color scribbles. It addresses three common issues in the scribble-based video colorization area: colorization vividness, temporal consistency, and color bleeding. To improve the colorization quality and strengthen the temporal consistency, we adopt two sequential sub-networks in SVCNet for precise colorization and temporal smoothing, respectively. The first stage includes a pyramid feature encoder to incorporate color scribbles with a grayscale frame, and a semantic feature encoder to extract semantics. The second stage finetunes the output from the first stage by aggregating the information of neighboring colorized frames (as short-range connections) and the first colorized frame (as a long-range connection). To alleviate the color bleeding artifacts, we learn video colorization and segmentation simultaneously. Furthermore, we set the majority of operations on a fixed small image resolution and use a Super-resolution Module at the tail of SVCNet to recover original sizes. It allows the SVCNet to fit different image resolutions at the inference. Finally, we evaluate the proposed SVCNet on DAVIS and Videvo benchmarks. The experimental results demonstrate that SVCNet produces both higher-quality and more temporally consistent videos than other well-known video colorization approaches. The codes and models can be found at https://github.com/zhaoyuzhi/SVCNet.
Yuzhi Zhao, Lai-Man Po, Kangcheng Liu, Wing Yin Yu, Pengfei Xian, Yujia Zhang 0002, Mengyang Liu
IEEE Trans. Image Process.3
2023 Synergizing Low Rank Representation and Deep Learning for Automatic Pavement Crack Detection
abstract
Due to the critical role of pavement crack detection for road maintenance and eventually ensuring safety, remarkable efforts have been devoted to this research area, and such a trend is further intensified for the coming unmanned vehicle era. However, such crack detection task still remains unexpectedly challenging in practice since the appearance of both cracks and the background are diverse and complex in real scenarios. In this work, we propose an automatic pavement crack detection method via synergizing low rank representation (LRR) and deep learning techniques. First, leveraging LRR which facilitates anomaly detection without making any specific assumption, we can easily discriminate most of the frames with cracks from the long sequence with a consistent pavement base, followed by a straightforward algorithm to localize the cracks. In order to achieve the intelligence of detecting cracks with different pavement basis under unconstrained imaging conditions, we resort to deep learning techniques and propose a deep convolutional neural network for crack detection leveraging on multi-level features and atrous spatial pyramid pooling (ASPP). We train this network based on the training data obtained in the previous stage in an end-to-end manner. Extensive experiments on a wide range of pavements demonstrate the high performance in terms of both accuracy and automaticity. Moreover, the dataset generated by us is much more extensive and challenging than public ones. We put it online athttps://gaozhinuswhu.comto benefit the community.
Zhi Gao 0005, Min Cao 0001, Ziyao Li, Kangcheng Liu, Ben M. Chen
IEEE Trans. Intell. Transp. Syst.5
2022 Weakly Supervised 3D Scene Segmentation with Region-Level Boundary Awareness and Instance Discrimination
Kangcheng Liu, Yuzhi Zhao, Qiang Nie, Zhi Gao 0005, Ben M. Chen
ECCV (28)1
2022 WeakLabel3D-Net: A Complete Framework for Real-Scene LiDAR Point Clouds Weakly Supervised Multi-Tasks Understanding
abstract
Existing state-of-the-art 3D point clouds understanding methods only perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework which simultaneously solves the downstream high-level understanding tasks, especially when labels are extremely limited. This work presents a general and simple framework to tackle point clouds understanding when labels are limited. We propose a novel unsupervised region expansion based clustering method for generating clusters. More importantly, we innovatively propose to learn to merge the over-divided clusters based on the local low-level geometric property similarities and the learned high-level feature similarities supervised by weak labels. Hence, the true weak labels guide pseudo labels merging taking both geometric and semantic feature correlations into consideration. Finally, the self-supervised data augmentation optimization module is proposed to guide the propagation of labels among semantically similar points within a scene. Experimental Results demonstrate that our framework has the best performance among the three most important weakly supervised point clouds understanding tasks including semantic segmentation, instance segmentation, and object detection even when limited points are labeled.
Kangcheng Liu, Yuzhi Zhao, Zhi Gao 0005, Ben M. Chen
ICRA1
2022 D-LC-Nets: Robust Denoising and Loop Closing Networks for LiDAR SLAM in Complicated Circumstances with Noisy Point Clouds
abstract
The current LiDAR SLAM (Simultaneous Localization and Mapping) system suffers greatly from low accuracy and limited robustness when faced with complicated circumstances. From our experiments, we find that current LiDAR SLAM systems have limited performance when the noise level in the obtained point clouds is large. Therefore, in this work, we propose a general framework to tackle the problem of denoising and loop closure for LiDAR SLAM in complex environments with many noises and outliers caused by reflective materials. Current approaches for point clouds denoising are mainly designed for small-scale point clouds and can not be extended to large-scale point clouds scenes. In this work, we firstly proposed a lightweight network for large-scale point clouds denoising. Subsequently, we have also designed an efficient loop closure network for place recognition in global optimization to improve the localization accuracy of the whole system. Finally, we have demonstrated by extensive experiments and benchmark studies that our method can have a significant boost on the localization accuracy of the LiDAR SLAM system when faced with noisy point clouds, with a marginal increase in computational cost.
Kangcheng Liu, Aoran Xiao, Jiaxing Huang 0001, Kaiwen Cui, Yun Xing 0001, Shijian Lu
IROS1
2022 Robust Industrial UAV/UGV-Based Unsupervised Domain Adaptive Crack Recognitions with Depth and Edge Awareness: From System and Database Constructions to Real-Site Inspections
abstract
The defect diagnosis of modern infrastructures is crucial to public safety. In this work, we propose a complete crack inspection system with three main components, including the autonomous system setup, the geographic-information-system-based 3D reconstruction, and the database construction as well as domain adaptive algorithms design. To fulfill the unsupervised domain adaptation (UDA) task of cracks recognition in infrastructural inspections, we propose a robust unsupervised domain adaptive learning strategy termed Crack-DA to increase the generalization capacity of the model in unseen test circumstances. Specifically, firstly, we propose leveraging the self-supervised depth information to help the learning of semantics. Secondly, we propose using the edge information to suppress the non-edge background objects and noises. Thirdly, we propose using the data augmentation-based consistency learning to increase the prediction robustness. Finally, we use the disparity in depth to evaluate the domain gap in semantics and explicitly consider the domain gap in the optimization of the network. Also, we propose a source database consisting of 11,298 crack images with detailed pixel-level labels for network training in domain adaptations. Extensive experiments on UAV-captured highway cracks and real-site UAV inspections of building cracks demonstrate the robustness and effectiveness of the proposed domain adaptive crack recognition approach.
Kangcheng Liu
ACM Multimedia1
2021 FG-Conv: Large-Scale LiDAR Point Clouds Understanding Leveraging Feature Correlation Mining and Geometric-Aware Modeling
abstract
This work presents a general deep learning framework for large-scale point clouds understanding without voxelizations, called FG-Conv, which achieves an accurate and real-time understanding of point clouds. Through our novel design combining feature level correlation mining and deformable convolutions based geometric aware modeling, the local feature relationships and geometric patterns can be captured. The attention mechanism is also adopted to enhance the global long-range feature correlations. Finally, the feature pyramid residual learning network is proposed to combine patterns at different resolutions in a memory-efficient way. Extensive experiments on real-world challenging datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and efficiency. Weakly supervised transfer learning demonstrates the generalization capacity of our methods.
Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen
ICRA1
2021 Legacy Photo Editing with Learned Noise Prior
abstract
There are quite a number of photographs captured under undesirable conditions in the last century. Thus, they are of-ten noisy, regionally incomplete, and grayscale formatted. Conventional approaches mainly focus on one point so that those restoration results are not perceptually sharp or clean enough. To solve these problems, we propose a noise prior learner NEGAN to simulate the noise distribution of real legacy photos using unpaired images. It mainly focuses on matching high-frequency parts of noisy images through discrete wavelet transform (DWT) since they include most of noise statistics. We also create a large legacy photo dataset for learning noise prior. Using learned noise prior, we can easily build valid training pairs by degrading clean images. Then, we propose an IEGAN framework performing image editing including joint denoising, inpainting and colorization based on the estimated noise prior. We evaluate the proposed system and compare it with state-of-the-art image enhancement methods. The experimental results demonstrate that it achieves the best perceptual quality. Please see the webpage https://github.com/zhaoyuzhi/Legacy-Photo-Editing-with-Learned-Noise-Prior for the codes and the proposed LP dataset.
Yuzhi Zhao, Lai-Man Po, Tingyu Lin 0002, Kangcheng Liu, Yujia Zhang 0002, Wing Yin Yu, Pengfei Xian, Jingjing Xiong
WACV5