Liang Du 0004

dblp:40/5548-4 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-7952-5736ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 UL-SLAM: A Universal Monocular Line-Based SLAM Via Unifying Structural and Non-Structural Constraints
abstract
Leveraging structural line features to complement sparse point features has been studied in recent years. However, this approach relies on a Manhattan world assumption and does not incorporate non-structural lines due to the triangulation degeneracy problem and tracking process instability. To address these problems, we propose a general line-based SLAM system that combines points, structural and non-structural lines. First, an efficient line matching algorithm for multi-scale is designed to obtain more accurate matching pairs through a divide-and-conquer approach. In addition, a novel line triangulation strategy utilizing spatial-temporal consistency and degeneracy identification is proposed to improve the quality of line generation in a sliding window. Finally, universal structural constraints based on measurement of the vanishing directions are implemented to complement the information missing from the Plücker line projection in local mapping optimization. Extensive experiments are conducted on the public EuRoC and TUM datasets as well as a self-collected dataset, and the results show that UL-SLAM achieves cutting-edge performance among recent state-of-the-art methods in both accuracy and speed. Ablation experiments also demonstrate that the integration of different line features can improve the robustness and accuracy of a visual SLAM system in challenging scenarios with low texture and weak illumination. Our implementation of the UL-SLAM will be open-sourced to benefit the community (https://github.com/jhch1995/UL-SLAM).Note to Practitioners—This article was motivated by the challenges of visual localization problems in human-made indoor scenes. Visual localization has been widely used in various robotic fields such as self-driving vehicles, augmented reality (AR), and virtual reality (VR). In real-world scenes, the localization accuracy will be significantly decreased because of the sparse and uncertain visual features in low-texture or weak illumination environments, which reduces the robustness of robot tracking. To address this problem, this article proposes a novel universal line-based SLAM system (UL-SLAM) that unifies structural and non-structural constraints within a general framework unrestricted by the strong global Manhattan world assumption. UL-SLAM can not only improve the accuracy of pose estimation due to the proposed methods for the line features but also achieve real-time performance. In addition, for 3D mapping construction, UL-SLAM can also enrich the geometric structure information of the indoor scenes. Extensive experiments are conducted on various indoor datasets for autonomous robots, and the results demonstrate the efficiency, accuracy, and robustness of the proposed system in different complex scenarios.
Haochen Jiang, Rui Qian 0004, Liang Du 0004, Jian Pu, Jianfeng Feng
IEEE Trans Autom. Sci. Eng.3
2024 DELTA: Dynamic Embedding Learning with Truncated Conscious Attention for CTR Prediction
Liang Du 0004, Hong Chen 0011, Zixun Sun, Xin Wang 0019, Wenwu Zhu 0001
CogSci2
2024 Stable Heterogeneous Treatment Effect Estimation across Out-of-Distribution Populations
abstract
Heterogeneous treatment effect (HTE) estimation is vital for understanding the change of treatment effect across individuals or subgroups. Most existing HTE estimation methods focus on addressing selection bias induced by imbalanced distributions of confounders between treated and control units, but ignore distribution shifts across populations. Thereby, their applicability has been limited to the in-distribution (ID) population, which shares a similar distribution with the training dataset. In real-world applications, where population distributions are subject to continuous changes, there is an urgent need for stable HTE estimation across out-of-distribution (OOD) populations, which, however, remains an open problem. As pioneers in resolving this problem, we propose a novel Stable Balanced Representation Learning with Hierarchical-Attention Paradigm (SBRL-HAP) framework, which consists of 1) Balancing Regularizer for eliminating selection bias, 2) Independence Regularizer for addressing the distribution shift issue, 3) Hierarchical-Attention Paradigm for coordination between balance and independence. In this way, SBRL- HAP regresses counterfactual outcomes using ID data, while ensuring the resulting HTE estimation can be successfully generalized to out-of-distribution scenarios, thereby enhancing the model's applicability in real-world settings. Extensive experiments conducted on synthetic and real-world datasets demonstrate the effectiveness of our SBRL-HAP in achieving stable HTE estimation across OOD populations, with an average 10% reduction in the error metric PEHE and 11% decrease in the ATE bias, compared to the SOTA methods.
Anpeng Wu, Kun Kuang 0001, Liang Du 0004, Zixun Sun
ICDE4
2024 RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction
abstract
Click-through rate (CTR) prediction is a critical task in recommendation systems, serving as the ultimate filtering step to sort items for a user. Most recent cutting-edge methods primarily focus on investigating complex implicit and explicit feature interactions; however, these methods neglect the spurious correlation issue caused by confounding factors, thereby diminishing the model’s generalization ability. We propose a CTR prediction framework that REmoves Spurious cORrelations in mulTilevel feature interactions, termed RE-SORT, which has two key components. I. A multilevel stacked recurrent (MSR) structure enables the model to efficiently capture diverse nonlinear interactions from feature spaces at different levels. II. A spurious correlation elimination (SCE) module further leverages Laplacian kernel mapping and sample reweighting methods to eliminate the spurious correlations concealed within the multilevel features, allowing the model to focus on the true causal features. Extensive experiments conducted on four challenging CTR datasets, our production dataset, and an online A/B test demonstrate that the proposed method achieves state-of-the-art performance in both accuracy and speed. The utilized codes, models, and dataset will be released at https://github.com/RE-SORT.
Songli Wu, Liang Du 0004, Yuai Wang, De-Chuan Zhan, Zixun Sun
UAI2
2024 Rethinking precision of pseudo label: Test-time adaptation via complementary learning
Longbin Zeng, Jiayi Han, Liang Du 0004, Weiyang Ding
Pattern Recognit. Lett.3
2022 Modify Self-Attention via Skeleton Decomposition for Effective Point Cloud Transformer
abstract
Although considerable progress has been achieved regarding the transformers in recent years, the large number of parameters, quadratic computational complexity, and memory cost conditioned on long sequences make the transformers hard to train and implement, especially in edge computing configurations. In this case, a dizzying number of works have sought to make improvements around computational and memory efficiency upon the original transformer architecture. Nevertheless, many of them restrict the context in the attention to seek a trade-off between cost and performance with prior knowledge of orderly stored data. It is imperative to dig deep into an efficient feature extractor for point clouds due to their irregularity and a large number of points. In this paper, we propose a novel skeleton decomposition-based self-attention (SD-SA) which has no sequence length limit and exhibits favorable scalability in long-sequence models. Due to the numerical low-rank nature of self-attention, we approximate it by the skeleton decomposition method while maintaining its effectiveness. At this point, we have shown that the proposed method works for the proposed approach on point cloud classification, segmentation, and detection tasks on the ModelNet40, ShapeNet, and KITTI datasets, respectively. Our approach significantly improves the efficiency of the point cloud transformer and exceeds other efficient transformers on point cloud tasks in terms of the speed at comparable performance.
Jiayi Han, Longbin Zeng, Liang Du 0004, Xiaoqing Ye, Weiyang Ding, Jianfeng Feng
AAAI3
2022 Repainting and Imitating Learning for Lane Detection
abstract
Current lane detection methods are struggling with the invisibility lane issue caused by heavy shadows, severe road mark degradation, and serious vehicle occlusion. As a result, discriminative lane features can be barely learned by the network despite elaborate designs due to the inherent invisibility of lanes in the wild. In this paper, we target at finding an enhanced feature space where the lane features are distinctive while maintaining a similar distribution of lanes in the wild. To achieve this, we propose a novel Repainting and Imitating Learning (RIL) framework containing a pair of teacher and student without any extra data or extra laborious labeling. Specifically, in the repainting step, an enhanced ideal virtual lane dataset is built in which only the lane regions are repainted while non-lane regions are kept unchanged, maintaining the similar distribution of lanes in the wild. The teacher model learns enhanced discriminative representation based on the virtual data and serves as the guidance for a student model to imitate. In the imitating learning step, through the scale-fusing distillation module, the student network is encouraged to generate features that mimic the teacher model both on the same scale and cross scales. Furthermore, the coupled adversarial module builds the bridge to connect not only teacher and student models but also virtual and real data, adjusting the imitating learning process dynamically. Note that our method introduces no extra time cost during inference and can be plug-and-play in various cutting-edge lane detection networks. Experimental results prove the effectiveness of the RIL framework both on CULane and TuSimple for four modern lane detection methods. The code and model will be available soon.
Minyue Jiang, Xiaoqing Ye, Liang Du 0004, Zhikang Zou, Wei Zhang 0197, Xiao Tan 0001, Errui Ding
ACM Multimedia4
2022 The devil is in the face: Exploiting harmonious representations for facial expression recognition
Jiayi Han, Liang Du 0004, Xiaoqing Ye, Li Zhang 0040, Jianfeng Feng
Neurocomputing2
2022 AGO-Net: Association-Guided 3D Point Cloud Object Detection Network
abstract
The human brain can effortlessly recognize and localize objects, whereas current 3D object detection methods based on LiDAR point clouds still report inferior performance for detecting occluded and distant objects: The point cloud appearance varies greatly due to occlusion, and has inherent variance in point densities along the distance to sensors. Therefore, designing feature representations robust to such point clouds is critical. Inspired by human associative recognition, we propose a novel 3D detection framework that associates intact features for objects via domain adaptation. We bridge the gap between the perceptual domain, where features are derived from real scenes with sub-optimal representations, and the conceptual domain, where features are extracted from augmented scenes that consist of non-occlusion objects with rich detailed information. A feasible method is investigated to construct conceptual scenes without external datasets. We further introduce an attention-based re-weighting module that adaptively strengthens the feature adaptation of more informative regions. The network's feature enhancement ability is exploited without introducing extra cost during inference, which is plug-and-play in various 3D detection frameworks. We achieve new state-of-the-art performance on the KITTI 3D detection benchmark in both accuracy and speed. Experiments on nuScenes and Waymo datasets also validate the versatility of our method.
Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Edward Johns, Errui Ding, Xiangyang Xue 0001, Jianfeng Feng
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection
abstract
The objective of this paper is to learn context- and depth- aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message propagation (DDMP) network to effectively integrate the multi-scale depth information with the image context; (ii) this is achieved by first adaptively sampling context-aware nodes in the image context and then dynamically predicting hybrid depth-dependent filter weights and affinity matrices for propagating information; (Hi) by augmenting a center-aware depth encoding (CDE) task, our method successfully alleviates the inaccurate depth prior; (iv) we thoroughly demonstrate the effectiveness of our proposed approach and show state-of-the-art results among the monocular-based approaches on the KITTI benchmark dataset. Particularly, we rank 1stin the highly competitive KITTI monocular 3D object detection track on the submission day (November 16th, 2020). Code and models are released at https: //github.com/fudan-zvg/DDMP
Li Wang 0033, Liang Du 0004, Xiaoqing Ye, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001, Jianfeng Feng, Li Zhang 0040
CVPR2
2021 The Devil is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection
abstract
Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. In this paper, we dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits to a deep excavation of reciprocal information underlying the entire task. We introduce a Dynamic Feature Reflecting Network, named DFR-Net, which contains two novel standalone modules: (i) the Appearance-Localization Feature Reflecting module (ALFR) that first separates task-specific features and then self-mutually reflects the reciprocal features; (ii) the Dynamic Intra-Trading module (DIT) that adaptively realigns the training processes of various sub-tasks via a self-learning manner. Extensive experiments on the challenging KITTI dataset demonstrate the effectiveness and generalization of DFR-Net. We rank 1stamong all the monocular 3D object detectors in the KITTI test set (till March 16th, 2021). The proposed method is also easy to be plug-and-play in many cutting-edge 3D detection frameworks at negligible cost to boost performance. The code will be made publicly available.
Zhikang Zou, Xiaoqing Ye, Liang Du 0004, Xianhui Cheng, Xiao Tan 0001, Li Zhang 0040, Jianfeng Feng, Xiangyang Xue 0001, Errui Ding
ICCV3
2021 AggNet for Self-supervised Monocular Depth Estimation: Go An Aggressive Step Furthe
abstract
Without appealing to exhaustive labeled data, self-supervised monocular depth estimation (MDE) plays a fundamental role in computer vision. Previous methods usually adopt a one-stage MDE network, which is insufficient to achieve high performance. In this paper, we dig deep into this task to propose an aggressive framework termed AggNet. The framework is based on a training-only progressive two-stage module to perform pseudo counter-surveillance as well as a simple yet effective dual-warp loss function between image pairs. In particular, we first propose a residual module, which follows the MDE network to learn a refined depth. The residual module takes both the initial depth generated from MDE and the initial color image as input to generate refined depth with residual depth learning. Then, the refined depth is leveraged to supervise the initial depth simultaneously during the training period. For inference, only the MDE network is retained to regress depth from a single image, which gains better performance without introducing extra computation. In addition to self-distillation loss, a simple yet effective dual-warp consistency loss is introduced to encourage the MDE network to keep depth consistency between stereo image pairs. Extensive experiments show that our AggNet achieves state-of-the-art performance on the KITTI and Make3D datasets.
Zhi Chen 0026, Xiaoqing Ye, Liang Du 0004, Wei Yang 0011, Liusheng Huang, Xiao Tan 0001, Zhenbo Shi, Fumin Shen, Errui Ding
ACM Multimedia3
2020 Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object Detection
abstract
Object detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point cloud data. Designing robust feature representation against such appearance changes is hence the key issue in a 3D object detection method. In this paper, we innovatively propose a domain adaptation like approach to enhance the robustness of the feature representation. More specifically, we bridge the gap between the perceptual domain where the feature comes from a real scene and the conceptual domain where the feature is extracted from an augmented scene consisting of non-occlusion point cloud rich of detailed information. This domain adaptation approach mimics the functionality of the human brain when proceeding object perception. Extensive experiments demonstrate that our simple yet effective approach fundamentally boosts the performance of 3D point cloud object detection and achieves the state-of-the-art results.
Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Jianfeng Feng, Zhenbo Xu, Errui Ding, Shilei Wen
CVPR1
2020 Monocular 3D Object Detection via Feature Domain Adaptation
Xiaoqing Ye, Liang Du 0004, Yifeng Shi, Xiao Tan 0001, Jianfeng Feng, Errui Ding, Shilei Wen
ECCV (9)2
2020 3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection
abstract
We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptively selects and fuses the reciprocal semantic and instance features from two tasks in a coupled manner. To further boost the performance of the instance segmentation task in our 3DCFS, we investigate a loss function that helps the model learn to balance the magnitudes of the output embedding dimensions during training, which makes calculating the Euclidean distance more reliable and enhances the generalizability of the model. Extensive experiments demonstrate that our 3DCFS outperforms state-of-the-art methods on benchmark datasets in terms of accuracy, speed and computational cost. Codes are available at: https://github.com/Biotan/3DCFS.
Liang Du 0004, Jingang Tan, Xiangyang Xue 0001, Hongkai Wen 0001, Jianfeng Feng, Jiamao Li
ICRA1
2019 SSF-DAN: Separated Semantic Feature Based Domain Adaptation Network for Semantic Segmentation
abstract
Despite the great success achieved by supervised fully convolutional models in semantic segmentation, training the models requires a large amount of labor-intensive work to generate pixel-level annotations. Recent works exploit synthetic data to train the model for semantic segmentation, but the domain adaptation between real and synthetic images remains a challenging problem. In this work, we propose a Separated Semantic Feature based domain adaptation network, named SSF-DAN, for semantic segmentation. First, a Semantic-wise Separable Discriminator (SS-D) is designed to independently adapt semantic features across the target and source domains, which addresses the inconsistent adaptation issue in the class-wise adversarial learning. In SS-D, a progressive confidence strategy is included to achieve a more reliable separation. Then, an efficient Class-wise Adversarial loss Reweighting module (CA-R) is introduced to balance the class-wise adversarial learning process, which leads the generator to focus more on poorly adapted classes. The presented framework demonstrates robust performance, superior to state-of-the-art methods on benchmark datasets.
Liang Du 0004, Jingang Tan, Hongye Yang, Jianfeng Feng, Xiangyang Xue 0001, Qibao Zheng, Xiaoqing Ye
ICCV1
2018 3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation
Xiaoqing Ye, Jiamao Li, Hexiao Huang, Liang Du 0004
ECCV (7)4