EDBT 2026 Demo / reviewers in the wild / expert
Liangji Fang
dblp:217/1931
· DBLP profile ↗
17ranked-venue papers
3as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Graph-DETR4D: Spatio-Temporal Graph Modeling for Multi-View 3D Object DetectionabstractMulti-View 3D object detection (MV3D) has made tremendous progress by leveraging multiple perspective features through surrounding cameras. Despite demonstrating promising prospects in various applications, accurately detecting objects through camera view in the 3D space is extremely difficult due to the ill-posed issue in monocular depth estimation. Recently, Graph-DETR3D presents a novel graph-based 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves competitive performance. Although it enriches the query representations with 2D image features through a learnable 3D graph, it still suffers from limited depth and velocity estimation abilities due to the adoption of a single-frame input setting. To solve this problem, we introduce a unified spatial-temporal graph modeling framework to fully leverage the multi-view imagery cues under the multi-frame inputs setting. Thanks to the flexibility and sparsity of the dynamic graph architecture, we lift the original 3D graph into the 4D space with an effective attention mechanism to automatically perceive imagery information at both spatial and temporal levels. Moreover, considering the main latency bottleneck lies in the image backbone, we propose a novel dense-sparse distillation framework for multi-view 3D object detection, to reduce the computational budget while sacrificing no detection accuracy, making it more suitable for real-world deployment. To this end, we propose Graph-DETR4D, a faster and stronger multi-view 3D object detection framework, built on top of Graph-DETR3D. Extensive experiments on nuScenes and Waymo benchmarks demonstrate the effectiveness and efficiency of Graph-DETR4D. Notably, our best model achieves 62.0% NDS on nuScenes test leaderboard. Code is available at https://github.com/zehuichen123/Graph-DETR4D. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Wu 0005, Feng Zhao 0004 |
IEEE Trans. Image Process. | 5 |
| 2023 | BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ICLR | 4 |
| 2022 | SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual RepresentationsabstractPre-training has become a standard paradigm in many computer vision tasks. However, most of the methods are generally designed on the RGB image domain. Due to the discrepancy between the two-dimensional image plane and the three-dimensional space, such pre-trained models fail to perceive spatial information and serve as sub-optimal solutions for 3D-related tasks. To bridge this gap, we aim to learn a spatial-aware visual representation that can describe the three-dimensional space and is more suitable and effective for these tasks. To leverage point clouds, which are much more superior in providing spatial information compared to images, we propose a simple yet effective 2D Image and 3D Point cloud Unsupervised pre-training strategy, called SimIPU. Specifically, we develop a multi-modal contrastive learning framework that consists of an intra-modal spatial perception module to learn a spatial-aware representation from point clouds and an inter-modal feature interaction module to transfer the capability of perceiving spatial information from the point cloud encoder to the image encoder, respectively. Positive pairs for contrastive losses are established by the matching algorithm and the projection matrix. The whole framework is trained in an unsupervised end-to-end fashion. To the best of our knowledge, this is the first study to explore contrastive learning pre-training strategies for outdoor multi-modal datasets, containing paired camera images and LIDAR point clouds. Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang, Bolei Zhou, Hang Zhao 0021 |
AAAI | 4 |
| 2022 | Deformable Feature Aggregation for Dynamic Multi-modal 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ECCV (8) | 4 |
| 2022 | Unsupervised Domain Adaptation for Monocular 3D Object Detection via Self-training
Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang |
ECCV (9) | 4 |
| 2022 | AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object DetectionabstractObject detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sources complementary and beneficial to each other. In this paper, we propose AutoAlign, an automatic feature fusion strategy for 3D object detection. Instead of establishing deterministic correspondence with camera projection matrix, we model the mapping relationship between the image and point clouds with a learnable alignment map. This map enables our model to automate the alignment of non-homogenous features in a dynamic and data-driven manner. Specifically, a cross-attention feature alignment module is devised to adaptively aggregate pixel-level image features for each voxel. To enhance the semantic consistency during feature alignment, we also design a self-supervised cross-modal feature interaction module, through which the model can learn feature aggregation with instance-level feature guidance. Extensive experimental results show that our approach can lead to 2.3 mAP and 7.0 mAP improvements on the KITTI and nuScenes datasets respectively. Notably, our best model reaches 70.9 NDS on the nuScenes testing leaderboard, achieving competitive performance among various state-of-the-arts. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004, Bolei Zhou, Hang Zhao 0021 |
IJCAI | 4 |
| 2022 | Graph-DETR3D: Rethinking Overlapping Regions for Multi-View 3D Object Detectionabstract3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. However, accurately detecting objects through perspective views in the 3D space is extremely difficult due to the lack of depth information. Recently, DETR3D introduces a novel 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves state-of-the-art performance. In this paper, with intensive pilot experiments, we quantify the objects located at different regions and find that the "truncated instances'' (i.e., at the border regions of each image) are the main bottleneck hindering the performance of DETR3D. Although it merges multiple features from two adjacent views in the overlapping regions, DETR3D still suffers from insufficient feature aggregation, thus missing the chance to fully boost the detection performance. In an effort to tackle the problem, we propose Graph-DETR3D to automatically aggregate multi-view imagery information through graph structure learning. It constructs a dynamic 3D graph between each object query and 2D feature maps to enhance the object representations, especially at the border regions. Besides, Graph-DETR3D benefits from a novel depth-invariant multi-scale training strategy, which maintains the visual depth consistency by simultaneously scaling the image size and the object depth. Extensive experiments on the nuScenes dataset demonstrate the effectiveness and efficiency of our Graph-DETR3D. Notably, our best model achieves 49.5 NDS on the nuScenes test leaderboard, achieving new state-of-the-art in comparison with various published image-view 3D object detectors. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ACM Multimedia | 4 |
| 2022 | Balanced Gradient Penalty Improves Deep Long-Tailed LearningabstractIn recent years, deep learning has achieved a great success in various image recognition tasks. However, the long-tailed setting over a semantic class plays a leading role in real-world applications. Common methods focus on optimization on balanced distribution or naive models. Few works explore long-tailed learning from a deep learning-based generalization perspective. The loss landscape on long-tailed learning is first investigated in this work. Empirical results show that sharpness-aware optimizers work not well on long-tailed learning. Because they do not take class priors into consideration, and they fail to improve performance of few-shot classes. To better guide the network and explicitly alleviate sharpness without extra computational burden, we develop a universal Balanced Gradient Penalty (BGP) method. Surprisingly, our BGP method does not need the detailed class priors and preserves privacy. Our new algorithm BGP, as a regularization loss, can achieve the state-of-the-art results on various image datasets (i.e., CIFAR-LT, ImageNet-LT and iNaturalist-2018) in the settings of different imbalance ratios. Dong Wang 0004, Liangji Fang, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
ACM Multimedia | 3 |
| 2021 | Multimodal Motion Prediction With Stacked TransformersabstractPredicting multiple plausible future trajectories of the nearby vehicles is crucial for the safety of autonomous driving. Recent motion prediction approaches attempt to achieve such multimodal motion prediction by implicitly regularizing the feature or explicitly generating multiple candidate proposals. However, it remains challenging since the latent features may concentrate on the most frequent mode of the data while the proposal-based methods depend largely on the prior knowledge to generate and select the proposals. In this work, we propose a novel transformer framework for multimodal motion prediction, termed as mmTransformer. A novel network architecture based on stacked transformers is designed to model the multimodality at feature level with a set of fixed independent proposals. A region-based training strategy is then developed to induce the multimodality of the generated proposals. Experiments on Argoverse dataset show that the proposed model achieves the state-of-the-art performance on motion prediction, substantially improving the diversity and the accuracy of the predicted trajectories. Demo video and code are available at https://decisionforce.github.io/mmTransformer. Jinghuai Zhang, Liangji Fang, Qinhong Jiang, Bolei Zhou |
CVPR | 3 |
| 2021 | CAT: Corner Aided Tracking With Deep Regression NetworkabstractSingle object tracking in visual media is an important yet challenging task. Various challenges, especially target scale variation, shape deformation and occlusion, can have large effects on the performances of trackers. Current deep regression based trackers only pay close attention to regression on the center key point of the tracking target, meanwhile employ the image pyramid based multi-scale testing method to deal with scale estimation. Such procedure can not properly handle the three challenges. We address these challenges in a principled way by the aid of auxiliary regressions on the four bounding box corners of the tracking target. In this work, we propose the novel Corner Aided Tracker with deep regression network, abbreviated as CAT. Different from RPN-based trackers, in CAT, four corners along with the center key point of the bounding box for tracking target are simultaneously obtained by five corresponding response maps. Furthermore, to robustly and accurately generate tight bounding boxes for the tracking target and collect reliable samples for online training of the network, we propose an adaptive key point selection method to select the subset of reliable key points and drop the unreliable ones, based on the qualities of their corresponding response maps as well as the constraints from shape, scale and location. We demonstrate that the regressed corners can help naturally locate the tracking target with tight bounding boxes. The challenges of scale variation, shape deformation and occlusion can be handled explicitly. The commonly used time-consuming image pyramid based multi-scale testing method can also be discarded. Extensive experiments on OTB2013, OTB2015, UAV123, LaSOT, VOT2016 and VOT2018 datasets are conducted to report new state-of-the-art performances and demonstrate the effectiveness of CAT. Shiquan Zhang, Xu Zhao 0001, Liangji Fang |
IEEE Trans. Multim. | 3 |
| 2020 | TPNet: Trajectory Proposal Network for Motion PredictionabstractMaking accurate motion prediction of the surrounding traffic agents such as pedestrians, vehicles, and cyclists is crucial for autonomous driving. Recent data-driven motion prediction methods have attempted to learn to directly regress the exact future position or its distribution from massive amount of trajectory data. However, it remains difficult for these methods to provide multimodal predictions as well as integrate physical constraints such as traffic rules and movable areas. In this work we propose a novel two-stage motion prediction framework, Trajectory Proposal Network (TPNet). TPNet first generates a candidate set of future trajectories as hypothesis proposals, then makes the final predictions by classifying and refining the proposals which meets the physical constraints. By steering the proposal generation process, safe and multimodal predictions are realized. Thus this framework effectively mitigates the complexity of motion prediction problem while ensuring the multimodal output. Experiments on four large-scale trajectory prediction datasets, i.e. the ETH, UCY, Apollo and Argoverse datasets, show that TPNet achieves the state-of-the-art results both quantitatively and qualitatively. Liangji Fang, Qinhong Jiang, Jianping Shi, Bolei Zhou |
CVPR | 1 |
| 2020 | EdgeStereo: An Effective Multi-task Learning Network for Stereo Matching and Edge Detection
Xiao Song 0002, Xu Zhao 0001, Liangji Fang, Hanwen Hu, Yizhou Yu |
Int. J. Comput. Vis. | 3 |
| 2019 | Small-objectness sensitive detection based on shifted single shot detector
Liangji Fang, Xu Zhao 0001, Shiquan Zhang |
Multim. Tools Appl. | 1 |
| 2019 | Discriminative representation combinations for accurate face spoofing detection
Xiao Song 0002, Xu Zhao 0001, Liangji Fang |
Pattern Recognit. | 3 |
| 2018 | Putting the Anchors Efficiently: Geometric Constrained Pedestrian Detection
Liangji Fang, Xu Zhao 0001, Xiao Song 0002, Shiquan Zhang, Ming Yang 0002 |
ACCV (5) | 1 |
| 2018 | EdgeStereo: A Context Integrated Residual Pyramid Network for Stereo Matching
Xiao Song 0002, Xu Zhao 0001, Hanwen Hu, Liangji Fang |
ACCV (5) | 4 |
| 2018 | Led: Localization-Quality Estimation Embedded DetectorabstractClassification subnetwork and box regression subnetwork are essential components in deep networks for object detection. However, we observe a contradiction that before NMS, some better localized detections do not correspond to higher classification confidences, and vice versa. This contradiction exists because classification confidences can not fully reflect the localization-quality (loc-quality) of each detection. In this work, we propose the Localization-quality Estimation embedded Detector abbreviated as LED, and a corresponding detection pipeline. In this detection pipeline, we first propose an accurate loc-quality estimation method for each detection, then combine the loc-quality with the corresponding classification confidence during inference to make each detection more reasonable and accurate. For efficiency, LED is designed as an one-stage network. Extensive experiments are conducted on Pascal VOC 2007 and KITTI car detection datasets to demonstrate the effectiveness of LED. Shiquan Zhang, Xu Zhao 0001, Liangji Fang, Haiping Fei |
ICIP | 3 |