Pengpeng Liang

dblp:132/4754 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2026 Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection
abstract
Unsupervised domain adaptation for LiDAR-based 3D object detection (3D UDA) based on the teacher-student architecture with pseudo labels has achieved notable improvements in recent years. Although it is quite popular to collect point clouds and images simultaneously, little attention has been paid to the usefulness of image data in 3D UDA when training the models. In this paper, we propose an approach named MMAssist that improves the performance of 3D UDA with multi-modal assistance. A method is designed to align 3D features between the source domain and the target domain by using image and text features as bridges. More specifically, we project the ground truth labels or pseudo labels to the images to get a set of 2D bounding boxes. For each 2D box, we extract its image feature from a pre-trained vision backbone. A large vision-language model (LVLM) is adopted to extract the box's text description, and a pre-trained text encoder is used to obtain its text feature. During the training of the model in the source domain and the student model in the target domain, we align the 3D features of the predicted boxes with their corresponding image and text features, and the 3D features and the aligned features are fused with learned weights for the final prediction. The features between the student branch and the teacher branch in the target domain are aligned as well. To enhance the pseudo labels, we use an off-the-shelf 2D object detector to generate 2D bounding boxes from images and estimate their corresponding 3D boxes with the aid of point cloud, and these 3D boxes are combined with the pseudo labels generated by the teacher model. Experimental results show that our approach achieves promising performance compared with state-of-the-art methods in three domain adaptation tasks on three popular 3D object detection datasets.
Shenao Zhao, Pengpeng Liang, Zhoufan Yang
AAAI2
2025 CurveFormer++: 3D Lane Detection by Curve Propagation With Temporal Curve Queries and Attention
abstract
In autonomous driving, accurate 3D lane detection using monocular cameras is important for downstream tasks. Recent CNN and Transformer approaches usually apply a two-stage model design. The first stage transforms the image feature from a front image into a bird’s-eye-view (BEV) representation. Subsequently, a sub-network processes the BEV feature to generate the 3D detection results. However, these approaches heavily rely on a challenging image feature transformation module from a perspective view to a BEV representation. In our work, we present CurveFormer++, a single-stage Transformer-based method that does not require the view transform module and directly infers 3D lane results from the perspective image features. Specifically, our approach models the 3D lane detection task as a curve propagation problem, where each lane is represented by a curve query with a dynamic and ordered anchor point set. By employing a Transformer decoder, the model can iteratively refine the 3D lane results. A curve cross-attention module is introduced to calculate similarities between image features and curve queries. To handle varying lane lengths, we employ context sampling and anchor point restriction techniques to compute more relevant image features. Furthermore, we apply a temporal fusion module that incorporates selected informative sparse curve queries and their corresponding anchor point sets to leverage historical information. In the experiments, we evaluate our approach on two publicly real-world datasets. The results demonstrate that our method provides outstanding performance compared with both CNN and Transformer based methods. We also conduct ablation studies to analyze the impact of each component.
Yifeng Bai, Zhirong Chen, Pengpeng Liang, Erkang Cheng
IEEE Trans. Intell. Transp. Syst.3
2024 Enhancing 3D Object Detection with 2D Detection-Guided Query Anchors
abstract
Multi-camera-based 3D object detection has made no-table progress in the past several years. However, we observe that there are cases (e.g. faraway regions) in which popular 2D object detectors are more reliable than state-of-the-art 3D detectors. In this paper, to improve the performance of query-based 3D object detectors, we present a novel query generating approach termed QAF2D, which infers 3D query anchors from 2D detection results. A 2D bounding box of an object in an image is lifted to a set of 3D anchors by associating each sampled point within the box with depth, yaw angle, and size candidates. Then, the validity of each 3D anchor is verified by comparing its projection in the image with its corresponding 2D box, and only valid anchors are kept and used to construct queries. The class information of the 2D bounding box associated with each query is also utilized to match the predicted boxes with ground truth for the set-based loss. The image feature extraction backbone is shared between the 3D detector and 2D detector by adding a small number of prompt parameters. We integrate QAF2D into three popular query-based 3D object detectors and carry out comprehensive evaluations on the nuScenes dataset. The largest improvement that QAF2D can bring about on the nuScenes validation subset is 2.3% NDS and 2.7% mAP. Code is available at https://github.com/nullmax-vision/QAF2D.
Haoxuanye Ji, Pengpeng Liang, Erkang Cheng
CVPR2
2024 A Broad Sparse Fine-Grained Image Classification Model Based on Dictionary Selection Strategy
abstract
When the fine-grained recognition problems of image classification processed, broad learning system (BLS) is more efficient in classification, but has difficulty in distinguishing features with large similarities. Sparse representation classification (SRC) is more capable of handling similarity features, but is more computationally expensive. To better use the BLS model to tackle the fine-grained recognition problem and improve the ability to handle similarity features, this article combines the advantages of BLS and SRC, and proposes a broad sparse fine-grained image classification model based on dictionary selection strategy, dictionary broad sparse representation classification (DBSRC). First, to solve the parameter selection problem of the BLS model, leave one out cross validation (LOO) is introduced to quickly find the better regularization parameters and build a BLS-LOO coarse-grained classification model. Then propose reliability criteria based on a threshold selection strategy for the selection of fine-grained images. Next, an adaptive dictionary selection strategy is designed based on the output of the BLS-LOO to construct a sparse subdictionary for each fine-grained image that is not distinguished by the BLS-LOO. Finally, a sparse subdictionary based SRC model is used to classify fine-grained images. Experimental results show that DBSRC achieves good classification performance on three image datasets with different complex dimensions, ImageNet, USPS, and Pavia, and has strong processing capability for fine-grained features.
Jianjie Zheng, Pengpeng Liang, Huimin Zhao 0002, Wu Deng 0001
IEEE Trans. Reliab.2
2023 CurveFormer: 3D Lane Detection by Curve Propagation with Curve Queries and Attention
abstract
3D lane detection is an integral part of au-tonomous driving systems. Previous CNN and Transformer-based methods usually first generate a bird's-eye-view (BEV) feature map from the front view image, and then use a sub-network with BEV feature map as input to predict 3D lanes. Such approaches require an explicit view transformation between BEV and front view, which itself is still a challenging problem. In this paper, we propose CurveFormer, a single-stage Transformer-based method that directly calculates 3D lane pa-rameters and can circumvent the difficult view transformation step. Specifically, we formulate 3D lane detection as a curve propagation problem by using curve queries. A 3D lane query is represented by a dynamic and ordered anchor point set. In this way, queries with curve representation in Transformer decoder iteratively refine the 3D lane detection results. Moreover, a curve cross-attention module is introduced to compute the similarities between curve queries and image features. Additionally, a context sampling module that can capture more relative image features of a curve query is provided to further boost the 3D lane detection performance. We evaluate our method for 3D lane detection on both synthetic and real-world datasets, and the experimental results show that our method achieves promising performance compared with the state-of-the-art approaches. The effectiveness of each component is validated via ablation studies as well.
Yifeng Bai, Zhirong Chen, Zhangjie Fu 0002, Lang Peng, Pengpeng Liang, Erkang Cheng
ICRA5
2023 CircleFormer: Circular Nuclei Detection in Whole Slide Images with Circle Queries and Attention
Hengxu Zhang, Pengpeng Liang, Zhiyong Sun 0002, Erkang Cheng
MICCAI (8)2
2023 BEVSegFormer: Bird's Eye View Semantic Segmentation From Arbitrary Camera Rigs
abstract
Semantic segmentation in bird's eye view (BEV) is an important task for autonomous driving. Though this task has attracted a large amount of research efforts, it is still challenging to flexibly cope with arbitrary (single or multiple) camera sensors equipped on the autonomous vehicle. In this paper, we present BEVSegFormer, an effective transformer-based method for BEV semantic segmentation from arbitrary camera rigs. Specifically, our method first encodes image features from arbitrary cameras with a shared backbone. These image features are then enhanced by a deformable transformer-based encoder. Moreover, we introduce a BEV transformer decoder module to parse BEV semantic segmentation results. An efficient multi-camera deformable attention unit is designed to carry out the BEV-to-image view transformation. Finally, the queries are reshaped according to the layout of grids in the BEV, and upsampled to produce the semantic segmentation result in a supervised manner. We evaluate the proposed algorithm on the public nuScenes dataset and a self-collected dataset. Experimental results show that our method achieves promising performance on BEV semantic segmentation from arbitrary camera rigs. We also demonstrate the effectiveness of each component via ablation study.
Lang Peng, Zhirong Chen, Zhangjie Fu 0002, Pengpeng Liang, Erkang Cheng
WACV4
2022 Traffic Context Aware Data Augmentation for Rare Object Detection in Autonomous Driving
abstract
Detection of rare objects (e.g., traffic cones, traffic barrels and traffic warning triangles) is an important perception task to improve the safety of autonomous driving. Training of such models typically requires a large number of annotated data which is expensive and time consuming to obtain. To address the above problem, an emerging approach is to apply data augmentation to automatically generate cost-free training samples. In this work, we propose a systematic study on simple Copy-Paste data augmentation for rare object detection in autonomous driving. Specifically, local adaptive instance-level image transformation is introduced to generate realistic rare object masks from source domain to the target domain. Moreover, traffic scene context is utilized to guide the placement of masks of rare objects. To this end, our data augmentation generates training data with high quality and realistic characteristics by leveraging both local and global consistency. In addition, we build a new dataset named NM10k consisting 10k training images, 4k validation images and the corresponding labels with a diverse range of scenarios in autonomous driving. Experiments on NM10k show that our method achieves promising results on rare object detection. We also present a thorough study to illustrate the effectiveness of our local-adaptive and global constraints based Copy-Paste data augmentation for rare object detection. The data, development kit and more information of NM10k dataset are available online at: https://nullmax-vision.github.io.
Naifan Li, Pengpeng Liang, Erkang Cheng
ICRA4
2022 Pseudo Segmentation for Semantic Information-Aware Stereo Matching
abstract
Stereo matching plays an important role in computer vision and robotics. Though substantial progress has been made on deep learning-based algorithms, the inherent semantic information within the ground truth of the training data for stereo matching has not been well explored. In this letter, we propose to use a pseudo segmentation sub-network to extract additional semantic information. More specifically, we divide the disparity label into groups and let each group correspond to a class for pseudo segmentation. To assist stereo matching with the semantic information obtained from pseudo segmentation, we inject the feature maps at the end of the pseudo segmentation sub-network into the cost volume that is used to infer the pixel-level disparity. To validate the effectiveness of the proposed approach, we select PSMNet (Chang and Chen, 2018)and GwcNet (Guoet al., 2019) as baselines and enhance them with the pseudo segmentation sub-network. Comprehensive experiments are carried out on the Scene Flow, KITTI 2015, and KITTI 2012 datasets, and the results show that our proposed method can improve the performance notably.
Shengyou Hua, Zhiyong Sun 0002, Pengpeng Liang, Erkang Cheng
IEEE Signal Process. Lett.4
2022 A Simple and Strong Baseline for Universal Targeted Attacks on Siamese Visual Tracking
abstract
Siamese trackers are shown to be vulnerable to adversarial attacks recently. However, the existing attack methods craft the perturbations for each video independently, which comes at a non-negligible computational cost. In this paper, we show the existence of universal perturbations that can enable the targeted attack, e.g., forcing a tracker to follow the ground-truth trajectory with specified offsets, to be video-agnostic and free from inference in a network. Specifically, we attack a tracker by adding a universal translucent perturbation to the template image and adding afake target, i.e., a small universal adversarial patch, into the search images adhering to the predefined trajectory, so that the tracker outputs the location and size of thefake targetinstead of the real target. Our approach allows perturbing a novel video to come at no additional cost except the mere addition operations – and not require gradient optimization or network inference. Experimental results on several datasets demonstrate that our approach can effectively fool the Siamese trackers in a targeted attack manner. We show that the proposed perturbations are not only universal across videos, but also generalize well across different trackers. Such perturbations are therefore doubly universal, both with respect to the data and the network architectures. Our code is available athttps://github.com/lizhenbang56/SiamAttack.
Zhenbang Li, Yaya Shi, Shaoru Wang, Bing Li 0001, Pengpeng Liang, Weiming Hu 0004
IEEE Trans. Circuits Syst. Video Technol.6
2021 Coarse-to-fine Semantic Localization with HD Map for Autonomous Driving in Structural Scenes
abstract
Robust and accurate localization is an essential component for robotic navigation and autonomous driving. The use of cameras for localization with high definition map (HD Map) provides an affordable localization sensor set. Existing methods suffer from pose estimation failure due to error prone data association or initialization with accurate initial pose requirement. In this paper, we propose a cost-effective vehicle localization system with HD map for autonomous driving that uses cameras as primary sensors. To this end, we formulate vision-based localization as a data association problem that maps visual semantics to landmarks in HD map. Specifically, system initialization is finished in a coarse to fine manner by combining coarse GPS (Global Positioning System) measurement and fine pose searching. In tracking stage, vehicle pose is refined by implicitly aligning the semantic segmentation result between image and landmarks in HD maps with photometric consistency. Finally, vehicle pose is computed by pose graph optimization in a sliding window fashion. We evaluate our method on two datasets and demonstrate that the proposed approach yields promising localization results in different driving scenarios. Additionally, our approach is suitable for both monocular camera and multi-cameras that provides flexibility and improves robustness for the localization system.
Minjie Lin, Heyang Guo, Pengpeng Liang, Erkang Cheng
IROS4
2021 Joint Spinal Centerline Extraction and Curvature Estimation with Row-Wise Classification and Curve Graph Network
Long Huo, Bin Cai 0006, Pengpeng Liang, Zhiyong Sun 0002, Chi Xiong, Chaoshi Niu, Erkang Cheng
MICCAI (5)3
2021 Learning local descriptors with multi-level feature aggregation and spatial context pyramid
Pengpeng Liang, Haoxuanye Ji, Erkang Cheng, Yumei Chai, Haibin Ling
Neurocomputing1
2021 Planar object tracking benchmark in the wild
Pengpeng Liang, Haoxuanye Ji, Yumei Chai, Chunyuan Liao, Haibin Ling
Neurocomputing1
2018 Planar Object Tracking in the Wild: A Benchmark
abstract
Planar object tracking is an actively studied problem in vision-based robotic applications. While several benchmarks have been constructed for evaluating state-of-the-art algorithms, there is a lack of video sequences captured in the wild rather than in constrained laboratory environment. In this paper, we present a carefully designed planar object tracking benchmark containing 210 videos of 30 planar objects sampled in the natural environment. In particular, for each object, we shoot seven videos involving various challenging factors, namely scale change, rotation, perspective distortion, motion blur, occlusion, out-of-view, and unconstrained. The ground truth is carefully annotated semi-manually to ensure the quality. Moreover, eleven state-of-the-art algorithms are evaluated on the benchmark using two evaluation metrics, with detailed analysis provided for the evaluation results. We expect the proposed benchmark to benefit future studies on planar object tracking.
Pengpeng Liang, Hu Lu, Chunyuan Liao, Haibin Ling
ICRA1
2016 Adaptive Objectness for Object Tracking
abstract
To exploit the reliable prior knowledge that the target object in tracking must be an object other than nonobject, in this letter, we propose to adapt objectness for visual object tracking. Instead of directly applying an existing objectness measure that is generic and handles various objects and environments, we adapt it to be compatible to the specific tracking sequence and object. More specifically, we use the newly proposed binarized normed gradient (BING) objectness as the base, and then train an object-adaptive objectness for each tracking task. The training is implemented by using an adaptive support vector machine that integrates information from the specific tracking target into the BING measure. We emphasize that the benefit of the proposed adaptive objectness, named ADOBING, is generic. To show this, we combine ADOBING with eight top performed trackers in recent evaluations. We run the ADOBING-enhanced trackers along with their base trackers on the CVPR2013 benchmark, and our methods consistently improve the base trackers both in overall performance and under all challenge factors. Noting that the way we integrate objectness in visual tracking is generic and straightforward, we expect even more improvement by using tracker-specific objectness.
Pengpeng Liang, Chunyuan Liao, Xue Mei, Haibin Ling
IEEE Signal Process. Lett.1
2015 Encoding Color Information for Visual Tracking: Algorithms and Benchmark
abstract
While color information is known to provide rich discriminative clues for visual inference, most modern visual trackers limit themselves to the grayscale realm. Despite recent efforts to integrate color in tracking, there is a lack of comprehensive understanding of the role color information can play. In this paper, we attack this problem by conducting a systematic study from both the algorithm and benchmark perspectives. On the algorithm side, we comprehensively encode 10 chromatic models into 16 carefully selected state-of-the-art visual trackers. On the benchmark side, we compile a large set of 128 color sequences with ground truth and challenge factor annotations (e.g., occlusion). A thorough evaluation is conducted by running all the color-encoded trackers, together with two recently proposed color trackers. A further validation is conducted on an RGBD tracking benchmark. The results clearly show the benefit of encoding color information for tracking. We also perform detailed analysis on several issues, including the behavior of various combinations between color model and visual tracker, the degree of difficulty of each sequence for tracking, and how different challenge factors affect the tracking performance. We expect the study to provide the guidance, motivation, and benchmark for future work on encoding color in visual tracking.
Pengpeng Liang, Erik Blasch, Haibin Ling
IEEE Trans. Image Process.1
2014 Blur-Resilient Tracking Using Group Sparsity
Pengpeng Liang, Yi Wu 0001, Xue Mei, Jingyi Yu 0001, Erik Blasch, Danil V. Prokhorov, Chunyuan Liao, Haitao Lang, Haibin Ling
ACCV (5)1
2013 Vehicle detection in wide area aerial surveillance using Temporal Context
Pengpeng Liang, Haibin Ling, Erik Blasch, Guna Seetharaman, Dan Shen 0004, Genshe Chen
FUSION1
2012 Multiple Kernel Learning for vehicle detection in wide area motion imagery
Pengpeng Liang, Gregory Teodoro, Haibin Ling, Erik Blasch, Genshe Chen, Li Bai 0002
FUSION1