Peisheng Qian

dblp:264/5492 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-9763-547XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 AIC3DOD: Advancing Indoor Class-Incremental 3D Object Detection with Point Transformer Architecture and Room Layout Constraints
abstract
Over the recent years, there has been a growing interest in class-incremental 3D object detection based on point clouds. However, the current state-of-the-art (SOTA) methods still fall short of practical adoption, mainly due to two key observations. Firstly, existing SOTA methods are limited by the capability of feature representation from the object detection model. Secondly, these methods overlook the importance of incorporating prior information or geometry constraints, which are crucial elements for 3D point cloud tasks. In this study, we strive to enhance the performance of class-incremental 3D object detection for indoor scenes by proposing AIC3DOD - Advancing Indoor Classincremental 3D Object Detection using the point transformer architecture with room layout constraints. Our approach employs a transformer architecture in our detection model and optimizes the class incremental step in the transformer architecture. Besides, AIC3DOD incorporates additional prior information, namely room layout, to impose physical constraints on detected objects, thereby enhancing overall object detection performance. Extensive experimental results on the ScanNet dataset demonstrate the effectiveness of our approach, showcasing our superior performance compared to other SOTA methods in the class-incremental 3D object detection task.
Zhongyao Cheng, Fang Wu 0009, Peisheng Qian, Ziyuan Zhao, Xulei Yang
WACV3
2025 SDCoT++: Improved Static-Dynamic Co-Teaching for Class-Incremental 3D Object Detection
abstract
Deep learning approaches have demonstrated high effectiveness in 3D object detection tasks. However, they often suffer from a notable drop in performance on the previously trained classes when learning new classes incrementally without revisiting the old data. This is the "catastrophic forgetting" phenomenon which impedes 3D object detection in real-world scenarios, where intelligent machines must continuously learn to detect previously unseen categories. Furthermore, frequent co-occurrences of old and new classes in scenes exacerbate catastrophic forgetting and cause model confusion. To address these challenges, we propose a novel static-dynamic co-teaching approach. Our framework involves a student model and two teacher models: a static teacher with fixed weights which imparts preserved old knowledge to the student, and a dynamic teacher with continuously updated weights which transfers underlying knowledge from new data to the student. To mitigate the issue of co-occurrence, we generate pseudo labels for base (i.e. old) classes from both static and dynamic sources during incremental learning. Additionally, to mitigate the negative impact of varying occurrence frequencies of classes on fixed thresholding during the selection of pseudo labels, we calibrate the probabilities of base classes to attain more balanced class probabilities. Moreover, our static-dynamic co-teaching framework is backbone-agnostic, making it compatible with different detection architectures. We demonstrate its backbone-agnostic nature by adapting three representative 3D object detectors: VoteNet, 3DETR and CAGroup3D. Extensive experiments showcase the superior performance of our proposed method compared to baseline approaches across indoor and outdoor benchmark datasets and applicability with different backbone models.
Na Zhao 0004, Peisheng Qian, Fang Wu 0009, Xun Xu 0002, Xulei Yang, Gim Hee Lee
IEEE Trans. Image Process.2
2023 COCO-TEACH: A Contrastive Co-Teaching Network For Incremental 3D Object Detection
abstract
Deep learning (DL) models for 3D object detection from point clouds have shown remarkable progress in various autonomous perception scenarios. However, the issue of catastrophic forgetting seriously hinders the deployment of these models in real-world applications where new classes are encountered over time. In order to address this issue, we present the Contrastive Co-Teaching Network (COCO-TEACH) framework for class-incremental 3D object detection. Our proposed framework consists of two teacher networks: a primary teacher network that detects old class objects in new data and provides them with pseudo-labels and an auxiliary teacher network that leverages the unlabelled objects in new data. The two teacher models transfer their learned knowledge to the target student model through a class-aware consistency loss. To enhance this transfer, a supervised contrastive loss is further incorporated into the loss function. We evaluate the performance of our proposed method against baseline methods through extensive experiments on two benchmark datasets. The results show that our proposed framework achieves state-of-the-art performance on incremental 3D object detection.
Zhongyao Cheng, Cen Chen 0002, Ziyuan Zhao, Peisheng Qian, Xiaoli Li 0001, Xulei Yang
ICIP4
2023 SemiGNN-PPI: Self-Ensembling Multi-Graph Neural Network for Efficient and Generalizable Protein-Protein Interaction Prediction
abstract
Protein-protein interactions (PPIs) are crucial in various biological processes and their study has significant implications for drug development and disease diagnosis. Existing deep learning methods suffer from significant performance degradation under complex real-world scenarios due to various factors, e.g., label scarcity and domain shift. In this paper, we propose a self-ensembling multi-graph neural network (SemiGNN-PPI) that can effectively predict PPIs while being both efficient and generalizable. In SemiGNN-PPI, we not only model the protein correlations but explore the label dependencies by constructing and processing multiple graphs from the perspectives of both features and labels in the graph learning process. We further marry GNN with Mean Teacher to effectively leverage unlabeled graph-structured PPI data for self-ensemble graph learning. We also design multiple graph consistency constraints to align the student and teacher graphs in the feature embedding space, enabling the student model to better learn from the teacher model by incorporating more relationships. Extensive experiments on PPI datasets of different scales with different evaluation settings demonstrate that SemiGNN-PPI outperforms state-of-the-art PPI prediction methods, particularly in challenging scenarios such as training with limited annotations and testing on unseen data.
Ziyuan Zhao, Peisheng Qian, Xulei Yang, Zeng Zeng, Cuntai Guan, Tam Wai Leong, Xiaoli Li 0001
IJCAI2
2022 DA-CIL: Towards Domain Adaptive Class-Incremental 3D Object Detection
Ziyuan Zhao, Mingxi Xu, Peisheng Qian, Ramanpreet Singh Pahwa, Richard Chang 0002
BMVC3
2022 Iterative Contrastive Learning for Single Image Raindrop Removal
abstract
Deep learning has achieved remarkable progress in computer vision and image analysis. However, raindrop removal from single image still remains challenging, due to a wide range of raindrop diversities and surface reflections. In this paper, we propose an iterative neural network with feedback strategy and contrastive learning for single image raindrop removal. First, we design an iterative feedback neural network to refine low-level representations with high-level information, i.e., the output of the previous iteration is used as input for the next iteration, together with the input image with raindrops. As a result, raindrops could be gradually removed through this feedback manner. Then, we deploy contrastive regularization to push the restored image from each iteration close to the clean images without raindrops, but away from rainy images with raindrops. Extensive experiments on two raindrop benchmark datasets demonstrate the effectiveness of the proposed approach in comparison with the state-of-the-art methods. The methodology in this work could be further extended to self-supervised contrastive learning to obtain robust feature representations with less labelled data.
Xulei Yang, Peisheng Qian, Li Wang 0057, Cen Chen 0001, Xiaoli Li 0001, Zeng Zeng
ICIP2
2022 MMGL: Multi-Scale Multi-View Global-Local Contrastive Learning for Semi-Supervised Cardiac Image Segmentation
abstract
With large-scale well-labeled datasets, deep learning has shown significant success in medical image segmentation. However, it is challenging to acquire abundant annotations in clinical practice due to extensive expertise requirements and costly labeling efforts. Recently, contrastive learning has shown a strong capacity for visual representation learning on unlabeled data, achieving impressive performance rivaling supervised learning in many domains. In this work, we propose a novel multi-scale multi-view global-local contrastive learning (MMGL) framework to thoroughly explore global and local features from different scales and views for robust contrastive learning performance, thereby improving segmentation performance with limited annotations. Extensive experiments on the MM-WHS dataset demonstrate the effectiveness of MMGL framework on semi-supervised cardiac image segmentation, outperforming the state-of-the-art contrastive learning methods by a large margin.
Ziyuan Zhao, Jinxuan Hu, Zeng Zeng, Xulei Yang, Peisheng Qian, Bharadwaj Veeravalli, Cuntai Guan
ICIP5
2022 Adaptive Mean-Residue Loss for Robust Facial Age Estimation
abstract
Automated facial age estimation has diverse real-world applications in multimedia analysis, e.g., video surveillance, and human-computer interaction. However, due to the randomness and ambiguity of the aging process, age assessment is challenging. Most research work over the topic regards the task as one of age regression, classification, and ranking problems, and cannot well leverage age distribution in representing labels with age ambiguity. In this work, we propose a simple yet effective loss function for robust facial age estimation via distribution learning, i.e., adaptive mean-residue loss, in which, the mean loss penalizes the difference between the estimated age distribution's mean and the ground-truth age, whereas the residue loss penalizes the entropy of age probability out of dynamic top-K in the distribution. Experimental results in the datasets FG-NET and CLAP2016 have validated the effectiveness of the proposed loss.
Ziyuan Zhao, Peisheng Qian, Yubo Hou, Zeng Zeng
ICME2
2022 Robust Traffic Prediction From Spatial-Temporal Data Based on Conditional Distribution Learning
abstract
Traffic prediction based on massive speed data collected from traffic sensors plays an important role in traffic management. However, it is still challenging to obtain satisfactory performance due to the complex and dynamic spatial-temporal correlations among the data. Recently, many research works have demonstrated the effectiveness of graph neural networks (GNNs) for spatial-temporal modeling. However, such models are restricted by conditional distribution during training, and may not perform well when the target is outside the primary region of interest in the distribution. In this article, we address this problem with a stagewise learning mechanism, in which we redefine speed prediction as a conditional distribution learning followed by speed regression. We first perform a conditional distribution learning for each observed speed class, and then obtain speed prediction by optimizing regression learning, based on the learned conditional distribution. To effectively learn the conditional distribution, we introduce a mean-residue loss, consisting of two parts: 1) a mean loss, which penalizes the differences between the mean of the estimated conditional distribution and the ground truth and 2) a residue loss, which penalizes residue errors of the long tails in the distribution. To optimize the subsequent regression based on distribution information, we combine the mean absolute error (MAE) as another part of the loss function. We also incorporate a GNN-based architecture with our proposed learning mechanism. Mean-residue loss is employed to supervise the hidden speed representation in the network at each time interval, followed by a shared layer to recalibrate the hidden temporal dependencies in the conditional distribution. The experimental results based on three public traffic datasets have demonstrated that the effectiveness of the proposed method outperforms state-of-the-art methods.
Zeng Zeng, Wei Zhao 0035, Peisheng Qian, Yingjie Zhou 0001, Ziyuan Zhao, Cen Chen 0002, Cuntai Guan
IEEE Trans. Cybern.3