Sanqing Qu

dblp:232/3041 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
32since 2021 · last 2026
0009-0000-0377-7120ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Temporal Awareness for Novel Class Discovery in 3D Point Cloud Segmentation
Tianpei Zou, Sanqing Qu, Alois C. Knoll, Guang Chen 0001
IV4
2026 Dynamic MAsk-Pruning Strategy for Source-Free Model Intellectual Property Protection
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
Int. J. Comput. Vis.2
2026 I2EKD: Efficient and Versatile Image-to-Event Knowledge Distillation
abstract
Recently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%.
Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 An Online-Training-Free Adaptor for Open Heterogeneous Collaborative Perception via Diffusion Model
abstract
Collaborative perception seeks to mitigate the limitations of single-vehicle perception, such as occlusions, by facilitating communication and information sharing among connected vehicles. However, most existing works assume a homogeneous scenario where all vehicles share identity sensor types and perception model architectures. In contrast, real-world systems often involve heterogeneous agents with diverse sensor configurations and independently developed models. In such settings, directly exchanging features without proper alignment can significantly degrade performance and hinder effective collaboration. While some methods have been proposed to address heterogeneity, they typically require retraining or access to internal model parameters, making them impractical for scalable deployment. To address these challenges, we propose DiffAlign, a plug-and-play adapter that enables feature alignment across heterogeneous agents in a training-free and model-agnostic manner. DiffAlign treats received BEV features as noisy latent representations and progressively refines them through a pretrained diffusion process. This alignment strategy does not require access to model internals or any retraining, which makes it both scalable and privacy-preserving while supporting diverse sensor modalities and perception backbones. Extensive experiments on simulated OPV2V and real-world V2V4Real datasets demonstrate that DiffAlign consistently improves detection performance in heterogeneous settings, improving CoBEVT by 132.01% and 91.95%, respectively. Our method provides a practical path toward scalable, generalizable, and deployment-ready collaborative perception.
Tianhang Wang, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Hu Cao, Alois C. Knoll, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptions
abstract
Collaborative perception enhances single-vehicle perception by integrating sensory data from multiple connected vehicles. However, existing studies often assume ideal conditions, overlooking resilience to real-world challenges, such as adverse weather and sensor malfunctions, which is critical for safe deployment. To address this gap, we introduce RCP-Bench, the first comprehensive benchmark designed to evaluate the robustness of collaborative detection models under a wide range of real-world corruptions. RCP-Bench includes three new datasets (i.e., OPV2V-C, V2XSet-C, and DAIR-V2X-C) that simulate six collaborative cases and 14 types of camera corruption resulting from external environmental factors, sensor failures, and temporal misalignments. Extensive experiments on 10 leading collaborative perception models reveal that, while these models perform well under ideal conditions, they are significantly affected by corruptions. To improve robustness, we propose two simple yet effective strategies, RCP-Drop and RCP-Mix, based on training regularization and feature augmentation. Additionally, we identify several critical factors influencing robustness, such as backbone architecture, camera number, feature fusion methods, and the number of connected vehicles. We hope that RCP-Bench, along with these strategies and insights, will stimulate future research toward developing more robust collaborative perception models. Our benchmark toolkit is available at https://github.com/LuckyDush/RCP-Bench.
Shihang Du, Sanqing Qu, Tianhang Wang, Yunwei Zhu, Fan Lu 0001, Guang Chen 0001
CVPR2
2025 Event-Aided Progressive Neural Radiance Fields Reconstruction Under Challenging Illumination
abstract
Scene reconstruction and novel view synthesis are extensively utilized in domains such as autonomous driving, facilitating the creation of more comprehensive datasets and accelerating algorithm development. Current techniques have achieved promising results in image-based synthesis. Nonethe-less, they encounter difficulties including limited dynamic range, motion blur, and suboptimal performance under challenging illumination. Furthermore, these techniques depend significantly on precise camera poses, which are difficult to acquire in practical situations. Event cameras, characterized by their high dynamic range and temporal resolution, demonstrate advantages in challenging illumination and rapid motion scenarios and have been utilized in several autonomous driving applications. This study is the first to combine image and event camera for neural radiance field (NeRF) reconstruction in driving scenarios. The proposed method leverages the unique properties of event camera and adopts a progressive reconstruction strategy to jointly optimize camera poses during training, reducing reliance on pose precision and enabling more robust and accurate scene reconstruction under challenging illumination. Experimental results demonstrate that the proposed approach attains enhanced performance in novel view synthesis, exhibiting a 3.5% increase in PSNR and a 2% increase in SSIM on the DSEC dataset compared to baseline method.
Zongtao Bu, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Guang Chen 0001
IV4
2025 Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition
abstract
Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread availability of RGB cameras and the robustness of point clouds in providing accurate spatial geometry and distance information. However, current methods rely on intermediate modalities that capture either the vertical or horizontal field of view, limiting their ability to fully exploit the complementary information from both sensors. In this work, we propose an innovative initial retrieval + re-rank method that effectively combines information from range (or RGB) images and Bird's Eye View (BEV) images. Our approach relies solely on a computationally efficient global descriptor similarity search process to achieve re-ranking. Additionally, we introduce a novel similarity label supervision technique to maximize the utility of limited training data. Specifically, we employ points average distance to approximate appearance similarity and incorporate an adaptive margin, based on similarity differences, into the vanilla triplet loss. Experimental results on the KITTI dataset demonstrate that our method significantly outperforms state-of-the-art approaches.
Jianyi Peng, Fan Lu 0001, Bin Li 0087, Sanqing Qu, Guang Chen 0001
IV5
2025 OOD-Barrier: Build a Middle-Barrier for Open-Set Single-Image Test Time Adaptation via Vision Language Models
abstract
In real-world environments, a well-designed model must be capable of handling dynamically evolving distributions, where both in-distribution (ID) and out-of-distribution (OOD) samples appear unpredictably and individually, making real-time adaptation particularly challenging. While open-set test-time adaptation has demonstrated effectiveness in adjusting to distribution shifts, existing methods often rely on batch processing and struggle to manage single-sample data stream in open-set environments. To address this limitation, we propose Open-IRT, a novel open-set Intermediate-Representation-based Test-time adaptation framework tailored for single-image test-time adaptation with vision-language models. Open-IRT comprises two key modules designed for dynamic, single-sample adaptation in open-set scenarios. The first is Polarity-aware Prompt-based OOD Filter module, which fully constructs the ID-OOD distribution, considering both the absolute semantic alignment and relative semantic polarity. The second module, Intermediate Domain-based Test-time Adaptation module, constructs an intermediate domain and indirectly decomposes the ID-OOD distributional discrepancy to refine the separation boundary during the test-time. Extensive experiments on a range of domain adaptation benchmarks demonstrate the superiority of Open-IRT. Compared to previous state-of-the-art methods, it achieves significant improvements on representative benchmarks, such as CIFAR-100C and SVHN — with gains of +8.45\% in accuracy, -10.80\% in FPR95, and +11.04\% in AUROC.
Boyang Peng, Sanqing Qu, Tianpei Zou, Fan Lu 0001, Siheng Chen, Yong Wu 0007, Guang Chen 0001
NeurIPS2
2025 Multimodal LiDAR-Camera Novel View Synthesis with Unified Pose-free Neural Fields
abstract
Pose-free Neural Radiance Field (NeRF) aims at novel view synthesis (NVS) without relying on accurate poses, exhibiting significant practical value. Image and LiDAR point cloud are two pivotal modalities in autonomous driving scenarios. While demonstrating impressive performance, single-modality pose-free NeRFs often suffer from local optima due to the limited geometric information provided by dense image textures or the sparse, textureless nature of point clouds. Although prior methods have explored the complementary strengths of both modalities, they have only leveraged inherently sparse point clouds for discrete, non-pixel-wise depth supervision, and are limited to NVS of images. As a result, a Multimodal Unified Pose-free framework remains notably absent. In light of this, we propose MUP, a pose-free framework for LiDAR-Camera joint NVS in large-scale scenes. This unified framework enables continuous depth supervision for image reconstruction using LiDAR-Fields rather than discrete point clouds. By leveraging multimodal inputs, pose optimization receives gradients from the rendering loss of point cloud geometry and image texture, thereby alleviating the issue of local optima commonly encountered in single-modality pose-free tasks. Moreover, to further guide pose optimization of NeRF, we propose a multimodal geometric optimizer that leverages geometric relations from point clouds and photometric regularization from adjacent image frames. Besides, to alleviate the domain gap between modalities, we propose a multimodal-specific coarse-to-fine training approach for unified, compact reconstruction. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate MUP's superiority in accomplishing geometry-aware, modality-consistent, and pose-free 3D reconstruction.
Weiyi Xue, Fan Lu 0001, Yunwei Zhu, Zehan Zheng, Sanqing Qu, Jiangtong Li, Haiyun Wei, Guang Chen 0001
NeurIPS5
2025 CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning
abstract
Constrained hybrid-action reinforcement learning (RL) promises to learn a safe policy within a parameterized action space, which is particularly valuable for safety-critical applications involving discrete-continuous hybrid action spaces. However, existing hybrid-action RL algorithms primarily focus on reward maximization, which faces significant challenges for tasks involving both cost constraints and hybrid action spaces. In this work, we propose a novel Constrained Hybrid-action Policy Optimization algorithm (CHPO) to address the problems of constrained hybrid-action RL. Concretely, we rethink the limitations of hybrid-action RL in handling safe tasks with parameterized action spaces and reframe the objective of constrained hybrid-action RL by introducing the concept of Constrained Parameterized-action Markov Decision Process (CPMDP). Subsequently, we present a constrained hybrid-action policy optimization algorithm to confront the constrained hybrid-action problems and conduct theoretical analyses demonstrating that the CHPO converges to the optimal solution while satisfying safety constraints. Finally, extensive experiments demonstrate that the CHPO achieves competitive performance across multiple experimental tasks.
Ao Zhou 0005, Jiayi Guan, Li Shen 0008, Fan Lu 0001, Sanqing Qu, Junqiao Zhao, Guang Chen 0001
NeurIPS5
2025 Deep reinforcement learning as an interaction agent to steer fragment-based 3D molecular generation for protein pockets
abstract
Designing high-affinity molecules for protein targets (especially novel protein families) is a crucial yet challenging task in drug discovery. Recently, there has been tremendous progress in structure-based 3D molecular generative models that incorporate structural information of protein pockets. However, the capacity for molecular representation learning and the generalization for capturing interaction patterns need substantial further developments. Here, we propose AMG, a framework that leverages deep reinforcement learning as a pocket-ligand interaction agent (IA) to gradually steer fragment-based 3D molecular generation targeting protein pockets. AMG is trained using a two-stage strategy to capture interaction features and explicitly optimize the IA. The framework also introduces a pair of separate encoders for pockets and ligands, coupled with a dedicated pre-training strategy. This enables AMG to enhance its generalization ability by leveraging a vast repository of undocked pockets and molecules, thus mitigating the constraints posed by the limited quantity and quality of available datasets. Extensive evaluations demonstrate that AMG significantly outperforms five state-of-the-art baselines in affinity performance while maintaining proper drug-likeness properties. Furthermore, visual analysis confirms the superiority of AMG at capturing 3D molecular geometrical features and interaction patterns within pocket-ligand complexes, indicating its considerable promise for various structure-based downstream tasks.
Sanqing Qu, Fan Lu 0001, Zhixin Tian, Alois C. Knoll, Guang Chen 0001, Shaorong Gao, Yanping Zhang 0007
Briefings Bioinform.3
2025 General Class-Balanced Multicentric Dynamic Prototype Pseudo-Labeling for Source-Free Domain Adaptation
Sanqing Qu, Guang Chen 0001, Jing Zhang 0037, Zhijun Li 0001, Wei He 0001, Dacheng Tao
Int. J. Comput. Vis.1
2025 GLC++: Source-Free Universal Domain Adaptation Through Global-Local Clustering and Contrastive Affinity Learning
abstract
Deep neural networks often exhibit sub-optimal performance under covariate and category shifts. Source-Free Domain Adaptation (SFDA) presents a promising solution to this dilemma, yet most SFDA approaches are restricted to closed-set scenarios. In this paper, we explore Source-Free Universal Domain Adaptation (SF-UniDA) aiming to accurately classify "known" data belonging to common categories and segregate them from target-private "unknown" data. We propose a novel Global and Local Clustering (GLC) technique, which comprises an adaptive one-vs-all global clustering algorithm to discern between target classes, complemented by a local k-NN clustering strategy to mitigate negative transfer. Despite the effectiveness, the inherent closed-set source architecture leads to uniform treatment of "unknown" data, impeding the identification of distinct "unknown" categories. To address this, we evolve GLC to GLC++, integrating a contrastive affinity learning strategy. We examine the superiority of GLC and GLC++ across multiple benchmarks and category shift scenarios. Remarkably, in the most challenging open-partial-set scenarios, GLC and GLC++ surpass GATE by 16.8% and 18.9% in H-score on VisDA, respectively. GLC++ enhances the novel category clustering accuracy of GLC by 4.1% in open-set scenarios on Office-Home. Furthermore, the introduced contrastive learning strategy not only enhances GLC but also significantly facilitates existing methodologies.
Sanqing Qu, Tianpei Zou, Florian Röhrbein, Cewu Lu, Guang Chen 0001, Dacheng Tao, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 MHCPP: A Motion-Based Historical Enhancement Collaborative Perception and Prediction Framework
Shihang Du, Sanqing Qu, Tianhang Wang, Fan Lu 0001, Guang Chen 0001, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.2
2024 MAP: MAsk-Pruning for Source-Free Model Intellectual Property Protection
abstract
Deep learning has achieved remarkable progress in various applications, heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment of models in authorized data domains, i.e., making models exclusive to certain target domains. Previous methods necessitate concurrent access to source training data and target unauthorized data when performing IP protection, making them risky and inefficient for decentralized private data. In this paper, we target a practical setting where only a well-trained source model is available and investigate how we can realize IP protection. To achieve this, we propose a novel MAsk Pruning (MAP) framework. MAP stems from an intuitive hypothesis, i.e., there are target-related parameters in a well-trained model, locating and pruning them is the key to IP protection. Technically, MAP freezes the source model and learns a target-specific binary mask to prevent unauthorized data usage while minimizing performance degradation on authorized data. Moreover, we introduce a new metric aimed at achieving a better balance between source and target performance degradation. To verify the effectiveness and versatility, we have evaluated MAP in a variety of scenarios, including vanilla source-available, practical source-free, and challenging data-free. Extensive experiments indicate that MAP yields new state-of-the-art performance. Code will be available at https://github.com/ispc-lab/MAP.
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR2
2024 LEAD: Learning Decomposition for Source-free Universal Domain Adaptation
abstract
Universal Domain Adaptation (UniDA) targets knowledge transfer in the presence of both covariate and label shifts. Recently, Source-free Universal Domain Adaptation (SF-UniDA) has emerged to achieve UniDA without access to source data, which tends to be more practical due to data protection policies. The main challenge lies in determining whether covariate-shifted samples belong to target-private unknown categories. Existing methods tackle this either through hand-crafted thresholding or by developing time-consuming iterative clustering strategies. In this paper, we propose a new idea of LEArning Decomposition (LEAD), which decouples features into source-known and-unknown components to identify target-private data. Technically, LEAD initially leverages the or-thogonal decomposition analysis for feature decomposition. Then, LEAD builds instance-level decision boundaries to adaptively identify target-private data. Extensive experiments across various UniDA scenarios have demonstrated the effectiveness and superiority of LEAD. Notably, in the OPDA scenario on VisDA dataset, LEAD outperforms GLC by 3.5% overall H-score and reduces 75% time to derive pseudo-labeling decision boundaries. Besides, LEAD is also appealing in that it is complementary to most existing methods. The code is available at https://github.com/ispc-lab/LEAD.
Sanqing Qu, Tianpei Zou, Lianghua He, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR1
2024 HGL: Hierarchical Geometry Learning for Test-Time Adaptation in 3D Point Cloud Segmentation
Tianpei Zou, Sanqing Qu, Zhijun Li 0001, Alois C. Knoll, Lianghua He, Guang Chen 0001, Changjun Jiang 0002
ECCV (55)2
2024 ELF-UA: Efficient Label-Free User Adaptation in Gaze Estimation
Yong Wu 0007, Yang Wang 0003, Sanqing Qu, Zhijun Li 0001, Guang Chen 0001
IJCAI3
2024 PCDepth: Pattern-based Complementary Learning for Monocular Depth Estimation by Best of Both Worlds
abstract
Event cameras can record scene dynamics with high temporal resolution, providing rich scene details for monocular depth estimation (MDE) even at low-level illumination. Therefore, existing complementary learning approaches for MDE fuse intensity information from images and scene details from event data for better scene understanding. However, most methods directly fuse two modalities at pixel level, ignoring that the attractive complementarity mainly impacts high-level patterns that only occupy a few pixels. For example, event data is likely to complement contours of scene objects. In this paper, we discretize the scene into a set of high-level patterns to explore the complementarity and propose a Pattern-based Complementary learning architecture for monocular Depth estimation (PCDepth). Concretely, PCDepth comprises two primary components: a complementary visual representation learning module for discretizing the scene into high-level patterns and integrating complementary patterns across modalities and a refined depth estimator aimed at scene reconstruction and depth prediction while maintaining an efficiency-accuracy balance. Through pattern-based complementary learning, PCDepth fully exploits two modalities and achieves more accurate predictions than existing methods, especially in challenging nighttime scenarios. Extensive experiments on MVSEC and DSEC datasets verify the effectiveness and superiority of our PCDepth. Remarkably, compared with state-of-the-art, PCDepth achieves a 37.9% improvement in accuracy in MVSEC nighttime scenarios.
Sanqing Qu, Fan Lu 0001, Zongtao Bu, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001
IROS2
2023 Modality-Agnostic Debiasing for Single Domain Generalization
abstract
Deep neural networks (DNNs) usually fail to generalize well to outside of distribution (OOD) data, especially in the extreme case of single domain generalization (single-DG) that transfers DNNs from single domain to multiple unseen domains. Existing single-DG techniques commonly devise various data-augmentation algorithms, and remould the multi-source domain generalization methodology to learn domain-generalized (semantic) features. Nevertheless, these methods are typically modality-specific, thereby being only applicable to one single modality (e.g., image). In contrast, we target a versatile Modality-Agnostic Debiasing (MAD) framework for single-DG, that enables generalization for different modalities. Technically, MAD introduces a novel two-branch classifier: a biased-branch encourages the classifier to identify the domain-specific (superficial) features, and a general-branch captures domain-generalized features based on the knowledge from biased-branch. Our MAD is appealing in view that it is pluggable to most single-DG models. We validate the superiority of our MAD in a variety of single-DG scenarios with different modalities, including recognition on 1D texts, 2D images, 3D point clouds, and semantic segmentation on 2D images. More remarkably, for recognition on 3D point clouds and semantic segmentation on 2D images, MAD improves DSU by 2.82% and 1.5% in accuracy and mIOU.
Sanqing Qu, Yingwei Pan, Guang Chen 0001, Ting Yao 0003, Changjun Jiang 0002, Tao Mei 0001
CVPR1
2023 Upcycling Models Under Domain and Category Shift
abstract
Deep neural networks (DNNs) often perform poorly in the presence of domain shift and category shift. How to upcycle DNNs and adapt them to the target task remains an important open problem. Unsupervised Domain Adaptation (UDA), especially recently proposed Source-free Domain Adaptation (SFDA), has become a promising technology to address this issue. Nevertheless, existing SFDA methods require that the source domain and target domain share the same label space, consequently being only applicable to the vanilla closed-set setting. In this paper, we take one step further and explore the Source-free Universal Domain Adaptation (SF-UniDA). The goal is to identify “known” data samples under both domain and category shift, and reject those “unknown” data samples (not present in source classes), with only the knowledge from standard pre-trained source model. To this end, we introduce an innovative global and local clustering learning technique (GLC). Specifically, we design a novel, adaptive one-vs-all global clustering algorithm to achieve the distinction across different target classes and introduce a local k-NN clustering strategy to alleviate negative transfer. We examine the superiority of our GLC on multiple benchmarks with different category shift scenarios, including partial-set, open-set, and open-partial-set DA. Remarkably, in the most challenging open-partial-set DA scenario, GLC outperforms UMAD by 14.8% on the VisDA benchmark. The code is available at https://github.com/ispc-lab/GLC.
Sanqing Qu, Tianpei Zou, Florian Röhrbein, Cewu Lu, Guang Chen 0001, Dacheng Tao, Changjun Jiang 0002
CVPR1
2023 TMA: Temporal Motion Aggregation for Event-based Optical Flow
abstract
Event cameras have the ability to record continuous and detailed trajectories of objects with high temporal resolution, thereby providing intuitive motion cues for optical flow estimation. Nevertheless, most existing learning-based approaches for event optical flow estimation directly remould the paradigm of conventional images by representing the consecutive event stream as static frames, ignoring the inherent temporal continuity of event data. In this paper, we argue that temporal continuity is a vital element of event-based optical flow and propose a novel Temporal Motion Aggregation (TMA) approach to unlock its potential. Technically, TMA comprises three components: an event splitting strategy to incorporate intermediate motion information underlying the temporal context, a linear lookup strategy to align temporally fine-grained motion features and a novel motion pattern aggregation module to emphasize consistent patterns for motion feature enhancement. By incorporating temporally fine-grained motion information, TMA can derive better flow estimates than existing methods at early stages, which not only enables TMA to obtain more accurate final predictions, but also greatly reduces the demand for a number of refinements. Extensive experiments on DSEC-Flow and MVSEC datasets verify the effectiveness and superiority of our TMA. Remarkably, compared to E-RAFT, TMA achieves a 6% improvement in accuracy and a 40% reduction in inference time on DSEC-Flow. Code will be available at https://github.com/ispc-lab/TMA.
Guang Chen 0001, Sanqing Qu, Yanping Zhang 0007, Zhijun Li 0001, Alois C. Knoll, Changjun Jiang 0002
ICCV3
2023 HRegNet: A Hierarchical Network for Efficient and Accurate Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale outdoor LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching, and novel similarity features are designed to incorporate them into the correspondence network, which significantly improves the registration performance. In addition, we design a consistency propagation strategy to effectively incorporate spatial consistency into the registration pipeline. The whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on three large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The source code of the proposed HRegNet is available at https://github.com/ispc-lab/HRegNet2.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 PSDC: A Prototype-Based Shared-Dummy Classifier Model for Open-Set Domain Adaptation
abstract
Open-set domain adaptation (OSDA) aims to achieve knowledge transfer in the presence of both domain shift and label shift, which assumes that there exist additional unknown target classes not presented in the source domain. To solve the OSDA problem, most existing methods introduce an additional unknown class to the source classifier and represent the unknown target instances as a whole. However, it is unreasonable to treat all unknown target instances as a group since these unknown instances typically consist of distinct categories and distributions. It is challenging to identify all unknown instances with only one additional class. In addition, most existing methods directly introduce marginal distribution alignment to alleviate distribution shift between the source and target domains, failing to learn discriminative class boundaries in the target domain since they ignore categorical discriminative information in the adaptation. To address these problems, in this article, we propose a novel prototype-based shared-dummy classifier (PSDC) model for the OSDA. Specifically, our PSDC introduces an auxiliary dummy classifier to calibrate the source classifier and simultaneously develops a weighted adaptation procedure to align class-wise prototypes for adaptation. We further design a pseudo-unknown learning algorithm to reduce the open-set risk. Extensive experiments on Office-31, Office-Home, and VisDA datasets show that the proposed PSDC can outperform existing methods and achieve the new state-of-the-art performance. The code will be made public.
Zhengfa Liu, Guang Chen 0001, Zhijun Li 0001, Yu Kang 0001, Sanqing Qu, Changjun Jiang 0002
IEEE Trans. Cybern.5
2023 RPP-Net: Rigid Constrained Point Cloud Prediction Network
abstract
Forecasting the future environment is an essential and fundamental capability of autonomous driving systems. In contrast to the widely studied video prediction, only a few literatures have explored LiDAR point cloud prediction. To achieve future point cloud generation, most existing methods are based on free-form 3D scene flow prediction. However, these simple scene flow prediction-based methods may cause distortions since the motions of 3D scenes can be seen as a combination of rigid (static background) and flow (dynamic foreground) motions. To address this issue, we propose a simple but effective rigid constrained point cloud prediction network named RPP-Net. The key component of the proposed method is the hybrid motion decoder, which generates motion masks and combines flow motions and rigid motions to produce hybrid motions without additional annotations and sensors. Besides, to reduce running time, we provide a hierarchical cell, which uses a feature encoder to extract deep features and an RPP-RNN module to capture temporal correlations across frames. To evaluate the effectiveness of our RPP-Net, we have conducted extensive experiments on both KITTI dataset and Argoverse dataset and the results show that our RPP-Net significantly outperforms existing methods and achieves a new state-of-the-art.
Tianpei Zou, Guang Chen 0001, Fan Lu 0001, Zhijun Li 0001, Sanqing Qu, Alois C. Knoll, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.5
2022 BMD: A General Class-Balanced Multicentric Dynamic Prototype Strategy for Source-Free Domain Adaptation
Sanqing Qu, Guang Chen 0001, Jing Zhang 0037, Zhijun Li 0001, Wei He 0001, Dacheng Tao
ECCV (34)1
2022 Neuromorphic Vision-Based Fall Localization in Event Streams With Temporal-Spatial Attention Weighted Network
abstract
Falling down is a serious problem for health and has become one of the major etiologies of accidental death for the elderly living alone. In recent years, many efforts have been paid to fall recognition based on wearable sensors or standard vision sensors. However, the prior methods have the risk of privacy leaks, and almost all these methods are based on video clips, which cannot localize where the falls occurred in long videos. For these reasons, in this article, the bioinspired vision sensor-based falls temporal localization framework is proposed. The bioinspired vision sensors, such as dynamic and active-pixel vision sensor (DAVIS) camera applied in this work responds to pixels' brightness change, and each pixel works independently and asynchronously compared to the standard vision sensors. This property makes it have a very high dynamic range and privacy preserving. First, to better represent event data, compared with the typical constant temporal window mechanism, an adaptive temporal window conversion mechanism is developed. The temporal localization framework follows a proven proposal and classification paradigm. Second, for the high-efficient and recall proposal generation, different from the traditional sliding window scheme, the event temporal density as the actionness score is set and the 1D-watershed algorithm to generate proposals is applied. In addition, we combine the temporal and spatial attention mechanism with our feature extraction network to temporally model the falls. Finally, to evaluate the performance of our framework, 30 volunteers are recruited to join the simulated fall experiments. According to the results of experiments, our framework can realize precise falls temporal localization and achieve the state-of-the-art performance.
Guang Chen 0001, Sanqing Qu, Zhijun Li 0001, Jiaxuan Dong, Min Liu 0031, Jörg Conradt
IEEE Trans. Cybern.2
2022 MoNet: Motion-Based Point Cloud Prediction Network
abstract
Predicting the future can significantly improve the safety of intelligent vehicles, which is a key component in autonomous driving. 3D point clouds can accurately model 3D information of surrounding environment and are crucial for intelligent vehicles to perceive the scene. Therefore, prediction of 3D point clouds has great significance for intelligent vehicles, which can be utilized for numerous further applications. However, due to point clouds are unordered and unstructured, point cloud prediction is challenging and has not been deeply explored in current literature. In this paper, we propose a novel motion-based neural network named MoNet. The key idea of the proposed MoNet is to integrate motion features between two consecutive point clouds into the prediction pipeline. The introduction of motion features enables the model to more accurately capture the variations of motion information across frames and thus make better predictions for future motion. In addition, content features are introduced to model the spatial content of individual point clouds. A recurrent neural network named MotionRNN is proposed to capture the temporal correlations of both features. Moreover, an attention-based motion align module is proposed to address the problem of missing motion features in the inference pipeline. Extensive experiments on two large-scale outdoor LiDAR point cloud datasets demonstrate the performance of the proposed MoNet. Moreover, we perform experiments on applications using the predicted point clouds and the results indicate the great application potential of the proposed method.
Fan Lu 0001, Guang Chen 0001, Zhijun Li 0001, Yinlong Liu, Sanqing Qu, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.6
2021 PointINet: Point Cloud Frame Interpolation Network
abstract
LiDAR point cloud streams are usually sparse in time dimension, which is limited by hardware performance. Generally, the frame rates of mechanical LiDAR sensors are 10 to 20 Hz, which is much lower than other commonly used sensors like cameras. To overcome the temporal limitations of LiDAR sensors, a novel task named Point Cloud Frame Interpolation is studied in this paper. Given two consecutive point cloud frames, Point Cloud Frame Interpolation aims to generate intermediate frame(s) between them. To achieve that, we propose a novel framework, namely Point Cloud Frame Interpolation Network (PointINet). Based on the proposed method, the low frame rate point cloud streams can be upsampled to higher frame rates. We start by estimating bi-directional 3D scene flow between the two point clouds and then warp them to the given time step based on the 3D scene flow. To fuse the two warped frames and generate intermediate point cloud(s), we propose a novel learning-based points fusion module, which simultaneously takes two warped point clouds into consideration. We design both quantitative and qualitative experiments to evaluate the performance of the point cloud frame interpolation method and extensive experiments on two large scale outdoor LiDAR datasets demonstrate the effectiveness of the proposed PointINet. Our code is available at https://github.com/ispc-lab/PointINet.git.
Fan Lu 0001, Guang Chen 0001, Sanqing Qu, Zhijun Li 0001, Yinlong Liu, Alois C. Knoll
AAAI3
2021 HRegNet: A Hierarchical Network for Large-scale Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale out-door LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching and novel similarity features are designed to in-corporate them into the correspondence network, which significantly improves the registration performance. Besides, the whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on two large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The project website is https://ispc-group.github.io/hregnet.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu
ICCV5
2021 A Novel Illumination-Robust Hand Gesture Recognition System With Event-Based Neuromorphic Vision Sensor
abstract
The hand gesture recognition system is a noncontact and intuitive communication approach, which, in turn, allows for natural and efficient interaction. This work focuses on developing a novel and robust gesture recognition system, which is insensitive to environmental illumination and background variation. In the field of gesture recognition, standard vision sensors, such as CMOS cameras, are widely used as the sensing devices in state-of-the-art hand gesture recognition systems. However, such cameras depend on environmental constraints, such as lighting variability and the cluttered background, which significantly deteriorates their performances. In this work, we propose an event-based gesture recognition system to overcome the detriment constraints and enhance the robustness of the recognition performance. Our system relies on a biologically inspired neuromorphic vision sensor that has microsecond temporal resolution, high dynamic range, and low latency. The sensor output is a sequence of asynchronous events instead of discrete frames. To interpret the visual data, we utilize a wearable glove as an interaction device with five high-frequency (>100 Hz) active LED markers (ALMs), representing fingers and palm, which are tracked precisely in the temporal domain using a restricted spatiotemporal particle filter algorithm. The latency of the sensing pipeline is negligible compared with the dynamics of the environment as the sensor's temporal resolution allows us to distinguish high frequencies precisely. We design an encoding process to extract features and adopt a lightweight network to classify the hand gestures. The recognition accuracy of our system is comparable to the state-of-the-art methods. To study the robustness of the system, experiments considering illumination and background variations are performed, and the results show that our system is more robust than the state-of-the-art deep learning-based gesture recognition systems. Note to Practitioners-This article addresses the robustness of the hand gesture recognition system that is important for gesture recognition-based applications. Existing methods rely on either the large-volume data to train a deep learning model or to restrict the applied environments (e.g., an ideal environment without dynamic background). However, a vision-based deep learning model requires large computational resources, while the ideal environment limits the practicality of the system. In this work, we introduce a biologically inspired neuromorphic vision sensor and an ALM glove and build a novel gesture recognition system to tackle the above issue. The neuromorphic vision sensor has a microsecond temporal resolution and a high dynamic range. With these properties, the sensing system of our prototype operates in a very low-latency space, which, in turn, ensures that our gesture recognition system is robust to illumination variance and dynamic background. Thus, this work is valuable to the research of illumination-robust gesture recognition systems. Preliminary experiments suggest that our system prototype is feasible, but it has not yet been incorporated into an online gesture recognition system nor tested with complex gestures. In future work, we will concentrate on the improvement of the signal processing methods that advance the current system to complex and practical applications.
Guang Chen 0001, Zhongcong Xu, Zhijun Li 0001, Huajin Tang, Sanqing Qu, Kejia Ren, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.5
2021 Pseudo-Image and Sparse Points: Vehicle Detection With 2D LiDAR Revisited by Deep Learning-Based Methods
abstract
Detecting and locating surrounding vehicles robustly and efficiently are essential capabilities for autonomous vehicles. Existing solutions often rely on vision-based methods or 3D LiDAR-based methods. These methods are either too expensive in both sensor pricing (3D LiDAR) and computation (camera and 3D LiDAR) or less robust in resisting harsh environment changes (camera). In this work, we revisit the LiDAR based approaches for vehicle detection with a less expensive 2D LiDAR by utilizing modern deep learning approaches. We aim at filling in the gap as few previous works conclude an efficient and robust vehicle detection solution in a deep learning way in 2D. To this end, we propose a learning based method with the input of pseudo-images, named Cascade Pyramid Region Proposal Convolution Neural Network (Cascade Pyramid RCNN), and a hybrid learning method with the input of sparse points, named Hybrid Resnet Lite. Experiments are conducted with our newly 2D LiDAR vehicle dataset recorded in complex traffic environments. Results demonstrate that the Cascade Pyramid RCNN outperforms state-of-the-art methods in accuracy while the proposed Hybrid Resnet Lite provides superior performance of the speed and lightweight model by hybridizing learning based and non-learning based modules. As few previous works conclude an efficient and robust vehicle detection solution with 2D LiDAR, our research fills in this gap and illustrates that even with limited sensing source from a 2D LiDAR, detecting obstacles like vehicles efficiently and robustly is still achievable.
Guang Chen 0001, Fa Wang, Sanqing Qu, Lu Xiong 0001, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.3