Guang Chen 0001

dblp:09/4891-1 · DBLP profile ↗
← Back
92ranked-venue papers
9as first author
77since 2021 · last 2026
0000-0002-7416-592XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 71 · 3 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 25 since 2021Systems, architecture and hardware · 14 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Intelligible Human-Robot Interaction: An Active Inference Approach to Occluded Pedestrian Scenarios
abstract
The sudden appearance of occluded pedestrians presents a critical safety challenge in autonomous driving. Conventional rule-based or purely data-driven approaches struggle with the inherent high uncertainty of these long-tail scenarios. To tackle this challenge, we propose a novel framework grounded in Active Inference, which endows the agent with a human-like, belief-driven mechanism. Our framework leverages a Rao-Blackwellized Particle Filter (RBPF) to efficiently estimate the pedestrian's hybrid state. To emulate human-like cognitive processes under uncertainty, we introduce a Conditional Belief Reset mechanism and a Hypothesis Injection technique to explicitly model beliefs about the pedestrian's multiple latent intentions. Planning is achieved via a Cross-Entropy Method (CEM) enhanced Model Predictive Path Integral (MPPI) controller, which synergizes the efficient, iterative search of CEM with the inherent robustness of MPPI. Simulation experiments demonstrate that our approach significantly reduces the collision rate compared to reactive, rule-based, and reinforcement learning (RL) baselines, while also exhibiting explainable and human-like driving behavior that reflects the agent's internal belief state.
Yuyao Huang 0001, Guang Chen 0001
HRI3
2026 Learning Temporal Awareness for Novel Class Discovery in 3D Point Cloud Segmentation
Tianpei Zou, Sanqing Qu, Alois C. Knoll, Guang Chen 0001
IV7
2026 Dynamic MAsk-Pruning Strategy for Source-Free Model Intellectual Property Protection
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
Int. J. Comput. Vis.7
2026 Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering
abstract
Video Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts.
Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Post-Refining LiDAR Point Cloud Registration via Diffusion-Based Correspondence Refinement
abstract
LiDAR point cloud registration is a fundamental problem in 3D computer vision. Recent learning-based methods have significantly improved the robustness and accuracy of LiDAR point cloud registration, while most of these methods are designed for global registration. In this paper, we focus on the post-refinement problem of LiDAR point cloud registration, which has been largely overlooked in previous learning-based approaches. To address the correspondence error between LiDAR point clouds, we formulate the transformation refinement problem as a correspondence optimization problem and propose DCR, a diffusion-based correspondence refinement model. DCR is built upon denoising diffusion models and recovers accurate correspondences from noisy initial correspondences estimated by arbitrary global registration methods. To achieve precise correspondence refinement, we employ a residual-based modeling scheme to constrain the output distribution and design a conditional generation paradigm to control the randomness of diffusion models. We combine DCR with both learning-based and traditional global registration methods and perform experiments on three large-scale outdoor LiDAR point cloud datasets to verify the performance. Extensive experimental results demonstrate DCR’s versatility and effectiveness in improving the accuracy of LiDAR point cloud registration.
Fan Lu 0001, Tianhang Wang, Bin Li 0087, Alois C. Knoll, Zhijun Li 0001, Guang Chen 0001
IEEE Trans Autom. Sci. Eng.8
2026 I2EKD: Efficient and Versatile Image-to-Event Knowledge Distillation
abstract
Recently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%.
Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.9
2026 An Online-Training-Free Adaptor for Open Heterogeneous Collaborative Perception via Diffusion Model
abstract
Collaborative perception seeks to mitigate the limitations of single-vehicle perception, such as occlusions, by facilitating communication and information sharing among connected vehicles. However, most existing works assume a homogeneous scenario where all vehicles share identity sensor types and perception model architectures. In contrast, real-world systems often involve heterogeneous agents with diverse sensor configurations and independently developed models. In such settings, directly exchanging features without proper alignment can significantly degrade performance and hinder effective collaboration. While some methods have been proposed to address heterogeneity, they typically require retraining or access to internal model parameters, making them impractical for scalable deployment. To address these challenges, we propose DiffAlign, a plug-and-play adapter that enables feature alignment across heterogeneous agents in a training-free and model-agnostic manner. DiffAlign treats received BEV features as noisy latent representations and progressively refines them through a pretrained diffusion process. This alignment strategy does not require access to model internals or any retraining, which makes it both scalable and privacy-preserving while supporting diverse sensor modalities and perception backbones. Extensive experiments on simulated OPV2V and real-world V2V4Real datasets demonstrate that DiffAlign consistently improves detection performance in heterogeneous settings, improving CoBEVT by 132.01% and 91.95%, respectively. Our method provides a practical path toward scalable, generalizable, and deployment-ready collaborative perception.
Tianhang Wang, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Hu Cao, Alois C. Knoll, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 Skill Information Representation Imitation Learning for Long-Horizon Dexterous Robot Micromanipulation of Deformable Cell
abstract
Robots performing collaborative long-horizon dexterity cell micromanipulation tasks are challenging and practically significant, such as peeling cell membranes, which is considered one of the most technically demanding procedures. The imitation learning (IL) approach is expected to address the challenges of multitask coupling and object modeling difficulties in long-horizon tasks. Existing IL algorithms suffer from compounding error as they perform only a simple mapping of the task environment space to the action space. In this article, we propose a skill information representation IL (SIRIL) algorithm for long-horizon dexterous robot micromanipulation tasks. First, SIRIL extracts the discrete latent codes of the expert video frames by the VQ-GAN encoder, and the distribution of the latent codes is modeled by an autoregressive transformer. SIRIL quantifies the representation of the expert's skill information by computing the log-likelihood of the latent discrete codes, which allows for the extraction of safe action constraints. Finally, SIRIL predicts actions by integrating actions from previous time steps, and actions that satisfy the safety constraints are executed, which effectively suppresses compounding error. Real experiments show that the SIRIL algorithm can complete the deformable zebrafish embryonic cells dexterous membrane stripping surgery. Ablation studies further confirmed SIRIL's high efficiency in various subtasks, including PushCell, GraspCell, and PeelCell, which achieves an average accuracy of 86.7% and a high final success rate of 64.7%, significantly outperforming existing algorithms. Code is available at https://github.com/zycrobot/SIRIL.
Youchao Zhang, Fanghao Wang, Guang Chen 0001, Alois C. Knoll, Yibin Ying, Mingchuan Zhou
IEEE Trans. Cybern.5
2025 RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptions
abstract
Collaborative perception enhances single-vehicle perception by integrating sensory data from multiple connected vehicles. However, existing studies often assume ideal conditions, overlooking resilience to real-world challenges, such as adverse weather and sensor malfunctions, which is critical for safe deployment. To address this gap, we introduce RCP-Bench, the first comprehensive benchmark designed to evaluate the robustness of collaborative detection models under a wide range of real-world corruptions. RCP-Bench includes three new datasets (i.e., OPV2V-C, V2XSet-C, and DAIR-V2X-C) that simulate six collaborative cases and 14 types of camera corruption resulting from external environmental factors, sensor failures, and temporal misalignments. Extensive experiments on 10 leading collaborative perception models reveal that, while these models perform well under ideal conditions, they are significantly affected by corruptions. To improve robustness, we propose two simple yet effective strategies, RCP-Drop and RCP-Mix, based on training regularization and feature augmentation. Additionally, we identify several critical factors influencing robustness, such as backbone architecture, camera number, feature fusion methods, and the number of connected vehicles. We hope that RCP-Bench, along with these strategies and insights, will stimulate future research toward developing more robust collaborative perception models. Our benchmark toolkit is available at https://github.com/LuckyDush/RCP-Bench.
Shihang Du, Sanqing Qu, Tianhang Wang, Yunwei Zhu, Fan Lu 0001, Guang Chen 0001
CVPR9
2025 Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering
abstract
Video Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition.
Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001
ICML8
2025 Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features
abstract
Point cloud completion is crucial for reconstructing accurate shapes in many 3D visual applications. Recent approaches incorporate images into the completion pipeline, introducing geometric clues and global constraints. However, their fusion processes often fail to reconstruct detailed parts and maintain global consistency simultaneously. Except for images, text is another important clue for recognizing the target's characteristics. Thus, in this work, we propose to combine multiple modalities including points, images and texts for point cloud completion. Specifically, inspired by recently pre-trained large language models, we generate the description texts for images by Visual Question Answering (VQA) models and introduce Visual-Textual Embedding (VTE) models to extract joint features of image-text pairs. Furthermore, we describe the edge geometric patterns by multi-scale edge convolution to guide the refinement of shapes in local areas. Then we adopt cross attention mechanism to effectively fuse multi-modal features and refine the coarse shape. Extensive experiments on commonly used benchmarks demonstrate our method's superior performance over previous uni-modal and cross-modal methods.
ChengKai Xia, Fan Lu 0001, Bin Li 0087, Alois C. Knoll, Guang Chen 0001
ICRA6
2025 Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review
abstract
Embodied multi-agent systems (EMAS) have attracted growing attention for their potential to address complex, real-world challenges in areas such as logistics and robotics. Recent advances in foundation models pave the way for generative agents capable of richer communication and adaptive problem-solving. This survey provides a systematic examination of how EMAS can benefit from these generative capabilities. We propose a taxonomy that categorizes EMAS by system architectures and embodiment modalities, emphasizing how collaboration spans both physical and virtual contexts. Central building blocks, perception, planning, communication, and feedback, are then analyzed to illustrate how generative techniques bolster system robustness and flexibility. Through concrete examples, we demonstrate the transformative effects of integrating foundation models into embodied, multi-agent frameworks. Finally, we discuss challenges and future directions, underlining the significant promise of EMAS to reshape the landscape of AI-driven collaboration.
Xian Wei, Guang Chen 0001, Hao Shen 0002, Bo Jin 0003
IJCAI3
2025 Event-Aided Progressive Neural Radiance Fields Reconstruction Under Challenging Illumination
abstract
Scene reconstruction and novel view synthesis are extensively utilized in domains such as autonomous driving, facilitating the creation of more comprehensive datasets and accelerating algorithm development. Current techniques have achieved promising results in image-based synthesis. Nonethe-less, they encounter difficulties including limited dynamic range, motion blur, and suboptimal performance under challenging illumination. Furthermore, these techniques depend significantly on precise camera poses, which are difficult to acquire in practical situations. Event cameras, characterized by their high dynamic range and temporal resolution, demonstrate advantages in challenging illumination and rapid motion scenarios and have been utilized in several autonomous driving applications. This study is the first to combine image and event camera for neural radiance field (NeRF) reconstruction in driving scenarios. The proposed method leverages the unique properties of event camera and adopts a progressive reconstruction strategy to jointly optimize camera poses during training, reducing reliance on pose precision and enabling more robust and accurate scene reconstruction under challenging illumination. Experimental results demonstrate that the proposed approach attains enhanced performance in novel view synthesis, exhibiting a 3.5% increase in PSNR and a 2% increase in SSIM on the DSEC dataset compared to baseline method.
Zongtao Bu, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Guang Chen 0001
IV6
2025 Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition
abstract
Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread availability of RGB cameras and the robustness of point clouds in providing accurate spatial geometry and distance information. However, current methods rely on intermediate modalities that capture either the vertical or horizontal field of view, limiting their ability to fully exploit the complementary information from both sensors. In this work, we propose an innovative initial retrieval + re-rank method that effectively combines information from range (or RGB) images and Bird's Eye View (BEV) images. Our approach relies solely on a computationally efficient global descriptor similarity search process to achieve re-ranking. Additionally, we introduce a novel similarity label supervision technique to maximize the utility of limited training data. Specifically, we employ points average distance to approximate appearance similarity and incorporate an adaptive margin, based on similarity differences, into the vanilla triplet loss. Experimental results on the KITTI dataset demonstrate that our method significantly outperforms state-of-the-art approaches.
Jianyi Peng, Fan Lu 0001, Bin Li 0087, Sanqing Qu, Guang Chen 0001
IV6
2025 MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive Clustering
abstract
Visual Place Recognition (VPR) enables robust localization through image retrieval based on learned descriptors. However, drastic appearance variations of images at the same place caused by viewpoint changes can lead to inconsistent supervision signals, thereby degrading descriptor learning. Existing methods either rely on manually defined cropping rules or labeled data for view differentiation, but they suffer from two major limitations: (1) reliance on labels or handcrafted rules restricts generalization capability; (2) even within the same view direction, occlusions can introduce feature ambiguity. To address these issues, we propose MutualVPR, a mutual learning framework that integrates unsupervised view self-classification and descriptor learning. We first group images by geographic coordinates, then iteratively refine the clusters using K-means to dynamically assign place categories without manual labeling. Specifically, we adopt a DINOv2-based encoder to initialize the clustering. During training, the encoder and clustering co-evolve, progressively separating drastic appearance variations of the same place and enabling consistent supervision. Furthermore, we find that capturing fine-grained image differences at a place enhances robustness. Experiments demonstrate that MutualVPR achieves state-of-the-art (SOTA) performance across multiple datasets, validating the effectiveness of our framework in improving view direction generalization, occlusion robustness.
Qiwen Gu, Xufei Wang, Junqiao Zhao, Siyue Tao, Tiantian Feng, Guang Chen 0001
NeurIPS7
2025 OOD-Barrier: Build a Middle-Barrier for Open-Set Single-Image Test Time Adaptation via Vision Language Models
abstract
In real-world environments, a well-designed model must be capable of handling dynamically evolving distributions, where both in-distribution (ID) and out-of-distribution (OOD) samples appear unpredictably and individually, making real-time adaptation particularly challenging. While open-set test-time adaptation has demonstrated effectiveness in adjusting to distribution shifts, existing methods often rely on batch processing and struggle to manage single-sample data stream in open-set environments. To address this limitation, we propose Open-IRT, a novel open-set Intermediate-Representation-based Test-time adaptation framework tailored for single-image test-time adaptation with vision-language models. Open-IRT comprises two key modules designed for dynamic, single-sample adaptation in open-set scenarios. The first is Polarity-aware Prompt-based OOD Filter module, which fully constructs the ID-OOD distribution, considering both the absolute semantic alignment and relative semantic polarity. The second module, Intermediate Domain-based Test-time Adaptation module, constructs an intermediate domain and indirectly decomposes the ID-OOD distributional discrepancy to refine the separation boundary during the test-time. Extensive experiments on a range of domain adaptation benchmarks demonstrate the superiority of Open-IRT. Compared to previous state-of-the-art methods, it achieves significant improvements on representative benchmarks, such as CIFAR-100C and SVHN — with gains of +8.45\% in accuracy, -10.80\% in FPR95, and +11.04\% in AUROC.
Boyang Peng, Sanqing Qu, Tianpei Zou, Fan Lu 0001, Siheng Chen, Yong Wu 0007, Guang Chen 0001
NeurIPS9
2025 Multimodal LiDAR-Camera Novel View Synthesis with Unified Pose-free Neural Fields
abstract
Pose-free Neural Radiance Field (NeRF) aims at novel view synthesis (NVS) without relying on accurate poses, exhibiting significant practical value. Image and LiDAR point cloud are two pivotal modalities in autonomous driving scenarios. While demonstrating impressive performance, single-modality pose-free NeRFs often suffer from local optima due to the limited geometric information provided by dense image textures or the sparse, textureless nature of point clouds. Although prior methods have explored the complementary strengths of both modalities, they have only leveraged inherently sparse point clouds for discrete, non-pixel-wise depth supervision, and are limited to NVS of images. As a result, a Multimodal Unified Pose-free framework remains notably absent. In light of this, we propose MUP, a pose-free framework for LiDAR-Camera joint NVS in large-scale scenes. This unified framework enables continuous depth supervision for image reconstruction using LiDAR-Fields rather than discrete point clouds. By leveraging multimodal inputs, pose optimization receives gradients from the rendering loss of point cloud geometry and image texture, thereby alleviating the issue of local optima commonly encountered in single-modality pose-free tasks. Moreover, to further guide pose optimization of NeRF, we propose a multimodal geometric optimizer that leverages geometric relations from point clouds and photometric regularization from adjacent image frames. Besides, to alleviate the domain gap between modalities, we propose a multimodal-specific coarse-to-fine training approach for unified, compact reconstruction. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate MUP's superiority in accomplishing geometry-aware, modality-consistent, and pose-free 3D reconstruction.
Weiyi Xue, Fan Lu 0001, Yunwei Zhu, Zehan Zheng, Sanqing Qu, Jiangtong Li, Haiyun Wei, Guang Chen 0001
NeurIPS9
2025 CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning
abstract
Constrained hybrid-action reinforcement learning (RL) promises to learn a safe policy within a parameterized action space, which is particularly valuable for safety-critical applications involving discrete-continuous hybrid action spaces. However, existing hybrid-action RL algorithms primarily focus on reward maximization, which faces significant challenges for tasks involving both cost constraints and hybrid action spaces. In this work, we propose a novel Constrained Hybrid-action Policy Optimization algorithm (CHPO) to address the problems of constrained hybrid-action RL. Concretely, we rethink the limitations of hybrid-action RL in handling safe tasks with parameterized action spaces and reframe the objective of constrained hybrid-action RL by introducing the concept of Constrained Parameterized-action Markov Decision Process (CPMDP). Subsequently, we present a constrained hybrid-action policy optimization algorithm to confront the constrained hybrid-action problems and conduct theoretical analyses demonstrating that the CHPO converges to the optimal solution while satisfying safety constraints. Finally, extensive experiments demonstrate that the CHPO achieves competitive performance across multiple experimental tasks.
Ao Zhou 0005, Jiayi Guan, Li Shen 0008, Fan Lu 0001, Sanqing Qu, Junqiao Zhao, Guang Chen 0001
NeurIPS9
2025 Deep reinforcement learning as an interaction agent to steer fragment-based 3D molecular generation for protein pockets
abstract
Designing high-affinity molecules for protein targets (especially novel protein families) is a crucial yet challenging task in drug discovery. Recently, there has been tremendous progress in structure-based 3D molecular generative models that incorporate structural information of protein pockets. However, the capacity for molecular representation learning and the generalization for capturing interaction patterns need substantial further developments. Here, we propose AMG, a framework that leverages deep reinforcement learning as a pocket-ligand interaction agent (IA) to gradually steer fragment-based 3D molecular generation targeting protein pockets. AMG is trained using a two-stage strategy to capture interaction features and explicitly optimize the IA. The framework also introduces a pair of separate encoders for pockets and ligands, coupled with a dedicated pre-training strategy. This enables AMG to enhance its generalization ability by leveraging a vast repository of undocked pockets and molecules, thus mitigating the constraints posed by the limited quantity and quality of available datasets. Extensive evaluations demonstrate that AMG significantly outperforms five state-of-the-art baselines in affinity performance while maintaining proper drug-likeness properties. Furthermore, visual analysis confirms the superiority of AMG at capturing 3D molecular geometrical features and interaction patterns within pocket-ligand complexes, indicating its considerable promise for various structure-based downstream tasks.
Sanqing Qu, Fan Lu 0001, Zhixin Tian, Alois C. Knoll, Guang Chen 0001, Shaorong Gao, Yanping Zhang 0007
Briefings Bioinform.7
2025 General Class-Balanced Multicentric Dynamic Prototype Pseudo-Labeling for Source-Free Domain Adaptation
Sanqing Qu, Guang Chen 0001, Jing Zhang 0037, Zhijun Li 0001, Wei He 0001, Dacheng Tao
Int. J. Comput. Vis.2
2025 GLC++: Source-Free Universal Domain Adaptation Through Global-Local Clustering and Contrastive Affinity Learning
abstract
Deep neural networks often exhibit sub-optimal performance under covariate and category shifts. Source-Free Domain Adaptation (SFDA) presents a promising solution to this dilemma, yet most SFDA approaches are restricted to closed-set scenarios. In this paper, we explore Source-Free Universal Domain Adaptation (SF-UniDA) aiming to accurately classify "known" data belonging to common categories and segregate them from target-private "unknown" data. We propose a novel Global and Local Clustering (GLC) technique, which comprises an adaptive one-vs-all global clustering algorithm to discern between target classes, complemented by a local k-NN clustering strategy to mitigate negative transfer. Despite the effectiveness, the inherent closed-set source architecture leads to uniform treatment of "unknown" data, impeding the identification of distinct "unknown" categories. To address this, we evolve GLC to GLC++, integrating a contrastive affinity learning strategy. We examine the superiority of GLC and GLC++ across multiple benchmarks and category shift scenarios. Remarkably, in the most challenging open-partial-set scenarios, GLC and GLC++ surpass GATE by 16.8% and 18.9% in H-score on VisDA, respectively. GLC++ enhances the novel category clustering accuracy of GLC by 4.1% in open-set scenarios on Office-Home. Furthermore, the introduced contrastive learning strategy not only enhances GLC but also significantly facilitates existing methodologies.
Sanqing Qu, Tianpei Zou, Florian Röhrbein, Cewu Lu, Guang Chen 0001, Dacheng Tao, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 ROSCOM: Robust Safe Reinforcement Learning on Stochastic Constraint Manifolds
abstract
Reinforcement Learning (RL) has demonstrated remarkable success across various domains. Nonetheless, a significant challenge in RL is to ensure safety, particularly when deploying it in safety-critical applications such as robotics and autonomous driving. In this work, we develop a robust and safe RL methodology grounded in manifold space. Initially, we construct a constrained manifold space, taking safety constraints into consideration. We then propose a robust safe RL approach, supported by theoretical analysis, based on the value at risk and conditional value at risk, in order to enhance the robustness of safety. Our methodology is designed to ensure safety within stochastic constraint environments. Following the theoretical analysis, we develop a practical, safe algorithm to search for a robust safe policy on stochastic constraint manifolds (ROSCOM). We evaluate the effectiveness of our approach through circular motion and air-hockey tasks. Our experiments demonstrate that ROSCOM outperforms existing baselines in terms of both reward and safety. Note to Practitioners—Real-world applications often involve inherent uncertainties, noise, and high-dimensional spaces. This complexity accentuates the urgency and challenge of ensuring safety in robot learning, especially when implementing RL in practical environments. To address this critical issue, we build a stochastic constraint manifold to delineate the safety space, thus establishing a rigorous framework for robot learning at each iteration. Compared with state-of-the-art baselines, our method can provide remarkable performance regarding safety and reward performance. For example, in an air hockey robot learning task, our method has demonstrated a remarkable 50% enhancement in safety performance compared to the ATACOM framework, while concurrently exhibiting superior reward performance. Moreover, in contrast to traditional algorithms, including CPO, PCPO, our method has achieved a 99% improvement in safety performance, coupled with significantly superior reward performance. These empirical insights render our approach not only theoretically sound but also practically efficacious, indicating its potential as a useful tool in real robot learning and beyond.
Shangding Gu, Puze Liu, Alap Kshirsagar, Guang Chen 0001, Jan Peters 0001, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.4
2025 MHCPP: A Motion-Based Historical Enhancement Collaborative Perception and Prediction Framework
Shihang Du, Sanqing Qu, Tianhang Wang, Fan Lu 0001, Guang Chen 0001, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.7
2025 Spreeze: High-Throughput Parallel Reinforcement Learning Framework
abstract
The promotion of large-scale applications of reinforcement learning (RL) requires efficient training computation. While existing parallel RL frameworks encompass a variety of RL algorithms and parallelization techniques, the excessively burdensome communication frameworks hinder the attainment of the hardware's limit for final throughput and training effects on a single desktop. In this article, we propose Spreeze, a lightweight parallel framework for RL that efficiently utilizes a single desktop hardware resource to approach the throughput limit. We asynchronously parallelize the experience sampling, network update, performance evaluation, and visualization operations, and employ multiple efficient data transmission techniques to transfer various types of data between processes. The framework can automatically adjust the parallelization hyperparameters based on the computing ability of the hardware device in order to perform efficient large-batch updates. Based on the characteristics of the “Actor-Critic” RL algorithm, our framework uses dual GPUs to independently update the network of actors and critics in order to further improve throughput. Simulation results show that our framework can achieve up to 15,000 Hz experience sampling and 370,000 Hz network update frame rate using only a personal desktop computer, which is an order of magnitude higher than other mainstream parallel RL frameworks, resulting in a 73% reduction of training time. Our work on fully utilizing the hardware resources of a single desktop computer is fundamental to enabling efficient large-scale distributed RL training.
Guang Chen 0001, Zhijun Li 0001, Shangding Gu, Changjun Jiang 0002
IEEE Trans. Parallel Distributed Syst.2
2024 POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement Learning
abstract
Multi-constraint offline reinforcement learning (RL) promises to learn policies that satisfy both cumulative and state- wise costs from offline datasets. This arrangement provides an effective approach for the widespread appli-cation of RL in high-risk scenarios where both cumulative and state-wise costs need to be considered simulta-neously. However, previously constrained offline RL algorithms are primarily designed to handle single-constraint problems related to cumulative cost, which faces challenges when addressing multi-constraint tasks that involve both cumulative and state-wise costs. In this work, we pro-pose a novel Primal policy Optimization with Conservative Estimation algorithm (POCE) to address the problem of multi-constraint offline RL. Concretely, we reframe the ob-jective of multi-constraint offline RL by introducing the con-cept of Maximum Markov Decision Processes (MMDP). Subsequently, we present a primal policy optimization al-gorithm to confront the multi-constraint problems, which improves the stability and convergence speed of model training. Furthermore, we propose a conditional Bell-man operator to estimate cumulative and state-wise Q-values, reducing the extrapolation error caused by out-of-distribution (OOD) actions. Finally, extensive experiments demonstrate that the POCE algorithm achieves competitive performance across multiple experimental tasks, particu-larly outperforming baseline algorithms in terms of safety. Our code is available at github. POCE.
Jiayi Guan, Li Shen 0008, Ao Zhou 0005, Lusong Li, Han Hu 0003, Xiaodong He 0001, Guang Chen 0001, Changjun Jiang 0002
CVPR7
2024 MAP: MAsk-Pruning for Source-Free Model Intellectual Property Protection
abstract
Deep learning has achieved remarkable progress in various applications, heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment of models in authorized data domains, i.e., making models exclusive to certain target domains. Previous methods necessitate concurrent access to source training data and target unauthorized data when performing IP protection, making them risky and inefficient for decentralized private data. In this paper, we target a practical setting where only a well-trained source model is available and investigate how we can realize IP protection. To achieve this, we propose a novel MAsk Pruning (MAP) framework. MAP stems from an intuitive hypothesis, i.e., there are target-related parameters in a well-trained model, locating and pruning them is the key to IP protection. Technically, MAP freezes the source model and learns a target-specific binary mask to prevent unauthorized data usage while minimizing performance degradation on authorized data. Moreover, we introduce a new metric aimed at achieving a better balance between source and target performance degradation. To verify the effectiveness and versatility, we have evaluated MAP in a variety of scenarios, including vanilla source-available, practical source-free, and challenging data-free. Extensive experiments indicate that MAP yields new state-of-the-art performance. Code will be available at https://github.com/ispc-lab/MAP.
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR7
2024 LEAD: Learning Decomposition for Source-free Universal Domain Adaptation
abstract
Universal Domain Adaptation (UniDA) targets knowledge transfer in the presence of both covariate and label shifts. Recently, Source-free Universal Domain Adaptation (SF-UniDA) has emerged to achieve UniDA without access to source data, which tends to be more practical due to data protection policies. The main challenge lies in determining whether covariate-shifted samples belong to target-private unknown categories. Existing methods tackle this either through hand-crafted thresholding or by developing time-consuming iterative clustering strategies. In this paper, we propose a new idea of LEArning Decomposition (LEAD), which decouples features into source-known and-unknown components to identify target-private data. Technically, LEAD initially leverages the or-thogonal decomposition analysis for feature decomposition. Then, LEAD builds instance-level decision boundaries to adaptively identify target-private data. Extensive experiments across various UniDA scenarios have demonstrated the effectiveness and superiority of LEAD. Notably, in the OPDA scenario on VisDA dataset, LEAD outperforms GLC by 3.5% overall H-score and reduces 75% time to derive pseudo-labeling decision boundaries. Besides, LEAD is also appealing in that it is complementary to most existing methods. The code is available at https://github.com/ispc-lab/LEAD.
Sanqing Qu, Tianpei Zou, Lianghua He, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR6
2024 LiDAR4D: Dynamic Neural Fields for Novel Space-Time View LiDAR Synthesis
abstract
Although neural radiance fields (NeRFs) have achieved triumphs in image novel view synthesis (NVS), LiDAR NVS remains largely unexplored. Previous LiDAR NVS methods employ a simple shift from image NVS methods while ignoring the dynamic nature and the large-scale reconstruction problem of LiDAR point clouds. In light of this, we propose LiDAR4D, a differentiable LiDAR-only frameworkfor novel space-time LiDAR view synthesis. In consideration of the sparsity and large-scale characteristics, we design a 4D hybrid representation combined with multi-planar and grid features to achieve effective reconstruction in a coarse-to-fine manner. Furthermore, we introduce geometric constraints derivedfrom point clouds to improve temporal consistency. For the realistic synthesis of LiDAR point clouds, we incorporate the global optimization of ray-drop prob-ability to preserve cross-region patterns. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate the superiority of our method in accomplishing geometry-aware and time-consistent dynamic reconstruction. Codes are available at https://github.com/ispc-labILiDAR4D.
Zehan Zheng, Fan Lu 0001, Weiyi Xue, Guang Chen 0001, Changjun Jiang 0002
CVPR4
2024 Embracing Events and Frames with Hierarchical Feature Refinement Network for Object Detection
Hu Cao, Zehua Zhang 0009, Yan Xia 0003, Jiahao Xia 0001, Guang Chen 0001, Alois C. Knoll
ECCV (84)6
2024 HGL: Hierarchical Geometry Learning for Test-Time Adaptation in 3D Point Cloud Segmentation
Tianpei Zou, Sanqing Qu, Zhijun Li 0001, Alois C. Knoll, Lianghua He, Guang Chen 0001, Changjun Jiang 0002
ECCV (55)6
2024 ELF-UA: Efficient Label-Free User Adaptation in Gaze Estimation
Yong Wu 0007, Yang Wang 0003, Sanqing Qu, Zhijun Li 0001, Guang Chen 0001
IJCAI5
2024 Lightweight Fisheye Object Detection Network with Transformer-based Feature Enhancement for Autonomous Driving
abstract
Fisheye cameras, offering a wide field of view (FOV) of 360◦, are extensively employed for surround-view perception in autonomous driving. Compared with the object detection on the standard images, it lacks studies for fisheye images. Moreover, efficient perception is crucial for autonomous vehicles with limited computational capability. In this work, we introduce a lightweight fisheye object detection network with transformer-based feature enhancement for autonomous driving. Specifically, we leverage ShuffleNet V2 as a feature extraction network to reduce computation complexity and develop a transformer-based feature enhancement module (TFEM) to integrate multi-level features. Notably, we observe that data augmentation methods like mix-up and mosaic, effective on standard images, do not yield positive results on fisheye images. The results on the WoodScape dataset demonstrate that our method can achieve better performance with fewer parameters and floating-point operations per second (FLOPs). Extending our evaluation to the Microsoft Common Objects in Context (MS COCO) dataset shows that the proposed method has excellent generalization capability.
Hu Cao, Yinlong Liu, Guang Chen 0001, Alois C. Knoll
IROS5
2024 3D Object Detection via Stereo Pyramid Transformers with Rich Semantic Feature Fusion
abstract
Camera-based 3D object detectors, prized for their broader applicability and cost-effectiveness compared to LiDAR sensors, still grapple with the inherently ill-posed nature of depth extraction from images. In this work, we present a novel approach that employs a transformer-based backbone and a fused geometry volume to bolster feature richness and elevate detection accuracy. Firstly, we propose the Stereo Pyramid Transformer backbone to extract features from stereo images, which can capture global information and establish cross-image semantic connections. Then, to tackle the challenge posed by small baseline binocular cameras, we propose to fuse stereo geometry volumes constructed by Stereo Plane Sweeping Volume (SPSV), Monocular Semantic Volume (MSV), and Lifted Volume (LV) to create finely detailed feature volumes. Through extensive experiments on both the KITTI and our datasets, our approach not only surpasses all existing transformer-based stereo 3D detection methods but also marks a significant milestone by achieving comparable performance with the leading CNN-based 3D detectors for the first time.
Rongqi Gu, Chu Yang, Yaohan Lu, Peigen Liu, Guang Chen 0001
IROS6
2024 PCDepth: Pattern-based Complementary Learning for Monocular Depth Estimation by Best of Both Worlds
abstract
Event cameras can record scene dynamics with high temporal resolution, providing rich scene details for monocular depth estimation (MDE) even at low-level illumination. Therefore, existing complementary learning approaches for MDE fuse intensity information from images and scene details from event data for better scene understanding. However, most methods directly fuse two modalities at pixel level, ignoring that the attractive complementarity mainly impacts high-level patterns that only occupy a few pixels. For example, event data is likely to complement contours of scene objects. In this paper, we discretize the scene into a set of high-level patterns to explore the complementarity and propose a Pattern-based Complementary learning architecture for monocular Depth estimation (PCDepth). Concretely, PCDepth comprises two primary components: a complementary visual representation learning module for discretizing the scene into high-level patterns and integrating complementary patterns across modalities and a refined depth estimator aimed at scene reconstruction and depth prediction while maintaining an efficiency-accuracy balance. Through pattern-based complementary learning, PCDepth fully exploits two modalities and achieves more accurate predictions than existing methods, especially in challenging nighttime scenarios. Extensive experiments on MVSEC and DSEC datasets verify the effectiveness and superiority of our PCDepth. Remarkably, compared with state-of-the-art, PCDepth achieves a 37.9% improvement in accuracy in MVSEC nighttime scenarios.
Sanqing Qu, Fan Lu 0001, Zongtao Bu, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001
IROS7
2024 RCDN: Towards Robust Camera-Insensitivity Collaborative Perception via Dynamic Feature-based 3D Neural Modeling
abstract
Collaborative perception is dedicated to tackling the constraints of single-agent perception, such as occlusions, based on the multiple agents' multi-view sensor inputs. However, most existing works assume an ideal condition that all agents' multi-view cameras are continuously available. In reality, cameras may be highly noisy, obscured or even failed during the collaboration. In this work, we introduce a new robust camera-insensitivity problem: how to overcome the issues caused by the failed camera perspectives, while stabilizing high collaborative performance with low calibration cost? To address above problems, we propose RCDN, a Robust Camera-insensitivity collaborative perception with a novel Dynamic feature-based 3D Neural modeling mechanism. The key intuition of RCDN is to construct collaborative neural rendering field representations to recover failed perceptual messages sent by multiple agents. To better model collaborative neural rendering field, RCDN first establishes a geometry BEV feature based time-invariant static field with other agents via fast hash grid modeling. Based on the static background field, the proposed time-varying dynamic field can model corresponding motion vector for foregrounds with appropriate positions. To validate RCDN, we create OPV2V-N, a new large-scale dataset with manual labelling under different camera failed scenarios. Extensive experiments conducted on OPV2V-N show that RCDN can be ported to other baselines and improve their robustness in extreme camera-insensitivity setting. Our code and datasets will be available soon.
Tianhang Wang, Fan Lu 0001, Zehan Zheng, Zhijun Li 0001, Guang Chen 0001, Changjun Jiang 0002
NeurIPS5
2024 GeoNLF: Geometry guided Pose-Free Neural LiDAR Fields
abstract
Although recent efforts have extended Neural Radiance Field (NeRF) into LiDAR point cloud synthesis, the majority of existing works exhibit a strong dependence on precomputed poses. However, point cloud registration methods struggle to achieve precise global pose estimation, whereas previous pose-free NeRFs overlook geometric consistency in global reconstruction. In light of this, we explore the geometric insights of point clouds, which provide explicit registration priors for reconstruction. Based on this, we propose Geometry guided Neural LiDAR Fields (GeoNLF), a hybrid framework performing alternately global neural reconstruction and pure geometric pose optimization. Furthermore, NeRFs tend to overfit individual frames and easily get stuck in local minima under sparse-view inputs. To tackle this issue, we develop a selective-reweighting strategy and introduce geometric constraints for robust optimization. Extensive experiments on NuScenes and KITTI-360 datasets demonstrate the superiority of GeoNLF in both novel view synthesis and multi-view registration of low-frequency large-scale point clouds.
Weiyi Xue, Zehan Zheng, Fan Lu 0001, Haiyun Wei, Guang Chen 0001, Changjun Jiang 0002
NeurIPS5
2024 HDMNet: A Hierarchical Matching Network with Double Attention for Large-scale Outdoor LiDAR Point Cloud Registration
abstract
Outdoor LiDAR point clouds are typically large-scale and complexly distributed. To achieve efficient and accurate registration, emphasizing the similarity among local regions and prioritizing global local-to-local matching is of utmost importance, subsequent to which accuracy can be enhanced through cost-effective fine registration. In this paper, a novel hierarchical neural network with double attention named HDMNet is proposed for large-scale outdoor LiDAR point cloud registration. Specifically, A novel feature consistency enhanced double-soft matching network is introduced to achieve two-stage matching with high flexibility while enlarging the receptive field with high efficiency in a patch-to-patch manner, which significantly improves the registration performance. Moreover, in order to further utilize the sparse matching information from deeper layer, we develop a novel trainable embedding mask to incorporate the confidence scores of correspondences obtained from pose estimation of deeper layer, eliminating additional computations. The high-confidence keypoints in the sparser point cloud of the deeper layer correspond to a high-confidence spatial neighborhood region in shallower layer, which will receive more attention, while the features of non-key regions will be masked. Extensive experiments are conducted on two large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HDMNet.
Weiyi Xue, Fan Lu 0001, Guang Chen 0001
WACV3
2024 A Review of Safe Reinforcement Learning: Methods, Theories, and Applications
abstract
Reinforcement Learning (RL) has achieved tremendous success in many complex decision-making tasks. However, safety concerns are raised during deploying RL in real-world applications, leading to a growing demand for safe RL algorithms, such as in autonomous driving and robotics scenarios. While safe control has a long history, the study of safe RL algorithms is still in the early stages. To establish a good foundation for future safe RL research, in this paper, we provide a review of safe RL from the perspectives of methods, theories, and applications. First, we review the progress of safe RL from five dimensions and come up with five crucial problems for safe RL being deployed in real-world applications, coined as "2H3W". Second, we analyze the algorithm and theory progress from the perspectives of answering the "2H3W" problems. Particularly, the sample complexity of safe RL algorithms is reviewed and discussed, followed by an introduction to the applications and benchmarks of safe RL algorithms. Finally, we open the discussion of the challenging problems in safe RL, hoping to inspire future research on this thread. To advance the study of safe RL algorithms, we release an open-sourced repository containing major safe RL algorithms at the link.
Shangding Gu, Long Yang 0004, Yali Du 0001, Guang Chen 0001, Florian Walter, Jun Wang 0012, Alois C. Knoll
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Absolute Pose Estimation With a Known Direction by Motion Decoupling
abstract
This paper develops an extremely robust solution for absolute pose estimation with known prior gravity direction by motion decoupling. Absolute pose estimation is a fundamental problem in computer vision, and recently the prior known vertical direction is commonly applied to help solve the pose estimation problem. In this paper, we explore the geometrical constraints of the absolute pose estimation with a known direction. We find that the rigid pose can be decoupled with the help of the known direction. Thereby, absolute pose estimation algorithms, which decouple rigid motion, are proposed. Notably, in real applications, there may be imperfect inputs, i.e., outliers, due to incorrect 2D-3D matches. Unfortunately, these outliers may lead to unacceptable results. To suppress the outliers, the decoupled absolute pose estimation problem is solved by branch-and-bound algorithm and globally voting, which can provide the optimal solution with provable guarantees. Moreover, in extreme case, the proposed method can solve absolute pose estimation problem without knowing the 2D-3D correspondences, which is also known as simultaneous camera pose correspondence estimation. To demonstrate the feasibility and the superiority of the proposed methods, comprehensive comparison experiment are conduced. The source code is available at https://github.com/Liu-Yinlong/algorithm-for-PnP-with-known-vertical-direction.
Yinlong Liu, Guang Chen 0001, Alois C. Knoll
IEEE Trans. Circuits Syst. Video Technol.2
2024 TTAGaze: Self-Supervised Test-Time Adaptation for Personalized Gaze Estimation
abstract
In this paper, we address the problem of personalized gaze estimation. Due to the anatomical differences between individuals, current personalized gaze models often rely on fine-tuning or fully-supervised methods with labeled calibration samples, which may not be practical in real-world applications. To tackle this limitation, we propose an approach called Self-Supervised Test-Time Adaptation for Personalized Gaze Estimation (TTAGaze), which enables adaptation with small unlabeled data at test time. Our goal is to develop a gaze estimation model specifically adapted to a target person using only a few unlabeled images. We call this setting as unsupervised few-shot personalized adaptation in gaze estimation, which is more aligned with real-world scenarios compared to existing approaches. Additionally, Our approach leverages self-supervised learning and meta-learning. The model consists of the main task (gaze estimation) and a self-supervised auxiliary task. During training, the two task are trained using a coupled method. At test time, adaptation is achieved by optimizing the self-supervised loss adapted to an unseen person with a few unlabeled data. The model parameters are learned via model-agnostic meta-learning (MAML) to facilitate effective unsupervised few-shot personalized adaptation in gaze estimation. Experimental results demonstrate that the proposed method outperforms alternative approaches on several widely-used benchmark datasets.
Yong Wu 0007, Guang Chen 0001, Linwei Ye, Yuanning Jia, Zhi Liu 0003, Yang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2024 Hybrid Residual Multiexpert Reinforcement Learning for Spatial Scheduling of High-Density Parking Lots
abstract
Industries, such as manufacturing, are accelerating their embrace of the metaverse to achieve higher productivity, especially in complex industrial scheduling. In view of the growing parking challenges in large cities, high-density vehicle spatial scheduling is one of the potential solutions. Stack-based parking lots utilize parking robots to densely park vehicles in the vertical stacks like container stacking, which greatly reduces the aisle area in the parking lot, but requires complex scheduling algorithms to park and take out the vehicles. The existing high-density parking (HDP) scheduling algorithms are mainly heuristic methods, which only contain simple logic and are difficult to utilize information effectively. We propose a hybrid residual multiexpert (HIRE) reinforcement learning (RL) approach, a method for interactive learning in the digital industrial metaverse, which efficiently solves the HDP batch space scheduling problem. In our proposed framework, each heuristic scheduling method is considered as an expert. The neural network trained by RL assigns the expert strategy according to the current parking lot state. Furthermore, to avoid being limited by heuristic expert performance, the proposed hierarchical network framework also sets up a residual output channel. Experiments show that our proposed algorithm outperforms various advanced heuristic methods and the end-to-end RL method in the number of vehicle maneuvers, and has good robustness to the parking lot size and the estimation accuracy of vehicle exit time. We believe that the proposed HIRE RL method can be effectively and conveniently applied to practical application scenarios, which can be regarded as a key step for RL to enter the practical application stage of the industrial metaverse.
Guang Chen 0001, Zhijun Li 0001, Wei He 0001, Shangding Gu, Alois C. Knoll, Changjun Jiang 0002
IEEE Trans. Cybern.2
2024 Safe Multiagent Learning With Soft Constrained Policy Optimization in Real Robot Control
abstract
Due to a lack of safety considerations, a wide range of multiagent reinforcement learning (MARL) applications are limited in real-world environments. Thus, ensuring MARL safety is essential and urgent in the domain. However, merely a few studies consider the safe MARL problem, and the investigation of real-world applications using safe MARL algorithms still needs to be improved. To fill this gap, we provide a framework with soft constrained policy optimization, in which we develop practical algorithms to address the problem in a cooperative game setting. First, the problem formulation of safe MARL is introduced. Second, the safe policy optimization of safe MARL algorithms based on soft constrained optimization is analyzed, and we further propose a safe learning framework for safe MARL. The framework can be plugged into MARL algorithms without manually fine-tuning safety bounds. Third, we investigate the sim-to-real problems, and conduct simulation and real-world experiments to evaluate the effectiveness of our algorithms. Finally, the comprehensive experimental results indicate that our method has significant benefits regarding the balance between reward and safety performance and outperforms several strong baselines.
Shangding Gu, Dianye Huang, Muning Wen, Guang Chen 0001, Alois C. Knoll
IEEE Trans. Ind. Informatics4
2024 SDPT: Semantic-Aware Dimension-Pooling Transformer for Image Segmentation
abstract
Image segmentation plays a critical role in autonomous driving by providing vehicles with a detailed and accurate understanding of their surroundings. Transformers have recently shown encouraging results in image segmentation. However, transformer-based models are challenging to strike a better balance between performance and efficiency. The computational complexity of the transformer-based models is quadratic with the number of inputs, which severely hinders their application in dense prediction tasks. In this paper, we present the semantic-aware dimension-pooling transformer (SDPT) to mitigate the conflict between accuracy and efficiency. The proposed model comprises an efficient transformer encoder for generating hierarchical features and a semantic-balanced decoder for predicting semantic masks. In the encoder, a dimension-pooling mechanism is used in the multi-head self-attention (MHSA) to reduce the computational cost, and a parallel depth-wise convolution is used to capture local semantics. Simultaneously, we further apply this dimension-pooling attention (DPA) to the decoder as a refinement module to integrate multi-level features. With such a simple yet powerful encoder-decoder framework, we empirically demonstrate that the proposed SDPT achieves excellent performance and efficiency on various popular benchmarks, including ADE20K, Cityscapes, and COCO-Stuff. For example, our SDPT achieves 48.6$\%$mIOU on the ADE20K dataset, which outperforms the current methods with fewer computational costs. The codes can be found at https://github.com/HuCaoFighting/SDPT.
Hu Cao, Guang Chen 0001, Hengshuang Zhao, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.2
2023 Modality-Agnostic Debiasing for Single Domain Generalization
abstract
Deep neural networks (DNNs) usually fail to generalize well to outside of distribution (OOD) data, especially in the extreme case of single domain generalization (single-DG) that transfers DNNs from single domain to multiple unseen domains. Existing single-DG techniques commonly devise various data-augmentation algorithms, and remould the multi-source domain generalization methodology to learn domain-generalized (semantic) features. Nevertheless, these methods are typically modality-specific, thereby being only applicable to one single modality (e.g., image). In contrast, we target a versatile Modality-Agnostic Debiasing (MAD) framework for single-DG, that enables generalization for different modalities. Technically, MAD introduces a novel two-branch classifier: a biased-branch encourages the classifier to identify the domain-specific (superficial) features, and a general-branch captures domain-generalized features based on the knowledge from biased-branch. Our MAD is appealing in view that it is pluggable to most single-DG models. We validate the superiority of our MAD in a variety of single-DG scenarios with different modalities, including recognition on 1D texts, 2D images, 3D point clouds, and semantic segmentation on 2D images. More remarkably, for recognition on 3D point clouds and semantic segmentation on 2D images, MAD improves DSU by 2.82% and 1.5% in accuracy and mIOU.
Sanqing Qu, Yingwei Pan, Guang Chen 0001, Ting Yao 0003, Changjun Jiang 0002, Tao Mei 0001
CVPR3
2023 Upcycling Models Under Domain and Category Shift
abstract
Deep neural networks (DNNs) often perform poorly in the presence of domain shift and category shift. How to upcycle DNNs and adapt them to the target task remains an important open problem. Unsupervised Domain Adaptation (UDA), especially recently proposed Source-free Domain Adaptation (SFDA), has become a promising technology to address this issue. Nevertheless, existing SFDA methods require that the source domain and target domain share the same label space, consequently being only applicable to the vanilla closed-set setting. In this paper, we take one step further and explore the Source-free Universal Domain Adaptation (SF-UniDA). The goal is to identify “known” data samples under both domain and category shift, and reject those “unknown” data samples (not present in source classes), with only the knowledge from standard pre-trained source model. To this end, we introduce an innovative global and local clustering learning technique (GLC). Specifically, we design a novel, adaptive one-vs-all global clustering algorithm to achieve the distinction across different target classes and introduce a local k-NN clustering strategy to alleviate negative transfer. We examine the superiority of our GLC on multiple benchmarks with different category shift scenarios, including partial-set, open-set, and open-partial-set DA. Remarkably, in the most challenging open-partial-set DA scenario, GLC outperforms UMAD by 14.8% on the VisDA benchmark. The code is available at https://github.com/ispc-lab/GLC.
Sanqing Qu, Tianpei Zou, Florian Röhrbein, Cewu Lu, Guang Chen 0001, Dacheng Tao, Changjun Jiang 0002
CVPR5
2023 NeuralPCI: Spatio-Temporal Neural Field for 3D Point Cloud Multi-Frame Non-Linear Interpolation
abstract
In recent years, there has been a significant increase in focus on the interpolation task of computer vision. Despite the tremendous advancement of video interpolation, point cloud interpolation remains insufficiently explored. Meanwhile, the existence of numerous nonlinear large motions in real-world scenarios makes the point cloud interpolation task more challenging. In light of these issues, we present NeuralPCI: an end-to-end 4D spatiotemporal Neural field for 3D Point Cloud Interpolation, which implicitly integrates multi-frame information to handle nonlinear large motions for both indoor and outdoor scenarios. Furthermore, we construct a new multi-frame point cloud interpolation dataset called NL-Drive for large nonlinear motions in autonomous driving scenes to better demonstrate the superiority of our method. Ultimately, NeuralPCI achieves state-of-the-art performance on both DHB (Dynamic Human Bodies) and NL-Drive datasets. Beyond the interpolation task, our method can be naturally extended to point cloud extrapolation, morphing, and auto-labeling, which indicates its substantial potential in other domains. Codes are available at https://github.com/ispc-lab/NeuralPCI.
Zehan Zheng, Danni Wu, Ruisi Lu, Fan Lu 0001, Guang Chen 0001, Changjun Jiang 0002
CVPR5
2023 TMA: Temporal Motion Aggregation for Event-based Optical Flow
abstract
Event cameras have the ability to record continuous and detailed trajectories of objects with high temporal resolution, thereby providing intuitive motion cues for optical flow estimation. Nevertheless, most existing learning-based approaches for event optical flow estimation directly remould the paradigm of conventional images by representing the consecutive event stream as static frames, ignoring the inherent temporal continuity of event data. In this paper, we argue that temporal continuity is a vital element of event-based optical flow and propose a novel Temporal Motion Aggregation (TMA) approach to unlock its potential. Technically, TMA comprises three components: an event splitting strategy to incorporate intermediate motion information underlying the temporal context, a linear lookup strategy to align temporally fine-grained motion features and a novel motion pattern aggregation module to emphasize consistent patterns for motion feature enhancement. By incorporating temporally fine-grained motion information, TMA can derive better flow estimates than existing methods at early stages, which not only enables TMA to obtain more accurate final predictions, but also greatly reduces the demand for a number of refinements. Extensive experiments on DSEC-Flow and MVSEC datasets verify the effectiveness and superiority of our TMA. Remarkably, compared to E-RAFT, TMA achieves a 6% improvement in accuracy and a 40% reduction in inference time on DSEC-Flow. Code will be available at https://github.com/ispc-lab/TMA.
Guang Chen 0001, Sanqing Qu, Yanping Zhang 0007, Zhijun Li 0001, Alois C. Knoll, Changjun Jiang 0002
ICCV2
2023 Urban Radiance Field Representation with Deformable Neural Mesh Primitives
abstract
Neural Radiance Fields (NeRFs) have achieved great success in the past few years. However, most current methods still require intensive resources due to ray marching-based rendering. To construct urban-level radiance fields efficiently, we design Deformable Neural Mesh Primitive (DNMP), and propose to parameterize the entire scene with such primitives. The DNMP is a flexible and compact neural variant of classic mesh representation, which enjoys both the efficiency of rasterization-based rendering and the powerful neural representation capability for photo-realistic image synthesis. Specifically, a DNMP consists of a set of connected deformable mesh vertices with paired vertex features to parameterize the geometry and radiance information of a local area. To constrain the degree of freedom for optimization and lower the storage budgets, we enforce the shape of each primitive to be decoded from a relatively low-dimensional latent space. The rendering colors are decoded from the vertex features (interpolated with rasterization) by a view-dependent MLP. The DNMP provides a new paradigm for urban-level scene representation with appealing properties: (1) High-quality rendering. Our method achieves leading performance for novel view synthesis in urban scenarios. (2) Low computational costs. Our representation enables fast rendering (2.07 ms/1K pixels) and low peak memory usage (110 MB/1K pixels). We also present a lightweight version that can run 33× faster than vanilla NeRFs, which is comparable to the highly-optimized Instant-NGP. Project page: https://dnmp.github.io/.
Fan Lu 0001, Guang Chen 0001, Hongsheng Li 0001, Kwan-Yee Lin, Changjun Jiang 0002
ICCV3
2023 UMC: A Unified Bandwidth-efficient and Multi-resolution based Collaborative Perception Framework
abstract
Multi-agent collaborative perception (MCP) has recently attracted much attention. It includes three key processes: communication for sharing, collaboration for integration, and reconstruction for different downstream tasks. Existing methods pursue designing the collaboration process alone, ignoring their intrinsic interactions and resulting in suboptimal performance. In contrast, we aim to propose a Unified Collaborative perception framework named UMC, optimizing the communication, collaboration, and reconstruction processes with the Multi-resolution technique. The communication introduces a novel trainable multi-resolution and selective-region (MRSR) mechanism, achieving higher quality and lower bandwidth. Then, a graph-based collaboration is proposed, conducting on each resolution to adapt the MRSR. Finally, the reconstruction integrates the multi-resolution collaborative features for downstream tasks. Since the general metric can not reflect the performance enhancement brought by MCP systematically, we introduce a brand-new evaluation metric that evaluates the MCP from different perspectives. To verify our algorithm, we conducted experiments on the V2X-Sim and OPV2V datasets. Our quantitative and qualitative experiments prove that the proposed UMC outperforms the state-of-the-art collaborative perception approaches.
Tianhang Wang, Guang Chen 0001, Zhengfa Liu, Alois C. Knoll, Changjun Jiang 0002
ICCV2
2023 Deep High-Level Policy Model Predictive Contour Control for Autonomous Racing
abstract
For autonomous racing vehicles, it is very important to perform fast overtaking maneuvers in a safe manner. To this aim, we design a high-level policy for model predictive contouring control (High-MPCC). Two decision variables will be obtained from the high-level policy using the observations of the ego and lead vehicle. Based on these two decision variables, the ego vehicle will determine whether to perform an overtaking maneuver when there is an opponent vehicle. These two decision variables will also be used as the control parameters of the lower model predictive contouring control (MPCC) controller to adjust the cost function of the MPCC controller. Then an overtaking maneuver or maintaining a safe distance from the lead vehicle is performed. To improve the real-time performance of the algorithm, the combination of deep neural network and policy search is introduced to form deep High-MPCC. Simulation results show that the proposed High-MPCC controller performs better than traditional MPCC, especially in the process of overtaking.
Minghao Zeng, Guang Chen 0001, Alois C. Knoll
IV3
2023 VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement Learning
abstract
Offline safe reinforcement learning (RL) algorithms promise to learn policies that satisfy safety constraints directly in offline datasets without interacting with the environment. This arrangement is particularly important in scenarios with high sampling costs and potential dangers, such as autonomous driving and robotics. However, the influence of safety constraints and out-of-distribution (OOD) actions have made it challenging for previous methods to achieve high reward returns while ensuring safety. In this work, we propose a Variational Optimization with Conservative Eestimation algorithm (VOCE) to solve the problem of optimizing safety policies in the offline dataset. Concretely, we reframe the problem of offline safe RL using probabilistic inference, which introduces variational distributions to make the optimization of policies more flexible. Subsequently, we utilize pessimistic estimation methods to estimate the Q-value of cost and reward, which mitigates the extrapolation errors induced by OOD actions. Finally, extensive experiments demonstrate that the VOCE algorithm achieves competitive performance across multiple experimental tasks, particularly outperforming state-of-the-art algorithms in terms of safety.
Jiayi Guan, Guang Chen 0001, Jiaming Ji, Long Yang 0004, Ao Zhou 0005, Zhijun Li 0001, Changjun Jiang 0002
NeurIPS2
2023 Real-time semantic segmentation in traffic scene using Cross Stage Partial-based encoder-decoder network
Liguo Zhou, Guang Chen 0001, Ruining Wang, Alois C. Knoll
Eng. Appl. Artif. Intell.2
2023 Residual encoding framework to compress DNN parameters for fast transfer
Liguo Zhou, Rui Song 0007, Guang Chen 0001, Andreas Festag, Alois C. Knoll
Knowl. Based Syst.3
2023 Sparse-to-Dense Matching Network for Large-Scale LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Previous learning-based methods for LiDAR point cloud registration can be categorized into two schemes: dense-to-dense matching methods and sparse-to-sparse matching methods. However, for large-scale outdoor LiDAR point clouds, solving dense point correspondences is time-consuming, whereas sparse keypoint matching easily suffers from keypoint detection error. In this paper, we propose SDMNet, a novel Sparse-to-Dense Matching Network for large-scale outdoor LiDAR point cloud registration. Specifically, SDMNet performs registration in two sequential stages: sparse matching stage and local-dense matching stage. In the sparse matching stage, we sample a set of sparse points from the source point cloud and then match them to the dense target point cloud using a spatial consistency enhanced soft matching network and a robust outlier rejection module. Furthermore, a novel neighborhood matching module is developed to incorporate local neighborhood consensus, significantly improving performance. The local-dense matching stage is followed for fine-grained performance, where dense correspondences are efficiently obtained by performing point matching in local spatial neighborhoods of high-confidence sparse correspondences. Extensive experiments on three large-scale outdoor LiDAR point cloud datasets demonstrate that the proposed SDMNet achieves state-of-the-art performance with high efficiency.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Yibing Zhan, Zhijun Li 0001, Dacheng Tao, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 HRegNet: A Hierarchical Network for Efficient and Accurate Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale outdoor LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching, and novel similarity features are designed to incorporate them into the correspondence network, which significantly improves the registration performance. In addition, we design a consistency propagation strategy to effectively incorporate spatial consistency into the registration pipeline. The whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on three large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The source code of the proposed HRegNet is available at https://github.com/ispc-lab/HRegNet2.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Robot Policy Improvement With Natural Evolution Strategies for Stable Nonlinear Dynamical System
abstract
Robot learning through kinesthetic teaching is a promising way of cloning human behaviors, but it has its limits in the performance of complex tasks with small amounts of data, due to compounding errors. In order to improve the robustness and adaptability of imitation learning, a hierarchical learning strategy is proposed: low-level learning comprises only behavioral cloning with supervised learning, and high-level learning constitutes policy improvement. First, the Gaussian mixture model (GMM)-based dynamical system is formulated to encode a motion from the demonstration. We then derive the sufficient conditions of the GMM parameters that guarantee the global stability of the dynamical system from any initial state, using the Lyapunov stability theorem. Generally, imitation learning should reason about the motion well into the future for a wide range of tasks; it is significant to improve the adaptability of the learning method by policy improvement. Finally, a method based on exponential natural evolution strategies is proposed to optimize the parameters of the dynamical system associated with the stiffness of variable impedance control, in which the exploration noise is subject to stability conditions of the dynamical system in the exploration space, thus guaranteeing the global stability. Empirical evaluations are conducted on manipulators for different scenarios, including motion planning with obstacle avoidance and stiffness learning.
Yingbai Hu, Guang Chen 0001, Zhijun Li 0001, Alois C. Knoll
IEEE Trans. Cybern.2
2023 PSDC: A Prototype-Based Shared-Dummy Classifier Model for Open-Set Domain Adaptation
abstract
Open-set domain adaptation (OSDA) aims to achieve knowledge transfer in the presence of both domain shift and label shift, which assumes that there exist additional unknown target classes not presented in the source domain. To solve the OSDA problem, most existing methods introduce an additional unknown class to the source classifier and represent the unknown target instances as a whole. However, it is unreasonable to treat all unknown target instances as a group since these unknown instances typically consist of distinct categories and distributions. It is challenging to identify all unknown instances with only one additional class. In addition, most existing methods directly introduce marginal distribution alignment to alleviate distribution shift between the source and target domains, failing to learn discriminative class boundaries in the target domain since they ignore categorical discriminative information in the adaptation. To address these problems, in this article, we propose a novel prototype-based shared-dummy classifier (PSDC) model for the OSDA. Specifically, our PSDC introduces an auxiliary dummy classifier to calibrate the source classifier and simultaneously develops a weighted adaptation procedure to align class-wise prototypes for adaptation. We further design a pseudo-unknown learning algorithm to reduce the open-set risk. Extensive experiments on Office-31, Office-Home, and VisDA datasets show that the proposed PSDC can outperform existing methods and achieve the new state-of-the-art performance. The code will be made public.
Zhengfa Liu, Guang Chen 0001, Zhijun Li 0001, Yu Kang 0001, Sanqing Qu, Changjun Jiang 0002
IEEE Trans. Cybern.2
2023 Cross-Modal Integration and Transfer Learning Using Fuzzy Logic Techniques for Intelligent Upper Limb Prosthesis
abstract
The integration and interaction of proprioception and exteroception in the human multisensory network facilitate high-level cognitive functionalities, such as cross-modal integration, recognition, and imagination for accurate evaluation and comprehensive understanding of the multimodal world. In this article, we propose a novel cross-modal integration framework for the upper limb prosthesis based on type-2 fuzzy logic system (FLS), which can facilitate the high-level cognitive and dexterous manipulation of the prosthesis by combing human's surface electromyography (sEMG) with the computer vision. First, a transfer learning approach is proposed to improve the decoding of human's intent and enhance the effectiveness of skill transition. Then the sEMG signals and image information are integrated to jointly determine the grasp posture of the bionic hand based on fuzzy decision strategy. Fusing multisensory data and using cross-modal integration, the system is capable of crossmodally recognizing multimodal information. In order to realize the prosthesis automatically reaching the target position under the guidance of computer vision, an interval type-2 fuzzy logic controller considering uncertain dynamic parameters and disturbance is designed. Experiments are performed in some typical 3C assembly scenarios, and results show that our proposed strategy can obviously improve the accuracy of grasp posture selection and trajectory tracking effect, which brings more potential job opportunities with hope to amputees, and provides a promising approach toward robotic sensing and perception.
Jin Huang 0002, Zhijun Li 0001, Haisheng Xia, Guang Chen 0001, Qingsheng Meng
IEEE Trans. Fuzzy Syst.4
2023 RPP-Net: Rigid Constrained Point Cloud Prediction Network
abstract
Forecasting the future environment is an essential and fundamental capability of autonomous driving systems. In contrast to the widely studied video prediction, only a few literatures have explored LiDAR point cloud prediction. To achieve future point cloud generation, most existing methods are based on free-form 3D scene flow prediction. However, these simple scene flow prediction-based methods may cause distortions since the motions of 3D scenes can be seen as a combination of rigid (static background) and flow (dynamic foreground) motions. To address this issue, we propose a simple but effective rigid constrained point cloud prediction network named RPP-Net. The key component of the proposed method is the hybrid motion decoder, which generates motion masks and combines flow motions and rigid motions to produce hybrid motions without additional annotations and sensors. Besides, to reduce running time, we provide a hierarchical cell, which uses a feature encoder to extract deep features and an RPP-RNN module to capture temporal correlations across frames. To evaluate the effectiveness of our RPP-Net, we have conducted extensive experiments on both KITTI dataset and Argoverse dataset and the results show that our RPP-Net significantly outperforms existing methods and achieves a new state-of-the-art.
Tianpei Zou, Guang Chen 0001, Fan Lu 0001, Zhijun Li 0001, Sanqing Qu, Alois C. Knoll, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.2
2022 Unsupervised Domain Adaptation for Nighttime Aerial Tracking
abstract
Previous advances in object tracking mostly reported on favorable illumination circumstances while neglecting performance at nighttime, which significantly impeded the development of related aerial robot applications. This work instead develops a novel unsupervised domain adaptation framework for nighttime aerial tracking (named UDAT). Specifically, a unique object discovery approach is provided to generate training patches from raw nighttime tracking videos. To tackle the domain discrepancy, we employ a Transformer-based bridging layer post to the feature extractor to align image features from both domains. With a Transformer day/night feature discriminator, the day-time tracking model is adversarially trained to track at night. Moreover, we construct a pioneering benchmark namely NAT2021 for unsupervised domain adaptive night-time tracking, which comprises a test set of 180 manually annotated tracking sequences and a train set of over 276k unlabelled nighttime tracking frames. Exhaustive experiments demonstrate the robustness and domain adaptability of the proposed framework in nighttime aerial tracking. The code and benchmark are available at https://github.com/vision4robotics/UDAT.
Junjie Ye 0004, Changhong Fu 0001, Guangze Zheng 0001, Danda Pani Paudel, Guang Chen 0001
CVPR5
2022 BMD: A General Class-Balanced Multicentric Dynamic Prototype Strategy for Source-Free Domain Adaptation
Sanqing Qu, Guang Chen 0001, Jing Zhang 0037, Zhijun Li 0001, Wei He 0001, Dacheng Tao
ECCV (34)2
2022 Image Grid Recognition and Regression for Fast and Accurate Face Detection
abstract
CNN-based face detection methods have achieved significant progress in recent years. However, for high performance face detection, there are still many challenging problems, e.g., the speed-accuracy balance and the performance degradation in adverse conditions. In this paper, by taking advantage of the characteristic of CNN, we propose an effective anchor generation and bounding-box regression method that can make a good balance between speed and accuracy, and also work well in bad conditions. The classic structure of CNN produces pyramid-like feature maps due to the pooling or other downscale operations. According to the size of a feature map, we divide the image into grids. Each grid corresponds to a point in the feature map. We make the corresponding feature point responsible for identifying the content of the grid. If this grid area belongs to the face area, it is a natural anchor for face bounding-box regression. Since this anchor is square, it is reasonable to use it to predict the face bounding-box which is square-like. The points in the lower-level feature map correspond to smaller grids, which are dedicated to predicting the bounding-boxes of smaller faces. The points in the higher-level feature maps correspond to larger grids, which are responsible for predicting the bounding-boxes of larger faces. Hence our method can effectively detect multi-scale faces. With this effectiveness, our method can achieve a high detection accuracy using fewer parameters which leads to a fast detection speed. The experiments demonstrate the effectiveness of our method.
Liguo Zhou, Guang Chen 0001, Alois C. Knoll
ICPR2
2022 Learning Local Event-based Descriptor for Patch-based Stereo Matching
abstract
Stereo matching is an indispensable function that enables machine vision system to obtain depth information of its environment. However, most of existing algorithms rely on conventional camera, which follows the frame-based scheme and has several shortcomings: low dynamic range, low temporal resolution and high power consumption. To address these issues, we propose two novel patch-based stereo matching methods that exploit the output from a pair of neuromorphic vision sensors. Compared to frame-based camera, neuromorphic vision sensor has independent pixels that generates events at the time intensity changes occur. Based on this unique output, we first construct event representations and present a novel encoding method, which integrates with attention mechanism to encode rich spatial-temporal information of event streams. Then, we design efficient and accuracy networks and propose corresponding loss to train them, which are used to extract event-based descriptors from representations. Finally, the disparity maps are calculated based on local features and refined by two simple smoothing methods. Extensive experiments on the Multi Vehicle Stereo Event Camera Dataset demonstrate the effectiveness of our methods.
Peigen Liu, Guang Chen 0001, Zhijun Li 0001, Huajin Tang, Alois C. Knoll
ICRA2
2022 Fast and Accurate Face Detection using Feature Pyramid with Grid Anchors
abstract
CNN-based face detection methods have achieved significant progress in recent years. However, making a good balance between time cost and detection accuracy is still a challenging problem. Those methods which can reach a very high detection accuracy always have complicated networks and rely on expensive GPUs for inference, while those methods which have shallow networks and can run on common devices always lose detection accuracy to a large extent. In this paper, we propose an effective anchor generation and bounding-box regression method which can improve the detection accuracy by modifying the detection head of the popular detection networks. With this effectiveness, we can reduce the trainable weights of the network to speed up the inference while maintaining high accuracy. As a result, our method can get a better speed-accuracy balance. In our method, we divide the input image into grids according to the sizes of the pyramid-like feature maps produced by CNN. In training, those grids close to the center of the ground-truth bounding-boxes are selected as anchors. After training, the regression mapping from the anchors to the ground-truth bounding-boxes can be acquired by the exponential transformation we designed. Our method explicitly and strictly use feature maps in different levels to detect faces of different sizes. The higher-level feature maps and larger grid anchors are responsible for detecting larger faces, while the lower-level feature maps and smaller grid anchors are dedicated to detecting smaller faces. Therefore, our method is effective for detecting multi-scale faces. The experiments on both GPU and CPU demonstrate that our method is effective. Our source code is publicly available on https://github.com/zhouliguo/GAFace.
Liguo Zhou, Guang Chen 0001, Alois C. Knoll
IJCNN2
2022 Gaussian Process based Model Predictive Control for Overtaking Scenarios at Highway Curves
abstract
Model predictive control (MPC) is a commonly applied vehicle control technique, but its performance depends highly on how accurate the model captures the vehicle dynamics. It is disreputable hard to get a precise vehicle model in complex situations. The unmodeled dynamic will cause the uncertainty of the prediction which brings the risk while overtaking. To address this issue, Gaussian process (GP) regression is employed to acquire the unexplored discrepancy between the nominal vehicle model and the real vehicle dynamics which can result in a more accurate model. To achieve safe overtaking at highway curves, the constraint conditions are carefully designed. The implementation of GP-based MPC including approximate uncertainty propagation and safety constraints ensures that the ego vehicle overtakes the obstacles without collision. Simulation results show that GP-based MPC has a strong adaptability to different scenarios and outperforms MPC in overtaking control.
Yulin Zhai, Guang Chen 0001, Alois C. Knoll
IV3
2022 Globally Optimal Linear Model Fitting with Unit-Norm Constraint
Yinlong Liu, Manning Wang, Guang Chen 0001, Alois C. Knoll, Zhijian Song
Int. J. Comput. Vis.4
2022 Globally Optimal Vertical Direction Estimation in Atlanta World
abstract
In man-made environments, most of the objects and structures are organized in the form of orthogonal and parallel planes. These planes can be approximated by an Atlanta world assumption, in which the normals of planes can be represented by Atlanta frames. The Atlanta world assumption has one vertical frame and multiple horizontal frames. Conventionally, given a set of inputs such as surface normals, the Atlanta frame estimation problem can be solved by a branch-and-bound (BnB) algorithm. However, the runtime of the BnB algorithm will increase greatly when the dimensionality (i.e., the number of horizontal frames) increases. In this paper, we estimate only the vertical direction, instead of all Atlanta frames at once. Accordingly, we propose a vertical direction estimation method by considering the relationship between the vertical frame and horizontal frames. Concretely, our approach employs a BnB algorithm to search the vertical direction, thereby guaranteeing global optimality without requiring prior knowledge of the number of Atlanta frames. In order to guarantee convergence, four novel bounds are investigated, by mapping a 3D hemisphere to a 2D region. We verify the feasibility of the proposed method using various challenging synthetic and real-world data.
Yinlong Liu, Guang Chen 0001, Alois C. Knoll
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Neuromorphic Vision-Based Fall Localization in Event Streams With Temporal-Spatial Attention Weighted Network
abstract
Falling down is a serious problem for health and has become one of the major etiologies of accidental death for the elderly living alone. In recent years, many efforts have been paid to fall recognition based on wearable sensors or standard vision sensors. However, the prior methods have the risk of privacy leaks, and almost all these methods are based on video clips, which cannot localize where the falls occurred in long videos. For these reasons, in this article, the bioinspired vision sensor-based falls temporal localization framework is proposed. The bioinspired vision sensors, such as dynamic and active-pixel vision sensor (DAVIS) camera applied in this work responds to pixels' brightness change, and each pixel works independently and asynchronously compared to the standard vision sensors. This property makes it have a very high dynamic range and privacy preserving. First, to better represent event data, compared with the typical constant temporal window mechanism, an adaptive temporal window conversion mechanism is developed. The temporal localization framework follows a proven proposal and classification paradigm. Second, for the high-efficient and recall proposal generation, different from the traditional sliding window scheme, the event temporal density as the actionness score is set and the 1D-watershed algorithm to generate proposals is applied. In addition, we combine the temporal and spatial attention mechanism with our feature extraction network to temporally model the falls. Finally, to evaluate the performance of our framework, 30 volunteers are recruited to join the simulated fall experiments. According to the results of experiments, our framework can realize precise falls temporal localization and achieve the state-of-the-art performance.
Guang Chen 0001, Sanqing Qu, Zhijun Li 0001, Jiaxuan Dong, Min Liu 0031, Jörg Conradt
IEEE Trans. Cybern.1
2022 NeuroIV: Neuromorphic Vision Meets Intelligent Vehicle Towards Safe Driving With a New Database and Baseline Evaluations
abstract
Neuromorphic vision sensors such as the Dynamic and Active-pixel Vision Sensor (DAVIS) using silicon retina are inspired by biological vision, they generate streams of asynchronous events to indicate local log-intensity brightness changes. Their properties of high temporal resolution, low-bandwidth, lightweight computation, and low-latency make them a good fit for many applications of motion perception in the intelligent vehicle. However, as a younger and smaller research field compared to classical computer vision, neuromorphic vision is rarely connected with the intelligent vehicle. For this purpose, we present three novel datasets recorded with DAVIS sensors and depth sensor for the distracted driving research and focus on driver drowsiness detection, driver gaze-zone recognition, and driver hand-gesture recognition. To facilitate the comparison with classical computer vision, we record the RGB, depth and infrared data with a depth sensor simultaneously. The total volume of this dataset has 27360 samples. To unlock the potential of neuromorphic vision on the intelligent vehicle, we utilize three popular event-encoding methods to convert asynchronous event slices to event-frames and adapt state-of-the-art convolutional architectures to extensively evaluate their performances on this dataset. Together with qualitative and quantitative results, this work provides a new database and baseline evaluations named NeuroIV in cross-cutting areas of neuromorphic vision and intelligent vehicle.
Guang Chen 0001, Fa Wang, Lin Hong, Jörg Conradt, Jieneng Chen, Zhenyan Zhang, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.1
2022 MoNet: Motion-Based Point Cloud Prediction Network
abstract
Predicting the future can significantly improve the safety of intelligent vehicles, which is a key component in autonomous driving. 3D point clouds can accurately model 3D information of surrounding environment and are crucial for intelligent vehicles to perceive the scene. Therefore, prediction of 3D point clouds has great significance for intelligent vehicles, which can be utilized for numerous further applications. However, due to point clouds are unordered and unstructured, point cloud prediction is challenging and has not been deeply explored in current literature. In this paper, we propose a novel motion-based neural network named MoNet. The key idea of the proposed MoNet is to integrate motion features between two consecutive point clouds into the prediction pipeline. The introduction of motion features enables the model to more accurately capture the variations of motion information across frames and thus make better predictions for future motion. In addition, content features are introduced to model the spatial content of individual point clouds. A recurrent neural network named MotionRNN is proposed to capture the temporal correlations of both features. Moreover, an attention-based motion align module is proposed to address the problem of missing motion features in the inference pipeline. Extensive experiments on two large-scale outdoor LiDAR point cloud datasets demonstrate the performance of the proposed MoNet. Moreover, we perform experiments on applications using the predicted point clouds and the results indicate the great application potential of the proposed method.
Fan Lu 0001, Guang Chen 0001, Zhijun Li 0001, Yinlong Liu, Sanqing Qu, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.2
2022 A Survey of the Four Pillars for Small Object Detection: Multiscale Representation, Contextual Information, Super-Resolution, and Region Proposal
abstract
Although great progress has been made in generic object detection by advanced deep learning techniques, detecting small objects from images is still a difficult and challenging problem in the field of computer vision due to the limited size, less appearance, and geometry cues, and the lack of large-scale datasets of small targets. Improving the performance of small object detection has a wider significance in many real-world applications, such as self-driving cars, unmanned aerial vehicles, and robotics. In this article, the first-ever survey of recent studies in deep learning-based small object detection is presented. Our review begins with a brief introduction of the four pillars for small object detection, includingmultiscale representation, contextual information, super-resolution, and region-proposal. Then, the collection of state-of-the-art datasets for small object detection is listed. The performance of different methods on these datasets is reported later. Moreover, the state-of-the-art small object detection networks are investigated along with a special focus on the differences and modifications to improve the detection performance comparing to generic object detection architectures. Finally, several promising directions and tasks for future work in small object detection are provided. Researchers can track up-to-date studies on this webpage available at:https://github.com/tjtum-chenlab/SmallObjectDetectionList.
Guang Chen 0001, Zhijun Li 0001, Zida Song, Yinlong Liu, Alois C. Knoll
IEEE Trans. Syst. Man Cybern. Syst.1
2021 PointINet: Point Cloud Frame Interpolation Network
abstract
LiDAR point cloud streams are usually sparse in time dimension, which is limited by hardware performance. Generally, the frame rates of mechanical LiDAR sensors are 10 to 20 Hz, which is much lower than other commonly used sensors like cameras. To overcome the temporal limitations of LiDAR sensors, a novel task named Point Cloud Frame Interpolation is studied in this paper. Given two consecutive point cloud frames, Point Cloud Frame Interpolation aims to generate intermediate frame(s) between them. To achieve that, we propose a novel framework, namely Point Cloud Frame Interpolation Network (PointINet). Based on the proposed method, the low frame rate point cloud streams can be upsampled to higher frame rates. We start by estimating bi-directional 3D scene flow between the two point clouds and then warp them to the given time step based on the 3D scene flow. To fuse the two warped frames and generate intermediate point cloud(s), we propose a novel learning-based points fusion module, which simultaneously takes two warped point clouds into consideration. We design both quantitative and qualitative experiments to evaluate the performance of the point cloud frame interpolation method and extensive experiments on two large scale outdoor LiDAR datasets demonstrate the effectiveness of the proposed PointINet. Our code is available at https://github.com/ispc-lab/PointINet.git.
Fan Lu 0001, Guang Chen 0001, Sanqing Qu, Zhijun Li 0001, Yinlong Liu, Alois C. Knoll
AAAI2
2021 HRegNet: A Hierarchical Network for Large-scale Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale out-door LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching and novel similarity features are designed to in-corporate them into the correspondence network, which significantly improves the registration performance. Besides, the whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on two large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The project website is https://ispc-group.github.io/hregnet.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu
ICCV2
2021 Residual Squeeze-and-Excitation Network with Multi-scale Spatial Pyramid Module for Fast Robotic Grasping Detection
abstract
This paper proposes an efficient, fully convolutional neural network to generate robotic grasps by using 300×300 depth images as input. Specifically, a residual squeeze-and-excitation network (RSEN) is introduced for deep feature extraction. Following the RSEN block, a multi-scale spatial pyramid module (MSSPM) is developed to obtain multi-scale contextual information. The outputs of each RSEN block and MSSPM are combined as inputs for hierarchical feature fusion. Then, the fused global features are upsampled to perform pixel-wise learning for grasping pose estimation. The experimental results on Cornell and Jacquard grasping datasets indicate that the proposed method has a fast inference speed of 5ms while achieving high grasp detection accuracy of 96.4% and 94.8% on Cornell and Jacquard, respectively, which strikes a balance between accuracy and running speed. Our method also gets a 90% physical grasp success rate with a UR5 robot arm.
Hu Cao, Guang Chen 0001, Zhijun Li 0001, Jianjie Lin, Alois C. Knoll
ICRA2
2021 A Novel Illumination-Robust Hand Gesture Recognition System With Event-Based Neuromorphic Vision Sensor
abstract
The hand gesture recognition system is a noncontact and intuitive communication approach, which, in turn, allows for natural and efficient interaction. This work focuses on developing a novel and robust gesture recognition system, which is insensitive to environmental illumination and background variation. In the field of gesture recognition, standard vision sensors, such as CMOS cameras, are widely used as the sensing devices in state-of-the-art hand gesture recognition systems. However, such cameras depend on environmental constraints, such as lighting variability and the cluttered background, which significantly deteriorates their performances. In this work, we propose an event-based gesture recognition system to overcome the detriment constraints and enhance the robustness of the recognition performance. Our system relies on a biologically inspired neuromorphic vision sensor that has microsecond temporal resolution, high dynamic range, and low latency. The sensor output is a sequence of asynchronous events instead of discrete frames. To interpret the visual data, we utilize a wearable glove as an interaction device with five high-frequency (>100 Hz) active LED markers (ALMs), representing fingers and palm, which are tracked precisely in the temporal domain using a restricted spatiotemporal particle filter algorithm. The latency of the sensing pipeline is negligible compared with the dynamics of the environment as the sensor's temporal resolution allows us to distinguish high frequencies precisely. We design an encoding process to extract features and adopt a lightweight network to classify the hand gestures. The recognition accuracy of our system is comparable to the state-of-the-art methods. To study the robustness of the system, experiments considering illumination and background variations are performed, and the results show that our system is more robust than the state-of-the-art deep learning-based gesture recognition systems. Note to Practitioners-This article addresses the robustness of the hand gesture recognition system that is important for gesture recognition-based applications. Existing methods rely on either the large-volume data to train a deep learning model or to restrict the applied environments (e.g., an ideal environment without dynamic background). However, a vision-based deep learning model requires large computational resources, while the ideal environment limits the practicality of the system. In this work, we introduce a biologically inspired neuromorphic vision sensor and an ALM glove and build a novel gesture recognition system to tackle the above issue. The neuromorphic vision sensor has a microsecond temporal resolution and a high dynamic range. With these properties, the sensing system of our prototype operates in a very low-latency space, which, in turn, ensures that our gesture recognition system is robust to illumination variance and dynamic background. Thus, this work is valuable to the research of illumination-robust gesture recognition systems. Preliminary experiments suggest that our system prototype is feasible, but it has not yet been incorporated into an online gesture recognition system nor tested with complex gestures. In future work, we will concentrate on the improvement of the signal processing methods that advance the current system to complex and practical applications.
Guang Chen 0001, Zhongcong Xu, Zhijun Li 0001, Huajin Tang, Sanqing Qu, Kejia Ren, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.1
2021 NeuroAED: Towards Efficient Abnormal Event Detection in Visual Surveillance With Neuromorphic Vision Sensor
abstract
Abnormal event detection is an important task in research and industrial applications, which has received considerable attention in recent years. Existing methods usually rely on standard frame-based cameras to record the data and process them with computer vision technologies. In contrast, this paper presents a novel neuromorphic vision based abnormal event detection system. Compared to the frame-based camera, neuromorphic vision sensors, such as Dynamic Vision Sensor (DVS), do not acquire full images at a fixed frame rate but rather have independent pixels that output intensity changes (called events) asynchronously at the time they occur. Thus, it avoids the design of the encryption scheme. Since events are triggered by moving edges on the scene, DVS is a natural motion detector for the abnormal objects and automatically filters out any temporally-redundant information. Based on this unique output, we first propose a highly efficient method based on the event density to select activated event cuboids and locate the foreground. We design a novel event-based multiscale spatio-temporal descriptor to extract features from the activated event cuboids for the abnormal event detection. Additionally, we build the NeuroAED dataset, the first public dataset dedicated to abnormal event detection with neuromorphic vision sensor. The NeuroAED dataset consists of four sub-datasets: Walking, Campus, Square, and Stair dataset. Experiments are conducted based on these datasets and demonstrate the high efficiency and accuracy of our method.
Guang Chen 0001, Peigen Liu, Zhengfa Liu, Huajin Tang, Lin Hong, Jinhu Dong, Jörg Conradt, Alois C. Knoll
IEEE Trans. Inf. Forensics Secur.1
2021 Pseudo-Image and Sparse Points: Vehicle Detection With 2D LiDAR Revisited by Deep Learning-Based Methods
abstract
Detecting and locating surrounding vehicles robustly and efficiently are essential capabilities for autonomous vehicles. Existing solutions often rely on vision-based methods or 3D LiDAR-based methods. These methods are either too expensive in both sensor pricing (3D LiDAR) and computation (camera and 3D LiDAR) or less robust in resisting harsh environment changes (camera). In this work, we revisit the LiDAR based approaches for vehicle detection with a less expensive 2D LiDAR by utilizing modern deep learning approaches. We aim at filling in the gap as few previous works conclude an efficient and robust vehicle detection solution in a deep learning way in 2D. To this end, we propose a learning based method with the input of pseudo-images, named Cascade Pyramid Region Proposal Convolution Neural Network (Cascade Pyramid RCNN), and a hybrid learning method with the input of sparse points, named Hybrid Resnet Lite. Experiments are conducted with our newly 2D LiDAR vehicle dataset recorded in complex traffic environments. Results demonstrate that the Cascade Pyramid RCNN outperforms state-of-the-art methods in accuracy while the proposed Hybrid Resnet Lite provides superior performance of the speed and lightweight model by hybridizing learning based and non-learning based modules. As few previous works conclude an efficient and robust vehicle detection solution with 2D LiDAR, our research fills in this gap and illustrates that even with limited sensing source from a 2D LiDAR, detecting obstacles like vehicles efficiently and robustly is still achievable.
Guang Chen 0001, Fa Wang, Sanqing Qu, Lu Xiong 0001, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.1
2020 Improving Low-Resolution Image Classification by Super-Resolution with Enhancing High-Frequency Content
abstract
With the prosperous development of Convolutional Neural Networks, currently they can perform excellently on visual understanding tasks when the input images are high quality and common quality images. However, large degradation in performance always occur when the input images are low quality images. In this paper, we propose a new super-resolution method in order to improve the classification performance for low-resolution images. In an image, the regions in which pixel values vary dramatically contain more abundant high frequency contents compared to other parts. Based on this fact, we design a weight map and integrate it with a super-resolution CNN training framework. During the process of training, this weight map can find out positions of the high frequency pixels in ground truth high-resolution images. After that, the pixel-level loss function takes effect only at these found positions to minimize the difference between reconstructed high-resolution images and ground truth high-resolution images. Compared with other state-of-the-art super-resolution methods, the experiment results show that our method can recover more high frequency contents in high-resolution image reconstructing, and better improve the classification accuracy after low-resolution image preprocessing.
Liguo Zhou, Guang Chen 0001, Mingyue Feng, Alois C. Knoll
ICPR2
2020 Hierarchical optimization Control of Redundant Manipulator for Robot-assisted Minimally Invasive Surgery
abstract
For the time varying optimization problem, the tracking error cannot converge to zero at the finite time because of the optimal solution changing over time. This paper proposes a novel varying parameter recurrent neural network (VPRNN) based hierarchical optimization of a 7-DoF surgical manipulator for Robot-Assisted Minimally Invasive Surgery (RAMIS), which guarantees task tracking, Remote Center of Motion (RCM) and manipulability index optimization. A theoretically grounded hierarchical optimization framework based is introduced to control multiple tasks based on their priority. Finally, the effectiveness of the proposed control strategy is demonstrated with both simulation and experimental results. The results show that the proposed VPRNN-based method can optimal three tasks at the same time and have better performance than previous work.
Yingbai Hu, Hang Su 0001, Guang Chen 0001, Giancarlo Ferrigno, Elena De Momi, Alois C. Knoll
IROS3
2020 RSKDD-Net: Random Sample-based Keypoint Detector and Descriptor
abstract
Keypoint detector and descriptor are two main components of point cloud registration. Previous learning-based keypoint detectors rely on saliency estimation for each point or farthest point sample (FPS) for candidate points selection, which are inefficient and not applicable in large scale scenes. This paper proposes Random Sample-based Keypoint Detector and Descriptor Network (RSKDD-Net) for large scale point cloud registration. The key idea is using random sampling to efficiently select candidate points and using a learning-based method to jointly generate keypoints and corresponding descriptors. To tackle the information loss of random sampling, we exploit a novel random dilation cluster strategy to enlarge the receptive field of each sampled point and an attention mechanism to aggregate the positions and features of neighbor points. Furthermore, we propose a matching loss to train the descriptor in a weakly supervised manner. Extensive experiments on two large scale outdoor LiDAR datasets show that the proposed RSKDD-Net achieves state-of-the-art performance with more than 15 times faster than existing methods. Our code is available at https://github.com/ispc-lab/RSKDD-Net.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Zhongnan Qu, Alois C. Knoll
NeurIPS2
2020 Ground Moving Vehicle Detection and Movement Tracking Based on the Neuromorphic Vision Sensor
abstract
Moving-objects detection is a critical ability for an autonomous vehicle. Facing the high detection requirements and the slow target-extraction problem of a common camera, this article proposes to utilize neuromorphic vision sensor (DVS) for detecting the moving objects and estimating their movement states. For a better detection work, the DVS image's noise points are filtered and the CTRV kinematics model is built in advance. In order to distinguish the overlapped or nearby bodies and get the accurate clustering number, this article proposes a 3-D improved K -means method. As the clustering centers can be influenced by the movement easily, the moving objects' clustering centers appear unstable, so the movement estimation also fluctuates. In order to obtain stable movement estimation, this article proposes a strong tracking center differential external Kalman filter (SCDEKF) to track the moving objects, and the method has higher accuracy and less computational load. In order to verify the advantages of proposed methods, a simulation environment was established in the Gazebo, and the common cameras were also added in simulation for comparison with the DVS sensor. The simulation results show that the 3-D improved K -means method with DVS can cluster the moving objects accurately, and the SCDEKF can provide more accurate movement estimation than the traditional EKF method. Finally, two experiments were conducted to prove the methods' superiority. The main contribution is to solve the exploitation problems faced by the short-term borne sensor and promote its application in transportation.
Guang Chen 0001, Xuesong Sun, Alois C. Knoll
IEEE Internet Things J.2
2020 Indirect and direct training of spiking neural networks for end-to-end control of a lane-keeping vehicle
Zhenshan Bing, Claus Meschede, Guang Chen 0001, Alois C. Knoll, Kai Huang 0001
Neural Networks3
2019 Mixed Frame-/Event-Driven Fast Pedestrian Detection
abstract
Pedestrian detection has attracted enormous research attention in the field of Intelligent Transportation System (ITS) due to that pedestrians are the most vulnerable traffic participants. So far, almost all pedestrian detection solutions are based on the conventional frame-based camera. However, they cannot perform very well in scenarios with bad light condition and high-speed motion. In this work, a Dynamic and Active Pixel Sensor (DAVIS), whose two channels concurrently output conventional gray-scale frames and asynchronous low-latency temporal contrast events of light intensity, was first used to detect pedestrians in a traffic monitoring scenario. Data from two camera channels were fed into Convolutional Neural Networks (CNNs) including three YOLOv3 models and three YOLO-tiny models to gather bounding boxes of pedestrians with respective confidence map. Furthermore, a confidence map fusion method combining the CNN-based detection results from both DAVIS channels was proposed to obtain higher accuracy. The experiments were conducted on a custom dataset collected on TUM campus. Benefiting from the high speed, low latency and wide dynamic range of the event channel, our method achieved higher frame rate and lower latency than those only using a conventional camera. Additionally, it reached higher average precision by using the fusion approach.
Zhuangyi Jiang, Kai Huang 0001, Walter Stechele, Guang Chen 0001, Zhenshan Bing, Alois C. Knoll
ICRA5
2019 Mobile Robot Learning from Human Demonstrations with Nonlinear Model Predictive Control
abstract
Learning by imitation is a powerful way that can reduce the complexly in searching space. It could help the mobile robot to acquire new skills from interaction with a human-being in natural way. In this paper, the dynamic movement primitives (DMPs) is utilized to imitate the trajectory from human walking. DMPs is a modified formulation of virtual spring-dampers (VSD) system that enjoys better fitting performance in learning. Further, while dealing with the trajectory tracking problem of mobile robots, a novel nonlinear model predictive control (MPC) approach is proposed for motion control. The nonlinear MPC scheme applies a new neural network named Varying-parameter Lagrangian Neural Network (VP-LNN) to solve a Quadratic Programming (QP) problem by iterating over a finite receding horizon. The new network of VP-LNN can converge to the global optimal solution. Thus, a new human-robot interaction (HRI) scheme for mobile robot is proposed, which can reduce the complexity in motion planning in various applications.
Yingbai Hu, Guang Chen 0001, Xiangyu Ning, Jinhu Dong, Alois C. Knoll
IROS2
2018 End to End Learning of Spiking Neural Network Based on R-STDP for a Lane Keeping Vehicle
abstract
Learning-based methods have demonstrated clear advantages in controlling robot tasks, such as the information fusion abilities, strong robustness, and high accuracy. Meanwhile, the on-board systems of robots have limited computation and energy resources, which are contradictory with state-of-the-art learning approaches. They are either too lightweight to solve complex problems or too heavyweight to be used for mobile applications. On the other hand, training spiking neural networks (SNNs) with biological plausibility has great potentials of performing fast computation and energy efficiency. However, the lack of effective learning rules for SNNs impedes their wide usage in mobile robot applications. This paper addresses the problem by introducing an end to end learning approach of spiking neural networks for a lane keeping vehicle. We consider the reward-modulated spike-timing-dependent-plasticity (R-STDP) as a promising solution in training SNNs, since it combines the advantages of both reinforcement learning and the well-known STDP. We test our approach in three scenarios that a Pioneer robot is controlled to keep lanes based on an SNN. Specifically, the lane information is encoded by the event data from a neuromorphic vision sensor. The SNN is constructed using R-STDP synapses in an all-to-all fashion. We demonstrate the advantages of our approach in terms of the lateral localization accuracy by comparing with other state-of-the-art learning algorithms based on SNNs.
Zhenshan Bing, Claus Meschede, Kai Huang 0001, Guang Chen 0001, Florian Röhrbein, Mahmoud Akl, Alois C. Knoll
ICRA4
2017 Event-Based Target Tracking Control for a Snake Robot Using a Dynamic Vision Sensor
Zhuangyi Jiang, Zhenshan Bing, Kai Huang 0001, Guang Chen 0001, Long Cheng 0007, Alois C. Knoll
ICONIP (6)4
2017 Towards autonomous locomotion: Slithering gait design of a snake-like robot for target observation and tracking
abstract
In this paper, a biologically inspired 3D slithering gait for a snake-like robot is designed and implemented for the purpose of target tracking. First, by balancing the forward speed and the stability of the robot, a straight slithering gait is modelled, under which the robot can march straight, fast, and stably. Then, for the purpose of steering, the straight slithering gait is modified into a biased slithering gait. The relationship between turning radius and gait parameters is analyzed by the resistive force theory. With the head composition algorithm, we investigate the orientation problem of the snake robot's head module to obtain stable visual information during the locomotion process. Finally, with the guidance of the vision sensor mounted in the head module, target tracking simulations and prototype experiments are conducted to demonstrate the practicality and effectiveness of the slithering gait in autonomous locomotion scenarios.
Zhenshan Bing, Long Cheng 0007, Kai Huang 0001, Zhuangyi Jiang, Guang Chen 0001, Florian Röhrbein, Alois C. Knoll
IROS5
2015 Combining unsupervised learning and discrimination for 3D action recognition
Guang Chen 0001, Daniel Clarke 0001, Manuel Giuliani, Andre Gaschler, Alois C. Knoll
Signal Process.1
2014 Action recognition using ensemble weighted multi-instance learning
abstract
This paper deals with recognizing human actions in depth video data. Current state-of-the-art action recognition methods use hand-designed features, which are difficult to produce and time-consuming to extend to new modalities. In this paper, we propose a novel, 3.5D representation of a depth video for action recognition. A 3.5D graph of the depth video consists of a set of nodes that are the joints of the human body. Each joint is represented by a set of spatio-temporal features, which are computed by an unsupervised learning approach. However, if occlusions occur, the 3D positions of the joints are noisy which increases the intra-class variations in action classes. To address this problem, we propose the Ensemble Weighted Multi-Instance Learning approach (EnwMi) for the action recognition task. It considers the class imbalance and intra-class variations. We formulate the action recognition task with depth videos as a weighted multi-instance problem. We further integrate an ensemble learning method into the weighted multi-instance learning framework. Our approach is evaluated on Microsoft Research Action3D dataset, and the results show that it outperforms state-of-the-art methods.
Guang Chen 0001, Manuel Giuliani, Daniel Clarke 0001, Andre Gaschler, Alois C. Knoll
ICRA1
2013 Learning to Track Multi-target Online by Boosting and Scene Layout
abstract
We address two principal difficulties of multi-target tracking in a real traffic scenario. Firstly, fast moving traffic scenarios lead to large displacements and complex interactions with occlusions and ambiguities. Secondly, the tracking application for real traffic scenarios has the online requirement. To surmount these difficulties, we propose an approach to track the multi-target online by Boosting and scene context reasoning. To this end, we use a two-stage system, where the first stage learns a non-linear classifier which is capable of generating the observation similarities. In the second stage, we demonstrate a novel relationship between observations and the scene layout parameters. Using a probabilistic formulation and the above relationship, our method has the unique ability to handle exceptions. To evaluate our method, we create three real traffic data sets, covering urban, rural, and highway conditions. We hope that these datasets will push forward the performance of tracking systems when being moved outside the laboratory to the real world.
Guang Chen 0001, Feihu Zhang, Daniel Clarke 0001, Alois C. Knoll
ICMLA (1)1
2013 Multiple vehicle cooperative localization under random finite set framework
abstract
This paper presents a new multiple vehicle cooperative localization approach based on Random Finite Set (RFS) theory. Assuming vehicles are equipped with proprioceptive and exteroceptive sensors to localize the positions, a solution based on RFS statistics is therefore proposed to consider the whole group behavior instead of each vehicle. For this, we rely on Probability Hypothesis Density (PHD) filtering. Compared to other methods, our approach presents a recursive filtering algorithm that provides dynamic estimation of multiple vehicle states. The proposed method addresses the current challenges in multiple vehicle cooperative localization domain such as communication bandwidth issue, data association uncertainty and the over-convergence problem. A comparative study based on simulations demonstrates the reliability and the feasibility of the proposed approach in large scale environments.
Feihu Zhang, Hauke Stähle, Guang Chen 0001, Christian Buckl, Alois C. Knoll
IROS3
2012 Visual odometry based on Random Finite Set Statistics in urban environment
abstract
This paper presents a novel approach for estimating the vehicle's trajectory in complex urban environments. In previous work, we presented a visual odometry solution that estimates frame-to-frame motion from a single camera based on Random Finite Set (RFS) Statistics. This paper extends that work by combining the stereo cameras and gyroscope sensor. We are among the first to apply RFS statistics to visual odometry in real traffic scenes. The method is based on two phases: a preprocessing phase to extract features from the image and transform the coordinates from the image space to vehicle coordinates; a tracking phase to estimate the egomotion vector of the camera. We consider features as a group target and use the Probability Hypothesis Density (PHD) filter to update the overall group state as the motion vector. Compared to other approaches, our method presents a recursive filtering algorithm that provides dynamic estimation of multiple-targets states in the presence of clutter and high association uncertainty. The experimental results show that this method exhibits good robustness under various scenarios.
Feihu Zhang, Guang Chen 0001, Hauke Stähle, Christian Buckl, Alois C. Knoll
Intelligent Vehicles Symposium2