Zhen Zhang 0019

dblp:19/5112-19 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-3786-4617ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer
abstract
Tracking a specific person in 3D scene is gaining momentum due to its numerous applications in robotics. Currently, most 3D trackers focus on driving scenarios with neglected jitter and uncomplicated surroundings, which results in their severe degeneration in complex environments, especially on jolting robot platforms (only 20-60% success rate). To improve the accuracy, a Point-Video-based Transformer Tracking model (PVTrack) is presented for robots. It is the first multi-modal 3D human tracking work that incorporates point clouds together with RGB videos to achieve information complementarity. Moreover, PVTrack proposes the Siamese Point-Video Transformer for feature aggregation to overcome dynamic environments, which captures more target-aware information through the hierarchical attention mechanism adaptively. Considering the violent shaking on robots and rugged terrains, a lateral Human-ware Proposal Network is designed together with an Anti-shake Proposal Compensation module. It alleviates the disturbance caused by complex scenes as well as the particularity of the robot platform. Experiments show that our method achieves state-of-the-art performance on both KITTI/Waymo datasets and a quadruped robot for various indoor and outdoor scenes.
Shuo Xin, Zhen Zhang 0019, Mengmeng Wang 0005, Xiaojun Hou, Yaowei Guo, Xiao Kang, Liang Liu 0007, Yong Liu 0007
ICRA2
2024 A Robotic-centric Paradigm for 3D Human Tracking Under Complex Environments Using Multi-modal Adaptation
abstract
The goal of this paper is to strike a feasible tracking paradigm that can make 3D human trackers applicable on robot platforms and enable more high-level tasks. Till now, two fundamental problems haven’t been adequately addressed. One is the computational cost lightweight enough for robotic deployment, and the other is the easily-influenced accuracy varied greatly in complex real environments. In this paper, a robotic-centric tracking paradigm called MATNet is proposed that directly matches the LiDAR point clouds and RGB videos through end-to-end learning. To improve the low accuracy of human tracking against disturbance, a coarse-to-fine Transformer along with target-ware augmentation is proposed by fusing RGB videos and point clouds through a pyramid encoding and decoding strategy. To better meet the real-time requirement of actual robot deployment, we introduce the parameter-efficient adaptation tuning that greatly shortens the model’s training time. Furthermore, we also propose a five-step Anti-shake Refinement strategy and have added human prior values to overcome the strong shaking on the robot plat-form. Extensive experiments confirm that MATNet significantly outperforms the previous state-of-the-art on both open-source datasets and large-scale robotic datasets.
Shuo Xin, Zhen Zhang 0019, Liang Liu 0007, Xiaojun Hou, Deye Zhu, Mengmeng Wang 0005, Yong Liu 0007
IROS2
2024 Learning Safe Locomotion for Quadrupedal Robots by Derived-Action Optimization
abstract
Deep reinforcement learning controllers with exteroception have enabled quadrupedal robots to traverse terrain robustly. However, most of these controllers heavily depend on complex reward functions and suffer from poor convergence. This work proposes a novel learning framework called derived-action optimization. The derived action is defined as a high-level representation of a policy and can be introduced into the reward function to guide decision-making behaviors. The proposed derived-action optimization method is applied to learn safer quadrupedal locomotion, achieving fast convergence and better performance. Specifically, we choose the foothold as the derived action and optimize the flatness of the terrain around the foothold to reduce potential sliding and collisions. Extensive experiments demonstrate the high safety and effectiveness of our method.
Deye Zhu, Chengrui Zhu, Zhen Zhang 0019, Shuo Xin, Yong Liu 0007
IROS3
2024 Dual-Feature Attention-Based Contrastive Prototypical Clustering for Multimodal Remote Sensing Data
abstract
The integrated use of multisource remote sensing (RS) data in Earth observation missions has garnered considerable attention. Hyperspectral images (HSIs) offer extensive spatial and spectral detail, whereas light detection and ranging (LiDAR) data provide elevation information. Therefore, the fusion of HSI and LiDAR data can enhance the accuracy (ACC) of image classification. However, contemporary supervised multimodal deep learning techniques depend heavily on extensive human-annotated training datasets. To address this challenge, we propose a contrastive prototypical clustering network enhanced with a dual-feature attention module. Specifically, two sets of enhanced modal views are constructed from the multimodal RS images for the subsequent contrastive learning. The proposed dual-feature attention module emphasizes channel and spatial attention separately for each modality, integrating both to adjust the feature representation across different channels and positions. By learning the importance weights of each channel and position, this module highlights the hierarchical structure and enhances the discriminative quality of the features. The learned features are utilized through an online clustering mechanism and a self-supervised training strategy that combines contrastive loss and cluster loss to achieve efficient and effective land cover classification. Extensive experiments on three widely used HSI and LiDAR datasets demonstrate that the proposed method outperforms current state-of-the-art approaches. The code for this method is openly available at:https://github.com/RogsDing/DFCPC.
Shufang Xu, Xinchen Ding, Zhen Zhang 0019, Hongmin Gao 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Contrastive Learning with Dynamic Weighting and Jigsaw Augmentation for Brain Tumor Classification in MRI
Guanghua Xiao, Jie Shen 0004, Zhe Chen 0004, Zhen Zhang 0019, Xiaomin Ge
Neural Process. Lett.5
2022 An Adaptive Hierarchical Concatenated Network With A Robust Loss Function For Image Denoising
Guanghua Xiao, Jie Shen 0004, Zhe Chen 0004, Zhen Zhang 0019
J. Grid Comput.5
2020 Underwater salient object detection by combining 2D and 3D visual features
Zhe Chen 0004, Hongmin Gao 0001, Zhen Zhang 0019, Helen Zhou, Xun Wang 0007
Neurocomputing3
2019 Background-foreground interaction for moving object detection in dynamic scenes
Zhe Chen 0004, Ruili Wang 0001, Zhen Zhang 0019
Inf. Sci.3