Zhijie Zheng 0004

dblp:42/10130-4 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0001-8849-4997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Through-Wall Human Pose Estimation by Mutual Information Maximizing Deeply Supervised Nets
abstract
This article proposes a three-dimensional (3D) human pose estimation method using through-wall radar (TWR) systems, which extends and supplements new applications in the era of the Internet of Things (IoT). TWR system can penetrate non-metallic obstacles and perceive wall-occlusive human targets, but the physical characteristics of radio frequency (RF) signals, such as poor imaging resolution and specularity effect, make the pose estimation process highly ill-posed. In this work, we propose a mutual information maximizing deeply-supervised network (MIMDSN), which aims to extract accurate and robust 3D human skeletons from TWR images. Inspired by past works, an optical system is attached to the TWR system to provide cross-modal pseudo labels. Based on a depth design philosophy of convolutional neural networks that meets radar resolution constraints, we design a resolution-guided pose estimation network for keypoint coordinate regression. To alleviate the ill-posed problem, supervising solely the network output is insufficient. The cross-modal supervision is not only built on predictions, but also on features of the network’s hidden layer. With the help of information theory, the mutual information between features and pseudo labels is maximized for feature alignment and discriminability enhancement. Experiments show competitive performance against state-of-the-art RF-based human pose estimation methods and can reconstruct accurate 3D skeletons in multi-target, low-visibility, and wall-occlusive scenes.
Zhijie Zheng 0004, Jun Pan 0005, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang
IEEE Internet Things J.1
2024 RadarFormer: End-to-End Human Perception With Through-Wall Radar and Transformers
abstract
For fine-grained human perception tasks such as pose estimation and activity recognition, radar-based sensors show advantages over optical cameras in low-visibility, privacy-aware, and wall-occlusive environments. Radar transmits radio frequency signals to irradiate the target of interest and store the target information in the echo signals. One common approach is to transform the echoes into radar images and extract the features with convolutional neural networks. This article introduces RadarFormer, the first method that introduces the self-attention (SA) mechanism to perform human perception tasks directly from radar echoes. It bypasses the imaging algorithm and realizes end-to-end signal processing. Specifically, we give constructive proof that processing radar echoes using the SA mechanism is at least as expressive as processing radar images using the convolutional layer. On this foundation, we design RadarFormer, which is a Transformer-like model to process radar signals. It benefits from the fast-/slow-time SA mechanism considering the physical characteristics of radar signals. RadarFormer extracts human representations from radar echoes and handles various downstream human perception tasks. The experimental results demonstrate that our method outperforms the state-of-the-art radar-based methods both in performance and computational cost and obtains accurate human perception results even in dark and occlusive environments.
Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang
IEEE Trans. Neural Networks Learn. Syst.1
2023 Unsupervised Human Contour Extraction From Through-Wall Radar Images Using Dual UNet
abstract
Through-wall radar (TWR) can image the target of interest and capture the human sensing information. However, the poor human interpretability of TWR images and the lack of effective supervision make the extraction of complete body contour intractable. This letter proposes dual UNet, an unsupervised human contour extraction method for TWR images. Specifically, the method adopts two UNets with the same structure. One serves as the encoder to convert the TWR images into the latent representation. Another serves as the decoder to reconstruct the latent representation into the original images. Reconstruction loss and smooth normalized cut loss are optimized together to offset the dependence on labels and supplement global segment constraints. After training and post-processing, the latent representation can be used as the result of contour extraction. Experimental results show that dual UNet stands out among unsupervised human contour extraction methods in both free space and wall-occlusive scenarios, opening the possibility of learning useful human sensing information from raw TWR images without manual annotations.
Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.1
2023 Fully Sparse Transformer 3-D Detector for LiDAR Point Cloud
abstract
The 3D object detector usually uses a framework similar to 2D detection and benefits from the advancements of 2D detection tasks. In these frameworks, it is necessary to make the unstructured, sparse point cloud features into dense grids to be compatible with popular 2D operators such as convolution and transformers, which also causes extra computational costs. In this paper, we propose a simple and efficient Fully Sparse TRansformer (FSTR) for LiDAR-based 3D object detection, which is able to combine with state-of-the-art sparse backbones to form a fully sparse, end-to-end, simple, and efficient detection framework. FSTR uses the sparse voxel feature from the sparse backbone as the input token without any custom operators. Further, we introduce the dynamic queries to provide a priori location and context of the foreground for the decoder and drop the high-confidence background tokens to further reduce redundant computations. We propose Gaussian denoising queries to speed up the decoder training and make it more adaptable to the distribution of sparse voxel features. Extensive experiments on the nuScenes benchmark and the Argoverse2 benchmark validate the effectiveness of the proposed method. FSTR outperforms all LiDAR real-time methods by 69.5 mAP and 72.9 NDS on the official benchmark of nuScenes dataset. On the long-range detection benchmark Argoverse2, the proposed method achieves a new state-of-art performance of 39.9 mAP which outperforms the existing LiDAR detectors, even the LiDAR-Camera detectors by a large margin (+9.4 mAP and +7.5mAP), showing the great advantage of the proposed method for long-range detection.
Diankun Zhang, Zhijie Zheng 0004, Haoyu Niu 0001, Xiaojun Liu 0004
IEEE Trans. Geosci. Remote. Sens.2
2022 Declutter-GAN: GPR B-Scan Data Clutter Removal Using Conditional Generative Adversarial Nets
abstract
Clutter removal in ground-penetrating radar (GPR) B-scan data has been widely studied in recent years. In this letter, we propose a novel data-driven clutter suppression method in GPR data based on conditional generative adversarial nets (cGANs). The proposed method learns a function that maps the cluttered data to the clutter-free data from the training set. The training set consists of pairs of cluttered data and corresponding clutter-free data. Different from the traditional method that only uses the simulation training set, we simulate the clutter-free data and add the real collected non-target data to the simulated clutter-free data as cluttered data, so that the trained network can generalize well to the real GPR data. The proposed method is compared with the subspace method, sparse representation-based method, and low-rank and sparse matrix decomposition (LRSD) methods on both simulation data and real collected data. The results show that the proposed method has higher performance in terms of computational complexity, clutter suppression results, and applicability than those state-of-the-art methods.
Zhi-Kang Ni, Jun Pan 0005, Zhijie Zheng 0004, Shengbo Ye, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.4
2022 Motion Compensation Method Based on MFDF of Moving Target for UWB MIMO Through-Wall Radar System
abstract
Ultrawideband (UWB) multiple-input–multiple-output (MIMO) radar is widely used for through-wall imaging (TWI) due to its excellent penetrability and large aperture. Multichannels in the MIMO radar system are usually time-division multiplexing based on microwave switches to reduce the complexity of the system in engineering. The switching process of the channel will bring time delay, which cannot be ignored in the TWI of the moving target. The switching time delay will cause the defocus and position shift of the TWI of the moving target. This letter proposes a motion compensation method based on multiframe data fusion (MFDF) used for correcting the echo of the through-wall moving target. A geometric model is established in the proposed method through the echo of the current frame and the next frame, and the compensated signal is obtained through the geometric solution. The proposed method is compared with before compensation and the traditional single-channel motion compensation algorithm (SCMCA) through simulation and experimental data verification. The visual images and quantitative results show that the proposed motion compensation method can obtain a good focus image of the through-wall moving target and reduce the positioning error.
Jun Pan 0005, Zhi-Kang Ni, Zhijie Zheng 0004, Shengbo Ye, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.4
2022 Human Posture Reconstruction for Through-the-Wall Radar Imaging Using Convolutional Neural Networks
abstract
Low imaging spatial resolution hinders through-the-wall radar imaging (TWRI) from reconstructing complete human postures. This letter mainly discusses a convolutional neural network (CNN)-based human posture reconstruction method for TWRI. The training process follows a supervision-prediction learning pipeline inspired by the cross-modal learning technique. Specifically, optical images and TWRI signals are collected simultaneously using a self-develop radar containing an optical camera. Then, the optical images are processed with a computer-vision-based supervision network to generate ground-truth human skeletons. Next, the same type of skeleton is predicted from corresponding TWRI signals using a prediction network. After training, the model shows complete predictions in wall-occlusive scenarios solely using TWRI signals. Experiments show comparable quantitative results with the state-of-the-art vision-based methods in nonwall-occlusive scenarios and accurate qualitative results with wall occlusion.
Zhijie Zheng 0004, Jun Pan 0005, Zhi-Kang Ni, Shengbo Ye, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.1
2022 Through-Wall Human Pose Reconstruction Based on Cross-Modal Learning and Self-Supervised Learning
abstract
Recent through-wall radar (TWR) systems can reconstruct the pose of human targets blocked by occlusion. They rely on the fusion of optical and radar data to avoid the painful annotation burden. However, the fusion process is not always reliable, especially for human joint coordinates that carry 3-D spatial information. Inspired by cross-modal learning and self-supervised learning, this letter proposes a two-stage 3-D human pose reconstruction method for TWR systems. In the cross-modal supervision stage, the pretrained optical model provides initial noisy labels extracted from optical images. In the self-supervision stage, supervised labels and the model weight are corrected circularly with radar images. The self-supervision enhances the robustness of the model and the reliability of labels. It can be directly extended to existing radar-based pose reconstruction methods, and hardly requires extra training time. Experiments show the model beats state of the art (SOTA) for reconstructing 3-D poses from TWR images and contains robust generalization in unseen wall-occlusive scenes.
Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.1
2022 Unsupervised Domain Adaptive 3-D Detection With Data Adaption From LiDAR Point Cloud
abstract
Existing unsupervised domain adaptive (UDA) 3D detection methods only address the domain gap caused by the prior size of 3D bounding boxes between different datasets, which ignore the difference in the distribution of point clouds. To address this challenge, we propose an unsupervised domain adaptive 3D detection by data adaption, which trains the model by transferring the source domain instances into the target domain scenes by adaptive point distribution. First, an instance transferring method is proposed for selecting and transferring suitable instances from the source domain into the target domain scene; Second, we propose an adaptive downsampling method to adjust the point cloud distribution of the transferred instances to approximate the points distribution of the target domain. Finally, our method trains the randomly initialized detector with the pseudo-instances in the target domain. To the best of our knowledge, we first address the UDA problem of the 3D detectors from the perspective of data. Extensive experiments on several popular datasets show that the proposed method outperforms the existing state-of-the-art methods by a large margin. Further experiments also show our approach is detector-agnostic and achieves consistent and significant gains on all types of 3D detectors.
Diankun Zhang, Zhijie Zheng 0004, Xiaojun Liu 0004
IEEE Trans. Geosci. Remote. Sens.3
2022 Recovering Human Pose and Shape From Through-the-Wall Radar Images
abstract
Although the through-the-wall radar imaging (TWRI) system working in the appropriate frequency band can penetrate the nonmetallic obstacles and sense the targets behind, its low imaging spatial resolution hinders the acquisition of more detailed information, such as human pose and shape. This article mainly discusses a deep learning-based human pose and shape recovery method from TWRI images. Inspired by cross-modal learning, the method follows a teacher–student learning pipeline that avoids the heavy cost of manual labeling. Specifically, a camera is attached to the self-develop radar system to simultaneously capture paired red-green-blue (RGB) images and TWRI images in a scenario without wall occlusion. A pose estimation framework (Hourglass) and a semantic segmentation framework (UNet) serve as the teacher network to convert the RGB images into the pose keypoints and the shape masks. By taking inspiration from the topological architecture of these frameworks, a student network radar pose shape network (RPSNet) is designed to extract the information from the corresponding radar images and predict the keypoints and masks that are close to the results above. Instead of learning two single-task objectives independently, multitasking learning is introduced to adaptatively learn common features. When applied to wall-occlusive scenarios, only the radar images are collected and fed into the student network for pose and shape recovery. The advantages of this method over computer vision-based methods for human recovery are demonstrated in scenarios both without and with wall occlusion.
Zhijie Zheng 0004, Jun Pan 0005, Zhi-Kang Ni, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang
IEEE Trans. Geosci. Remote. Sens.1