EDBT 2026 Demo / reviewers in the wild / expert
Diankun Zhang
dblp:305/3775
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-2004-0897ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Dingyuan Zhang, Hongwei Xie, Xiang Bai |
ICCV | 2 |
| 2024 | Through-Wall Human Pose Estimation by Mutual Information Maximizing Deeply Supervised NetsabstractThis article proposes a three-dimensional (3D) human pose estimation method using through-wall radar (TWR) systems, which extends and supplements new applications in the era of the Internet of Things (IoT). TWR system can penetrate non-metallic obstacles and perceive wall-occlusive human targets, but the physical characteristics of radio frequency (RF) signals, such as poor imaging resolution and specularity effect, make the pose estimation process highly ill-posed. In this work, we propose a mutual information maximizing deeply-supervised network (MIMDSN), which aims to extract accurate and robust 3D human skeletons from TWR images. Inspired by past works, an optical system is attached to the TWR system to provide cross-modal pseudo labels. Based on a depth design philosophy of convolutional neural networks that meets radar resolution constraints, we design a resolution-guided pose estimation network for keypoint coordinate regression. To alleviate the ill-posed problem, supervising solely the network output is insufficient. The cross-modal supervision is not only built on predictions, but also on features of the network’s hidden layer. With the help of information theory, the mutual information between features and pseudo labels is maximized for feature alignment and discriminability enhancement. Experiments show competitive performance against state-of-the-art RF-based human pose estimation methods and can reconstruct accurate 3D skeletons in multi-target, low-visibility, and wall-occlusive scenes. Zhijie Zheng 0004, Jun Pan 0005, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang |
IEEE Internet Things J. | 3 |
| 2024 | Enhancing concealed object detection in Active Millimeter Wave Images using wavelet transformabstractIn contemporary security detection systems, the utilization of millimeter-wave radar has assumed a central role owing to its non-contact and innocuous nature. This study addresses the challenging issues of detecting low-resolution and small targets in active millimeter wave (AMMW) images concealed detection. We identify a prevalent drawback in existing detectors, specifically the adoption of strided convolution or pooling layers, leading to a loss of crucial details that adversely impacts the detection rate. To overcome this limitation, we introduce a novel convolutional structure, termed Wavelet-Conv, which maintains information integrity while segregating features from high and low frequency bands, effectively replacing the unfavorable design. Furthermore, we harness the wavelet transform to enhance channel and spatial attention mechanisms, enabling more effective utilization of frequency band features and ensuring interpretability in the computational process. In this pursuit, we integrate the proposed Wavelet-Conv and Wavelet-Attention modules into the YOLOv8 framework, culminating in a unified model, termed Wavelet-YOLO. Through rigorous experimental validation on two AMMW datasets, our approach exhibits superior performance by significantly enhancing the recall and mean average precision (mAP) of small targets in AMMW images, while maintaining competitive inference speed. Extensive experiments demonstrate the outperformance of our proposed method over existing state-of-the-art approaches. Weixian Tan, Wei Xu 0018, Pingping Huang, Diankun Zhang |
Signal Process. | 7 |
| 2024 | RadarFormer: End-to-End Human Perception With Through-Wall Radar and TransformersabstractFor fine-grained human perception tasks such as pose estimation and activity recognition, radar-based sensors show advantages over optical cameras in low-visibility, privacy-aware, and wall-occlusive environments. Radar transmits radio frequency signals to irradiate the target of interest and store the target information in the echo signals. One common approach is to transform the echoes into radar images and extract the features with convolutional neural networks. This article introduces RadarFormer, the first method that introduces the self-attention (SA) mechanism to perform human perception tasks directly from radar echoes. It bypasses the imaging algorithm and realizes end-to-end signal processing. Specifically, we give constructive proof that processing radar echoes using the SA mechanism is at least as expressive as processing radar images using the convolutional layer. On this foundation, we design RadarFormer, which is a Transformer-like model to process radar signals. It benefits from the fast-/slow-time SA mechanism considering the physical characteristics of radar signals. RadarFormer extracts human representations from radar echoes and handles various downstream human perception tasks. The experimental results demonstrate that our method outperforms the state-of-the-art radar-based methods both in performance and computational cost and obtains accurate human perception results even in dark and occlusive environments. Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Unsupervised Human Contour Extraction From Through-Wall Radar Images Using Dual UNetabstractThrough-wall radar (TWR) can image the target of interest and capture the human sensing information. However, the poor human interpretability of TWR images and the lack of effective supervision make the extraction of complete body contour intractable. This letter proposes dual UNet, an unsupervised human contour extraction method for TWR images. Specifically, the method adopts two UNets with the same structure. One serves as the encoder to convert the TWR images into the latent representation. Another serves as the decoder to reconstruct the latent representation into the original images. Reconstruction loss and smooth normalized cut loss are optimized together to offset the dependence on labels and supplement global segment constraints. After training and post-processing, the latent representation can be used as the result of contour extraction. Experimental results show that dual UNet stands out among unsupervised human contour extraction methods in both free space and wall-occlusive scenarios, opening the possibility of learning useful human sensing information from raw TWR images without manual annotations. Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Fully Sparse Transformer 3-D Detector for LiDAR Point CloudabstractThe 3D object detector usually uses a framework similar to 2D detection and benefits from the advancements of 2D detection tasks. In these frameworks, it is necessary to make the unstructured, sparse point cloud features into dense grids to be compatible with popular 2D operators such as convolution and transformers, which also causes extra computational costs. In this paper, we propose a simple and efficient Fully Sparse TRansformer (FSTR) for LiDAR-based 3D object detection, which is able to combine with state-of-the-art sparse backbones to form a fully sparse, end-to-end, simple, and efficient detection framework. FSTR uses the sparse voxel feature from the sparse backbone as the input token without any custom operators. Further, we introduce the dynamic queries to provide a priori location and context of the foreground for the decoder and drop the high-confidence background tokens to further reduce redundant computations. We propose Gaussian denoising queries to speed up the decoder training and make it more adaptable to the distribution of sparse voxel features. Extensive experiments on the nuScenes benchmark and the Argoverse2 benchmark validate the effectiveness of the proposed method. FSTR outperforms all LiDAR real-time methods by 69.5 mAP and 72.9 NDS on the official benchmark of nuScenes dataset. On the long-range detection benchmark Argoverse2, the proposed method achieves a new state-of-art performance of 39.9 mAP which outperforms the existing LiDAR detectors, even the LiDAR-Camera detectors by a large margin (+9.4 mAP and +7.5mAP), showing the great advantage of the proposed method for long-range detection. Diankun Zhang, Zhijie Zheng 0004, Haoyu Niu 0001, Xiaojun Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Through-Wall Human Pose Reconstruction Based on Cross-Modal Learning and Self-Supervised LearningabstractRecent through-wall radar (TWR) systems can reconstruct the pose of human targets blocked by occlusion. They rely on the fusion of optical and radar data to avoid the painful annotation burden. However, the fusion process is not always reliable, especially for human joint coordinates that carry 3-D spatial information. Inspired by cross-modal learning and self-supervised learning, this letter proposes a two-stage 3-D human pose reconstruction method for TWR systems. In the cross-modal supervision stage, the pretrained optical model provides initial noisy labels extracted from optical images. In the self-supervision stage, supervised labels and the model weight are corrected circularly with radar images. The self-supervision enhances the robustness of the model and the reliability of labels. It can be directly extended to existing radar-based pose reconstruction methods, and hardly requires extra training time. Experiments show the model beats state of the art (SOTA) for reconstructing 3-D poses from TWR images and contains robust generalization in unseen wall-occlusive scenes. Zhijie Zheng 0004, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Analysis of Aeromagnetic Swing Noise and Corresponding Compensation MethodabstractAeromagnetic noise compensation is a vital part of aerial survey measurement, and its compensation effect directly determines the quality of aeromagnetic survey data. At present, the commonly used compensation model is the T-L model, and the least squares method is used to solve for the coefficients. However, the noise source modeled in the T-L model is incomplete. Since the tail boom cannot be completely rigid, tail-boom swing is an unavoidable problem in aeromagnetic measurement. This kind of swing is the most obvious when the aircraft is maneuvering, and it will significantly interfere with the measurement data of the sensor. In this article, two causes of the swing noise are analyzed, and the nonlinear relationship between the swing displacement and the noise is derived. Since it is difficult to express the nonlinear relationship with mathematical forms to compensate for the aeromagnetic data, we propose a new compensation method that uses a 1-D convolutional neural network to perform secondary compensation on the data already compensated by the T-L model in order to remove the effect of tail-boom swing. The flight experiment data show that the proposed method can significantly improve the quality of aeromagnetic data. Compared with the T-L method, the improve ratio is increased by 60%–100%. It shows that the proposed method has a remarkable compensation effect for aeromagnetic noise. Diankun Zhang, Xiaojun Liu 0004, Wanhua Zhu, Ling Huang 0007, Guangyou Fang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Unsupervised Domain Adaptive 3-D Detection With Data Adaption From LiDAR Point CloudabstractExisting unsupervised domain adaptive (UDA) 3D detection methods only address the domain gap caused by the prior size of 3D bounding boxes between different datasets, which ignore the difference in the distribution of point clouds. To address this challenge, we propose an unsupervised domain adaptive 3D detection by data adaption, which trains the model by transferring the source domain instances into the target domain scenes by adaptive point distribution. First, an instance transferring method is proposed for selecting and transferring suitable instances from the source domain into the target domain scene; Second, we propose an adaptive downsampling method to adjust the point cloud distribution of the transferred instances to approximate the points distribution of the target domain. Finally, our method trains the randomly initialized detector with the pseudo-instances in the target domain. To the best of our knowledge, we first address the UDA problem of the 3D detectors from the perspective of data. Extensive experiments on several popular datasets show that the proposed method outperforms the existing state-of-the-art methods by a large margin. Further experiments also show our approach is detector-agnostic and achieves consistent and significant gains on all types of 3D detectors. Diankun Zhang, Zhijie Zheng 0004, Xiaojun Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Recovering Human Pose and Shape From Through-the-Wall Radar ImagesabstractAlthough the through-the-wall radar imaging (TWRI) system working in the appropriate frequency band can penetrate the nonmetallic obstacles and sense the targets behind, its low imaging spatial resolution hinders the acquisition of more detailed information, such as human pose and shape. This article mainly discusses a deep learning-based human pose and shape recovery method from TWRI images. Inspired by cross-modal learning, the method follows a teacher–student learning pipeline that avoids the heavy cost of manual labeling. Specifically, a camera is attached to the self-develop radar system to simultaneously capture paired red-green-blue (RGB) images and TWRI images in a scenario without wall occlusion. A pose estimation framework (Hourglass) and a semantic segmentation framework (UNet) serve as the teacher network to convert the RGB images into the pose keypoints and the shape masks. By taking inspiration from the topological architecture of these frameworks, a student network radar pose shape network (RPSNet) is designed to extract the information from the corresponding radar images and predict the keypoints and masks that are close to the results above. Instead of learning two single-task objectives independently, multitasking learning is introduced to adaptatively learn common features. When applied to wall-occlusive scenarios, only the radar images are collected and fed into the student network for pose and shape recovery. The advantages of this method over computer vision-based methods for human recovery are demonstrated in scenarios both without and with wall occlusion. Zhijie Zheng 0004, Jun Pan 0005, Zhi-Kang Ni, Diankun Zhang, Xiaojun Liu 0004, Guangyou Fang |
IEEE Trans. Geosci. Remote. Sens. | 5 |