Dezhi Zheng

dblp:155/4239 · DBLP profile ↗
← Back
44ranked-venue papers
2as first author
43since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 17 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quantum-Based Broadband Integrated Sensing and Communication with Rydberg Atomic Receiver
Minze Chen, Tianqi Mao 0001, Zhiao Zhu, Zhaocheng Wang 0001, Dezhi Zheng
ICC6
2026 Broadband Scanning-Free Rydberg Atomic Communications: Hybrid Noise Modeling and Detector Design
Tianqi Mao 0001, Minze Chen, Meng Hua, Dezhi Zheng
IWCMC6
2026 MPANet: Motion Pattern Aggregation Network for Gait Recognition
Chunshui Cao, Guoren Wang, Dezhi Zheng
Int. J. Comput. Vis.5
2026 Pseudo-Random TDM-MIMO FMCW-Based Millimeter-Wave Sensing and Communication Integration for UAV Swarm
Zhen Gao 0001, Ziwei Wan, Tuan Li, Chunli Zhu, Guanghui Wen, Dezhi Zheng, Dusit Niyato
IEEE Internet Things J.9
2026 Rydberg Atomic Receivers for Multi-Band Communications and Sensing
abstract
Harnessing multi-level electron transitions, Rydberg Atomic REceivers (RAREs) can detect wireless signals across a wide range of frequency bands, from Megahertz to Terahertz. This capability enables multi-band wireless communications and sensing (CommunSense). Existing research on multi-band RAREs primarily focuses on experimental demonstrations, lacking a tractable model to mathematically characterize their mechanisms. This issue leaves the multi-band RARE as a black box and poses challenges in its practical applications. To fill in this gap, this paper investigates the underlying mechanism of multi-band RAREs and explores their optimal performance. For the first time, an analytical transfer function with a closed-form expression for multi-band RAREs is derived by solving the quantum response of Rydberg atoms. It shows that a multi-band RARE simultaneously serves as amulti-band atomic mixerfor down-converting multi-band signals and amulti-band atomic amplifierthat reflects its sensitivity to each band. Further analysis of the atomic amplifier unveils that the intrinsic gain at each frequency band can be decoupled into aglobal gainterm and aRabi attentionterm. The former determines the overall sensitivity of a RARE to all frequency bands of wireless signals. The latter influences the allocation of the overall sensitivity to each frequency band, representing a unique attention mechanism of multi-band RAREs. The optimal design of the global gain is provided to maximize the overall sensitivity of multi-band RAREs. Subsequently, the optimal Rabi attentions are also derived to maximize the practical multi-band CommunSense performance. An experiment platform is built to validate the effectiveness of the derived transfer function, and numerical results confirm the superiority of multi-band RAREs.
Mingyao Cui, Qunsong Zeng, Minze Chen, Zhanwei Wang, Tianqi Mao 0001, Dezhi Zheng, Kaibin Huang
IEEE Trans. Wirel. Commun.6
2025 DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis
abstract
Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code is available at https://github.com/CVI-SZU/DEGSTalk.
Kaijun Deng, Dezhi Zheng, Jindong Xie, Jinbao Wang 0001, Weicheng Xie 0001, LinLin Shen, Siyang Song
ICASSP2
2025 Dual Encoders for Diffusion-based Image Inpainting
abstract
Current diffusion-based inpainting models struggle to preserve unmasked regions or generate highly coherent content. Additionally, it is hard for them to generate meaningful content for 3D inpainting. To tackle these challenges, we design a plug-and-play branch that runs through the entire generation process to enhance existing models. Specifically, we utilize dual encoders - a Convolutional Neural Network (CNN) encoder and the pre-trained Variational AutoEncoder (VAE) encoder, to encode masked images. The latent code and the feature map from the dual encoders are fed to diffusion models simultaneously. In addition, we apply Zero-padded initialization to solve the problem of mode collapse caused by this branch. Experiments on BrushBench and EditBench demonstrate that models with our plug-and-play branch can improve the coherence of inpainting, and our model achieves new state-of-the-art results.
Dezhi Zheng, Kaijun Deng, Jinbao Wang 0001, LinLin Shen
ICASSP1
2025 Unknown Pixel Mask Based Fine-tuning of 2D Inpainting Models for Unbounded 3D Scene Generation from a Single Image
abstract
Conventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency.
Dezhi Zheng, Kaijun Deng, Xianxu Hou, Jinbao Wang 0001, LinLin Shen
ACM Multimedia1
2025 Quantum Demodulation of QAM Signals at a Rydberg atomic Homodyne Receiver
abstract
Facing the trend of 5G/6G communications toward high spectral efficiency and low power consumption, RF receivers urgently need to overcome the triple challenges of high sensitivity, terahertz (THz) coverage, and system integration. Conventional approaches are limited by electronic thermal noise, band correlation and antenna perturbation effects. In this context, sensors based on the Rydberg atom become a promising alternative, with his large electric dipole moment bringing extremely high sensitivity, inherent frequency selectivity, and the potential to build low-power integrated photonic platforms. This paper describes a Rydberg atomic RF sensor using a quantum coherent mechanism to receive and demodulate quadrature AM signals. The core innovations include: (1) a scheme for 4QAM signal splitting and demodulation within the atomic resonance region via precise signal control; and (2) dynamic, real-time control of the demodulation channels (I/Q paths) by exploiting the phase difference between the local oscillator (LO) and signal (SIG) fields. The proposed architecture can support demodulation of quadrature AM signals, demonstrating its great potential as a versatile next-generation communications platform.
Zhiao Zhu, Zhongxiang Li, Dezhi Zheng, Chun Hu, WeiDong Dai, Minze Chen
PIMRC3
2025 LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
Dezhi Zheng, Lei Wang 0018, Liping xiang, Kaijun Deng, Xiaowen Fu, LinLin Shen, Jinbao Wang 0001
PRCV (10)2
2025 Fast Autonomous Exploration in Complex Environments via the Farthest Cluster Representative and Dynamic Information Gain
Junhui Kuang, Dezhi Zheng, Youyi Huang, Jianye Fu, Shubin Cai, Yinghui Pan, Zhong Ming 0001
WASA (3)2
2025 Monocular thermal SLAM with neural radiance fields for 3D scene reconstruction
Yuzhen Wu, Lingxue Wang, Lian Zhang 0001, Ming-Kun Chen, Wenqu Zhao, Dezhi Zheng
Neurocomputing6
2025 ISCDFuse: Interval sampling correlation driven visual state space models for multimodal image fusion
Lian Zhang 0001, Lingxue Wang, Yuzhen Wu, Ming-Kun Chen, Dezhi Zheng
Neurocomputing5
2025 ACNTrack: Agent cross-attention guided Multimodal Multi-Object Tracking with Neural Kalman Filter
Lian Zhang 0001, Lingxue Wang, Yuzhen Wu, Ming-Kun Chen, Dezhi Zheng
Neurocomputing5
2025 Spatial Frequency Modulation for Semantic Segmentation
abstract
High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency components are vulnerable to aliasing or distortion when propagating through downsampling layers such as strided-convolution. Here, we propose a novel Spatial Frequency Modulation (SFM) that modulates high-frequency features to a lower frequency before downsampling and then demodulates them back during upsampling. Specifically, we implement modulation through adaptive resampling (ARS) and design a lightweight add-on that can densely sample the high-frequency areas to scale up the signal, thereby lowering its frequency in accordance with the Frequency Scaling Property. We also propose Multi-Scale Adaptive Upsampling (MSAU) to demodulate the modulated feature and recover high-frequency information through non-uniform upsampling This module further improves segmentation by explicitly exploiting information interaction between densely and sparsely resampled areas at multiple scales. Both modules can seamlessly integrate with various architectures, extending from convolutional neural networks to transformers. Feature visualization and analysis demonstrate that our method effectively alleviates aliasing while successfully retaining details after demodulation. As a result, the proposed approach considerably enhances existing state-of-the-art segmentation models (e.g., Mask2Former-Swin-T +1.5 mIoU, InternImage-T +1.4 mIoU on ADE20 K). Furthermore, ARS also enhances the performance of powerful Deformable Convolution (+0.8 mIoU on Cityscapes) by maintaining relative positional order during non-uniform sampling. Finally, we validate the broad applicability and effectiveness of SFM by extending it to image classification, adversarial robustness, instance segmentation, and panoptic segmentation tasks.
Ying Fu 0001, Lin Gu 0003, Dezhi Zheng, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 UniRTL: A universal RGBT and low-light benchmark for object tracking
Lian Zhang 0001, Lingxue Wang, Yuzhen Wu, Ming-Kun Chen, Dezhi Zheng, Liangcai Cao, Bangze Zeng
Pattern Recognit.5
2025 Distributed Robust Event-Triggered Platooning Control of Connected Vehicles With Uncertain Dynamics: A Neuro-adaptive Approach
abstract
This article aims to address the distributed robust platooning control of connected automated vehicles (CAVs) with general unknown uncertain dynamics. Despite recent progress in this area, achieving the objective of distributed robust platooning control for CAVs with limited communication resources and uncertain dynamics is an outstanding problem. To solve such a problem, a new Zeno-free event-triggered scheme is successfully established to determine whether the vehicle's state should be sampled and transmitted among the interacting vehicles. An adaptive law for updating the weighting matrix for the neural network approximator is designed, where the relative state variables are utilized only at triggered instants. Moreover, such a neuro-adaptive approach incorporates a low-pass filter structure to effectively mitigate undesirable high-frequency oscillations that may arise with the application of high-gain learning rates. Following this, a new class of distributed event-based neuro-adaptive control protocols is meticulously designed to guarantee the uniform ultimate boundedness of spacing error, relative velocity, and relative acceleration of the whole platoon. Finally, simulation examples with different scenarios are conducted, and it is interesting to find that the proposed protocol has a lower average communication rate than traditional ones without the low-pass filter structure.
Guanghui Wen, Ying Wan 0002, Jialing Zhou, Dezhi Zheng, C. L. Philip Chen
IEEE Trans. Ind. Informatics4
2025 Consensus Tracking of Disturbed Second-Order Multiagent Systems With Actuator Attacks: Reinforcement-Learning-Based Approach
abstract
This article is devoted to solving the leaderless and leader-following consensus tracking problems for a class of disturbed second-order multiagent systems (MASs) under the influence of actuator attacks. To achieve this, a two-step control strategy is developed, where the effects of disturbances and actuator attacks on the achievement of consensus tracking are addressed in distinct stages. In the first step, a reference system model is constructed for each agent. Upon which a sliding mode control (SMC) protocol is constructed and utilized to resolve the consensus tracking problem of disturbed second-order MASs in the absence of actuator attacks, facilitating the design of a baseline control term for the MASs under consideration. In the second step, a secure control policy is trained using an off-policy soft actor-critic algorithm, aiming at achieving secure consensus tracking in the presence of actuator attacks. Both numerical simulations and a multipendulum consensus example verify that the designed control structure has better control performance than using only the SMC method and also effectively improves the training efficiency over the traditional reinforcement learning (RL) alone method.
Guanghui Wen, Junjie Fu, Zhexin Luo, Dezhi Zheng, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Spatial-Temporal Fusion Network for Unsupervised Ultrasound Video Object Segmentation
abstract
Automatic tracking and segmentation of lesions in ultrasound videos could assist in early diagnosis and treatment plan development. However, this task is quite challenging due to problems such as low visual saliency of the lesions and large variation between adjacent frames. In this paper, we develop a Spatial-Temporal Fusion Network (STFNet) for unsupervised ultrasound video object segmentation. First, an Edge Blur Enhancement Module is designed to extract and preserve the edge details of the target objects in ultrasound frames for spatial feature enhancement. Then, a Dynamic Alignment Module is developed to correct the inter-frame inconsistencies by aligning target objects from adjacent frames with those in the current frame for temporal feature enhancement. To incorporate both the spatial and temporal information, we further implement a mixed training strategy. These innovations collectively refine the model’s learning process and substantially boost segmentation accuracy. Extensive evaluations on real lymphoma ultrasound video data demonstrate the competitive segmentation results of STFNet. Specifically, compared with the best results among seven competing baselines, STFNet achieves the best scores in terms of region similarity ${\mathcal{J}}$, contour accuracy ${\mathcal{F}}$ as well as their average ${\mathcal{J}}\& {\mathcal{F}}$.
Dezhi Zheng, Qiao Pan, Dehua Chen, Jianwen Su
BIBM2
2024 Frequency-Adaptive Dilated Convolution for Semantic Segmentation
abstract
Dilated convolution, which expands the receptive field by inserting gaps between its consecutive elements, is widely employed in computer vision. In this study, we propose three strategies to improve individual phases of dilated convolution from the perspective of spectrum analysis. Departing from the conventional practice of fixing a global dilation rate as a hyperparameter, we introduce Frequency-Adaptive Dilated Convolution (FADC), which dynamically adjusts dilation rates spatially based on local frequency components. Subsequently, we design two plug-in modules to directly enhance effective bandwidth and receptive field size. The Adaptive Kernel (AdaKern) module decomposes convolution weights into low-frequency and high-frequency components, dynamically adjusting the ratio between these components on a per-channel basis. By increasing the high-frequency part of convolution weights, AdaKern captures more high-frequency components, thereby improving effective bandwidth. The Frequency Selection (FreqSelect) module optimally balances high- and low-frequency components in feature representations through spatially variant reweighting. It suppresses high frequencies in the background to encourage FADC to learn a larger dilation, thereby increasing the receptive field for an expanded scope. Extensive experiments on segmentation and object detection consistently validate the efficacy of our approach. The code is made publicly available at https://github.com/ying-fu/FADC.
Lin Gu 0003, Dezhi Zheng, Ying Fu 0001
CVPR3
2024 Learning Visual Prompt for Gait Recognition
abstract
Gait, a prevalent and complex form of human motion, plays a significant role in the field of long-range pedestrian retrieval due to the unique characteristics inherent in individual motion patterns. However, gait recognition in real-world scenarios is challenging due to the limitations of capturing comprehensive cross-viewing and crossclothing data. Additionally, distractors such as occlusions, directional changes, and lingering movements further complicate the problem. The widespread application of deep learning techniques has led to the development of various potential gait recognition methods. However, these methods utilize convolutional networks to extract shared information across different views and attire conditions. Once trained, the parameters and non-linear function become constrained to fixed patterns, limiting their adaptability to various distractors in real-world scenarios. In this paper, we present a unified gait recognition framework to extract global motion patterns and develop a novel dynamic transformer to generate representative gait features. Specifically, we develop a trainable part-based prompt pool with numerous key-value pairs that can dynamically select prompt templates to incorporate into the gait sequence, thereby providing task-relevant shared knowledge information. Furthermore, we specifically design dynamic attention to extract robust motion patterns and address the length generalization issue. Extensive experiments on four widely recognized gait datasets, i.e., Gait3D, GREW, OUMVLP, and CASIA-B, reveal that the proposed method yields substantial improvements compared to current state-of-the-art approaches.
Ying Fu 0001, Chunshui Cao, Saihui Hou, Yongzhen Huang, Dezhi Zheng
CVPR6
2024 Multicarrier Waveform Design for mmWave/THz Integrated Sensing and Communication
abstract
Integrated sensing and communication (ISAC) is recognized as one of key enabling technologies for the Meta-verse. To enhance both communication data rate and sensing accuracy, the exploitation of millimeter wave (mmWave) and terahertz (THz) frequencies becomes mandatory due to huge amount of spectrum resources. To combat the severe path-loss at mmWave/THz band, large-scale antenna arrays are usually employed to form directional beams. However, the ultra-broad bandwidth induces undesirable beam squint (BS) effects, where the beams from different subcarriers point to diverse angles, leading to the communication performance loss. Fortunately, this BS effect can be leveraged to facilitate the multi-angle super-resolution sensing with minimal beam sweeping overhead. Against this background, we propose a novel multicarrier waveform design methodology for the BS-assisted ISAC systems, which optimizes the frequency resource allocation between sensing and communications to reach a good dual-functional performance trade-off. To mitigate mutual interference, both functions are assigned non-overlapping subcarriers. The subcarrier assignment design is formulated as a mixed integer programming problem to maximize the communication throughput while ensuring the required sensing range and resolution, which involves high complexity to get an exact solution. To this end, we propose a two-stage iterative update algorithm to obtain a quasi-optimal solution with low computational complexity. Numerical results demonstrate that our proposed methodology achieves high-rate communication and high-resolution sensing simultaneously with relatively low overhead.
Fan Zhang 0071, Tianqi Mao 0001, Ruiqi Liu 0002, Leyi Zhang, Dezhi Zheng, Zhaocheng Wang 0001
IWCMC5
2024 FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-Training
abstract
While significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP.
Yudong Li 0001, Xianxu Hou, Dezhi Zheng, LinLin Shen, Zhe Zhao 0006
ACM Multimedia3
2024 Temperature compensation for humidity sensors using ISSA-BP neural network
abstract
High-precision humidity detection is essential in various fields. However, the sensor's output signal is often affected by complex environmental temperature changes. To address these challenges, this paper propose an improved sparrow search algorithm based on the back propagation neural network (ISSA-BP). This method optimizes the initialization process of the traditional algorithm and introduces an edge evolution strategy to enhance its iterative update process, significantly improving both efficiency and accuracy. Experimental results demonstrate that the proposed algorithm reduces the mean absolute percentage error (MAPE) from 4.35% to 0.92% and decreases the convergence time from 1.25 s to 0.65 s, compared to traditional methods.
Hechu Zhang, Wei Li 0181, Shuai Wang 0049, Dezhi Zheng
MobiCom6
2024 A classification and quantitative assessment method for internal and external surface defects in pipelines based on ASTC-Net
Mengtian Qiao, Chun Hu, Yufei Cheng, Dezhi Zheng
Adv. Eng. Informatics6
2024 Multi-UAV Collaborative Surveillance Network Recovery via Deep Reinforcement Learning
abstract
As a typical nonterrestrial network (NTN)-enabled Internet of Things (IoT), the multi-Unmanned aerial vehicle (UAV) collaborative surveillance network boasts efficient capabilities in information collection and transmission. However, manufacturing techniques and environmental conditions can lead to UAV failures, thereby impacting network performance. To recover the performance of the multi-UAV collaborative surveillance network, the effective movement of multiple UAVs is under investigation in order to improve target coverage and data backhaul efficiency. In this article, we present a novel multiagent deep reinforcement learning-based algorithm to accomplish network recovery. The proposed algorithm employs a multihead attention network to facilitate coupled multiobjective learning and overcome the limitations imposed by local information. Additionally, a stable learning method is introduced to address the difficult convergence problem caused by dynamic topology changes due to UAV motion. Experimental results show that the proposed algorithm can generate feasible multi-UAV motion strategies, effectively facilitating network recovery and improving the performance of the multi-UAV collaborative surveillance network in different scenarios.
Tao Wang 0151, Jingjing Wang 0001, Wenbo Du 0001, Dezhi Zheng, Shuai Wang 0049
IEEE Internet Things J.5
2024 Category-Level Band Learning-Based Feature Extraction for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a classical task in remote sensing image analysis. With the development of deep learning, schemes based on deep learning have gradually become the mainstream of HSI classification. However, existing HSI classification schemes either lack the exploration of category-specific information in the spectral bands and the intrinsic value of information contained in features at different scales, or are unable to extract multiscale spatial information and global spectral properties simultaneously. To solve these problems, in this article, we propose a novel HSI classification framework named CL-MGNet, which can fully exploit the category-specific properties in spectral bands and obtain features with multiscale spatial information and global spectral properties. Specifically, we first propose a spectral weight learning (SWL) module with a category consistency loss to achieve the enhancement of information in important bands and the mining of category-specific properties. Then, a multiscale backbone is proposed to extract the spatial information at different scales and the cross-channel attention via multiscale convolution and a grouping attention module. Finally, we employ an attention multilayer perceptron (attention-MLP) block to exploit the global spectral properties of HSI, which is helpful for the final fully connected layer to obtain the classification result. The experimental results on five representative hyperspectral remote sensing datasets demonstrate the superiority of our method.
Ying Fu 0001, Hongrong Liu, Yunhao Zou, Shuai Wang 0049, Zhongxiang Li, Dezhi Zheng
IEEE Trans. Geosci. Remote. Sens.6
2023 Dynamic Aggregated Network for Gait Recognition
abstract
Gait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exterior factors in real scenes, such as carrying conditions, wearing overcoats, and diverse viewing angles. Recently, various deep learning-based gait recognition methods have achieved promising results, but they tend to extract one of the salient features using fixed-weighted convolutional networks, do not well consider the relationship within gait features in key regions, and ignore the aggregation of complete motion patterns. In this paper, we propose a new perspective that actual gait features include global motion patterns in multiple key regions, and each global motion pattern is composed of a series of local motion patterns. To this end, we propose a Dynamic Aggregation Network (DANet) to learn more discriminative gait features. Specifically, we create a dynamic attention mechanism between the features of neighboring pixels that not only adaptively focuses on key regions but also generates more expressive local motion patterns. In addition, we develop a selfattention mechanism to select representative local motion patterns and further learn robust global motion patterns. Extensive experiments on three popular public gait datasets, i.e., CASIA-B, OUMVLP, and Gait3D, demonstrate that the proposed method can provide substantial improvements over the current state-of-the-art methods.1
Ying Fu 0001, Dezhi Zheng, Chunshui Cao, Xuecai Hu, Yongzhen Huang
CVPR3
2023 Fine-grained Unsupervised Domain Adaptation for Gait Recognition
abstract
Gait recognition has emerged as a promising technique for the long-range retrieval of pedestrians, providing numerous advantages such as accurate identification in challenging conditions and non-intrusiveness, making it highly desirable for improving public safety and security. However, the high cost of labeling datasets, which is a prerequisite for most existing fully supervised approaches, poses a significant obstacle to the development of gait recognition. Recently, some unsupervised methods for gait recognition have shown promising results. However, these methods mainly rely on a fine-tuning approach that does not sufficiently consider the relationship between source and target domains, leading to the catastrophic forgetting of source domain knowledge. This paper presents a novel perspective that adjacent-view sequences exhibit overlapping views, which can be leveraged by the network to gradually attain cross-view and cross-dressing capabilities without pre-training on the labeled source domain. Specifically, we propose a fine-grained Unsupervised Domain Adaptation (UDA) framework that iteratively alternates between two stages. The initial stage involves offline clustering, which transfers knowledge from the labeled source domain to the unlabeled target domain and adaptively generates pseudo-labels according to the expressiveness of each part. Subsequently, the second stage encompasses online training, which further achieves cross-dressing capabilities by continuously learning to distinguish numerous features of source and target domains. The effectiveness of the proposed method is demonstrated through extensive experiments conducted on widely-used public gait datasets.
Ying Fu 0001, Dezhi Zheng, Yunjie Peng, Chunshui Cao, Yongzhen Huang
ICCV3
2023 Instance Segmentation in the Dark
Ying Fu 0001, Kaixuan Wei, Dezhi Zheng, Felix Heide
Int. J. Comput. Vis.4
2023 Next-Generation URLLC With Massive Devices: A Unified Semi-Blind Detection Framework for Sourced and Unsourced Random Access
abstract
This paper proposes a unified semi-blind detection framework for sourced and unsourced random access (RA), which enables next-generation ultra-reliable low-latency communications (URLLC) with a massive number of devices. Specifically, the active devices transmit their uplink access signals in a grant-free manner to realize ultra-low access latency. Meanwhile, the base station aims to achieve ultra-reliable data detection under severe inter-device interference without exploiting explicit channel state information (CSI). We first propose an efficient transmitter design, where a small amount of reference information (RI) is embedded in the access signal to resolve the inherent ambiguities incurred by the unknown CSI. At the receiver, we further develop a successive interference cancellation-based semi-blind detection scheme, where a bilinear generalized approximate message passing algorithm is utilized for joint channel and signal estimation (JCSE), while the embedded RI is exploited for ambiguity elimination. Particularly, a rank selection approach and a RI-aided initialization strategy are incorporated to reduce the algorithmic computational complexity and to enhance the JCSE reliability, respectively. Besides, four enabling techniques are integrated to satisfy the stringent latency and reliability requirements of massive URLLC. Numerical results demonstrate that the proposed semi-blind detection framework offers a better scalability-latency-reliability tradeoff than the state-of-the-art detection schemes dedicated to sourced or unsourced RA.
Malong Ke, Zhen Gao 0001, Dezhi Zheng, Derrick Wing Kwan Ng, H. Vincent Poor
IEEE J. Sel. Areas Commun.4
2023 Quasi-Synchronous Random Access for Massive MIMO-Based LEO Satellite Constellations
abstract
Low earth orbit (LEO) satellite constellation-enabled communication networks are expected to be an important part of many Internet of Things (IoT) deployments due to their unique advantage of providing seamless global coverage. In this paper, we investigate the random access problem in massive multiple-input multiple-output-based LEO satellite systems, where the multi-satellite cooperative processing mechanism is considered. Specifically, at edge satellite nodes, we conceive a training sequence padded multi-carrier system to overcome the issue of imperfect synchronization, where the training sequence is utilized to detect the devices’ activity and estimate their channels. Considering the inherent sparsity of terrestrial-satellite links and the sporadic traffic feature of IoT terminals, we utilize the orthogonal approximate message passing-multiple measurement vector algorithm to estimate the delay coefficients and user terminal activity. To further utilize the structure of the receive array, a two-dimensional estimation of signal parameters via rotational invariance technique is performed for enhancing channel estimation. Finally, at the central server node, we propose a majority voting scheme to enhance activity detection by aggregating backhaul information from multiple satellites. Moreover, multi-satellite cooperative linear data detection and multi-satellite cooperative Bayesian dequantization data detection are proposed to cope with perfect and quantized backhaul, respectively. Simulation results verify the effectiveness of our proposed schemes in terms of channel estimation, activity detection, and data detection for quasi-synchronous random access in satellite systems.
Keke Ying, Zhen Gao 0001, Sheng Chen 0001, Dezhi Zheng, Symeon Chatzinotas, Björn Ottersten 0001, H. Vincent Poor
IEEE J. Sel. Areas Commun.5
2023 Distributed Nash Equilibrium Seeking in Consistency-Constrained Multicoalition Games
abstract
The distributed Nash equilibrium (NE) seeking problem for multicoalition games has attracted increasing attention in recent years, but the research mainly focuses on the case without agreement demand within coalitions. This article considers a class of networked games among multiple coalitions where each coalition contains multiple agents that cooperate to minimize the sum of their costs, subject to the demand of reaching an agreement on their state values. Furthermore, the underlying network topology among the agents does not need to be balanced. To achieve the goal of NE seeking within such a context, two estimates are constructed for each agent, namely, an estimate of partial derivatives of the cost function and an estimate of global state values, based on which, an iterative state updating law is elaborately designed. Linear convergence of the proposed algorithm is demonstrated. It is shown that the consistency-constrained multicoalition games investigated in this article put the well-studied networked games among individual players and distributed optimization in a unified framework, and the proposed algorithm can easily degenerate into solutions to these problems.
Jialing Zhou, Yuezu Lv, Guanghui Wen, Jinhu Lü 0001, Dezhi Zheng
IEEE Trans. Cybern.5
2023 Integrated Sensing and Communication With mmWave Massive MIMO: A Compressed Sampling Perspective
abstract
Integrated sensing and communication (ISAC) has opened up numerous game-changing opportunities for realizing future wireless systems. In this paper, we propose an ISAC processing framework relying on millimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) systems. Specifically, we provide a compressed sampling (CS) perspective to facilitate ISAC processing, which can not only recover the high-dimensional channel state information or/and radar imaging information, but also significantly reduce pilot overhead. First, an energy-efficient widely spaced array (WSA) architecture is tailored for the radar receiver, which enhances the angular resolution of radar sensing at the cost of angular ambiguity. Then, we propose an ISAC frame structure for time-varying ISAC systems considering different timescales. The pilot waveforms are judiciously designed by taking into account both CS theories and hardware constraints induced by hybrid beamforming (HBF) architecture. Next, we design the dedicated dictionary for WSA that serves as a building block for formulating the ISAC processing as sparse signal recovery problems. The orthogonal matching pursuit with support refinement (OMP-SR) algorithm is proposed to effectively solve the problems in the existence of the angular ambiguity. We also provide a framework for estimating the Doppler frequencies during payload data transmission to guarantee communication performances. Simulation results demonstrate the good performances of both communications and radar sensing under the proposed ISAC framework.
Zhen Gao 0001, Ziwei Wan, Dezhi Zheng, Shufeng Tan, Christos Masouros, Derrick Wing Kwan Ng, Sheng Chen 0001
IEEE Trans. Wirel. Commun.3
2022 Riemannian Geometric Instance Filtering for Transfer Learning in Brain-Computer Interfaces
abstract
Due to the inter-subject variability of Electroencephalogram(EEG) signals, a long calibration time is required to collect a large number of labeled trials to calibrate classifier parameters before using the Brain-computer Interface(BCI). This challenge greatly limits the practical roll-out of BCIs. To address this problem, we propose a novel instance-based transfer learning framework named Riemannian Geometric Instance Filtering (RGIF) to reduce calibration time without sacrificing accuracy. A new inter-subject similarity metric based on Riemannian geometry is proposed to measure the similarity between a few trials from the target subject and adequate trials from source subjects. The classification model for the target subject is then trained with the help of abundant trials from similar source subjects with high similarity to the target subject. We evaluate our method on two open-source EEG datasets. The results show that our approach improves significantly compared with other baselines. Furthermore, compared with using all source subjects data, our method reduces the training time by at least half and achieves slightly better accuracy.
Qianxin Hui, Yang Li 0104, Susu Xu, Shuailei Zhang, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng
SenSys9
2022 C-RIDGE: Indoor CO2 Data Collection System for Large Venues Based on prior Knowledge
abstract
CO2 concentration data with high resolution in large venues is highly required during indoor sport events for in-time environment adjustment to guarantee the athlete performances and audience experience. However, the limited battery energy of the wireless sensors cannot support high data resolution and long time coverage simultaneously. Besides, there also lacks effective embedded methods to clean anomaly data caused by the human and environmental factors probably occurring in large venues. Thus, in this paper, we propose C-RIDGE, a low-power sensing system for high resolution CO2 data collection in large venues. Based on prior knowledge, firstly, an adaptive sampling rate adjustment policy is developed for lower energy consumption to extend the time coverage of data. Secondly, CO2 physical property (CPP) aided data cleaning algorithm is designed to improve data quality as well, using Pearson Correlation Coefficient (PCC) and standard deviation with sliding windows. C-RIDGE has been deployed in one venue during a world-class event. The experiments and collected data have shown the system power consumption can be reduced by 36.1%, with measurement error less than 10.2%. The outliers and anomaly trends can also be detected and calibrated effectively via CPP algorithm. The dataset is available at https://doi.org/10.5281/zenodo.7160830.
Yuxuan Liu 0010, Xiaolei Qu, Dezhi Zheng, Xinlei Chen
SenSys5
2022 Non-Acoustic Speech Sensing System Based on Flexible Piezoelectric
abstract
Speech is one of the most important biological signals to complement human-human and human-computer interaction. Traditional speech datasets were collected by air microphones, but using these datasets in noisy environments such as factories is practically challenging. Therefore, speech recognition in noisy environments poses higher requirements. The non-acoustic speech dataset plays a significant role in robust speech recognition under high background noise. Existing datasets suffered from dull sound, low intelligibility and poor recognition accuracy due to hardware and computer technology limitations. This paper presents a non-acoustic speech sensing system based on flexible piezoelectric. The system collected vibration signals from the jaws of six males and five females, and the corpus contained ten different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and capable of achieving 93.7% recognition accuracy by calculation. With the aforementioned benefits, this dataset is an essential tool for studying human-computer interaction in high-noise environments, analyzing human acoustic properties, and aiding medical rehabilitation.
Shiji Yuan, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng, Shangchun Fan
SenSys6
2022 A Wearable Low-Power Collaborative Sensing System for High-Quality SSVEP-BCI Signal Acquisition
abstract
The brain–computer interface (BCI) technology improves the communication efficiency between people and Internet of Things (IoT) devices. BCI based on the steady-state visual evoked potential (SSVEP-BCI) is the preferred scheme for controlling devices because of its convenient operation, low training requirement, and high information transmission rate (ITR). Most signal acquisition devices for BCIs are used for medical diagnosis and scientific research and utilize multiple channels and wet electrodes to obtain high-quality signals. However, the practicability, wearability, and cost of the signal acquisition devices for real-life applications need to be considered, resulting in new requirements for the acquisition mode, the number of electrodes, power consumption, and signal processing methods. This article presents a wearable low-power collaborative sensing system based on a time mask window canonical correlation analysis method (TMW-CCA). An 8-array spring dry electrode signal acquisition device based on a flexible circuit board is designed to address the shortcomings of traditional wet electrode acquisition devices, such as high-power consumption, discomfort, and being unsuitable for long-time use. The proposed TMW-CCA method, which uses a dry electrode sensor to evaluate the time domain’s signal quality dynamically, exhibits 12.5% higher steady-state visual evoked potential recognition accuracy and 40% lower average power consumption (only 740 mW) than the benchmark.
Rui Na, Dezhi Zheng, Ying Sun 0012, Mingzhe Han, Shuai Wang 0049, Shuailei Zhang, Qianxin Hui, Xinlei Chen, Jun Zhang 0007, Chun Hu
IEEE Internet Things J.2
2022 Trajectory Design for UAV-Based Internet of Things Data Collection: A Deep Reinforcement Learning Approach
abstract
In this article, we investigate an unmanned aerial vehicle (UAV)-assisted Internet of Things (IoT) system in a sophisticated 3-D environment, where the UAV’s trajectory is optimized to efficiently collect data from multiple IoT ground nodes. Unlike existing approaches focusing only on a simplified 2-D scenario and the availability of perfect channel state information (CSI), this article considers a practical 3-D urban environment with imperfect CSI, where the UAV’s trajectory is designed to minimize data collection completion time subject to practical throughput and flight movement constraints. Specifically, inspired by the state-of-the-art deep reinforcement learning approaches, we leverage the twin-delayed deep deterministic policy gradient (TD3) to design the UAV’s trajectory and we present a TD3-based trajectory design for completion time minimization (TD3-TDCTM) algorithm. In particular, we set an additional information, i.e., the merged pheromone, to represent the state information of the UAV and environment as a reference of reward which facilitates the algorithm design. By taking the service statuses of the IoT nodes, the UAV’s position, and the merged pheromone as input, the proposed algorithm can continuously and adaptively learn how to adjust the UAV’s movement strategy. By interacting with the external environment in the corresponding Markov decision process, the proposed algorithm can achieve a near-optimal navigation strategy. Our simulation results show the superiority of the proposed TD3-TDCTM algorithm over three conventional nonlearning-based baseline methods.
Yang Wang 0154, Zhen Gao 0001, Jun Zhang 0007, Xianbin Cao 0001, Dezhi Zheng, Yue Gao 0001, Derrick Wing Kwan Ng, Marco Di Renzo
IEEE Internet Things J.5
2022 Ultralow-Power Sensing Framework for Internet of Things: A Smart Gas Meter as a Case
abstract
Gas serves as one of the most indispensable energy sources for industrial production and household life. In order to improve the gas utilization efficiency, smart gas meters for the Internet of Things (IoT) has been designed to achieve two-way communication and remote control functions. However, the existing smart gas meter does not consider further reduction of power consumption. Thus, we designed an ultralow-power sensing framework for IoT and applied it to the smart gas meter. Based on a thorough analysis of the metering system, we propose an ultralow power system framework and a low-power peripheral management solution to reach ultralow power consumption. Also, we propose a cooperative sensing scheme to achieve stability and accuracy of gas volume detection. Finally, we implement a real smart gas manage system to evaluate our low-power smart gas meter solution. Verified by a comparative test, the proposed gas meter successfully reduces the power consumption by at least 37% compared to the baseline. The experimental results verify the innovative design and confirm that the proposed gas meter features ultralow power consumption, high precision, and high reliability.
Ziteng Wang 0007, Chun Hu, Dezhi Zheng, Xinlei Chen
IEEE Internet Things J.3
2022 TCACNet: Temporal and channel attention convolutional network for motor imagery classification of EEG-based BCI
abstract
Brain–computer interface (BCI) is a promising intelligent healthcare technology to improve human living quality across the lifespan, which enables assistance of movement and communication, rehabilitation of exercise and nerves, monitoring sleep quality, fatigue and emotion. Most BCI systems are based on motor imagery electroencephalogram (MI-EEG) due to its advantages of sensory organs affection, operation at free will and etc. However, MI-EEG classification, a core problem in BCI systems, suffers from two critical challenges: the EEG signal’s temporal non-stationarity and the nonuniform information distribution over different electrode channels. To address these two challenges, this paper proposes TCACNet, a temporal and channel attention convolutional network for MI-EEG classification. TCACNet leverages a novel attention mechanism module and a well-designed network architecture to process the EEG signals. The former enables the TCACNet to pay more attention to signals of task-related time slices and electrode channels, supporting the latter to make accurate classification decisions. We compare the proposed TCACNet with other state-of-the-art deep learning baselines on two open source EEG datasets. Experimental results show that TCACNet achieves 11.4% and 7.9% classification accuracy improvement on two datasets respectively. Additionally, TCACNet achieves the same accuracy as other baselines with about 50% less training data. In terms of classification accuracy and data efficiency, the superiority of the TCACNet over advanced baselines demonstrates its practical value for BCI systems.
Rongye Shi, Qianxin Hui, Susu Xu, Shuai Wang 0049, Rui Na, Ying Sun 0012, Wenbo Ding 0001, Dezhi Zheng, Xinlei Chen
Inf. Process. Manag.9
2022 Data-Driven Deep Learning Based Hybrid Beamforming for Aerial Massive MIMO-OFDM Systems With Implicit CSI
abstract
In an aerial hybrid massive multiple-input multiple-output (MIMO) and orthogonal frequency division multiplexing (OFDM) system, how to design a spectral-efficient broadband multi-user hybrid beamforming with a limited pilot and feedback overhead is challenging. To this end, by modeling the key transmission modules as an end-to-end (E2E) neural network, this paper proposes a data-driven deep learning (DL)-based unified hybrid beamforming framework for both the time division duplex (TDD) and frequency division duplex (FDD) systems with implicit channel state information (CSI). For TDD systems, the proposed DL-based approach jointly models the uplink pilot combining and downlink hybrid beamforming modules as an E2E neural network. While for FDD systems, we jointly model the downlink pilot transmission, uplink CSI feedback, and downlink hybrid beamforming modules as an E2E neural network. Different from conventional approaches separately processing different modules, the proposed solution simultaneously optimizes all modules with the sum rate as the optimization object. Therefore, by perceiving the inherent property of air-to-ground massive MIMO-OFDM channel samples, the DL-based E2E neural network can establish the mapping function from the channel to the beamformer, so that the explicit channel reconstruction can be avoided with reduced pilot and feedback overhead. Besides, practical low-resolution phase shifters (PSs) introduce the quantization constraint, leading to the intractable gradient backpropagation when training the neural network. To mitigate the performance loss caused by the phase quantization error, we adopt the transfer learning strategy to further fine-tune the E2E neural network based on a pre-trained network that assumes the ideal infinite-resolution PSs. Numerical results show that our DL-based schemes have considerable advantages over state-of-the-art schemes.
Zhen Gao 0001, Minghui Wu 0002, Chun Hu, Feifei Gao 0001, Guanghui Wen, Dezhi Zheng, Jun Zhang 0007
IEEE J. Sel. Areas Commun.6
2022 Joint Activity and Blind Information Detection for UAV-Assisted Massive IoT Access
abstract
International audience
Li Qiao 0001, Jun Zhang 0007, Zhen Gao 0001, Dezhi Zheng, Md. Jahangir Hossain 0002, Yue Gao 0001, Derrick Wing Kwan Ng, Marco Di Renzo
IEEE J. Sel. Areas Commun.4
2019 A novel pattern with high-level commands for encoding motor imagery-based brain computer interface
Shuailei Zhang, Shuai Wang 0049, Dezhi Zheng, Mengxi Dai
Pattern Recognit. Lett.3