Yang Hu 0006

dblp:43/4685-6 · DBLP profile ↗
← Back
93ranked-venue papers
7as first author
67since 2021 · last 2026
0000-0003-0379-1525ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-author · 31 since 2021Computer networks · 37 · 1 first-author · 25 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Towards Practical Reliability Assessment of RF-based Heartbeat Sensing with Multi-Domain Analysis
Hanqin Gong, Jinbo Chen 0001, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
ISCAS6
2026 Adversarially Regularized Latent Flow for Enhanced Conditional Video Generation
Jinduo Wang, Binquan Wang, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
ISCAS5
2026 Contactless Premature Ventricular Contractions Diagnosis With Radio Frequency Signals Based on Deep Auxiliary Learning
abstract
Cardiovascular diseases dominate global mortality, with premature ventricular contractions (PVCs) among the most prevalent arrhythmias. PVCs are paroxysmal and sporadic. If undetected, they may escalate to lethal arrhythmias or sudden cardiac death, underscoring the imperative for early and accurate diagnosis. Conventional monitoring via electrocardiogram (ECG) necessitates prolonged skin-electrode contact, provoking discomfort and potential allergic reactions, whereas photoplethysmography (PPG) is intrinsically susceptible to ambient light and motion artifacts. These limitations restrict both modalities to short-term or intermittent use. Radio-frequency (RF) sensing offers a contactless alternative for unobtrusive, long-term cardiac surveillance. The cardiac-induced mechanical displacements, however, are orders of magnitude smaller than respiratory and body-motion artifacts, rendering PVC-related signatures in RF echoes highly susceptible to noise and consequently difficult to extract and classify. To handle this, we propose a PVC diagnosis model based on RF signals, which constructs an auxiliary learning framework, thereby improving the accuracy of PVC diagnosis. To further enhance the effectiveness of the auxiliary framework, we designed a residual-assisted Mixture of Experts architecture, which effectively alleviates the gradient conflict between the auxiliary task and the main task. We conducted experiments on a large-scale dataset (7,015 subjects) in a hospital outpatient setting, where our system achieved the best performance, demonstrating the superiority of our approach.
Xilong Yuan, Zehan Guo, Yang Hu 0006, Yan Chen 0007
IEEE Internet Things J.4
2026 mmGuard: A Countermeasure Against Physical Adversarial Attacks on mmWave Radar Sensing
abstract
Physical adversarial attacks (PAAs) pose a serious security threat to millimeter-wave (mmWave) radar systems used in safety-critical applications such as autonomous driving and security checking. These attacks, manipulating radar signals via specially crafted materials, are proven feasible and can cause severe sensing failures; however, effective defenses remain unexplored due to the difficulty of distinguishing adversarial examples from normal environmental objects. This paper presents mmGuard, a physics-based defense framework that addresses this challenge by exploiting a fundamental insight: the engineering process that makes materials adversarial inevitably creates detectable physical signatures. We identify three key domains where adversarial examples show artificial nature: spatial phase discontinuities, anomalous radar cross-section patterns, and violations of natural physico-kinematic relationships. mmGuard systematically captures these signatures through multi-domain feature extraction, enhances their discriminability via neural refinement, and enables efficient per-object attack detection and mitigation compatible with automotive radar update rates. To enable evaluation, we introduce mmAD, comprising over 110,000 annotated radar frames with diverse adversarial examples across realistic deployment scenarios. Experimental results demonstrate that mmGuard achieves over 90% detection accuracy while exhibiting strong in-distribution performance, with few-shot adaptation enabling calibration to unseen settings Case studies further validate that mmGuard can reliably defend against PAAs in real-world settings.
Ruixu Geng, Dongheng Zhang, Jianyang Wang, Qian Liang 0001, Rui Zhang 0120, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Inf. Forensics Secur.9
2026 Authentication With Passports for Deep RF Sensing Model Protection
abstract
As RF sensing increasingly moves toward real-world deployments, critical concerns emerge around unauthorized model usage and access control. Existing protection approaches-such as watermarking-are typically reactive, task-specific, and ineffective at runtime. This paper presents AuthRF (Authentication with passports for RF sensing models), a novel signal-level passport mechanism that proactively enforces access control by mapping a user-specific passport to phase-compensation weights in the signal processing pipeline. Valid passports yield coherent phase alignment and high-fidelity representations, while invalid or forged ones induce phase distortion that significantly degrades model performance. This design effectively deters unauthorized access, supports scalable multi-user authentication, and enables personalized service provisioning through controlled passport variation. We evaluate AuthRF on six representative RF sensing tasks using both WiFi and radar signals. Experimental results demonstrate its robust protection capabilities and seamless integration with existing sensing pipelines, positioning AuthRF as a practical foundation for secure and commercial-grade RF sensing deployment.
Ruiyuan Song, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Inf. Forensics Secur.6
2026 Radar HRV Monitoring With Physiological Prior Inspired Deep Neural Networks
abstract
Radar sensing has emerged as a promising solution for the contactless monitoring of Heart Rate Variability (HRV), a crucial indicator of the cardiovascular and autonomic nervous systems. However, due to signal noise and interference that easily obscure heartbeat details, along with variations in heartbeat across different physiological conditions, existing methods remain restricted to laboratory settings with healthy subjects and fail in real-world scenarios involving more complex physiological conditions. In this study, we propose a physiological prior-inspired deep learning framework for robust radar-based HRV monitoring. Specifically, we leverage the prior that internal heartbeats drive movements across the entire torso surface and design a hybrid deep neural network to model the spatio-temporal relationship between full-body radio reflections and heartbeats, effectively mitigating interference. Then, we incorporate the cardiac motion's self-similarity prior to establish a signal augmentation strategy, effectively remodeling the HRV distribution and enhancing performance across diverse physiological conditions. We build and validate our method on a large-scale dataset comprising 7,150 outpatients with complex physiological conditions in real-world scenarios. The experimental results demonstrate that our method achieves a mean IBI error of 19.21 ms, an RMSSD error of 16.23 ms, an SDSD error of 16.70 ms, and a pNN50 error of 7.28%. We further validate the performance by classifying five common cardiac conditions based on HRV results, demonstrating performance comparable to ECG-based methods. These results highlight the great potential of our approach for accurate, contactless HRV monitoring in real-world applications.
Jinbo Chen 0001, Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE J. Biomed. Health Informatics5
2026 WN-Sleep: Modeling Whole-Night Data for Improved Sleep Staging Classification
abstract
Sleep staging, crucial for diagnosing sleep disorders, requires precise recognition of physiological signals within 30-second epochs, a task fundamentally different from managing long-term semantic dependencies in natural language processing (NLP). Our model aims to refine the integration of local and global features for more accurate sleep stage classification. Following the American Academy of Sleep Medicine (AASM) guidelines, it focuses on rigorous intra-epoch feature extraction to ensure reliable identification of sleep stages. Moreover, our approach incorporates a global perspective by analyzing whole-night data, which is essential for handling transitional periods and ambiguities. Existing sequential modeling techniques often overlook the unique requirements of sleep staging, leading to performance declines when epochs extend beyond approximately 200. Our model addresses this by structurally processing local and global information and carefully balancing detailed intra-epoch analysis with an overarching view of sleep cycles through a gating mechanism. This gate mechanism selectively integrates long-term dependencies, optimizing the balance between local accuracy and global context. This approach represents a significant advancement over existing models, offering more accurate, reliable, and clinically relevant sleep staging. Extensive experiments on the SHHS, SleepEDF-20, and SleepEDF-78 datasets demonstrate that our method outperforms state-of-the-art approaches.
Gaohan Ye, Lingjie Shu, Yu Pu, Beilei Wang, Dong Zhang 0015, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE J. Biomed. Health Informatics10
2026 Contactless Arrhythmia Detection via Diversity-Invariant Contrastive mmWave Sensing
abstract
Arrhythmias are prevalent cardiac disorders affecting millions worldwide. By analyzing cardiac motion modulated in mmWave reflections, mmWave sensing is emerging as a promising contactless revolution in arrhythmia detection compared to conventional contact-based methods. However, the fundamental bottleneck of existing mmWave sensing methods is their restriction to controlled laboratory settings with small-scale cohorts, limiting generalization to real-world populations. This limitation arises because mmWave signals undergo complex signal transformations during propagation, resulting in an explosion of signal diversity across large populations in real-world scenarios. Such diversity significantly complicates the direct recognition of arrhythmia. In this paper, we theoretically analyze the mechanism and impact of mmWave cardiac signal diversity. Leveraging the inherent transformation properties of mmWave signals, we propose a Diversity-Invariant Contrastive mmWave Sensing framework, which learns invariant features robust to complex signal transformations encountered in real-world scenarios. We evaluate our method in a practical, clinically-oriented scenario involving a large-scale population of 7,338 subjects, achieving an average F1-score of 0.8241 across four common arrhythmias. These results demonstrate that our method effectively bridges the diversity gap, representing a significant step toward practical clinical deployment of contactless arrhythmia detection via mmWave sensing.
Xinmeng Cai, Jinbo Chen 0001, Yuqin Yuan, Guixin Xu, Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.8
2026 Physio-BP: Physiologically-Guided Dual Radar System for Continuous Blood Pressure Estimation
abstract
Continuous blood pressure (BP) estimation remains a significant challenge in non-contact hemodynamic monitoring, where existing methods struggle to strike a balance between clinical accuracy and usability. We present Physio-BP, a physiologically-guided dual-radar system enabling calibration-free, continuous tracking of systolic (SBP) and diastolic blood pressure (DBP) by integrating spatially resolved vascular dynamics. Unlike conventional single-waveform radar approaches, Physio-BP synchronously captures mechanical signals from the heart and pulse waves from the carotid artery, allowing precise measurement of pulse arrival time (PAT) and inter-beat intervals (IBI). To enhance signal quality, we introduce a novel joint signal selection algorithm that optimizes feature extraction across dual radars, effectively addressing the inherent signal-to-noise ratio (SNR) disparities between the chest and neck monitoring sites. To validate our system, an extensive dataset exceeding 50 hours was collected under diverse conditions, including different seasons, times of day, body postures, and radar devices, ensuring comprehensive coverage. Experimental results on the dataset confirm that Physio-BP enables continuous and precise SBP and DBP tracking, a capability not achieved by current baseline methods, thereby advancing non-contact sensing for hemodynamic monitoring.
Zhenzhen Cao, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Mob. Comput.3
2026 Automatic Phase Calibration for High-Resolution mmWave Sensing via Ambient Radio Anchors
abstract
Millimeter-wave (mmWave) radar systems with large array have pushed radar sensing into a new era, thanks to their high angular resolution. However, our long-term experiments indicate that array elements exhibit phase drift over time and require periodic phase calibration to maintain high-resolution, creating an obstacle for practical high-resolution mmWave sensing with large array. Unfortunately, existing calibration methods are inadequate for periodic recalibration, either because they rely on artificial references or fail to provide sufficient precision. To address this challenge, we introduce AutoCalib, the first framework designed to automatically and accurately calibrate high-resolution mmWave radars by identifying Ambient Radio Anchors (ARAs)—naturally existing objects in ambient environments that offer stable phase references. AutoCalib achieves calibration by first generating spatial spectrum templates based on theoretical electromagnetic characteristics. It then employs a pattern-matching and scoring mechanism to accurately detect these anchors and select the optimal one for calibration. Extensive experiments across 11 environments demonstrate that AutoCalib is capable of identifying ARAs that existing methods miss due to their focus on strong reflectors. AutoCalib's calibration performance approaches corner reflectors (74% phase error reduction) while outperforming existing methods by 83%. Beyond radar calibration, AutoCalib effectively supports other phase-dependent applications like handheld imaging, delivering 96% of corner reflector calibration performance without artificial references.
Ruixu Geng, Dongheng Zhang, Binquan Wang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Mob. Comput.8
2026 RF$^{2}$2 Transformer: Refocusing Transformer for RF Sensing
abstract
The learning-based RF sensing methods typically involve signal processing to transform the RF signals into spectrograms, which are then input into neural networks. However, this approach is suboptimal due to the potential degradation caused by the inherent multi-path interference, resulting in blurred RF spectrograms. To tackle this challenge, we introduce a novelRefocusingTransformerbackbone customized forRFsensing (RF$^{2}$Transformer). Rather than attempting to precisely model interference using the traditional signal processing techniques, the RF$^{2}$Transformer utilizes a self-compensation mechanism to treat the interference-induced phase shifts as learnable parameters. This mechanism learns and compensates for the phase shift caused by interference in the complex feature space to obtain refocused and high-quality feature maps of RF spectrograms, thereby improving the performance of downstream tasks. We show that the RF$^{2}$Transformer is general for various RF sensing tasks by evaluating it on six typical RF sensing tasks using two general RF signals (WiFi and radar). Experimental results indicate that the RF$^{2}$Transformer takes an important step toward learning-based solutions for RF sensing.
Ruiyuan Song, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Mob. Comput.7
2026 MaskSense: Motion-Robust Dynamic IBI Estimation via Deep RF Masked Learning
abstract
Although radio frequency (RF) sensing offers a promising approach for monitoring human cardiac activity, it traditionally requires subjects to remain in a steady-state throughout the monitoring process to avoid motion artifacts—an inherently impractical constraint that hampers real-world adoption. However, existing methods either yield incorrect estimates from motion-affected segments or discard them entirely, leading to fragmented data and incomplete observations. To address the challenge, we introduce MaskSense, a novel framework designed to address motion interference in long-term monitoring. The key insight is that latent patterns within dynamic inter-beat interval (IBI) sequences allow for accurate heartbeat reconstruction from incomplete observations. Leveraging this insight, MaskSense treats motion-affected periods as ”masked” and steady-state periods as ”unmasked”, and employs a contrastive-learning-assisted masked modeling architecture to reconstruct the masked information. Our 400-hour evaluation with 18 participants confirms that MaskSense effectively recovers IBIs in the presence of motion artifacts, paving the way for more natural and unobtrusive RF-based cardiac activity monitoring.
Jianyang Wang, Ruixu Geng, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Mob. Comput.5
2025 Passive Non-Line-of-Sight Imaging with Parallel Encoder
abstract
Passive non-line-of-sight (NLOS) imaging has developed rapidly in recent years. However, existing models generally suffer from low-quality reconstruction due to the severe loss of information during the projection process. In this paper, we introduce ParaEncodeNet, an NLOS imaging method for reconstructing high-quality, complex hidden scenes. Our approach utilizes a reconstruction network with parallel encoder to bridge the distribution gap between projection images and hidden images. The parallel encoder employs a codebook pretrained on a natural image dataset to construct a discrete prior, enabling the efficient encoding of projection images into hidden images. Moreover, we apply pixel-level constraints to the projection images to further reduce noise and distortion during reconstruction. Extensive experiments on a large-scale passive NLOS dataset have effectively demonstrated the superiority of our method over existing approaches, achieving a 1.2 dB increase in the Peak Signal-to-Noise Ratio (PSNR) metric. This validates the effectiveness and robustness of our proposed model in improving reconstruction quality and handling complex scenes.
Xiaolong Du, Ruixu Geng, Yan Chen 0007, Yang Hu 0006
ICASSP5
2025 Spatial Alignment and Temporal Matching Adapter for Video-Radar Remote Physiological Measurement
Qian Liang 0001, Ruixu Geng, Jinbo Chen 0001, Yan Chen 0007, Yang Hu 0006
ICCV6
2025 RFMamba: Frequency-Aware State Space Model for RF-Based Human-Centric Perception
abstract
Human-centric perception with radio frequency (RF) signals has recently entered a new era of end-to-end processing with Transformers. Considering the long-sequence nature of RF signals, the State Space Model (SSM) has emerged as a superior alternative due to its effective long-sequence modeling and linear complexity. However, integrating SSM into RF-based sensing presents unique challenges including the fundamentally different signal representation, distinct frequency responses in different scenarios, and incomplete capture caused by specular reflection. To address this, we carefully devise a dual-branch SSM block that is characterized by adaptively grasping the most informative frequency cues and the assistant spatial information to fully explore the human representations from radar echoes. Based on these two branchs, we further introduce an SSM-based network for handling various downstream human perception tasks, named RFMamba. Extensive experimental results demonstrate the superior performance of our proposed RFMamba across all three downstream tasks. To the best of our knowledge, RFMamba is the first attempt to introduce SSM into RF-based human-centric perception.
Rui Zhang 0120, Ruixu Geng, Ruiyuan Song, Hanqin Gong, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
ICLR7
2025 NLOS-R2: Alternate Reconstruction and Recognition for Non-Line-of-Sight Understanding
abstract
Passive non-line-of-sight (NLOS) imaging aims to recover hidden scenes from indirect reflections. While reconstruction has been extensively studied, the high-level tasks for understanding hidden scenes, such as recognition, remain insufficiently explored despite their importance for practical applications. Direct classifying using either the projection or reconstructed images yields limited performance due to the severe image degradation. In this paper, we propose NLOS-R2, an alternate reconstruction-recognition framework that leverages the complementary nature of both tasks to enhance NLOS scene understanding. By iteratively optimizing reconstruction and recognition networks, our framework effectively improves recognition accuracy while maintaining reconstruction quality. To enable systematic evaluation, we introduce the first large-scale multi-class passive NLOS dataset, containing 42 classes and 50,400 projection and hidden image pairs. Extensive experiments demonstrate that our approach achieves 52.88% recognition accuracy, significantly outperforming existing methods. The code and dataset are available at https://github.com/ustceewy/NLOS-R2.
Ruixu Geng, Xiaolong Du, Yan Chen 0007, Yang Hu 0006
ICME6
2025 RFinger: Environmental Fingerprint Embedding for Harmless mmWave Dataset Ownership Verification
abstract
The rapid evolution of millimeter-wave radar sensing technology has given rise to a proliferation of open-source radar datasets, creating an urgent need for innovative digital copyright protection techniques. However, conventional image and audio watermarking techniques are inadequate for radar copyright protection due to radar signals' sparsity, vulnerability and complexity. In this paper, we present RFinger, an ownership verification framework for static indoor millimeter-wave radar datasets. Our approach encodes environmental information extracted from radar signals into digital watermarks, strategically embedding these within carefully selected data frames to establish robust verification credentials. We develop statistical hypothesis testing metrics to detect unauthorized access to RFinger-protected data in black-box setting. Our strategic watermark design ensures that the unauthorized models exhibit distinctly anomalous performance on verification data compared to legitimate models. Through experiments on two large millimeter-wave radar datasets, we have validated that our designed strategy provides high watermark retrieval accuracy without compromising downstream tasks.
Zixin Shang, Jiamu Li, Yang Hu 0006, Yan Chen 0007
WISEC4
2025 OSense: Omni-Directional Heartbeat Sensing With Radio Signal
abstract
By analyzing cardiac motion modulated in body reflections, radio signals offer a novel contactless pathway for heartbeat sensing, attracting growing research attention. However, current studies overlooked the unique signal interaction when radio signals are incident at non-normal directions to the torso surface. This missing component significantly limits the effectiveness and results in unreliable performance in practical usage, where normal sensing direction cannot always be guaranteed. In this paper, we aim to answer the questions of what causes this performance degradation and how to solve it. Specifically, we analyze the signal interaction using a fine-grained thoracic motion model and reveal that non-stationary interference, caused by physiologically-driven spatial variation of the body surface, is the key to the problem. Correspondingly, we propose OSense, a framework based on a time-domain optimization method to cancel the non-stationary interference. This framework can be seamlessly integrated into various heartbeat sensing tasks. We validate OSense using a commercial Frequency Modulated Continuous Wave radar across multiple downstream tasks. The results demonstrate that our method effectively eliminates interference and enables direction-robust heartbeat sensing, highlighting its potential for practical cardiac monitoring using radio signals.
Hanqin Gong, Jinbo Chen 0001, Guixin Xu, Jianwen Tong, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Internet Things J.7
2025 RF-URL 2.0: A General Unsupervised Representation Learning Method for RF Sensing
abstract
The major challenge in learning-based RF sensing is acquiring high-quality large-scale annotated datasets. Unlike visual datasets, RF signals are inherently non-intuitive and non-interpretable, making their annotation both time-consuming and labor-intensive. To address this challenge, we propose RF-URL 2.0, a novel unsupervised representation learning (URL) framework for RF sensing, which enables pre-training on easily collected, large-scale unannotated RF datasets to make downstream tasks solve easier. Existing URL techniques, such as contrastive learning, are primarily designed for natural images and are prone to learn shortcuts rather than meaningful information when applied to RF signals. RF-URL 2.0 is the first framework to overcome these limitations by constructing positive and negative pairs through well-established RF signal processing algorithms. Besides, it introduces a novel signal-model-driven augmentation technique, which augments signal representations by identifying and perturbing physically meaningful parameters of signal processing models. Moreover, the RF-URL 2.0 is carefully designed to take into account the heterogeneity characteristics of different RF signal processing representations. We show the universality of RF-URL 2.0 in three typical RF sensing tasks using two general RF devices (WiFi and radar), including human gesture recognition, 3D pose estimation, and silhouette generation. Extensive experiments on the HIBER and WiDAR 3.0 datasets demonstrate that RF-URL 2.0 takes a significant step toward learning-based solutions for RF sensing.
Ruiyuan Song, Dongheng Zhang, Cong Yu 0011, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Attacking mmWave Imaging With Neural Meta-Material Rendering
abstract
Millimeter-wave (mmWave) radar imaging has shown remarkable potential in critical applications. While previous researches have explored attacks on high-level radar perception, the vulnerability of low-level radar imaging to adversarial attacks remains largely unexplored. In this work, we introduce mmHide, the first general attack framework on mmWave radar imaging that utilizes neural rendering of meta-materials to hide imaging targets (e.g., handguns). mmHide’s novelty lies in its three-fold approach: (1) an implicit neural rendering network that efficiently represents and optimizes complex 3D meta-material structures, (2) an explicit differentiable forward imaging model that provides physical constraints, and (3) a self-supervised learning strategy that iteratively refines the meta-material design. This unique combination enables mmHide to create an “invisible cloak” for target objects while maintaining plausible imaging results. Extensive real-world experiments demonstrate mmHide’s effectiveness in significantly reducing target visibility while preserving background similarity. A user study confirms its high success rate in deceiving human observers, outperforming existing methods. These findings not only showcase the potential of our approach but also underscore the urgent need for robust defense mechanisms in mmWave imaging systems.
Ruixu Geng, Dongheng Zhang, Jiamu Li, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Inf. Forensics Secur.7
2025 Passive Non-Line-of-Sight Imaging With Light Transport Modulation
abstract
Passive non-line-of-sight (NLOS) imaging has witnessed rapid development in recent years, due to its ability to image objects that are out of sight. The light transport condition plays an important role in this task since changing the conditions will lead to different imaging models. Existing learning-based NLOS methods usually train independent models for different light transport conditions, which is computationally inefficient and impairs the practicality of the models. In this work, we propose NLOS-LTM, a novel passive NLOS imaging method that effectively handles multiple light transport conditions with a single network. We achieve this by inferring a latent light transport representation from the projection image and using this representation to modulate the network that reconstructs the hidden image from the projection image. We train a light transport encoder together with a vector quantizer to obtain the light transport representation. To further regulate this representation, we jointly learn both the reconstruction network and the reprojection network during training. A set of light transport modulation blocks is used to modulate the two jointly trained networks in a multi-scale way. Extensive experiments on a large-scale passive NLOS dataset demonstrate the superiority of the proposed method. The code is available at https://github.com/JerryOctopus/NLOS-LTM.
Ruixu Geng, Xiaolong Du, Yan Chen 0007, Houqiang Li, Yang Hu 0006
IEEE Trans. Image Process.6
2025 IFNet: Deep Imaging and Focusing for Handheld SAR With Millimeter-Wave Signals
abstract
Recent advancements have showcased the potential of handheld millimeter-wave (mmWave) imaging, which applies synthetic aperture radar (SAR) principles in portable settings. However, existing studies addressing handheld motion errors either rely on costly tracking devices or employ simplified imaging models, leading to impractical deployment or limited performance. In this paper, we present IFNet, a novel deep unfolding network that combines the strengths of signal processing models and deep neural networks to achieve robust imaging and focusing for handheld mmWave systems. We first formulate the handheld imaging model by integrating multiple priors about mmWave images and handheld phase errors. Furthermore, we transform the optimization processes into an iterative network structure for improved and efficient imaging performance. Extensive experiments demonstrate that IFNet effectively compensates for handheld phase errors and recovers high-fidelity images from severely distorted signals. In comparison with existing methods, IFNet can achieve at least 11.89 dB improvement in average peak signal-to-noise ratio (PSNR) and 64.91% improvement in average structural similarity index measure (SSIM) on a real-world dataset.
Dongheng Zhang, Ruixu Geng, Jincheng Wu, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.5
2025 Unleashing the Potential of Self-Supervised RF Learning With Group Shuffle
abstract
Self-supervised learning (SSL) is a powerful approach that learns general semantic representations from large-scale unlabeled data to make downstream tasks solve easier, offering significant potential in enhancing downstream performance and alleviating the appetite for large-scale annotated data. However, existing SSL techniques, predominantly designed for natural images, may be prone to shortcuts when applied to RF signals. This study presents surprising empirical findings showing that SSL can indeed learn meaningful RF representations by employing simple group shuffle (GS) and asymmetry augmentation techniques. The GS augmentation is inspired by blind calibration tasks in Time-Interleaved Analog-to-Digital Converters (TIADC). By treating the original RF signal as a composite output from sub-ADCs, GS augmentation enriches RF signals while preserving their global semantics. We also provide a theoretical validation of the GS augmentation’s singular value consistency. Notably, we observe that the shortcut is essentially a domain gap between the pre-trained and the downstream task models. This issue can be mitigated by an asymmetry augmentation technique, which maximizes the similarity between an original RF signal and its augmented version, rather than between two augmentations of the same RF signal. By integratinggroupshuffle andasymmetryaugmentation (GSAA) into an existing contrastive learning framework, we develop an effective contrastive learning approach for RF signals. Our evaluations, spanning seven downstream RF sensing tasks across two general RF devices (WiFi and radar), strongly demonstrate that GSAA plays a significant role in advancing SSL-based solutions in RF sensing.
Ruiyuan Song, Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.6
2025 Learning-Based Tracking-Before-Detect for Unconstrained Indoor Human Tracking Using RF Signal
abstract
Human tracking plays a crucial role in various wireless sensing applications. However, recent advancements have primarily focused on constrained experimental scenarios with less interference, often involving a few individuals performing actions in an empty space without obstacles. In empirical unconstrained scenarios, such as daily office scenes, severe interference and attenuation caused by chaotic environments is inevitable which results in dramatic performance degradation. In this paper, we introduce TBDNet, which incorporates tracking-before-detect (TBD) from conventional signal processing into learning-based models, achieving impressive tracking performance in unconstrained scenarios. TBDNet follows first-track-then-detect pipeline. It maps input heatmap sequence into high-level frame-wise features to adapt the time-varying intensity distribution and motion pattern of targets. After that, the temporal information is accumulated in feature space to obtain trace proposals. We then predict the accurate positions and probability of traces at each timestamp. To assess the efficiency of TBDNet, we collect and release the first RF-UNIT (RF-based Unconstrained Indoor Tracking) dataset, which comprises 4,030,880 radar heatmaps and the corresponding tracking annotations under 6 different scenarios. To our knowledge, RF-UNIT is the first dataset for RF-based human tracking in unconstrained scenes. We anticipate that TBDNet and the RF-UNIT dataset will significantly contribute to the advancement of RF-based sensing technologies.
Dongheng Zhang, Zixin Shang, Yuqin Yuan, Hanqin Gong, Binquan Wang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.9
2025 UMIMO: Universal Unsupervised Learning for Mmwave Radar Sensing With MIMO Array Synthesis
abstract
Millimeter-wave (mmWave) radar sensing powered by deep learning is now emerging in numerous applications, which are predominantly trained in a supervised manner. However, due to the non-interpretable nature of mmWave signals, labeling the radar data has always been a difficult task. While there have been investigations on unsupervised pre-training for mmWave radar sensing, these methods are tailored to specific signal representations. In this paper, we propose UMIMO, an unsupervised learning framework combining the hardware nature of MIMO radar and deep learning techniques to resolve the challenge raised by the insufficient labeled data. UMIMO leverages the antenna arrays synthesized from multiple transmitting and receiving antennas in mmWave radar to construct positive samples for contrastive learning. To achieve this, we propose the constraints on angular resolution and grating lobes to generate effective signal representations with different synthetic arrays. We conduct experiments using UMIMO on three tasks: contactless ECG monitoring, 3D human pose estimation, and human silhouette generation. All experimental results demonstrate that UMIMO can effectively improve the performance of learning-based mmWave radar sensing in an unsupervised manner.
Dongheng Zhang, Ruiyuan Song, Jinbo Chen 0001, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Mob. Comput.8
2024 Continual Learning for Remote Physiological Measurement: Minimize Forgetting and Simplify Inference
Qian Liang 0001, Yan Chen 0007, Yang Hu 0006
ECCV (36)3
2024 Enabling Orientation-Free Mmwave-Based Vital Sign Sensing with Multi-Domain Signal Analysis
abstract
Contactless vital signs estimation using mmWave radar has gained significant attention. However, existing studies are built upon the radar being directed facing the thorax to capture fine-grained vital signs, ignoring the angle variation between the radar and thorax in practical deployment. In this paper, we propose a spatial-temporal optimization model to estimate the human body orientations between the radar and thorax through extracting the multi-domain features of reflected signal. By aligning the signal variation captured from different angles, we can realize orientation-free vital sign sensing. The system achieves an average angle estimation error of 13.1°, and a 14.8% discrepancy reduction in terms of the mean absolute error of the signal captured at different angles.
Hanqin Gong, Dongheng Zhang, Jinbo Chen 0001, Guixin Xu, Yuqin Yuan, Yang Hu 0006, Yan Chen 0007
ICASSP7
2024 SIMFALL: A Data Generator for RF-Based Fall Detection
abstract
Fall detection using Radio Frequency (RF) signals with deep learning has exhibited significant promise in recent years. However, the costly collection of RF data with falls has hampered the performance of existing methods. While there has been approaches which can generate RF signals using various simulation methods, they rely on human-body modeling based on other modalities. Moreover, the realism of the generated signals is insufficient because these approaches cannot accurately capture the human radar cross section (RCS). In this paper, we propose SimFall, which generates simulated data for RF-based fall detection without overhead for data collection. SimFall first simulates the fall process by manipulating the human body mesh based on practical fall model. Then a grid shooting and bouncing ray (SBR) method is utilized to calculate the accurate RCS. Finally, SimFall computes the original signal and transforms it into different forms that reveal the features of falls. The experimental results demonstrate that the data produced by SimFall effectively enhances the accuracy of the RF-based fall detection network.
Jiamu Li, Dongheng Zhang, Jianyang Wang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICASSP7
2024 IFNet: Imaging and Focusing Network for handheld mmWave Devices
abstract
Recent advancements have showcased the potential of hand-held millimeter-wave (mmWave) imaging, which applies synthetic aperture radar (SAR) principles in portable settings. However, existing studies addressing handheld motion errors either rely on costly tracking devices or employ simplified imaging models, leading to impractical deployment or limited performance. In this paper, we present IFNet, a novel deep unfolding network that combines the strengths of signal processing models and deep neural networks to achieve imaging and focusing for handheld mmWave systems. By integrating multiple priors and mapping the optimization processes into an iterative network structure, IFNet effectively compensates for phase errors and recovers high-fidelity images from severely distorted signals. Extensive experiments demonstrate that IFNet outperforms state-of-the-art methods, both qualitatively and quantitatively.
Dongheng Zhang, Ruixu Geng, Jincheng Wu, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICASSP5
2024 Contactless Radar Heart Rate Variability Monitoring Via Deep Spatio-Temporal Modeling
abstract
Radar sensing has been a promising solution for contactless monitoring of Heart Rate Variability (HRV), an essential indicator of the cardiovascular and autonomic nervous systems. However, existing works neglect heartbeat-driven body surface motions spreading across the entire body with spatial variations, which limits their accuracy in identifying fine-grid consecutive heartbeat timings and overall HRV performance. In this paper, we propose to exploit the entire body reflections and model the inherent spatial-temporal relationship between these reflections and heartbeats by deep neural network for contactless HRV monitoring. Specifically, a hybrid convolution-transformer-based network is designed to convert the complex multi-dimensional spatial-temporal modeling problem into an efficient sequence modeling process. Experimental results demonstrate its superiority over the baseline method, achieving the median IBI estimation error of 12ms (w.r.t. 98.47% accuracy), RMSDD error of 7.3ms, SDRR error of 2.9ms, pNN50 error of 5.5%.
Jinbo Chen 0001, Dongheng Zhang, Changwei Wu, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICASSP6
2024 RoFi: Robust WiFi Intrusion Detection via Distribution Matching
abstract
Intrusion detection acts as a key to in-home security, where WiFi-based systems have gained wide attention due to the ubiquitous nature of WiFi signals. While existing methods achieve impressive performance in specific environments, they are susceptible to environmental changes, especially for complex scenarios where outdoor human activities can be mistaken as intrusions. In this paper, we propose RoFi, a robust WiFi intrusion detection system which can handle more complex scenarios. It achieves this by exploring the distribution of autocorrelation function (ACF) of Channel State Information (CSI) when intrusion occurs, where likelihood ratio testing is employed to discriminate intrusion and non-intrusion scenarios, eliminating the variance of different environments. Without complex calibration, RoFi achieves an accuracy of over 97.5% in practical deployment, outperforming existing methods.
Dongheng Zhang, Fengquan Zhan, Xuecheng Xie, Yang Hu 0006, Yan Chen 0007
ICASSP6
2024 Diffradar: High-Quality Mmwave Radar Perception With Diffusion Probabilistic Model
abstract
Millimeter-wave (mmWave) radar has gained increasing attention in environmental perception due to its robustness under low-light conditions. However, existing methods fail to address the challenges of multipath interference and low angle resolution. In this paper, we introduce DiffRadar which leverages the diffusion probabilistic model (DPM) for high-quality mmWave environmental sensing. To adapt DPM for radar signals that lack pix-level structural information, we design a contour encoder to capture intrinsic scene features that enable the DPM to learn a robust representation from radar data. Then the DPM decoder utilizes this high-level semantic information to effectively reconstruct real-world scene distribution. Extensive experiments have demonstrated that our approach surpasses state-of-the-art methods in various complex scenarios.
Jincheng Wu, Ruixu Geng, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
ICASSP6
2024 AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase Calibration
abstract
Recent advancements in WiFi indoor localization have demonstrated the potential for achieving decimeter-level accuracy based on Angle of Arrival (AoA). However, existing commercial WiFi Access Points (APs) suffer from phase offset across different antennas, which significantly degrade the performance of AoA-based methods in practical deployment. Previous work either relied on labor-intensive manual calibration or involved inaccurate and non-robust automatic calibration. In this paper, we propose AutoCali, an accurate and robust automatic phase offset calibration system. The key insight is to utilize the binary nature of phase offsets and the property that triangulation exhibits higher convergence when the correct combination of phase offsets is employed. Extensive experiments demonstrate that AutoCali outperforms state-of-the-art methods by 22.1% in median localization error for simple scenarios and by 37.1% for complex multipath scenarios.
Pengfei Yin, Dongheng Zhang, Guanzhong Wang, Yang Hu 0006, Yan Chen 0007
ICASSP6
2024 Learning-Based Tracking-before-Detect for RF-Based Unconstrained Indoor Human Tracking
Dongheng Zhang, Zixin Shang, Yuqin Yuan, Hanqin Gong, Binquan Wang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IJCAI9
2024 PRISM: Pre-training RF Signals in Sparsity-aware Masked Autoencoders
abstract
This paper introduces a novel paradigm for learning-based RF sensing, termed Pre-training RF signals In Sparsity-aware Masked autoencoders (PRISM), which shifts the RF sensing paradigm from supervised training on limited annotated datasets to unsupervised pre-training on large-scale unannotated datasets, followed by fine-tuning with a small annotated dataset. PRISM leverages a carefully designed sparsity-aware masking strategy to predict missing contents by masking a portion of RF signals, resulting in an efficient pre-training framework that significantly reduces computation and memory resources. This addresses the major challenges posed by large-scale and high-dimensional RF datasets, where memory consumption and computation speed are critical factors. We demonstrate PRISM’s excellent generalization performance across diverse RF sensing tasks by evaluating it on three typical scenarios: human silhouette segmentation, 3D pose estimation, and gesture recognition, involving two general RF devices, radar and WiFi. The experimental results provide strong evidence for the effectiveness of PRISM as a robust learning-based solution for large-scale RF sensing applications.
Ruiyuan Song, Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
INFOCOM5
2024 RPM 2.0: RF-Based Pose Machines for Multi-Person 3D Pose Estimation
abstract
Advanced human sensing technologies based on radio frequency (RF) signals have gained widespread attention in recent years. However, due to the sparsity and incompleteness of RF signals, fine-grained RF-based multi-person 3D pose estimation has progressed more slowly. In this paper, we present RF-based Pose Machine (RPM 2.0) for multi-person 3D pose estimation using RF signals. Specifically, we first develop a lightweight anchor-free detector module to locate and crop regions of interest from horizontal and vertical RF signals. Afterward, we treat the horizontal and vertical millimeter-wave radars as “RF cameras” with different viewing angles and propose a Multi-view Fusion Network to unproject the RF signals into a unified latent feature space, and then calculate the correlation for weighted fusion. Finally, a Spatio-Temporal Attention Network is designed to reconstruct the multi-person 3D skeleton sequences, in which the spatial attention module is proposed to recover invisible body parts using non-local correlations among joints and the temporal attention module refines the 3D pose sequences using temporal coherency learned from frame queries. We evaluate the performance of the proposed RPM 2.0 and state-of-the-art methods on a large-scale dataset with multi-person 3D pose labels and corresponding radar signals. The experimental results show that RPM 2.0 outperforms all of the baseline methods, which locates multi-person 3D key points with an average error of$73 mm$and generalizes well in new data such as occlusion, low illumination.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Circuits Syst. Video Technol.5
2024 DREAM-PCD: Deep Reconstruction and Enhancement of mmWave Radar Pointcloud
abstract
Millimeter-wave (mmWave) radar pointcloud offers attractive potential for 3D sensing, thanks to its robustness in challenging conditions such as smoke and low illumination. However, existing methods failed to simultaneously address the three main challenges in mmWave radar pointcloud reconstruction: specular information lost, low angular resolution, and severe interference. In this paper, we propose DREAM-PCD, a novel framework specifically designed for real-time 3D environment sensing that combines signal processing and deep learning methods into three well-designed components to tackle all three challenges: Non-Coherent Accumulation for dense points, Synthetic Aperture Accumulation for improved angular resolution, and Real-Denoise Multiframe network for interference removal. By leveraging causal multiple viewpoints accumulation and the "real-denoise" mechanism, DREAM-PCD significantly enhances the generalization performance and real-time capability. We also introduce RadarEyes, the largest mmWave indoor dataset with over 1,000,000 frames, featuring a unique design incorporating two orthogonal single-chip radars, Lidar, and camera, enriching dataset diversity and applications. Experimental results demonstrate that DREAM-PCD surpasses existing methods in reconstruction quality, and exhibits superior generalization and real-time capabilities, enabling high-quality real-time reconstruction of radar pointcloud under various parameters and scenarios. We believe that DREAM-PCD, along with the RadarEyes dataset, will significantly advance mmWave radar perception in future real-world applications.
Ruixu Geng, Dongheng Zhang, Jincheng Wu, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Image Process.6
2024 SBRF: A Fine-Grained Radar Signal Generator for Human Sensing
abstract
While deep learning-based RF perception has received significant attention in recent years, the requirement for massive labeled RF data has hindered its further advancement. Despite existing efforts in synthesizing signals, they fail to accurately calculate the Radar Cross Section (RCS) of the target, leading to less practicality of the synthesized signals. In this paper, we introduce Simulated Body Radio Frequency (SBRF), a novel signal synthesis framework for calculating more realistic RCS by combining ray tracing with electromagnetic computation. SBRF involves three key components: a grid-based Shooting and Bouncing Ray (SBR) algorithm to calculate fine-grained human body RCS, a novel ray partitioning algorithm to improve the efficiency of ray tracing, and a coordinate transformation method to sense moving targets. Furthermore, we also design unique data augmentation techniques to improve the efficiency and generalizability of signal synthesis. Extensive experimental evaluations conducted on two publicly available datasets, involving wide-scale activity recognition and fine-grained gesture recognition, demonstrate the effectiveness of SBRF-generated signals in improving RF perception performance and alleviating the challenge of RF data collection.
Jiamu Li, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.7
2024 Robust WiFi Respiration Sensing in the Presence of Interfering Individual
abstract
WiFi-based respiration sensing technology has gained increasing attention due to its contactless sensing capabilities and utilization of existing WiFi devices. However, existing studies are limited to certain scenarios without addressing the motion interference from other individuals. In this paper, we tackle the challenge of robust respiration sensing in the presence of other individuals. Specifically, through an in-depth examination of the correlation between respiratory signals and spatial beam patterns, we develop a respiratory-energy based approach to evaluate the diverse impact of dynamic interference on respiratory signals. When significant interference is detected, we employ a convex-optimization-based beam control strategy, which exploits the inherent characteristics of human respiration, to adaptively adjust the spatial beam pattern. This approach enables a robust and precise gain adjustment between the target and interfering individual, effectively mitigating the impact of interference. Experimental results demonstrate that our approach can reduce the mean absolute error (MAE) of respiration detection by up to 32% compared to state-of-the-art methods, significantly enhancing the accuracy and robustness of WiFi-based respiration sensing.
Xuecheng Xie, Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.4
2024 iSense: Enabling Radar Sensing Under Mutual Device Interference
abstract
Millimeter-wave (mmWave) radar has been widely used in wireless sensing due to its non-contact nature, privacy preservation, and immunity to adverse lighting conditions. However, with more and more mmWave radars working in the same frequency band, mutual device interference among them is inevitable and has become a serious problem. The device interference reduces the signal-to-interference-plus-noise ratio (SINR) and significantly degrades the detection performance. Existing works mainly focus on the vital sign monitoring in different practical scenarios (e.g., device movement, human movement, multi-person interference, and in-car scenario), and the vital sign monitoring in the presence of mutual device interference is still not well resolved. In this paper, we propose a novel interference mitigation framework, iSense, to enable radar vital sign sensing under device interference. By exploiting one-way propagation characteristic of device interference, iSense can effectively detect and suppress the interference. We evaluate iSense under a variety of complex device interference scenarios, including different distances, angles, and numbers of aggressor radars, as well as the impact of different environments. Experimental results show that the accuracy of respiration and heartbeat estimation of iSense can reach over 99.2% and 98.6%, indicating that iSense takes an important step towards the practical development of radar sensing.
Dongheng Zhang, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.4
2024 RPM: RF-Based Pose Machines
abstract
Radio-frequency (RF) based human sensing technologies, due to their great practical value in various applications and privacy-preserving nature, have gained tremendous attention in recent years. However, without fully exploiting the characteristics of radio signals, the performance of existing methods are still limited. First, RF features of the moving human body have different representations in dimensions such as channel and scale, which is challenging when performing feature fusion. Besides, the human body is specularly reflective with respect to the radar, which means the human body cannot be fully captured by a single RF snapshot. Therefore, the radar signal reflected by the human body is sparse and incomplete, which is difficult to extract high-quality features for 3D human pose estimation. In this paper, we present the RF-based Pose Machines (RPM), a novel framework which can generate 3D skeletons from RF signals. Considering the characteristics of RF signals, RPM includes several modules to overcome the challenges. Firstly, a Feature Fusion Network (FFN) is designed to effectively fuse radio signals from horizontal and vertical planes based on the channels' correlation and maintain high-quality feature via a multi-scale fusion block. A Spatio-Temporal Attention network is then designed to reconstruct 3D skeletons from the sparse and incomplete RF signals. Specifically, a spatial attention module is designed to model non-local relationships among joints and reconstruct body parts that a single RF snapshot cannot capture. Afterwards, a temporal attention module is proposed to refine 3D pose based on temporal coherency learned from frame queries. To evaluate the performance of our RPM framework, we construct a large-scale dataset of synchronized 3d skeletons and RF signals, RFSkeleton3D. Our experimental results show that RPM locates 3D key points of the human body with an average error of$5.71 cm$and maintains its performance in new environments with occlusion or bad illumination. The dataset and codes will be made in public.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.5
2024 MobiRFPose: Portable RF-Based 3D Human Pose Camera
abstract
Existing RF-based human pose estimation methods usually require intensive computations and cannot meet the real-time processing and portability requirements for mobile devices. To tackle the limitation, in this article, we introduce a lightweight RF-based pose estimation model, i.e., MobiRFPose, to construct the portable RF-based pose camera. Different from traditional optical-based cameras, the RF-based camera does not capture visual information, which means the privacy-preserving characteristic. Specifically, we only utilize a horizontal antenna array to transceive RF signals, then estimate the human locations on the RF signal heatmap and crop the human location regions, and finally estimate the fine-grained human poses based on the cropped small RF signal heatmaps. To evaluate the performance, we compare MobiRFPose with state-of-the-art methods. Experimental results demonstrate that MobiRFPose can achieve accurate 3D human pose estimation with fewer parameters and computations. We also test the trained MobiRFPose model using mobile computing devices, where the model structures and parameters only take up 268 KB and 3226 KB of disk space, and MobiRFPose can achieve 66 FPS processing speed. The pose estimation error is 11.05 cm in the case of a single person and 11.29 cm in the case of multiple people. All experimental results indicate that our proposed method can construct a portable RF camera to estimate human poses accurately.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.6
2023 Fast 3D Human Pose Estimation Using RF Signals
abstract
Existing deep learning-based wireless sensing models usually require intensive computation. In this paper, we introduce a lightweight RF-based 3D human pose estimation model, i.e., Fast RFPose, to enable real-time human pose estimation. Specifically, Fast RFPose first estimates the human locations in the RF heatmap and crops the human location regions, then estimates the fine-grained human poses based on the cropped small RF heatmaps. In the experiments, we build a radio system and a multi-view camera system to acquire the RF signals and the ground-truth human poses, and compare Fast RFPose with state-of-the-art methods. Experimental results demonstrate that Fast RFPose outperforms the alternative methods. Besides, we further deploy the trained Fast RFPose model on a laptop with a CPU and Fast RFPose can achieve 66 FPS processing speed, which means it can meet the real-time running requirements in mobile devices.
Cong Yu 0011, Yudong Zhang 0001, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
ICASSP6
2023 RF-based Multi-view Pose Machine for Multi-Person 3D Pose Estimation
abstract
In this paper, we present RF-based Multi-view Pose machine (RF-MvP) for multi-person 3D pose estimation using RF signals. Specifically, we first develop a lightweight anchor-free detector module to locate and crop regions of interest from horizontal and vertical RF signals. Afterward, we propose a Multi-view Fusion Network to unproject the RF signals from the horizontal and vertical millimeter-wave radars into a unified latent space, and then calculate the correlation for weighted fusion. Finally, a Spatio-Temporal Attention Network is designed to reconstruct the multi-person 3D skeleton sequences, in which the spatial attention module is proposed to recover invisible body parts using non-local correlations among joints and the temporal attention module refines the 3D pose sequences using temporal coherency learned from frame queries. We evaluate the performance of the proposed RF-MvP and state-of-the-art methods on a large-scale dataset with multi-person 3D pose labels and corresponding radar signals. The experimental results show that RF-MvP outperforms all of the baseline methods, which locates multi-person 3D key points with an average error of 73mm and generalizes well in new data such as occlusion, low illumination.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICME5
2023 RF-Search: Searching Unconscious Victim in Smoke Scenes with RF-enabled Drone
abstract
Toxic gases inhalation is the most common cause of death in fire scenes, which can make people unconscious and unable to save themselves. Hence, discovering the unconscious victims is crucial to improve their survival rate. In this paper, we propose RF-Search, a victim searching system with RF device mounted on the drone. The challenge mainly comes from the fact that drone motion would overwhelm the subtle vital signs utilized for victim identification. To resolve this problem, we have noted that the physical signature of drone motion has been encoded in stationary object reflections. Leveraging this unique physical signature, we propose to identify the unconscious victim through the spatio-temporal correlation between signals reflected from the victim and the surrounding stationary objects. To extract respiration information of the victim, we propose a motion segmentation module and a motion compensation module to suppress the signal variation caused by drone movement. Extensive experiments have demonstrated that our system could achieve an accuracy of 92.5% for victim identification.
Dongheng Zhang, Ruiyuan Song, Binquan Wang, Yang Hu 0006, Yan Chen 0007
MobiCom5
2023 Robust Respiration Sensing with WiFi
abstract
The past decade has witnessed emerging applications of breath monitoring using off-the-shelf WiFi devices owing to their low-cost, non-intrusive, and privacy-friendly characteristics. While existing works have achieved promising results in certain scenarios, the performance degradation introduced by the interfering person who moves around the target user has not been fully investigated, which hinders practical applications of WiFi-based breath sensing. In this paper, we propose a robust respiration sensing system with WiFi which could achieve accurate respiration sensing under strong interference. To achieve this, we first design a 2-D Capon beamformer to maximize the signal-to-interference-plus-noise ratio (SINR). Then, the interfering user’s trajectory is estimated through spatial-temporal processing. Finally, we design a respiration extracting algorithm based on the constraint of the interferer’s trajectory and breath energy to find the optimal position to extract breath signals. Extensive experimental results show that the proposed framework can reduce the Mean Absolute Error (MAE) of breath rate estimation by up to 48% compared with the existing state-of-the-art methods, which demonstrates the superior robustness and effectiveness of our system.
Xuecheng Xie, Dongheng Zhang, Jinbo Chen 0001, Yang Hu 0006, Qibin Sun, Yan Chen 0007
WCNC5
2023 Unsupervised Domain Adaptation for WiFi Gesture Recognition
abstract
Human gesture recognition with WiFi signals has attained acclaim due to the omnipresence, privacy protection, and broad coverage nature of WiFi signals. These gesture recognition systems rely on neural networks trained with a large number of labeled data. However, the recognition model trained with data under certain conditions would suffer from significant performance degradation when applied in practical deployment, which limits the application of gesture recognition systems. In this paper, we propose UDAWiGR, an unsupervised domain adaptation framework for WiFi-based gesture recognition aiming to enhance the performance of the recognition model in new conditions by making effective use of the unlabeled data from new conditions. We first propose a pseudo-labeling method with confidence control constraint to utilize unlabeled data for model training. We then utilize consistency regularization to align the output distribution for enhancing the robustness of neural network under signal perturbations. Furthermore, we propose a cross-match loss to combine the pseudo-labeling and consistency regularization, which makes the whole framework simple yet effective. Extensive experiments demonstrate that the proposed framework could achieve 4.35% accuracy improvement comparing with the state-of-the-art methods on public dataset.
Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
WCNC3
2023 Unsupervised Domain Adaptation for RF-Based Gesture Recognition
abstract
Human gesture recognition with radio frequency (RF) signals has attained acclaim due to the omnipresence, privacy protection, and broad coverage nature of RF signals. These gesture recognition systems rely on neural networks trained with a large number of labeled data. However, the recognition model trained with data under certain conditions would suffer from significant performance degradation when applied in practical deployment, which limits the application of gesture recognition systems. In this article, we propose an unsupervised domain adaptation framework for RF-based gesture recognition aiming to enhance the performance of the recognition model in new conditions by making effective use of the unlabeled data from new conditions. We first propose pseudo labeling and consistency regularization to utilize unlabeled data for model training and eliminate the feature discrepancies in different domains. Then we propose a confidence constraint loss to enhance the effectiveness of pseudo labeling, and design two corresponding data augmentation methods based on the characteristic of the RF signals to strengthen the performance of the consistency regularization, which can make the framework more effective and robust. Furthermore, we propose a cross-match loss to integrate the pseudo labeling and consistency regularization, which makes the whole framework simple yet effective. Extensive experiments demonstrate that the proposed framework could achieve 4.35% and 2.25% accuracy improvement comparing with the state-of-the-art methods on public WiFi data set and millimeter wave (mmWave) radar data set, respectively.
Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Internet Things J.4
2023 RFPose-OT: RF-based 3D human pose estimation via optimal transport theory
abstract
This paper introduces a novel framework, i.e., RFPose-OT, to enable three-dimensional (3D) human pose estimation from radio frequency (RF) signals. Different from existing methods that predict human poses from RF signals at the signal level directly, we consider the structure difference between the RF signals and the human poses, propose a transformation of the RF signals to the pose domain at the feature level based on the optimal transport (OT) theory, and generate human poses from the transformed features. To evaluate RFPose-OT, we build a radio system and a multi-view camera system to acquire the RF signal data and the ground-truth human poses. The experimental results in a basic indoor environment, an occlusion indoor environment, and an outdoor environment demonstrate that RFPose-OT can predict 3D human poses with higher precision than state-of-the-art methods.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
Frontiers Inf. Technol. Electron. Eng.6
2023 Towards Domain-Independent and Real-Time Gesture Recognition Using mmWave Signal
abstract
Human gesture recognition using millimeter-wave (mmWave) signals provides attractive applications including smart home and in-car interfaces. While existing works achieve promising performance under controlled settings, practical applications are still limited due to the need of intensive data collection, extra training efforts when adapting to new domains, and poor performance for real-time recognition. In this paper, we propose DI-Gesture, a domain-independent and real-time mmWave gesture recognition system. Specifically, we first derive signal variations corresponding to human gestures with spatial-temporal processing. To enhance the robustness of the system and reduce data collecting efforts, we design a data augmentation framework for mmWave signals based on correlations between signal patterns and gesture variations. Furthermore, a spatial-temporal gesture segmentation algorithm is employed for real-time recognition. Extensive experimental results show DI-Gesture achieves an average accuracy of 97.92%, 99.18%, and 98.76% for new users, environments, and locations, respectively. We also evaluate DI-Gesture in challenging scenarios like real-time recogntion and sensing at extreme angles, all of which demonstrates the superior robustness and effectiveness of our system.
Dongheng Zhang, Jinbo Chen 0001, Jinwei Wan, Dong Zhang 0015, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.6
2023 Learning Fashion Compatibility With Context Conditioning Embedding
abstract
Fashion compatibility predictions have obtained a lot of attention recently. Mining the compatibility between fashion items in an outfit is different from learning the visual similarity, since this relationship is more delicate. Decomposing the outfit compatibility into pairwise item matching is a popular way to treat the problem. However, in most existing methods, the items are matched without considering the context, i.e, the remaining items in the outfit. Recent efforts have been made to learn the underlying high order relationships among items by treating the outfit as a whole. These models could be sensitive to the properties of different datasets, and the item representations in these models are not as compact as those in the pairwise models. In this paper, we propose a context conditioning embedding approach to learn compact representations that preserve the shared information among items under the existence of contextual items. We use two different spaces, the general and the contextual spaces, to embed items, where the representation in the contextual space contains information from the context. We employ mutual information maximization for model learning, which is shown to be more appropriate for the problem. With extensive experiments, we show that our model achieves superior performance than other state-of-the-art methods.
Yang Hu 0006, Cong Yu 0011, Yan Chen 0007, Bing Zeng 0001
IEEE Trans. Multim.2
2023 Personalized Fashion Recommendation With Discrete Content-Based Tensor Factorization
abstract
Fashion outfit recommendation has attracted lots of attention recently. The problem becomes even more interesting and challenging when considering users’ personalized fashion preferences. Although existing works have successfully improved the recommendation accuracy, the efficiency issue of computation and storage is still under-investigated and often ignored. In this paper, we propose a discrete content-based tensor factorization model that maps items and user to binary codes for efficient fashion recommendation. We introduce a probabilistic perspective for learning to hash, where the binary codes are sampled from a set of underlying Bernoulli variables. To demonstrate the effectiveness of our model, we collect a large-scale outfit dataset together with user label information from a fashion-focused social website. Extensive experiments on our dataset show that the proposed model outperforms other state-of-the-art methods.
Yang Hu 0006, Cong Yu 0011, Yunchao Jiang, Yan Chen 0007, Bing Zeng 0001
IEEE Trans. Multim.2
2023 Radio-Assisted Human Detection
abstract
In this paper, we propose a radio-assisted human detection framework by incorporating radio information into the state-of-the-art detection methods, including anchor-based one-stage detectors and two-stage detectors. We extract the radio localization and identifier information from the radio signals to assist the human detection, due to which the problem of false positives and false negatives can be greatly alleviated. For both detectors, we use the confidence score revision based on the radio localization to improve the detection performance. For two-stage detection methods, we propose to utilize the region proposals generated from radio localization rather than relying on region proposal network (RPN). Moreover, with the radio identifier information, a non-max suppression method with the radio localization constraint has also been proposed to further suppress the false detections and reduce miss detections. Experiments on the simulative Microsoft COCO dataset and Caltech pedestrian datasets show that the mean average precision (mAP) and the miss rate of the state-of-the-art detection methods can be improved with the aid of radio information. Finally, we conduct experiments in real-world scenarios to demonstrate the feasibility of our proposed method in practice.
Chengrun Qiu, Dongheng Zhang, Yang Hu 0006, Houqiang Li, Qibin Sun, Yan Chen 0007
IEEE Trans. Multim.3
2023 RFMask: A Simple Baseline for Human Silhouette Segmentation With Radio Signals
abstract
Human silhouette segmentation, which is originally defined in computer vision, has achieved promising results for understanding human activities. However, the physical limitation makes existing systems based on optical cameras suffer from severe performance degradation under low illumination, smoke, and/or opaque obstruction conditions. To overcome such limitations, in this paper, we propose to utilize the radio signals, which can traverse obstacles and are unaffected by the lighting conditions to achieve silhouette segmentation. The proposed RFMask framework is composed of three modules. It first transforms RF signals captured by millimeter wave radar on two planes into spatial domain and suppress interference with the signal processing module. Then, it locates human reflections on RF frames and extract features from surrounding signals with human detection module. Finally, the extracted features from RF frames are aggregated with an attention based mask generation module. To verify our proposed framework, we collect a dataset containing804,760radio frames and402,380camera frames with human activities under various scenes. Experimental results show that the proposed framework can achieve impressive human silhouette segmentation even under the challenging scenarios (such as low light and occlusion scenarios) where traditional optical-camera-based methods fail. To the best of our knowledge, this is the first investigation towards segmenting human silhouette based on millimeter wave signals. We hope that our work can serve as a baseline and inspire further research that perform vision tasks with radio signals. The dataset and codes will be made in public.
Dongheng Zhang, Chunyang Xie, Cong Yu 0011, Jinbo Chen 0001, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.6
2023 RFGAN: RF-Based Human Synthesis
abstract
This paper demonstrates human synthesis based on the Radio Frequency (RF) signals, which leverages the fact that RF signals can record human movements with the signal reflections off the human body. Different from existing RF sensing works that can only perceive humans roughly, this paper aims to generate fine-grained optical human images by introducing a novel cross-modal RFGAN model. Specifically, we first build a radio system equipped with horizontal and vertical antenna arrays to transceive RF signals. Since the reflected RF signals are processed as obscure signal projection heatmaps on the horizontal and vertical planes, we design a RF-Extractor with RNN in RFGAN for RF heatmap encoding and combining to obtain the human activity information. Then we inject the information extracted by the RF-Extractor and RNN as the condition into GAN using the proposed RF-based adaptive normalizations. Finally, we train the whole model in an end-to-end manner. To evaluate our proposed model, we create two cross-modal datasets (RF-Walk&RF-Activity) that contain thousands of optical human activity frames and corresponding RF signals. Experimental results show that the RFGAN can generate target human activity frames using RF signals. To the best of our knowledge, this is the first work to generate optical images based on RF signals.
Cong Yu 0011, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.5
2022 Learning Token-Based Representation for Image Retrieval
abstract
In image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match kernel. However, the complexity of these approaches is non-trivial with large memory footprint, which limits their capability to jointly perform feature learning and aggregation. To generate compact global representations while maintaining regional matching capability, we propose a unified framework to jointly learn local feature representation and aggregation. In our framework, we first extract local features using CNNs. Then, we design a tokenizer module to aggregate them into a few visual tokens, each corresponding to a specific visual pattern. This helps to remove background noise, and capture more discriminative regions in the image. Next, a refinement block is introduced to enhance the visual tokens with self-attention and cross-attention. Finally, different visual tokens are concatenated to generate a compact global representation. The whole framework is trained end-to-end with image-level labels. Extensive experiments are conducted to evaluate our approach, which outperforms the state-of-the-art methods on the Revisited Oxford and Paris datasets.
Min Wang 0019, Wengang Zhou 0001, Yang Hu 0006, Houqiang Li
AAAI4
2022 DI-Gesture: Domain-Independent and Real-Time Gesture Recognition with Millimeter-Wave Signals
abstract
Human gesture recognition using millimeter wave (mmWave) signals provides attractive applications including smart home and in-car interfaces. While existing works achieve promising performance under controlled settings, practical applications are still limited due to the need for intensive data collection, extra training efforts when adapting to new domains (i.e. environments, persons and locations) and poor performance for real-time recognition. In this paper, we propose DI-Gesture, a domain-independent and real-time mmWave gesture recognition system. Specifically, we first derive the signal variation corresponding to human gestures with spatial-temporal processing. To enhance the robustness of the system and reduce data collecting efforts, we design a data augmentation framework based on the correlation between signal patterns and gesture variations. Furthermore, we propose a dynamic window mechanism to perform gesture segmentation automatically and accurately, thus enabling real-time recognition. Finally, we build a lightweight neural network to extract spatial-temporal information from the data for gesture classification. Extensive experimental results show DI-Gesture achieves an average accuracy of 97.92%, 99.18% and 98.76% for new users, environments and locations, respectively. In real-time scenario, the accuracy of DI-Gesture reaches over 97% with an average inference time of 2.87ms, which demonstrates the superior robustness and effectiveness of our system.
Dongheng Zhang, Jinbo Chen 0001, Jinwei Wan, Dong Zhang 0015, Yang Hu 0006, Qibin Sun, Yan Chen 0007
GLOBECOM6
2022 Contactless Blood Pressure Monitoring with mmWave Radar
abstract
The monitoring of blood pressure is critical for the prevention, diagnosis and treatment of cardiovascular diseases. However, existing methods require physical contact between human body and sensor, which are not suitable for long-term monitoring. In this paper, we propose a contactless blood pressure monitoring system, mmBP, using millimeter wave radar. Specifically, we first separate signals reflected from different spatial locations by coherently combining the signals on different antennas. Then, we locate and extract the arterial pulse using convolutional neural network (CNN) assisted template matching with location tracking. Finally, we design an encoder-decoder neural network to derive the blood pressure information from the extracted signal. Experimental results on 20 subjects show that the measurement deviation rate is 9.00% and 3.69% for systolic and diastolic blood pressure, which demonstrates the feasibility and effectiveness of the proposed system.
You Ran, Dongheng Zhang, Jinbo Chen 0001, Yang Hu 0006, Yan Chen 0007
GLOBECOM4
2022 Real-Time Fall Detection Using Mmwave Radar
abstract
Fall is a severe health threat for elders’ health care. While existing systems could achieve promising performance under specific scenarios, the required computing resources are usually not affordable, which is not applicable for real-time detection. In this paper, we propose mmFall, a real time fall detection system using millimeter wave signal which can achieve impressive accuracy with low computation complexity. Specifically, we first extract the signal variation corresponding to human activity with spatial-temporal processing. To enhance the system performance and robustness, we perform data augmentation by shifting, flipping, extracting and interpolating the signal. Finally, we design a light-weight convolutional neural network to achieve real-time fall detection. Extensive experimental results demonstrate that the pro-posed system could achieve state-of-the-art performance with limited computation complexity.
Dongheng Zhang, Jinbo Chen 0001, Dong Zhang 0015, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICASSP7
2022 Accurate Human Pose Estimation using RF Signals
abstract
Radio-frequency (RF) based human sensing technologies, due to their great practical value in various applications and privacy-preserving nature, have gained tremendous attention in recent years. However, without fully exploiting the characteristics of radio signals, the performance of existing methods are still limited. First, RF features of the moving human body have different representations in dimensions such as channel and scale, which is challenging when performing feature fusion. Besides, the human body is specularly reflective with respect to the radar, which means the human body cannot be fully captured by a single RF snapshot. Therefore, the radar signal reflected by the human body is sparse and incomplete, which is difficult to extract high-quality features for 3D human pose estimation. In this paper, we present the RF-based Pose Machines (RPM), a novel framework which can generate 3D skeletons from RF signals. Considering the characteristics of RF signals, RPM includes several modules to overcome the challenges. Firstly, a Multidimensional Feature Fusion (MFF) backbone is designed to effectively fuse radio signals based on the channels' correlation and maintain high-quality feature via a multi-scale fusion block. A Spatio-Temporal Attention network is then designed to reconstruct 3D skeletons by modeling the non-local spatio-temporal relationships. To evaluate the performance of our RPM framework, we construct a large-scale dataset of synchronized 3D skeletons and RF signals, RFSkeleton3D. Our experimental results show that RPM locates 3D key points of the human body with an average error of 5.71cm and maintains its performance in new environments with occlusion or bad illumination. The dataset and codes will be made in public.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
MMSP5
2022 WiFi-Based Human Pose Image Generation
abstract
This paper tackles a new challenge: how to generate human pose images from wireless signals? Although the optical camera can capture optical images, it is easily restricted by bad lighting. The wireless signals do not rely on visible lights. However, the low-resolution characteristics make previous works can only generate a rough skeleton of human posture, missing a lot of detailed visual information, such as background, appearance, etc. Since the visual information usually maintains unchanged for a period and the wireless signals can capture the movements of the human, in this paper, we propose a framework to generate the target human pose images by combining the wireless signals with an initial optical image. We utilize multiple wireless devices to collect the WiFi signals and a camera to capture the initial optical image. Then a data preprocessing component is designed to preprocess the wireless and vision data. Finally, a deep learning model learns to generate the human pose images from the processed wireless signals and the initial optical image. We conduct experiments to evaluate our proposed framework and results show that it achieves higher accuracy than the state-of-the-art WiFi-based pose estimation method and better visual quality than the state-of-the-art human generation method.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Houqiang Li, Qibin Sun, Yan Chen 0007
MMSP5
2022 RF-URL: unsupervised representation learning for RF sensing
abstract
The major obstacle for learning-based RF sensing is to obtain a high-quality large-scale annotated dataset. However, unlike visual datasets that can be easily annotated by human workers, RF signal is non-intuitive and non-interpretable, which causes the annotation of RF signals time-consuming and laborious. To resolve the rapacious appetite of annotated data, we propose a novel unsupervised representation learning (URL) framework for RF sensing, RF-URL, to learn a pre-training model on large-scale unannotated RF datasets that can be easily collected. RF-URL utilizes a contrastive framework to mind the gap between signal-processing-based RF sensing and learning-based RF sensing. By constructing positive and negative pairs through different signal processing representations, RF-URL seamlessly integrates the existing RF signal processing algorithms into the learning-based networks. Moreover, the RF-URL is carefully designed to take into account the asymmetric characteristics of different RF signal processing representations. We show that RF-URL is universal to a variety of RF sensing tasks by evaluating RF-URL in three typical RF sensing tasks (human gesture recognition, 3D pose estimation and silhouette generation) based on two general RF devices (WiFi and radar). All experimental results strongly demonstrate that RF-URL takes an important step towards learning-based solutions for large-scale RF sensing applications.
Ruiyuan Song, Dongheng Zhang, Cong Yu 0011, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
MobiCom7
2022 Hierarchical Dynamic Programming Module for Human Pose Refinement
abstract
We observed that remarkable and impressive performance on image-based human pose estimation have been achieved by deep Convolutional Neural Networks (CNN). Nevertheless, directly applying these image-based models on videos is not only computionally intensive, but also may cause jitter and loss. The main reason is that the image-based models purely focus on the local features of individual frames and totally ignore the temporal information among adjacent frames. Some existing methods are proposed to address the temporal coherency issue. However, these methods need to be designed carefully and cannot be combined with existing image-based methods. In this paper, we propose a simple yet effective module to refine the estimated pose by exploiting the temporal coherency among the heatmaps of adjacent frames, which can be easily inserted into image-based networks as a plug-in. We show that the temporal coherency issue among the heatmap frames could be re-formulated as a graph path selection optimization problem. Moreover, to speed up the refinement process, we propose a hierarchical graph optimization to achieve the refinement from coarse to fine. Experimental results on two large-scale video pose estimation benchmarks show that our module can improve the performance with little speed loss when combined with image-based methods as an efficient plug-in.
Chunyang Xie, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Circuits Syst. Video Technol.3
2022 Passive Non-Line-of-Sight Imaging Using Optimal Transport
abstract
Passive non-line-of-sight (NLOS) imaging has drawn great attention in recent years. However, all existing methods are in common limited to simple hidden scenes, low-quality reconstruction, and small-scale datasets. In this paper, we propose NLOS-OT, a novel passive NLOS imaging framework based on manifold embedding and optimal transport, to reconstruct high-quality complicated hidden scenes. NLOS-OT converts the high-dimensional reconstruction task to a low-dimensional manifold mapping through optimal transport, alleviating the ill-posedness in passive NLOS imaging. Besides, we create the first large-scale passive NLOS imaging dataset, NLOS-Passive, which includes 50 groups and more than 3,200,000 images. NLOS-Passive collects target images with different distributions and their corresponding observed projections under various conditions, which can be used to evaluate the performance of passive NLOS imaging algorithms. It is shown that the proposed NLOS-OT framework achieves much better performance than the state-of-the-art methods on NLOS-Passive. We believe that the NLOS-OT framework together with the NLOS-Passive dataset is a big step and can inspire many ideas towards the development of learning-based passive NLOS imaging. Codes and dataset are publicly available (https://github.com/ruixv/NLOS-OT).
Ruixu Geng, Yang Hu 0006, Cong Yu 0011, Houqiang Li, Heng-Yu Zhang, Yan Chen 0007
IEEE Trans. Image Process.2
2021 Personalized Outfit Recommendation With Learnable Anchors
abstract
The multimedia community has recently seen a tremendous surge of interest in the fashion recommendation problem. A lot of efforts have been made to model the compatibility between fashion items. Some have also studied users’ personal preferences for the outfits. There is, however, another difficulty in the task that hasn’t been dealt with carefully by previous work. Users that are new to the system usually only have several (less than 5) outfits available for learning. With such a limited number of training examples, it is challenging to model the user’s preferences reliably. In this work, we propose a new solution for personalized outfit recommendation that is capable of handling this case. We use a stacked self-attention mechanism to model the high-order interactions among the items. We then embed the items in an outfit into a single compact representation within the outfit space. To accommodate the variety of users’ preferences, we characterize each user with a set of anchors, i.e. a group of learnable latent vectors in the outfit space that are the representatives of the outfits the user likes. We also learn a set of general anchors to model the general preference shared by all users. Based on this representation of the outfits and the users, we propose a simple but effective strategy for the new user profiling tasks. Extensive experiments on large scale real-world datasets demonstrate the performance of our proposed method.
Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
CVPR2
2021 SpeedNet: Indoor Speed Estimation With Radio Signals
abstract
Indoor human speed estimation is critical to in-home health monitoring of elderly people since it can provide the moving status of the human. Contactless indoor speed estimation with radio signals is challenging due to the complicated relationship between the speed of moving human and radio signals. In this article, we propose an indoor speed estimation framework, SpeedNet, to estimate the speed from the radio signals. Specifically, SpeedNet first extracts the dominant path signal reflected from the human through the beamforming technique. Then, SpeedNet obtains the doppler frequency shift (DFS) corresponding to the moving human by analyzing the short-time Fourier transform (STFT) spectrogram of the dominant path signal. Finally, SpeedNet trains a deep neural network composed of convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) to utilize the spatial and temporal features of DFS to estimate the speed of moving human. The experimental results show that SpeedNet can estimate the human moving speed with an average accuracy of 96.33% in a typical indoor environment, which is better than the state-of-the-art approaches.
Yan Chen 0007, Hongyu Deng, Dongheng Zhang, Yang Hu 0006
IEEE Internet Things J.4
2021 MTrack: Tracking Multiperson Moving Trajectories and Vital Signs With Radio Signals
abstract
In this article, we propose a human sensing system with radio signals, MTrack, for in-home healthcare, which is capable of tracking the trajectories of moving persons and vital signs of static persons under the multiperson scenarios. To achieve this, we implement a multiantenna wideband system that can provide high-resolution Angle of Arrival (AoA) and Time of Flight (ToF). A 2-D beamformer is utilized to transform the raw radio signals into the AoA-ToF domain. To track the trajectories of moving persons, we leverage the movement of persons to cancel static multipaths and propose a path selection algorithm to estimate the locations of human and suppress the interferences from dynamic multipaths. To track the vital signs of static persons, we utilize the breath of static persons to eliminate static multipaths and propose a correlation-based algorithm to eliminate dynamic multipaths. Extensive experiments show that the proposed MTrack system is capable of tracking multiple moving persons with subdecimeter level accuracy, and can estimate the breath and heartbeat rate of static persons with the median accuracy of 99.8% and 98.46%, respectively.
Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Internet Things J.2
2020 Estimating Indoor Human Speed via Radio Signals
abstract
Indoor human speed estimation, which can provide the moving status of human, is attracting considerable critical attention, especially in the field of in-home health monitoring of elderly people. Since the relationship between the moving of human and radio signal is very complicated, indoor speed estimation via radio signals is non-trivial and challenging. To address the challenge, in this paper, we propose a SpeedNet framework to estimate the speed of moving human from the radio signals. Specifically, SpeedNet first utilizes the beamforming technique to extract the dominant path signal reflected from individuals. Then, with short time Fourier transform (STFT), SpeedNet analyzes the spectrogram of the dominant path signal and obtains the doppler frequency shift (DFS) that corresponds to the moving human. Finally, SpeedNet exploits the spatial and temporal features of the DFS through a deep neural network, which consists of convolutional neural networks (CNN) and long short-term memory networks (LSTM), to estimate the speed of moving human. Extensive experiments show that compared with the state-of-the-art approaches, SpeedNet can achieve much better speed estimation performance with a mean absolute percentage error (MAPE) of 3.67% in a typical indoor environment.
Hongyu Deng, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
GLOBECOM3
2020 WiFi Vision: Sensing, Recognition, and Detection With Commodity MIMO-OFDM WiFi
abstract
Indoor human sensing, recognition, and detection, as key enablers of building smart environments, such as smart home, smart retail, and smart museum, have gained tremendous attention in recent years. Compared with traditional vision-based and wearable sensor-based solutions, radio-frequency (RF)-based approaches are more desirable with the contactless and nonline-of-sight nature. Among all RF-based approaches, WiFi-based approaches have been the focus of many researchers because of the ubiquitous availability and cost efficiency. In this article, we present a survey of recent advances in WiFi vision problems, i.e., sensing, recognition, and detection by utilizing the channel state information (CSI) of the commodity WiFi devices. We focus on nine key applications of smart environments, including WiFi imaging, vital sign monitoring, human identification, gesture recognition, gait recognition, daily activity recognition, fall detection, human detection, and indoor positioning. Such a survey can help readers have an overall understanding of sensing, recognition, and detection with commodity WiFi, and thus expedite the development of smart environments.
Ying He 0013, Yan Chen 0007, Yang Hu 0006, Bing Zeng 0001
IEEE Internet Things J.3
2020 MUcast: Linear Uncoded Multiuser Video Streaming With Channel Assignment and Power Allocation Optimization
abstract
Multiuser video transmission, where the server transmits videos to multiple users that require different contents at the same time, becomes more and more popular with the development of wireless communication technology. One key problem in multiuser video transmission is how to optimally allocate system resources such as transmission power and channels to multiple users to achieve the best system performance. To resolve the problem, in this paper, we propose an uncoded multiuser video streaming system, which exploits diversities of video contents and channel conditions of multiple users. We first solve the channel assignment problem with known power allocation by taking into account the intra-block energy diffusion and inter-block energy aliasing. Then, with the obtained channel assignment, we derive a closed-form solution to the multiuser power allocation optimization problem. Finally, we conduct simulations to evaluate the proposed uncoded multiuser video streaming system by comparing with three other approaches, and the simulation results show that the proposed method can achieve the best system performance.
Chaofan He, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Residual Carrier Frequency Offset Estimation and Compensation for Commodity WiFi
abstract
The various offsets existed in the commodity WiFi devices greatly limit the use of ubiquitous WiFi signals for indoor applications. In this paper, we focus on the estimation and compensation of the residual carrier frequency offset (CFO) for the commodity WiFi devices. Specifically, we consider a distorted channel state information (CSI) model by taking into consideration various CSI errors such as packet detection delay (PDD) and CFO. We propose a multiscale sparse recovery algorithm to get rid of the effect of PDD and extract the carrier frequency component out of CSI. Then, we formulate the residual CFO estimation as a spectrum estimation problem and utilize the MUSIC algorithm to estimate the residual CFO. Real experiments and numerical simulations are conducted to evaluate the performance of the proposed method. The experimental results and simulation results show that the residual CFO is time-varying, and compared with existing methods, the proposed method can better estimate and compensate the residual CFO, and thus achieve better results.
Yan Chen 0007, Yang Hu 0006, Bing Zeng 0001
IEEE Trans. Mob. Comput.3
2019 Learning Binary Code for Personalized Fashion Recommendation
abstract
With the rapid growth of fashion-focused social networks and online shopping, intelligent fashion recommendation is now in great needs. Recommending fashion outfits, each of which is composed of multiple interacted clothing and accessories, is relatively new to the field. The problem becomes even more interesting and challenging when considering users' personalized fashion style. Another challenge in a large-scale fashion outfit recommendation system is the efficiency issue of item/outfit search and storage. In this paper, we propose to learn binary code for efficient personalized fashion outfits recommendation. Our system consists of three components, a feature network for content extraction, a set of type-dependent hashing modules to learn binary codes, and a matching block that conducts pairwise matching. The whole framework is trained in an end-to-end manner. We collect outfit data together with user label information from a fashion-focused social website for the personalized recommendation task. Extensive experiments on our datasets show that the proposed framework outperforms the state-of-the-art methods significantly even with a simple backbone.
Yang Hu 0006, Yunchao Jiang, Yan Chen 0007, Bing Zeng 0001
CVPR2
2019 Exploiting Channel Assignment and Power Allocation for Linear Uncoded Multiuser Video Streaming
abstract
Multiuser video transmission, where the server transmits videos to multiple users that request different contents at the same time, becomes more and more popular with the development of wireless communication technology. One key problem in multiuser video transmission is how to optimally allocate the system resources such as transmission power and channels to multiple users to achieve the best system performance. To resolve the problem, in this paper, we propose an uncoded multiuser video streaming system, which exploits diversities of video contents and channel conditions of multiple users. We first solve the channel assignment problem with known power allocation by taking into account the intra-block energy diffusion and inter-block energy aliasing. Then, with the obtained channel assignment, we derive a closed-form solution to the multiuser power allocation optimization problem. Finally, we conduct simulations to evaluate the proposed uncoded multiuser video streaming system by comparing with three other approaches, and simulation results show that the proposed method can achieve the best system performance.
Chaofan He, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001
ICC2
2019 Estimating and Compensating Residual Carrier Frequency Offset for Commodity WiFi
abstract
The various offsets existed on the commodity WiFi devices greatly limit the use of ubiquitous WiFi signals for indoor applications. In this paper, we focus on the estimation and compensation of the residual carrier frequency offset (CFO) for the commodity WiFi devices. Specifically, we introduce a distorted channel state information (CSI) model by taking into consideration various CSI errors such as packet detection delay (PDD) and CFO. We propose a multiscale sparse recovery algorithm to get rid of the effect of PDD and extract the carrier frequency component out of CSI. Then, we formulate the residual CFO estimation as a spectrum estimation problem and propose to utilize the MUSIC algorithm to estimate the residual CFO. Real experiments are conducted to evaluate the performance of the proposed method. The experimental results show that the residual CFO is time-varying, and compared with existing methods, the proposed method can better estimate and compensate the residual CFO, and thus achieve better results.
Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
ICC2
2019 Personalized Fashion Design
abstract
Fashion recommendation is the task of suggesting a fashion item that fits well with a given item. In this work, we propose to automatically synthesis new items for recommendation. We jointly consider the two key issues for the task, i.e., compatibility and personalization. We propose a personalized fashion design framework with the help of generative adversarial training. A convolutional network is first used to map the query image into a latent vector representation. This latent representation, together with another vector which characterizes user's style preference, are taken as the input to the generator network to generate the target item image. Two discriminator networks are built to guide the generation process. One is the classic real/fake discriminator. The other is a matching network which simultaneously models the compatibility between fashion items and learns users' preference representations. The performance of the proposed method is evaluated on thousands of outfits composited by online users. The experiments show that the items generated by our model are quite realistic. They have better visual quality and higher matching degree than those generated by alternative methods.
Cong Yu 0011, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
ICCV2
2019 Deep Deterministic Policy Gradient (DDPG)-Based Energy Harvesting Wireless Communications
abstract
To overcome the difficulties of charging the wireless sensors in the wild with conventional energy supply, more and more researchers have focused on the sensor networks with renewable generations. Considering the uncertainty of the renewable generations, an effective energy management strategy is necessary for the sensors. In this paper, we propose a novel energy management algorithm based on the reinforcement learning. By utilizing deep deterministic policy gradient (DDPG), the proposed algorithm is applicable for the continuous states and realizes the continuous energy management. We also propose a state normalization algorithm to help the neural network initialize and learn. With only one day's real solar data and the simulative channel data for training, the proposed algorithm shows excellent performance in the validation with about 800 days length of real solar data. Compared with the state-of-the-art algorithms, the proposed algorithm achieves better performance in terms of long-term average net bit rate.
Chengrun Qiu, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
IEEE Internet Things J.2
2019 BreathTrack: Tracking Indoor Human Breath Status via Commodity WiFi
abstract
In this paper, we propose a contact-free breath tracking system, BreathTrack, to track the status of breath using the off-the-shelf WiFi devices. BreathTrack exploits the phase variation of the channel state information (CSI) to track human breath. To resolve the phase distortions introduced by the hardware imperfection of the commodity WiFi chips, BreathTrack utilizes both the hardware and software correction methods. The time-invariant PLL phase offset is calibrated by the hardware correction using cables and splitters, while the time-varying carrier frequency offset, sampling frequency offset and packet detection delay are removed by the software corrections using the phase difference between the CSI at the receiver antennas and that at the reference antenna connected from the transmitter. Moreover, BreathTrack utilizes the sparse recovery method to find the dominant path in the multipath indoor environment and derive the corresponding complex attenuation coefficient. Then, the phase variation of the complex attenuation coefficient is utilized to extract the detailed breath status and the breath rate. Extensive experiments are conducted to show that BreathTrack could estimate the breath rate with the median accuracy of over 99% in most scenarios, and could track the detailed status of breath directly using the raw phase variation.
Dongheng Zhang, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
IEEE Internet Things J.2
2019 Joint Power Allocation and Channel Assignment for NOMA With Deep Reinforcement Learning
abstract
Non-orthogonal multiple access (NOMA) has been considered as a significant candidate technique for the next generation wireless communication to support high throughput and massive connectivity. It allows different users to be multiplexed on one channel through applying superposition coding at the transmitter and successive interference cancellation (SIC) at the receiver. To fully utilize the benefit of the NOMA technique, the key problem is how to optimally allocate resources, such as power and channels, to users to maximize the system performance. There have been some existing works on the power allocation for the single-carrier NOMA system. However, how to optimally assign channels in the multi-carrier NOMA system is still unclear. In this paper, we propose a deep reinforcement learning framework to allocate resources to users in a near optimal way. Specifically, we exploit an attention-based neural network (ANN) to perform the channel assignment. Simulation results show that the proposed framework can achieve better system performance, compared with the state-of-the-art approaches.
Chaofan He, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
IEEE J. Sel. Areas Commun.2
2019 Lyapunov Optimized Resource Management for Multiuser Mobile Video Streaming
abstract
Buffering techniques have been commonly used in mobile video streaming systems to handle the bandwidth fluctuation and to mitigate the impact of the stochastic characteristic of wireless channels on a mobile user's quality of experience. However, it has been shown by measurement study that users tend to abort when watching videos with mobile devices, which results in a significant wastage of the video data in the buffer. Therefore, one important problem in mobile video streaming is how to manage the buffer at each mobile user. On the other hand, mobile users generally share the wireless media to download the video data, i.e., mobile users compete with each other for the bandwidth to download the video data. Thus, another important problem in mobile video streaming is how to allocate bandwidth among mobile users. In this paper, we propose to optimize the resource management, i.e., to design buffer management strategy at each mobile user and bandwidth allocation strategy among mobile users, for the multiuser mobile video streaming systems. Specifically, we optimize the long-term average total cost of data wastage and quality of experience of mobile users with certain constraints. By introducing virtual queues and employing the Lyapunov optimization theory, we transform the original optimization problem into the drift-plus-penalty minimization problem. Then, we adopt the primal decomposition to decouple the relationship among different mobile users, which decomposes the problem into a master problem with multiple subproblems. A one-dimension full search algorithm is applied to find the global optimal solution to each subproblem, and the subgradient descent algorithm is utilized to update the solution to the master problem. Finally, simulations are conducted to show that the proposed algorithm is effective for the buffer management and bandwidth allocation in a multiuser mobile video streaming system.
Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Breath Status Tracking Using Commodity WiFi
abstract
In this paper, we propose a contact-free breath tracking system, BreathTrack, to track the status of breath using the off-the-shelf WiFi devices by exploiting the phase variation of the channel state information(CSI). BreathTrack utilizes a reference antenna connected from the transmitter to resolve the phase distortions introduced by the hardware imperfection. Moreover, BreathTrack utilizes the sparse recovery method to find the dominant path in the multipath indoor environment and derive the corresponding complex attenuation coefficient. Then, the phase variation of the complex attenuation coefficient is utilized to extract the detailed breath status and the breath rate. Extensive experiments are conducted to show that BreathTrack could estimate the breath rate with the median accuracy of over 99% in most scenarios, and could track the detailed status of breath directly using the raw phase variation.
Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
GLOBECOM2
2018 Lyapunov Optimized Cooperative Communications With Stochastic Energy Harvesting Relay
abstract
Energy harvesting (EH) wireless communications have become more and more popular due to its capability of effectively reducing the battery replacement time. In this paper, we focus on decode-and-forward-based cooperative wireless communications with an EH relay. The stochastic characteristics of the harvestable energy and the wireless channel make it difficult to optimally manage the energy harvested at the relay for better communication performance. We formulate such a problem as an optimization by minimizing the long-term average symbol error rate (SER) subject to a battery constraint. Based on the Lyapunov optimization theory, we utilize the virtual queue technique and transform the optimization problem into a drift-plus-penalty minimization. Then, we conduct theoretic analysis of the optimal energy strategy derived by the proposed scheme. Specifically, we show the convexity of the driftplus-penalty minimization problem, and prove that the virtual queue is bounded and the battery constraint is always satisfied. We also show that the proposed optimal energy management strategy is limited by an upper bound that is independent of the operation time index, derive the closed-form expression for the asymptotical average SER, and analyze the corresponding diversity order and EH gain. Finally, simulation results using the real solar irradiance data show that the proposed algorithm can achieve much better performance in terms of both average SER and diversity order, compared with the Markov decision process-based method.
Chengrun Qiu, Yang Hu 0006, Yan Chen 0007
IEEE Internet Things J.2
2018 Lyapunov Optimization for Energy Harvesting Wireless Sensor Communications
abstract
With the development and popularity of the renewable energy harvesting devices, the energy harvesting wireless sensor communications that can make use of the energy harvested from the nearby environments have gained more and more attentions. One key problem in the energy harvesting wireless sensor communications is the transmission strategy management, i.e., how to manage the transmission strategy at each time slot to optimize the transmission performance. In this paper, we propose to use Lyapunov optimization theory to maximize the expected good bits per packet transmission for the source node in an energy harvesting wireless communication system. Considering the channel and battery states, we adapt the transmission power and modulation type to achieve such a goal. The problem is formulated as an optimization where the objective function is the long-term average good bits per packet transmission and the constraints are the bounded long-term average battery level and bit error rate. To solve the optimization, we introduce virtual queues and employ the Lyapunov optimization theory to transform the optimization with long-term average format into optimizing the drift-plus-penalty problem. The drift-plus-penalty is further upper bounded with variables only related to current time slot, which greatly simplifies the optimization problem. Theoretic analysis is also conducted to show that the optimal solution is limited by an upper bound that is independent of the operation time index. Finally, simulation results with real solar irradiance data show that the proposed algorithm can achieve much better performance than existing approaches based on Markov decision process and water-filling.
Chengrun Qiu, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
IEEE Internet Things J.2
2018 MCast: High-Quality Linear Video Transmission With Time and Frequency Diversities
abstract
Uncoded linear video transmission has recently attracted people's attention due to its capacity to provide robust and scalable transmission. However, in reality, with the fluctuation of the wireless channels, the received quality may not be good enough. In such a case, the data may need to be transmitted multiple times to exploit both the time and frequency diversities to improve the received quality. Such a problem has never been investigated in the literature of uncoded video transmission. To resolve the problem, in this paper, we propose a framework, named MCast, to utilize the time and frequency diversities to achieve high-quality linear video transmission. We study how to optimally allocate the power and assign the channels at each time slot to the source data such that the overall performance is maximized. Specifically, we first derive a closed-form optimal power allocation solution for any given channel assignment. With the optimal power allocation, we then propose a suboptimal channel assignment scheme, where we sort the channels with their gains and assign the channels one-by-one to the corresponding block that can reduce the most reconstruction error. Finally, we compare MCast system with four other systems that are based on Softcast and Parcast, and simulation results show that MCast system can achieve better performance in terms of both PSNR performance and visual quality.
Chaofan He, Huiying Wang, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001
IEEE Trans. Image Process.3
2018 Lyapunov-Optimized Two-Way Relay Networks With Stochastic Energy Harvesting
abstract
Energy harvesting wireless cooperative communications have become more and more popular in recent years. In this paper, we consider a two-way relay cooperative network, where the relay is an energy harvesting node and uses decode-and-forward (DF) or amplify-and-forward (AF) cooperation protocol. We formulate the energy management problem in such a network as an optimization problem, which minimizes the long-term average outages with a long-term average battery constraint. We then apply Lyapunov optimization to transform the long-term optimization problem into the drift-plus-penalty. We also conduct theoretic analysis of the proposed energy management strategy. Specifically, we prove that the long-term average battery constraint can be guaranteed with the proposed strategy and derive an upper bound of the average outages with the proposed strategy. Furthermore, we analyze theoretically the diversity order and energy harvesting gain of the proposed strategy with the DF and AF protocols. Finally, simulation results using real-solar irradiance data measured by the solar site in Elizabeth City University show that the outage performance and the diversity order of our algorithms are better than the existing MDP-based method.
Yang Hu 0006, Chengrun Qiu, Yan Chen 0007
IEEE Trans. Wirel. Commun.1
2017 Sampling for Approximate Maximum Search in Factorized Tensor
abstract
Factorization models have been extensively used for recovering the missing entries of a matrix or tensor. However, directly computing all of the entries using the learned factorization models is prohibitive when the size of the matrix/tensor is large. On the other hand, in many applications, such as collaborative filtering, we are only interested in a few entries that are the largest among them. In this work, we propose a sampling-based approach for finding the top entries of a tensor which is decomposed by the CANDECOMP/PARAFAC model. We develop an algorithm to sample the entries with probabilities proportional to their values. We further extend it to make the sampling proportional to the $k$-th power of the values, amplifying the focus on the top ones. We provide theoretical analysis of the sampling algorithm and evaluate its performance on several real-world data sets. Experimental results indicate that the proposed approach is orders of magnitude faster than exhaustive computing. When applied to the special case of searching in a matrix, it also requires fewer samples than the other state-of-the-art method.
Yang Hu 0006, Bing Zeng 0001
IJCAI2
2015 Collaborative Fashion Recommendation: A Functional Tensor Factorization Approach
abstract
With the rapid expansion of online shopping for fashion products, effective fashion recommendation has become an increasingly important problem. In this work, we study the problem of personalized outfit recommendation, i.e. automatically suggesting outfits to users that fit their personal fashion preferences. Unlike existing recommendation systems that usually recommend individual items, we suggest sets of items, which interact with each other, to users. We propose a functional tensor factorization method to model the interactions between user and fashion items. To effectively utilize the multi-modal features of the fashion items, we use a gradient boosting based method to learn nonlinear functions to map the feature vectors from the feature space into some low dimensional latent space. The effectiveness of the proposed algorithm is validated through extensive experiments on real world user data from a popular fashion-focused social network.
Yang Hu 0006, Xi Yi, Larry Davis 0001
ACM Multimedia1
2009 Scale-Invariant Visual Language Modeling for Object Categorization
abstract
In recent years, “bag-of-words” models, which treat an image as a collection of unordered visual words, have been widely applied in the multimedia and computer vision fields. However, their ignorance of the spatial structure among visual words makes them indiscriminative for objects with similar word frequencies but different word spatial distributions. In this paper, we propose a visual language modeling method (VLM), which incorporates the spatial context of the local appearance features into the statistical language model. To represent the object categories, models with different orders of statistical dependencies have been exploited. In addition, the multilayer extension to the VLM makes it more resistant to scale variations of objects. The model is effective and applicable to large scale image categorization. We train scale invariant visual language models based on the images which are grouped by Flickr tags, and use these models for object categorization. Experimental results show they achieve better performance than single layer visual language models and “bag-of-words” models. They also achieve comparable performance with 2-D MHMM and SVM-based methods, while costing much less computational time.
Lei Wu 0017, Yang Hu 0006, Mingjing Li, Nenghai Yu, Xian-Sheng Hua 0001
IEEE Trans. Multim.2
2008 Multiple-instance ranking: Learning to rank images for image retrieval
abstract
We study the problem of learning to rank images for image retrieval. For a noisy set of images indexed or tagged by the same keyword, we learn a ranking model from some training examples and then use the learned model to rank new images. Unlike previous work on image retrieval, which usually coarsely divide the images into relevant and irrelevant images and learn a binary classifier, we learn the ranking model from image pairs with preference relations. In addition to the relevance of images, we are further interested in what portion of the image is of interest to the user. Therefore, we consider images represented by sets of regions and propose multiple-instance rank learning based on the max margin framework. Three different schemes are designed to encode the multiple-instance assumption. We evaluate the performance of the multiple-instance ranking algorithms on real-word images collected from Flickr - a popular photo sharing service. The experimental results show that the proposed algorithms are capable of learning effective ranking models for image retrieval.
Yang Hu 0006, Mingjing Li, Nenghai Yu
CVPR1
2008 Maximum Margin Clustering with Pairwise Constraints
abstract
Maximum margin clustering (MMC), which extends the theory of support vector machine to unsupervised learning, has been attracting considerable attention recently. The existing approaches mainly focus on reducing the computational complexity of MMC. The accuracy of these methods, however, has not always been guaranteed. In this paper, we propose to incorporate additional side-information, which is in the form of pairwise constraints, into MMC to further improve its performance. A set of pairwise loss functions are introduced into the clustering objective function which effectively penalize the violation of the given constraints. We show that the resulting optimization problem can be easily solved via constrained concave-convex procedure (CCCP). Moreover, for constrained multi-class MMC, we present an efficient cutting-plane algorithm to solve the sub-problem in each iteration of CCCP. The experiments demonstrate that the pairwise constrained MMC algorithms considerably outperform the unconstrained MMC algorithms and two other clustering algorithms that exploit the same type of side-information.
Yang Hu 0006, Jingdong Wang 0001, Nenghai Yu, Xian-Sheng Hua 0001
ICDM1
2008 Efficient near-duplicate image detection by learning from examples
abstract
In this paper, we propose a novel scheme for near-duplicate image detection, which is an important problem in variety of applications. While in general content based image retrieval, an image could be similar to the query image in infinitely various ways, the ways in which near-duplicate images deviate from the reference image are very limited. Based on this observation, we proposed to use examplar near-duplicate images, which can be obtained automatically, to improve the performance of near-duplicate image retrieval. We first use examplar near-duplicates to learn an effective distance measure and incorporate the learned metric into locality-sensitive hashing to achieve fast retrieval. We then use examplar near-duplicates to automatically expand the query to further improve the retrieval accuracy. The experimental results validate the effectiveness of the proposed algorithms.
Yang Hu 0006, Mingjing Li, Nenghai Yu
ICME1
2008 Video Error Concealment Using Spatio-Temporal Boundary Matching and Partial Differential Equation
abstract
Error concealment techniques are very important for video communication since compressed video sequences may be corrupted or lost when transmitted over error-prone networks. In this paper, we propose a novel two-stage error concealment scheme for erroneously received video sequences. In the first stage, we propose a novel spatio-temporal boundary matching algorithm (STBMA) to reconstruct the lost motion vectors (MV). A well defined cost function is introduced which exploits both spatial and temporal smoothness properties of video signals. By minimizing the cost function, the MV of each lost macroblock (MB) is recovered and the corresponding reference MB in the reference frame is obtained using this MV. In the second stage, instead of directly copying the reference MB as the final recovered pixel values, we use a novel partial differential equation (PDE) based algorithm to refine the reconstruction. We minimize, in a weighted manner, the difference between the gradient field of the reconstructed MB in current frame and that of the reference MB in the reference frame under given boundary condition. A weighting factor is used to control the regulation level according to the local blockiness degree. With this algorithm, the annoying blocking artifacts are effectively reduced while the structures of the reference MB are well preserved. Compared with the error concealment feature implemented in the H.264 reference software, our algorithm is able to achieve significantly higher PSNR as well as better visual quality.
Yan Chen 0007, Yang Hu 0006, Oscar C. Au, Houqiang Li, Chang Wen Chen
IEEE Trans. Multim.2
2007 Image Search Result Clustering and Re-Ranking via Partial Grouping
abstract
Image search result clustering has become an active research topic. However, due to the limitations of current image search engines, the search result always exhibits partial clustering character, which makes the traditional clustering assumption unreasonable. In this paper, we apply Bregman bubble clustering (BBC), which clusters only a fraction of the whole data set, to image search result clustering. We show that relevant and irrelevant images are less mixed in the clusters produced by BBC. Therefore, we are able to incorporate a cluster based relevance feedback scheme to the clustering result and improve the relevance ranking of the search result according to user's feedback. Experiments on animal images from Flickr demonstrate the effectiveness of our clustering and re-ranking algorithms.
Yang Hu 0006, Nenghai Yu, Zhiwei Li 0006, Mingjing Li
ICME1
2007 Dual-Space Pyramid Matching for Medical Image Classification
Yang Hu 0006, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma
MMM (1)1