Cong Yu 0011

dblp:259/5253 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0001-6744-021XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 RF-URL 2.0: A General Unsupervised Representation Learning Method for RF Sensing
abstract
The major challenge in learning-based RF sensing is acquiring high-quality large-scale annotated datasets. Unlike visual datasets, RF signals are inherently non-intuitive and non-interpretable, making their annotation both time-consuming and labor-intensive. To address this challenge, we propose RF-URL 2.0, a novel unsupervised representation learning (URL) framework for RF sensing, which enables pre-training on easily collected, large-scale unannotated RF datasets to make downstream tasks solve easier. Existing URL techniques, such as contrastive learning, are primarily designed for natural images and are prone to learn shortcuts rather than meaningful information when applied to RF signals. RF-URL 2.0 is the first framework to overcome these limitations by constructing positive and negative pairs through well-established RF signal processing algorithms. Besides, it introduces a novel signal-model-driven augmentation technique, which augments signal representations by identifying and perturbing physically meaningful parameters of signal processing models. Moreover, the RF-URL 2.0 is carefully designed to take into account the heterogeneity characteristics of different RF signal processing representations. We show the universality of RF-URL 2.0 in three typical RF sensing tasks using two general RF devices (WiFi and radar), including human gesture recognition, 3D pose estimation, and silhouette generation. Extensive experiments on the HIBER and WiDAR 3.0 datasets demonstrate that RF-URL 2.0 takes a significant step toward learning-based solutions for RF sensing.
Ruiyuan Song, Dongheng Zhang, Cong Yu 0011, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 RPM 2.0: RF-Based Pose Machines for Multi-Person 3D Pose Estimation
abstract
Advanced human sensing technologies based on radio frequency (RF) signals have gained widespread attention in recent years. However, due to the sparsity and incompleteness of RF signals, fine-grained RF-based multi-person 3D pose estimation has progressed more slowly. In this paper, we present RF-based Pose Machine (RPM 2.0) for multi-person 3D pose estimation using RF signals. Specifically, we first develop a lightweight anchor-free detector module to locate and crop regions of interest from horizontal and vertical RF signals. Afterward, we treat the horizontal and vertical millimeter-wave radars as “RF cameras” with different viewing angles and propose a Multi-view Fusion Network to unproject the RF signals into a unified latent feature space, and then calculate the correlation for weighted fusion. Finally, a Spatio-Temporal Attention Network is designed to reconstruct the multi-person 3D skeleton sequences, in which the spatial attention module is proposed to recover invisible body parts using non-local correlations among joints and the temporal attention module refines the 3D pose sequences using temporal coherency learned from frame queries. We evaluate the performance of the proposed RPM 2.0 and state-of-the-art methods on a large-scale dataset with multi-person 3D pose labels and corresponding radar signals. The experimental results show that RPM 2.0 outperforms all of the baseline methods, which locates multi-person 3D key points with an average error of$73 mm$and generalizes well in new data such as occlusion, low illumination.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Circuits Syst. Video Technol.4
2024 SBRF: A Fine-Grained Radar Signal Generator for Human Sensing
abstract
While deep learning-based RF perception has received significant attention in recent years, the requirement for massive labeled RF data has hindered its further advancement. Despite existing efforts in synthesizing signals, they fail to accurately calculate the Radar Cross Section (RCS) of the target, leading to less practicality of the synthesized signals. In this paper, we introduce Simulated Body Radio Frequency (SBRF), a novel signal synthesis framework for calculating more realistic RCS by combining ray tracing with electromagnetic computation. SBRF involves three key components: a grid-based Shooting and Bouncing Ray (SBR) algorithm to calculate fine-grained human body RCS, a novel ray partitioning algorithm to improve the efficiency of ray tracing, and a coordinate transformation method to sense moving targets. Furthermore, we also design unique data augmentation techniques to improve the efficiency and generalizability of signal synthesis. Extensive experimental evaluations conducted on two publicly available datasets, involving wide-scale activity recognition and fine-grained gesture recognition, demonstrate the effectiveness of SBRF-generated signals in improving RF perception performance and alleviating the challenge of RF data collection.
Jiamu Li, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
IEEE Trans. Mob. Comput.4
2024 RPM: RF-Based Pose Machines
abstract
Radio-frequency (RF) based human sensing technologies, due to their great practical value in various applications and privacy-preserving nature, have gained tremendous attention in recent years. However, without fully exploiting the characteristics of radio signals, the performance of existing methods are still limited. First, RF features of the moving human body have different representations in dimensions such as channel and scale, which is challenging when performing feature fusion. Besides, the human body is specularly reflective with respect to the radar, which means the human body cannot be fully captured by a single RF snapshot. Therefore, the radar signal reflected by the human body is sparse and incomplete, which is difficult to extract high-quality features for 3D human pose estimation. In this paper, we present the RF-based Pose Machines (RPM), a novel framework which can generate 3D skeletons from RF signals. Considering the characteristics of RF signals, RPM includes several modules to overcome the challenges. Firstly, a Feature Fusion Network (FFN) is designed to effectively fuse radio signals from horizontal and vertical planes based on the channels' correlation and maintain high-quality feature via a multi-scale fusion block. A Spatio-Temporal Attention network is then designed to reconstruct 3D skeletons from the sparse and incomplete RF signals. Specifically, a spatial attention module is designed to model non-local relationships among joints and reconstruct body parts that a single RF snapshot cannot capture. Afterwards, a temporal attention module is proposed to refine 3D pose based on temporal coherency learned from frame queries. To evaluate the performance of our RPM framework, we construct a large-scale dataset of synchronized 3d skeletons and RF signals, RFSkeleton3D. Our experimental results show that RPM locates 3D key points of the human body with an average error of$5.71 cm$and maintains its performance in new environments with occlusion or bad illumination. The dataset and codes will be made in public.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.4
2024 MobiRFPose: Portable RF-Based 3D Human Pose Camera
abstract
Existing RF-based human pose estimation methods usually require intensive computations and cannot meet the real-time processing and portability requirements for mobile devices. To tackle the limitation, in this article, we introduce a lightweight RF-based pose estimation model, i.e., MobiRFPose, to construct the portable RF-based pose camera. Different from traditional optical-based cameras, the RF-based camera does not capture visual information, which means the privacy-preserving characteristic. Specifically, we only utilize a horizontal antenna array to transceive RF signals, then estimate the human locations on the RF signal heatmap and crop the human location regions, and finally estimate the fine-grained human poses based on the cropped small RF signal heatmaps. To evaluate the performance, we compare MobiRFPose with state-of-the-art methods. Experimental results demonstrate that MobiRFPose can achieve accurate 3D human pose estimation with fewer parameters and computations. We also test the trained MobiRFPose model using mobile computing devices, where the model structures and parameters only take up 268 KB and 3226 KB of disk space, and MobiRFPose can achieve 66 FPS processing speed. The pose estimation error is 11.05 cm in the case of a single person and 11.29 cm in the case of multiple people. All experimental results indicate that our proposed method can construct a portable RF camera to estimate human poses accurately.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.1
2023 Fast 3D Human Pose Estimation Using RF Signals
abstract
Existing deep learning-based wireless sensing models usually require intensive computation. In this paper, we introduce a lightweight RF-based 3D human pose estimation model, i.e., Fast RFPose, to enable real-time human pose estimation. Specifically, Fast RFPose first estimates the human locations in the RF heatmap and crops the human location regions, then estimates the fine-grained human poses based on the cropped small RF heatmaps. In the experiments, we build a radio system and a multi-view camera system to acquire the RF signals and the ground-truth human poses, and compare Fast RFPose with state-of-the-art methods. Experimental results demonstrate that Fast RFPose outperforms the alternative methods. Besides, we further deploy the trained Fast RFPose model on a laptop with a CPU and Fast RFPose can achieve 66 FPS processing speed, which means it can meet the real-time running requirements in mobile devices.
Cong Yu 0011, Yudong Zhang 0001, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
ICASSP1
2023 RF-based Multi-view Pose Machine for Multi-Person 3D Pose Estimation
abstract
In this paper, we present RF-based Multi-view Pose machine (RF-MvP) for multi-person 3D pose estimation using RF signals. Specifically, we first develop a lightweight anchor-free detector module to locate and crop regions of interest from horizontal and vertical RF signals. Afterward, we propose a Multi-view Fusion Network to unproject the RF signals from the horizontal and vertical millimeter-wave radars into a unified latent space, and then calculate the correlation for weighted fusion. Finally, a Spatio-Temporal Attention Network is designed to reconstruct the multi-person 3D skeleton sequences, in which the spatial attention module is proposed to recover invisible body parts using non-local correlations among joints and the temporal attention module refines the 3D pose sequences using temporal coherency learned from frame queries. We evaluate the performance of the proposed RF-MvP and state-of-the-art methods on a large-scale dataset with multi-person 3D pose labels and corresponding radar signals. The experimental results show that RF-MvP outperforms all of the baseline methods, which locates multi-person 3D key points with an average error of 73mm and generalizes well in new data such as occlusion, low illumination.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
ICME4
2023 RFPose-OT: RF-based 3D human pose estimation via optimal transport theory
abstract
This paper introduces a novel framework, i.e., RFPose-OT, to enable three-dimensional (3D) human pose estimation from radio frequency (RF) signals. Different from existing methods that predict human poses from RF signals at the signal level directly, we consider the structure difference between the RF signals and the human poses, propose a transformation of the RF signals to the pose domain at the feature level based on the optimal transport (OT) theory, and generate human poses from the transformed features. To evaluate RFPose-OT, we build a radio system and a multi-view camera system to acquire the RF signal data and the ground-truth human poses. The experimental results in a basic indoor environment, an occlusion indoor environment, and an outdoor environment demonstrate that RFPose-OT can predict 3D human poses with higher precision than state-of-the-art methods.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
Frontiers Inf. Technol. Electron. Eng.1
2023 Learning Fashion Compatibility With Context Conditioning Embedding
abstract
Fashion compatibility predictions have obtained a lot of attention recently. Mining the compatibility between fashion items in an outfit is different from learning the visual similarity, since this relationship is more delicate. Decomposing the outfit compatibility into pairwise item matching is a popular way to treat the problem. However, in most existing methods, the items are matched without considering the context, i.e, the remaining items in the outfit. Recent efforts have been made to learn the underlying high order relationships among items by treating the outfit as a whole. These models could be sensitive to the properties of different datasets, and the item representations in these models are not as compact as those in the pairwise models. In this paper, we propose a context conditioning embedding approach to learn compact representations that preserve the shared information among items under the existence of contextual items. We use two different spaces, the general and the contextual spaces, to embed items, where the representation in the contextual space contains information from the context. We employ mutual information maximization for model learning, which is shown to be more appropriate for the problem. With extensive experiments, we show that our model achieves superior performance than other state-of-the-art methods.
Yang Hu 0006, Cong Yu 0011, Yan Chen 0007, Bing Zeng 0001
IEEE Trans. Multim.3
2023 Personalized Fashion Recommendation With Discrete Content-Based Tensor Factorization
abstract
Fashion outfit recommendation has attracted lots of attention recently. The problem becomes even more interesting and challenging when considering users’ personalized fashion preferences. Although existing works have successfully improved the recommendation accuracy, the efficiency issue of computation and storage is still under-investigated and often ignored. In this paper, we propose a discrete content-based tensor factorization model that maps items and user to binary codes for efficient fashion recommendation. We introduce a probabilistic perspective for learning to hash, where the binary codes are sampled from a set of underlying Bernoulli variables. To demonstrate the effectiveness of our model, we collect a large-scale outfit dataset together with user label information from a fashion-focused social website. Extensive experiments on our dataset show that the proposed model outperforms other state-of-the-art methods.
Yang Hu 0006, Cong Yu 0011, Yunchao Jiang, Yan Chen 0007, Bing Zeng 0001
IEEE Trans. Multim.3
2023 RFMask: A Simple Baseline for Human Silhouette Segmentation With Radio Signals
abstract
Human silhouette segmentation, which is originally defined in computer vision, has achieved promising results for understanding human activities. However, the physical limitation makes existing systems based on optical cameras suffer from severe performance degradation under low illumination, smoke, and/or opaque obstruction conditions. To overcome such limitations, in this paper, we propose to utilize the radio signals, which can traverse obstacles and are unaffected by the lighting conditions to achieve silhouette segmentation. The proposed RFMask framework is composed of three modules. It first transforms RF signals captured by millimeter wave radar on two planes into spatial domain and suppress interference with the signal processing module. Then, it locates human reflections on RF frames and extract features from surrounding signals with human detection module. Finally, the extracted features from RF frames are aggregated with an attention based mask generation module. To verify our proposed framework, we collect a dataset containing804,760radio frames and402,380camera frames with human activities under various scenes. Experimental results show that the proposed framework can achieve impressive human silhouette segmentation even under the challenging scenarios (such as low light and occlusion scenarios) where traditional optical-camera-based methods fail. To the best of our knowledge, this is the first investigation towards segmenting human silhouette based on millimeter wave signals. We hope that our work can serve as a baseline and inspire further research that perform vision tasks with radio signals. The dataset and codes will be made in public.
Dongheng Zhang, Chunyang Xie, Cong Yu 0011, Jinbo Chen 0001, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.4
2023 RFGAN: RF-Based Human Synthesis
abstract
This paper demonstrates human synthesis based on the Radio Frequency (RF) signals, which leverages the fact that RF signals can record human movements with the signal reflections off the human body. Different from existing RF sensing works that can only perceive humans roughly, this paper aims to generate fine-grained optical human images by introducing a novel cross-modal RFGAN model. Specifically, we first build a radio system equipped with horizontal and vertical antenna arrays to transceive RF signals. Since the reflected RF signals are processed as obscure signal projection heatmaps on the horizontal and vertical planes, we design a RF-Extractor with RNN in RFGAN for RF heatmap encoding and combining to obtain the human activity information. Then we inject the information extracted by the RF-Extractor and RNN as the condition into GAN using the proposed RF-based adaptive normalizations. Finally, we train the whole model in an end-to-end manner. To evaluate our proposed model, we create two cross-modal datasets (RF-Walk&RF-Activity) that contain thousands of optical human activity frames and corresponding RF signals. Experimental results show that the RFGAN can generate target human activity frames using RF signals. To the best of our knowledge, this is the first work to generate optical images based on RF signals.
Cong Yu 0011, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE Trans. Multim.1
2022 Accurate Human Pose Estimation using RF Signals
abstract
Radio-frequency (RF) based human sensing technologies, due to their great practical value in various applications and privacy-preserving nature, have gained tremendous attention in recent years. However, without fully exploiting the characteristics of radio signals, the performance of existing methods are still limited. First, RF features of the moving human body have different representations in dimensions such as channel and scale, which is challenging when performing feature fusion. Besides, the human body is specularly reflective with respect to the radar, which means the human body cannot be fully captured by a single RF snapshot. Therefore, the radar signal reflected by the human body is sparse and incomplete, which is difficult to extract high-quality features for 3D human pose estimation. In this paper, we present the RF-based Pose Machines (RPM), a novel framework which can generate 3D skeletons from RF signals. Considering the characteristics of RF signals, RPM includes several modules to overcome the challenges. Firstly, a Multidimensional Feature Fusion (MFF) backbone is designed to effectively fuse radio signals based on the channels' correlation and maintain high-quality feature via a multi-scale fusion block. A Spatio-Temporal Attention network is then designed to reconstruct 3D skeletons by modeling the non-local spatio-temporal relationships. To evaluate the performance of our RPM framework, we construct a large-scale dataset of synchronized 3D skeletons and RF signals, RFSkeleton3D. Our experimental results show that RPM locates 3D key points of the human body with an average error of 5.71cm and maintains its performance in new environments with occlusion or bad illumination. The dataset and codes will be made in public.
Chunyang Xie, Dongheng Zhang, Cong Yu 0011, Yang Hu 0006, Qibin Sun, Yan Chen 0007
MMSP4
2022 WiFi-Based Human Pose Image Generation
abstract
This paper tackles a new challenge: how to generate human pose images from wireless signals? Although the optical camera can capture optical images, it is easily restricted by bad lighting. The wireless signals do not rely on visible lights. However, the low-resolution characteristics make previous works can only generate a rough skeleton of human posture, missing a lot of detailed visual information, such as background, appearance, etc. Since the visual information usually maintains unchanged for a period and the wireless signals can capture the movements of the human, in this paper, we propose a framework to generate the target human pose images by combining the wireless signals with an initial optical image. We utilize multiple wireless devices to collect the WiFi signals and a camera to capture the initial optical image. Then a data preprocessing component is designed to preprocess the wireless and vision data. Finally, a deep learning model learns to generate the human pose images from the processed wireless signals and the initial optical image. We conduct experiments to evaluate our proposed framework and results show that it achieves higher accuracy than the state-of-the-art WiFi-based pose estimation method and better visual quality than the state-of-the-art human generation method.
Cong Yu 0011, Dongheng Zhang, Chunyang Xie, Yang Hu 0006, Houqiang Li, Qibin Sun, Yan Chen 0007
MMSP1
2022 RF-URL: unsupervised representation learning for RF sensing
abstract
The major obstacle for learning-based RF sensing is to obtain a high-quality large-scale annotated dataset. However, unlike visual datasets that can be easily annotated by human workers, RF signal is non-intuitive and non-interpretable, which causes the annotation of RF signals time-consuming and laborious. To resolve the rapacious appetite of annotated data, we propose a novel unsupervised representation learning (URL) framework for RF sensing, RF-URL, to learn a pre-training model on large-scale unannotated RF datasets that can be easily collected. RF-URL utilizes a contrastive framework to mind the gap between signal-processing-based RF sensing and learning-based RF sensing. By constructing positive and negative pairs through different signal processing representations, RF-URL seamlessly integrates the existing RF signal processing algorithms into the learning-based networks. Moreover, the RF-URL is carefully designed to take into account the asymmetric characteristics of different RF signal processing representations. We show that RF-URL is universal to a variety of RF sensing tasks by evaluating RF-URL in three typical RF sensing tasks (human gesture recognition, 3D pose estimation and silhouette generation) based on two general RF devices (WiFi and radar). All experimental results strongly demonstrate that RF-URL takes an important step towards learning-based solutions for large-scale RF sensing applications.
Ruiyuan Song, Dongheng Zhang, Cong Yu 0011, Chunyang Xie, Yang Hu 0006, Yan Chen 0007
MobiCom4
2022 Passive Non-Line-of-Sight Imaging Using Optimal Transport
abstract
Passive non-line-of-sight (NLOS) imaging has drawn great attention in recent years. However, all existing methods are in common limited to simple hidden scenes, low-quality reconstruction, and small-scale datasets. In this paper, we propose NLOS-OT, a novel passive NLOS imaging framework based on manifold embedding and optimal transport, to reconstruct high-quality complicated hidden scenes. NLOS-OT converts the high-dimensional reconstruction task to a low-dimensional manifold mapping through optimal transport, alleviating the ill-posedness in passive NLOS imaging. Besides, we create the first large-scale passive NLOS imaging dataset, NLOS-Passive, which includes 50 groups and more than 3,200,000 images. NLOS-Passive collects target images with different distributions and their corresponding observed projections under various conditions, which can be used to evaluate the performance of passive NLOS imaging algorithms. It is shown that the proposed NLOS-OT framework achieves much better performance than the state-of-the-art methods on NLOS-Passive. We believe that the NLOS-OT framework together with the NLOS-Passive dataset is a big step and can inspire many ideas towards the development of learning-based passive NLOS imaging. Codes and dataset are publicly available (https://github.com/ruixv/NLOS-OT).
Ruixu Geng, Yang Hu 0006, Cong Yu 0011, Houqiang Li, Heng-Yu Zhang, Yan Chen 0007
IEEE Trans. Image Process.4
2019 Personalized Fashion Design
abstract
Fashion recommendation is the task of suggesting a fashion item that fits well with a given item. In this work, we propose to automatically synthesis new items for recommendation. We jointly consider the two key issues for the task, i.e., compatibility and personalization. We propose a personalized fashion design framework with the help of generative adversarial training. A convolutional network is first used to map the query image into a latent vector representation. This latent representation, together with another vector which characterizes user's style preference, are taken as the input to the generator network to generate the target item image. Two discriminator networks are built to guide the generation process. One is the classic real/fake discriminator. The other is a matching network which simultaneously models the compatibility between fashion items and learns users' preference representations. The performance of the proposed method is evaluated on thousands of outfits composited by online users. The experiments show that the items generated by our model are quite realistic. They have better visual quality and higher matching degree than those generated by alternative methods.
Cong Yu 0011, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001
ICCV1