Biyun Sheng

dblp:158/1357 · DBLP profile ↗
← Back
36ranked-venue papers
17as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NeuroForensics: Unmasking Covert Backdoor Attack via Endogenous Defense
Feirun Huang, Biyun Sheng, Lu Zhao 0001, Yanmeng Wang, Jian Zhou 0009
IWQoS3
2026 CA-PFL: Client-adaptive Parameter-efficient Fine-tuning for Personalized Federated Learning
Daixin Song, Biyun Sheng, Jian Zhou 0009, Mang Ye, Fu Xiao 0001
WWW4
2026 Self-Supervised Wi-Fi Activity Recognition via Iterative Pseudo-Labeling
abstract
With the rapid advancement of the Internet of Things (IoT) and smart environments, WiFi-based human activity recognition has made significant progress, providing non-intrusive and flexible sensing capabilities. However, most existing methods still depend on manual annotations, which severely limits their scalability and practicality in real-world scenarios where labeled data is scarce or unavailable. To address this, we propose LISAR, a novel two-stage iterative self-supervised WiFi activity recognition framework that learns from unlabeled data. Specifically, we introduce a contrastive pre-training approach by training with the proposed composite loss function, co-InfoNCE Loss, and a physics-informed data augmentation strategy to learn discriminative representations from unlabeled data. In addition, we design a Pseudo-label Confidence-guided Iterative Self-supervised Learning module, PC-ISL. Through iterative updates, this module refines pseudo-labels and enables efficient data utilization, thereby enhancing the model’s discrimination on unlabeled data. Extensive experiments demonstrate that LISAR achieves high-accuracy activity recognition while relying strictly on unlabeled data for representation learning, significantly out-performing state-of-the-art self-supervised methods.
Biyun Sheng, Linqing Gui, Fu Xiao 0001
IEEE Internet Things J.2
2026 Toward Personalized Location Privacy Trading for Mobile Crowd Sensing
abstract
With the commercialization of private data, location privacy trading in Mobile Crowd Sensing (MCS) has become a fascinating research topic. In consideration of location-dependent sensing tasks, mobile workers take risks at location privacy disclosure when reporting their actual locations. Existing work fail to take workers' diverse privacy protection and trading into account. This paper proposes a novel trading framework with personalized differential privacy guarantee, referred to asLeaper, to bridge the gap between location privacy protection and task allocation efficiency. In particular,Leaperoutputs a personalized obfuscated range for each worker and further obfuscates his location based on a perturbation set within this range by incorporating differential privacy and$k$-anonymity techniques, and thus improves the efficiency of task allocation. Moreover,Leaperquantifies each worker's location privacy loss and compensates him with reasonable payment by running auction in a cost-effective way. Through real-world datasets, our evaluations and analysis demonstrate thatLeaperindeed guarantees all desired properties of personalized differential privacy, truthfulness, individual rationality and budget feasibility.
Chen Lan, Yuanyuan Yang 0001, Fu Xiao 0001, Yanmin Zhu 0006, Jian Zhou 0009, Biyun Sheng
IEEE Trans. Dependable Secur. Comput.7
2026 Error-Correction Enabled Contactless Sedentary Behavior Detection via WiFi Sensing
abstract
Sedentary lifestyle has become a major health risk in modern society. Long sitting time can be detected by accurate recognition of sitting and standing (sit-stand) activities. WiFi-based sitting time detection has the remarkable advantage of low cost, noncontact, and privacy-protection. However, accurate recognition of sit-stand activities via WiFi signal is still facing two challenges. The first and also tougher challenge is inevitable mistakes in recognition results of traditional machine learning methods, while the second challenge is the difficulty of accurate activity segmentation before activity recognition. To the best of the authors' knowledge, few work addresses the above challenges, particularly the first challenge. A new contactless sitting time detection system is designed accordingly. The system first accurately segments all activities and removes in-seat activities. Then, the mistakes in sit-stand activity recognition are effectively corrected by a new recognition error correction method. The proposed method first creates and updates a correction benchmark that can satisfy both successive correlation and waveform symmetry between sit-stand activities. The recognition results of traditional machine learning methods are then corrected based on the latest correction benchmark. Extensive experiment results demonstrate that compared to related work, the designed system has much better accuracy on both sit-stand activity recognition and sitting time estimation. The experiment results also demonstrate the robustness of the designed system.
Linqing Gui, Chuanyue Xie, Biyun Sheng, Fu Xiao 0001
IEEE Trans. Hum. Mach. Syst.4
2026 ADGTrace: Achieving Adaptive Trajectory Synthesis With Generated Data
abstract
User trajectory publication has promoted various location-based applications like user travel recommendation. However, possible privacy leakages have hindered more inclusive trajectory data analysis and utilization. Privacy-preserving trajectory synthesis is a popular approach to address the above privacy issues. Existing methods unavoidably produce low trajectory utility since they usually apply perturbed versions of human moving patterns. Worse still, they cannot adaptively adjust this synthesis according to the varying granularity demands of different users. This paper proposes a novel adaptive trajectory synthesis framework with generated data, namelyADGTrace. Our model achieves privacy preservation without introducing additional noise while maintaining high adaptation.ADGTracedirectly synthesizes artificial trajectories that share the similar patterns with real ones through agenerative and selectiveoptimization process. Additionally, we present a grid granularity alignment strategy to achieve adaptive trajectory synthesis, satisfying varying user demands. Extensive experiments on real-world datasets demonstrate the superiority ofADGTraceover the state-of-the art methods under various utility metrics, maintaining strong attack resilience.
Chen Lan, Biyun Sheng, Jian Zhou 0009, Yuanyuan Yang 0001, Yanmin Zhu 0006, Fu Xiao 0001
IEEE Trans. Mob. Comput.3
2026 AceNet: Attention-Guided Context Enhancement for Imbalanced Action Recognition via RF Signals
abstract
Although radio frequency (RF)-based activity recognition has made significant progress in recent years, the sensing performance will be significantly degraded under class imbalance conditions, especially when minority and majority classes share semantically similar local motion patterns. Traditional data augmentation approaches in the original sample space may cause semantic deviation and meanwhile bring high computational cost. Instead, in this work we turn to address the issue at feature level, in which we focus on how to distinguish highly similar actions and mitigate imbalance-induced decision boundary bias. To tackle these challenges, we present attention-guided context enhancement network (AceNet), which designs a discriminative feature extractor and develops a feature-augmentation based classifier refinement strategy. Specifically, an attention-guided mechanism is presented to dynamically select the most distinctive temporal segments, and a hierarchical Transformer structure is then proposed to characterize both inner-segment micro-dynamics and inter-segment contextual relationships. Moreover, AceNet synthesizes features via Synthetic Minority Over-sampling Technique (SMOTE) to balance feature distribution for each category and then refine the classifier parameters to mitigate class imbalance bias. Comprehensive experiments on two public datasets with different RF modalities demonstrate that AceNet significantly outperforms existing approaches under various levels of data scarcity at a low cost.
Biyun Sheng, Yiping Zuo, Jian Zhou 0009, Fu Xiao 0001
IEEE Trans. Mob. Comput.1
2026 AE-IPP: Adversarial Example Enabled Identity Privacy Preserving With mmWave Signals
abstract
Despite convenience and reliability of mmWave-based action recognition, it still raises privacy concerns on identity leakage threat since human behaviors could meanwhile expose massive user information in real-world applications. Existing solutions attempt to send anonymized features extracted from mmWave signals; however, features not only reduce the application flexibility but also increase the privacy disclosure risk due to original data reconstruction. Instead, in this paper we propose a de-identification system, AE-IPP, which customizes learned noises into the raw data to generate adversarial examples for identity privacy and action utility balance. In other words, the noises are sample-specific perturbations that are automatically learned for each sample through our presented network. To achieve the performance balance and ensure robustness to other models, we are faced with two challenges, including the decoupling of action and identity information and the transferability of models. To this end, AE-IPP focuses on respective attention areas by leveraging task-specific gradients and designs a dynamic attention mechanism to update the attention weights according to the final optimization objective. Moreover, we present a multidirectional perturbation strategy to improve the model generalization capabilities, enabling robust de-identification. Extensive experiments on mmWave datasets demonstrate the superiority of our method over state-of-the-art approaches.
Biyun Sheng, Wangquan Qin, Jun Li 0033, Li Lu 0008, Tie Qiu 0001, Fu Xiao 0001
IEEE Trans. Mob. Comput.1
2026 Toward Adaptive Person Re-Identification via mmWave Radar Point Clouds
abstract
MmWave radar-based person re-identification (ReID), namely catching a specified person from the database, demonstrates enormous prospects for practical applications such as public security and intelligent surveillance. Towards adaptive ReID across different scenarios, we attempt to explore abundant spatial-temporal features from point clouds for walking individuals. Most existing approaches fail to describe fine-grained 3D spatial properties associated with gaits and neglect the impacts of walking speed changes on perception. To address the two problems, we extract gait features by integrating the anchor-based orientation descriptor (AOD) and multi-scale gait catch (MGC) modules into a ReID system. Specifically, AOD automatically learns virtual anchors, selects anchor-centered neighbors from eight different subspaces and designs orientation-driven feature aggregation to elaborately describe 3D local space. Then, MGC adopts the sub-sampling strategy on AOD results to estimate multiple temporal resolution pathways reflecting relative speeds, on which sequence-specific and cross-sequence dependencies are respectively characterized by our self-attention (SA) and hierarchical attention (HA) to generate more discriminative gait representations. For evaluation, we collect mmWave ReID datasets at three scenes, and comprehensive experiments illustrate that we can maintain superior ReID performances over 90.0% Top-1 accuracy under different scenario settings. Our code and dataset are available athttps://github.com/dpjqw195/ReID-AOD-MGC
Biyun Sheng, Pengju Ding, Fu Xiao 0001, Tie Qiu 0001
IEEE Trans. Netw.1
2025 Hierarchical Matching Game for Multiple User Association in Fully Decoupled Networks
abstract
In fully decoupled networks with separate uplink/downlink (UL/DL) base station (BS) deployments and high user mobility, ensuring efficient UL/DL user association remains a critical challenge. The dynamic environment and complex channel conditions necessitate state perception for optimal association strategies, while the densification of nodes demands scalable solutions to handle increased combinatorial complexity in UL and DL transmissions. This paper introduces a novel framework leveraging unmanned aerial vehicle (UAV) sensing-assisted to predict user mobility and channel dynamics, combined with multiple association mechanism to address dense node interactions. Accordingly, a joint optimization problem is formulated, where Kriging-based prediction is adopted to assist the user association for both UL and DL. To solve it, a hierarchical matching game is developed to decompose the joint problem into decoupled UL and DL games. Particularly, a low-complexity Kriging prediction-based hierarchical matching algorithm is designed to obtain the solution. Simulation results in dynamic network scenarios demonstrate that the effectiveness of proposed approach and the superiority is validated by comparisons.
Chen Dai, Haotong Cao, Biyun Sheng, Wael Bazzi, Shahid Mumtaz
GLOBECOM3
2025 Correlation-Aware Multi-Similarity Learning for Federated Human Activity Recognition
abstract
Centralized training for Human Activity Recognition (HAR) typically relies heavily on vast amounts of aggregated data, compromising user privacy. Federated learning (FL) for HAR offers a solution to protect local data privacy. However, existing FL methodologies often fail to fully capture the heterogeneity of user data and the latent correlations among user models, resulting in suboptimal performance and limited robustness. This paper proposes a Correlation-Aware Multi-Similarty Learning Method for Federated HAR, namely MultiSim. Our approach enhances model accuracy with an effective inter-user knowledge learning while protecting data privacy. MultiSim first constructs multiple similarity metrics, and then makes model feature fusion cunningly by the above metrics to learn inherent user similarity profiles. Additionally, we introduce a novel clustering-based FL framework by isolating malicious nodes, thereby mitigating the impact of adversarial attacks. Extensive evaluations on two realworld HAR datasets demonstrate the superiority of MultiSim over other state-of-the-art FL methods under accuracy and robustness. These findings demonstrate MultiSim's potential as a robust and effective solution for HAR.
Jinming Ju, Tianyang Zhou, Biyun Sheng, Jian Zhou 0009, Weibei Fan, Fu Xiao 0001
IWQoS4
2025 mmReID: Person Reidentification Based on Commodity Millimeter-Wave Radar
abstract
Person reidentification (Re-ID) plays an increasingly important role in the development of smart cities and public security systems. Typical person Re-ID is generally used to query and retrieve pedestrians across cameras in the form of images or videos. For the sake of privacy protection and invariability to resolution, light, and occlusion, Re-ID performs excellent prospect by using millimeter-wave radars. Existing radio frequency (RF) person Re-ID approaches either suffer from relatively unreliable accuracy due to the sparse characteristics of point clouds or additionally rely on other sensing task (e.g., 3-D skeleton prediction) to avoid overfitting. In this article, we present mmReID, an RF person Re-ID system which integrates frequency-modulated continuous wave (FMCW) mmWave radar time-velocity micro-Doppler imaging heatmaps with different frequencies into the proposed dual-stream multilayer feature fusion network named ConvSnet. ConvSnet effectively fuses shallow and deep features at different levels, and adopts the attention module with intramodal aggregation to extract the contextual relevance of velocity features. In addition, we construct and publish a dataset mmReIData based on mmWave radar for person Re-ID, consisting of sampling RF data from 41 pedestrians. Experimental results show that our proposed ConvSnet network achieves the best performance against other state-of-the-art networks both in person identification and Re-ID tasks. Further ablation studies indicate the effectiveness of each component of the ConvSnet network.
Chong Han 0002, Biyun Sheng, Jian Guo 0006
IEEE Internet Things J.3
2025 A Blockchain-Based Secure and Fair Online Incentive Mechanism for Crowdsensed Data Trading
abstract
With the development of blockchain technology, Blockchain-based Crowdsensed Data Trading (BCDT) has emerged as an attractive data exchange paradigm. Although it addresses security issues in data transactions, most recent research primarily focuses on offline scenarios, overlooking the critical importance of enabling real-time online data trading, where it suffers from dynamic worker participation and potential malicious attacks. In this paper, we propose a Blockchain-based Secure and Fair Online Incentive Mechanism (BSFOIM), which primarily incorporates a smart contract called BSFOIMToken, designed to function in online scenarios. In particular, we first introduce a multi-stage auction combined with a time discount factor in BSFOIM to quantify the contribution of workers in completing sensing tasks. Meanwhile, to ensure sensing data quality and worker selection fairness, we propose a Fairness-based Truth Discovery Mechanism (FTDM) with two core modules: a fine-grained reputation system to identify reliable workers and filter out malicious ones, and an upper confidence bound algorithm to optimize worker selection and avoid local optima. Finally, we implement these functions in BSFOIMToken and deploy a prototype on the Ethereum blockchain, demonstrating its practicality and robust performance. Rigorous theoretical and comprehensive experimental tests have proven their adherence to truthfulness, budget feasibility and individual rationality.
Biyun Sheng, Juan Li 0011, Jian Zhou 0009, Haiping Huang, Mang Ye, Fu Xiao 0001
IEEE Trans. Inf. Forensics Secur.3
2025 I Sense You Fast: Simultaneous Action and Identity Inference by Slimming Multi-Branch RadarNet
abstract
With the increasing connection between internet and human society, millimeter-wave radar based action recognition and user authentication exhibit remarkable prospects in security scenarios. Existing solutions usually focus on one of the tasks and mainly emphasize accuracy without reducing the inference time. In this paper, we propose a dual-task based Polymorphic Lightweight (PolyLite) RadarNet framework, in which the shared features are fed into two split streams for different tasks under joint supervision. The polymorphic concept here means that the trained network with parallel designs can be slimmed as a single-branch structure for inference. By this design strategy, we can not only efficiently extract spatial-temporal features during the training stage but also largely improve the response speed for simultaneously testing human activities and identities. Specifically, we design triple-view (TRIview) video-like data as the input by successively concatenating the range-velocity and range-angle matrices. Then a PolyLite module with linear and lightweight designs in each branch is integrated into our RadarNet framework to learn discriminative representations. Experimental results demonstrate that our approach is able to reach the accuracy over 98${\%}$within 0.21ms inference time. Especially, untrained intruders can also be successfully identified by a simple matching computation. Our code is available athttps://github.com/MagicalLiHua/PolyLite-RadarNet.
Biyun Sheng, Linqing Gui, Fu Xiao 0001
IEEE Trans. Mob. Comput.1
2025 mmZeAR: Zero-Effort Cross-Category Action Recognition With mmWave Radar
abstract
Despite the widespread application of radio frequency (RF) signal-based human action recognition, traditional solutions can only recognize seen categories and the perception scope is restrained by the limited activity classes. When a novel category emerges, the model needs to be optimized again on additionally collected samples at the cost of computation and labor burden. To address this challenge, we develop the mmZeAR system, which learns semantic knowledge from available vision data as class attributes and then transforms the classification into a matching problem. Specifically, we build the attribute space by fusing the coarse-grained video classification features and fine-grained angle change features of 3D joint skeletons. Then we design an efficient feature extraction backbone named TriSqN, which integrates triple radar heatmaps into the final representations by sufficiently exploring the heterogeneous and complementary characteristics. Finally, a projection network is developed between semantic attributes and radar features to construct indirect relationships between samples and labels. By implementing mmZeAR on millimeter wave (mmWave) radar signal datasets, our extensive experiments have demonstrated its remarkable recognition accuracy in novel category recognition with zero effort and achieved state-of-the-art performance.
Biyun Sheng, Jiabin Li, Yiping Zuo, Li Lu 0008, Fu Xiao 0001
IEEE Trans. Mob. Comput.1
2024 UWTracking: Passive Human Tracking Under LOS/NLOS Scenarios Using IR-UWB Radar
abstract
Passive human tracking plays a critical role in the field of ubiquitous sensing, offering customized services such as real-time location tracking for vital sign monitoring and motion detection. Traditional contact-free tracking systems are primarily designed for Line-of-Sight (LOS) scenarios, requiring a direct path between the radio device and the target. However, in Non-Line-of-Sight (NLOS) scenarios, where obstacles obstruct this direct path, these systems suffer from sensing failures and are unable to accurately obtain the motion trajectory of the sensing target. In this paper, we propose the UWTracking system, which utilizes the Commercial Off-the-Shelf (COTS) Impulse Radio-Ultra Wideband (IR-UWB) radars to enable precise indoor passive human tracking in both LOS and NLOS scenarios. To effectively capture the motion information of a moving target in NLOS scenarios, we present the Reconstructed Distributed- Doppler Frequency Shift (RD-DFS) features. We then binarize the RD-DFS features and design the Distance Extraction Algorithm (DEA) to obtain the target's distance in both scenarios. Subsequently, the Circle Intersection Method with Distance Stretching (CIM-DS) algorithm is developed to determine the indoor position of the sensing target, and the Scanning Angle and Velocity Particle Filter (SAV-PF) algorithm facilitates high-precision trajectory tracking. We implement a prototype of UWTracking system and conduct extensive evaluations to showcase its trajectory tracking performance under various scenarios. The results demonstrate that UWTracking achieves effective real-time tracking, with a median tracking error of 17.65 cm in the LOS scenario and 23.34 cm in the NLOS scenario, outperforming the state-of-the-art trajectory tracking systems based on COTS IR-UWB Radars.
Dongzi Wang 0001, Linqing Gui, Biyun Sheng, Fu Xiao 0001, Jinsong Han
IEEE Trans. Mob. Comput.4
2024 RoSeFi: A Robust Sedentary Behavior Monitoring System With Commodity WiFi Devices
abstract
Sedentary behaviors are shown to be hazardous to human health. Detecting sedentary behaviors in a ubiquitous way can be realized by the promising WiFi sensing technique. The accurate detection of sedentary behaviors is determined by the accurate recognition of sit-stand postural transition (SPT). However, according to our findings, SPT recognition errors are inevitable even with advanced machine-learning methods, because different SPTs may result in a similar change in WiFi channel state information (CSI). To effectively reduce SPT recognition errors, in this paper we propose RoSeFi, a robust sedentary behavior monitoring system. We first classify the errors in SPT recognition results into two categories: the errors violating SPT's consistency and the errors violating SPTs' symmetry. To correct the above errors, we reveal two inherent features in the CSI data of SPTs, i.e., contextual association and waveform mirror symmetry. Then a novel metric named WMSF is defined to quantify the degree of waveform mirror symmetry between two SPTs' CSI data. Integrating the above features, the problem of recognition error correction can be modeled as a constrained nonlinear optimization problem (CNOP). To solve the problem, we design a unified error detection/correction scheme, named UEDC, which converts the CNOP into a sequence decoding problem in Hidden Markov Model (HMM). A tailored Viterbi algorithm combined with WMSF is proposed to detect and correct the errors simultaneously. The experimental results show that RoseFi reduces 60-82% SPT recognition errors, gains 15-20% relative improvement in the accuracy of SPT recognition, and eventually reduces the sedentary time estimation errors by 10%-20%, compared with typical existing systems. In addition, our error correction method can be adapted to most existing machine learning based human action recognition methods, effectively improving their performance.
Cheng Peng 0019, Linqing Gui, Biyun Sheng, Fu Xiao 0001
IEEE Trans. Mob. Comput.3
2024 CDFi: Cross-Domain Action Recognition Using WiFi Signals
abstract
Contactless WiFi based human action recognition exhibits remarkable prospects in the fields such as human-computer interaction and smart home. However, domain dependency restricts its generalization into the real-world deployment. Since it is expensive to label enough new data for retaining a model, it is beneficial to explore few-shot learning for cross-domain sensing with limited target labels. Nevertheless, there are two challenges to be addressed. The first challenge is how to select a suitable dataset from a series of available source domains to prevent negative transfer. The second is to mine action-related characteristics by the feature learning model for the following effective knowledge transfer. In order to tackle the above challenges, we present a cross-domain sensing framework named CDFi, which consists of Nearest Neighbor based Domain Selector (NNDS) and Fine-to-Coarse-Grained Transformer Network (FCGTN). NNDS is proposed to evaluate the source-target domain similarities by measurements among local and global feature distributions. Besides, FCGTN embeds convolution map based hierarchical transformer structures and the modified linear layer into an end-to-end deep network, which can quickly adapt to the unseen domain by few samples. Comprehensive experiments show that CDFi can effectively realize cross-domain action recognition, and achieve about 4 cross-scene cases, respectively, compared to the state-of-the-art.
Biyun Sheng, Fu Xiao 0001, Linqing Gui
IEEE Trans. Mob. Comput.1
2024 LiteWiSys: A Lightweight System for WiFi-based Dual-task Action Perception
abstract
As two important contents in WiFi-based action perception, detection and recognition require localizing motion regions from the entire temporal sequences and classifying the corresponding categories. Existing approaches, though yielding reasonably acceptable performances, are suffering from two major drawbacks: heavy empirical dependency and large computational complexity. In order to solve these issues, we develop LiteWiSys in this article, a lightweight system in an end-to-end deep learning manner to simultaneously detect and recognize WiFi-based human actions. Specifically, we assign different attentions on sub-carriers, which are then compressed to reduce noise and information redundancy. Then, LiteWiSys integrates deep separable convolution and a channel shuffle mechanism into a multi-scale convolutional backbone structure. By feature channel split, two network branches are obtained and further trained with a joint loss function for dual tasks. We collect different datasets at multi-scenes and conduct experiments to evaluate the performance of LiteWiSys. In comparison to existing WiFi sensing systems, LiteWiSys achieves promising precision with lower complexity.
Biyun Sheng, Jiabin Li, Linqing Gui, Fu Xiao 0001
ACM Trans. Sens. Networks1
2023 Efficient Respiration Rate Estimation Based on MIMO mmWave Radar
Ling Deng, Biyun Sheng, Linqing Gui, Fu Xiao 0001
ICA3PP (3)3
2023 DyLiteRADHAR: Dynamic Lightweight Slowfast Network for Human Activity Recognition Using MMWAVE Radar
abstract
Millimeter-wave radar based human activity recognition (RADHAR) exhibits remarkable prospects in the field of device-free sensing. However, most existing RADHAR systems only focus on performance improvement, failing to simultaneously lighten the network parameters. In this paper, we propose a dynamic lightweight SlowFast network named DyLiteRADHAR, which can efficiently extract spatial-temporal features and largely reduce the resource consumption for human activity recognition. Specifically, we design triple-view signal maps (TRIview) as the input by successively concatenating the range-velocity, range-azimuth and range-elevation matrices. Then dynamic lightweight network is presented to learn discriminative representations which integrates dynamic convolution and lightweight shuffle net structure into the SlowFast framework. Experimental results demonstrate that the proposed approach DyLiteRADHAR is able to achieve superiority performance with limited computation complexity.
Biyun Sheng, Fu Xiao 0001, Linqing Gui
ICASSP1
2023 MMHeart: An Efficient Heartbeat Monitoring System Based on MIMO mmWave Radar
abstract
Heart rate provides aln important reference for human physical conditions and psychological changes. MmWave-based heart rate estimation has increasingly attracted attention in recent years due to its non-intrusiveness and cost-effectiveness. However, when the subject locates far away from the mmWave radar and also deviates from it, the low accuracy of heart rate estimation becomes a major concern. This paper presents MMHeart, a new heart rate estimation and heartbeat waveform reconstruction system based on MIMO mmWave radar. In order to effectively improve the accuracy of heart rate estimation, MMHeart first calculates appropriate range bins based on positioning results, then estimates candidate heart rates in all channels, removes abnormal candidates based on spectrum kurtosis, and finally estimates the heart rate by clustering the remaining candidates. Then in order to reconstruct more accurate heartbeat waveform, MMHeart first segments the signal by trough detection, then resamples heartbeat waveform template, and finally fine-tunes the start and end points of each segment. Our extensive experiments show that in long-range and large-deviation scenarios, MMHeart can improve the accuracy of heart rate estimation by at least 54.1% compared to mmEGC and PiVimo, while it can improve the accuracy of cardiac cycle duration by 52.2% compared to mmEGC.
Linqing Gui, Ling Deng, Cheng Peng 0019, Biyun Sheng, Fu Xiao 0001
MSN5
2023 MuAt-Va: Multi-Attention and Video-Auxiliary Network for Device-Free Action Recognition
abstract
With the growing popularity of Internet of Things (IoT) systems, device-free action recognition begins to attract extensive attention due to its friendly feasibility in broad applications, such as human–computer interaction and smart elderly care. Considering abundant information in the vision modality, existing methods adopt the cross-model methods for performance enhancement. However, the dependency of synchronous multimodal data in the collection and recognition stage brings into the vision weaknesses, such as sensitivity to occlusion and privacy invasion. In this article, we integrate multi-attention structure and auxiliary video information into a novel end-to-end deep learning framework named MuAt-Va, in which video soft labels learned in advance are utilized to teach the multi-attention WiFi feature training process without vision information involved during the test. Specifically, in order to enlarge the application scope and reduce the data cost, we beforehand acquire videos under a satisfactory condition only once, and then leverage teacher–student mechanism to guide the WiFi stream. Instead of straightforwardly concatenating multiantenna channel state information (CSI) from homogeneous wireless signals as previous works, we design a CSI subcarrier-wise, temporal-wise, and view-wise attention module to assign different weights on the basis of data characteristics for the sensing task. Our experiments with multiple subjects data in two scenes demonstrate that MuAt-Va can accurately recognize human actions with more superior performances.
Biyun Sheng, Chaorun Sun, Fu Xiao 0001, Linqing Gui
IEEE Internet Things J.1
2023 Context-Aware Faster RCNN for CSI-Based Human Action Perception
abstract
With the widespread deployment of commercial wireless devices, researchers begin to focus on device-free sensing tasks. In the field of action perception, existing WiFi-based sensing works mostly follow the framework in which action instances of channel state information (CSI) are first extracted and then classified. As for the part of human action detection, a majority of works adopt threshold based sliding window or frame-by-frame detection methods. However, it is hard for the former approach to set a reasonable threshold for all samples. As for the latter, it costs a relatively substantial amount of labor to label each moment of the time sequences. In order to overcome the above problems, we design an end-to-end context-aware faster region-based convolutional neural networks (RCNN) framework named Wisense to simultaneously detect the temporal boundaries as well as classify the actions. More specifically, Wisense consists of backbone net, region proposal net (RPN), pooling layer, and the prediction net, which directly regresses the action location along the time axis and classifies the action types. For the sake of wireless signal temporal detection, we transform the input into 1-D feature map and extract multiscale 1-D anchors. Besides, in order to sufficiently mine the context information, we extend the boundaries of region proposals and further establish the temporal pyramid features. Experimental results conducted in three indoor scenes validate the effectiveness of our proposed Wisense.
Biyun Sheng, Fu Xiao 0001, Linqing Gui
IEEE Trans. Hum. Mach. Syst.1
2023 BreatheBand: A Fine-grained and Robust Respiration Monitor System Using WiFi Signals
abstract
Respiration is a vital indicator of the state of the human body. Monitoring human respiration enables the realization of a variety of intelligent applications, including smart medical and sleep monitoring. Traditional methods that are dependent upon wearable devices are more costly and inconvenient for users. Recent studies have evidenced that low-cost commodity WiFi devices can be used to accomplish contactless respiration monitoring. In this article, we present BreatheBand, a fine-grained and robust respiration monitoring system based on commercial WiFi signals. We first remove the time-varying phase shift in the channel state information (CSI) by developing the Multi-antenna CSI–Subpopulation Genetic algorithm. Then we separate human respiratory components from WiFi signals by employing subcarrier selection and Independent Component Analysis. Next, applying a Mixed Cluster Gaussian–Hidden Markov Model, we generate a respiration signal resembling that of wearable devices. Finally, we integrate the BreatheBand system into commercial WiFi infrastructure. The results show that the BreatheBand’s respiration signal is remarkably identical to the signal collected by the wearable device in various scenarios. In particular, the mean absolute error of the BreatheBand’s respiration rate is approximately 0.1 bpm, outperforming state-of-the-art algorithms.
Wenyang Yuan, Linqing Gui, Biyun Sheng, Fu Xiao 0001
ACM Trans. Sens. Networks4
2022 Cross-scene passive human activity recognition using commodity WiFi
Yuanrun Fang, Fu Xiao 0001, Biyun Sheng, Letian Sha
Frontiers Comput. Sci.3
2022 Inner Knowledge-based Img2Doc Scheme for Visual Question Answering
abstract
Visual Question Answering (VQA) is a research topic of significant interest at the intersection of computer vision and natural language understanding. Recent research indicates that attributes and knowledge can effectively improve performance for both image captioning and VQA. In this article, an inner knowledge-based Img2Doc algorithm for VQA is presented. The inner knowledge is characterized as the inner attribute relationship in visual images. In addition to using an attribute network for inner knowledge-based image representation, VQA scheme is associated with a question-guided Doc2Vec method for question–answering. The attribute network generates inner knowledge-based features for visual images, while a novel question-guided Doc2Vec method aims at converting natural language text to vector features. After the vector features are extracted, they are combined with visual image features into a classifier to provide an answer. Based on our model, the VQA problem is resolved by textual question answering. The experimental results demonstrate that the proposed method achieves superior performance on multiple benchmark datasets.
Qun Li 0002, Fu Xiao 0001, Bir Bhanu, Biyun Sheng, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.4
2021 TS-Net: Device-Free Action Recognition with Cross-Modal Learning
Biyun Sheng, Linqing Gui, Fu Xiao 0001
WASA (1)1
2020 Multilayer deep features with multiple kernel learning for action recognition
Biyun Sheng, Fu Xiao 0001, Wankou Yang
Neurocomputing1
2020 WiReader: Adaptive Air Handwriting Recognition Based on Commercial WiFi Signal
abstract
In recent years, with the rapid development of the Internet-of-Things (IoT) technologies, many intelligent sensing applications have emerged, which realize contactless sensing and human-computer interaction (HCI). Handwriting recognition is the communication link between the human and computer. Previous handwriting recognition applications are usually founded on images and sensors, which require significant device overhead and are device dependent. Recently, the revolution of the wireless signal sensing technology has laid the foundation for the intelligent handwriting recognition technology without devices. In this article, we propose WiReader, an adaptive air handwriting recognition system based on wireless signals. WiReader utilizes ubiquitous commercial WiFi devices to process the collected channel state information (CSI), segments the data in combination with activity factors, and then transforms the original signal using the CSI-Ratio model. In order to address the problem of feature extraction caused by handwriting, we utilize the cumulative principal components and multilayer wavelet transform for the transformed signal. Finally, the energy feature matrix is generated and combines with long short-term memory (LSTM) to realize the recognition of different handwriting actions. Extensive real-world experiments show that WiReader achieves an average recognition accuracy of 90.64% leading other applications in three scenarios and has strong robustness to user location, user diversity, and different scenarios.
Fu Xiao 0001, Biyun Sheng, Huan Fei, Shui Yu 0001
IEEE Internet Things J.3
2020 Deep Spatial-Temporal Model Based Cross-Scene Action Recognition Using Commodity WiFi
abstract
With the popularization of Internet-of-Things (IoT) systems, passive action recognition on channel state information (CSI) has attracted much attention. Most conventional work under the machine-learning framework utilizes handcrafted features (e.g., statistic features) that are unable to sufficiently describe the sequence data and heavily rely on designers' experiences. Therefore, how to automatically learn abundant spatial-temporal information from CSI data is a topic worthy of study. In this article, we propose a deep learning framework that integrates spatial features learned from the convolutional neural network (CNN) into the temporal model multilayer bidirectional long short-term memory (Bi-LSTM). Specifically, CSI streams are segmented into a series of patches, from which spatial features are extracted by our designed CNN structure. Considering long-term dependencies between adjacent sequences, the fully connected layer of CNN for each patch is taken as the Bi-LSTM sequential input to further capture temporal features. Our model is appealing in that it can simultaneously learn temporal dynamics and convolutional perceptual representations. To the best of our knowledge, this is the first work to explore deep spatial-temporal features for CSI-based action recognition. Furthermore, in order to solve the problem that the trained model fully fails with environmental changes, we use the off-the-shelf model as the pretrained model and fine-tune it in the new scenario. The transfer method is able to realize cross-scene action recognition with low computational consumption and satisfactory accuracy. We carry out experiments on indoor data and the experimental results validate the effectiveness of our algorithm.
Biyun Sheng, Fu Xiao 0001, Letian Sha
IEEE Internet Things J.1
2020 Discriminative Multi-View Subspace Feature Learning for Action Recognition
abstract
Although deep features have achieved the state-of-the-art performance in action recognition recently, the hand-crafted shallow features still play a critical role in characterizing human actions for taking advantage of visual contents in an intuitive way such as edge features. Therefore, the shallow features can serve as auxiliary visual cues supplementary to deep representations. In this paper, we propose a discriminative subspace learning model (DSLM) to explore the complementary properties between the hand-crafted shallow feature representations and the deep features. As for the RGB action recognition, this is the first work attempting to mine multi-level feature complementaries by the multi-view subspace learning scheme. To sufficiently capture the complementary information among heterogeneous features, we construct the DSLM by integrating the multi-view reconstruction error and classification error into an unified objective function. To be specific, we first use Fisher Vector to encode improved dense trajectories (iDT+FV) for shallow representations and two-stream convolutional neural network models (T-CNN) for generating deep features. Moreover, the presented DSLM algorithm projects multi-level features onto a shared discriminative subspace with the complementary information and discriminating capacity simultaneously incorporated. Finally, the action types of test samples are identified by the margins from the learned compact representations to the decision boundary. The experimental results on three datasets demonstrate the effectiveness of the proposed method.
Biyun Sheng, Jun Li 0033, Fu Xiao 0001, Qun Li 0002, Wankou Yang, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Crowd Counting via Weighted VLAD on a Dense Attribute Feature Map
abstract
Crowd counting is an important task in computer vision, which has many applications in video surveillance. Although the regression-based framework has achieved great improvements for crowd counting, how to improve the discriminative power of image representation is still an open problem. Conventional holistic features used in crowd counting often fail to capture semantic attributes and spatial cues of the image. In this paper, we propose integrating semantic information into learning locality-aware feature (LAF) sets for accurate crowd counting. First, with the help of a convolutional neural network, the original pixel space is mapped onto a dense attribute feature map, where each dimension of the pixelwise feature indicates the probabilistic strength of a certain semantic class. Then, LAF built on the idea of spatial pyramids on neighboring patches is proposed to explore more spatial context and local information. Finally, the traditional vector of locally aggregated descriptor (VLAD) encoding method is extended to a more generalized form weighted-VLAD (W-VLAD) in which diverse coefficient weights are taken into consideration. Experimental results validate the effectiveness of our presented method.
Biyun Sheng, Chunhua Shen, Guosheng Lin, Jun Li 0033, Wankou Yang, Changyin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Filtered shallow-deep feature channels for pedestrian detection
Biyun Sheng, Qichang Hu, Jun Li 0033, Wankou Yang, Baochang Zhang 0001, Changyin Sun 0001
Neurocomputing1
2016 Discriminative low-rank dictionary learning for face recognition
Hoangvu Nguyen, Wankou Yang, Biyun Sheng, Changyin Sun 0001
Neurocomputing3
2015 Action recognition using direction-dependent feature pairs and non-negative low rank sparse model
Biyun Sheng, Wankou Yang, Changyin Sun 0001
Neurocomputing1