Tianyue Zheng

dblp:205/2918 · DBLP profile ↗
← Back
53ranked-venue papers
15as first author
51since 2021 · last 2026
0000-0002-2826-6498ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 44 · 11 first-author · 43 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dimension-Independent Separate Channel Estimation for RIS-Aided Communications
Tianyue Zheng, Jieao Zhu, Linglong Dai
ICC2
2026 A Low-Complexity Unified Error Correction Transformer
Yongli Yan, Jieao Zhu, Tianyue Zheng, Linglong Dai
ICC3
2026 Stereo-Fi: Free-Form 3D Reconstruction via Generatively Co-Trained Inverse RF Rendering
abstract
This paper presents Stereo-Fi, the first system for high-fidelity 3D reconstruction of free-form objects using radio frequency (RF) signals from an unconstrained, handheld sensor. Conventional RF methods rely on simplified physics and constrained data acquisition, which degrades geometric fidelity and generality. Removing these constraints reframes reconstruction as an unconstrained, yet fundamentally ill-posed, optimization problem. To this end, Stereo-Fi establishes a new framework that solves this ill-posed problem by iteratively co-training a physics-based and a learned generative model. To this end, Stereo-Fi integrates three innovations. First, it introduces a novel hybrid representation, rfMold, that combines the benefits of explicit and implicit models for a robust optimization foundation. Second, it employs an inverse rendering optimizer, rfChisel, to jointly refine scene geometry and sensor trajectory while effectively navigating the non-convex loss landscape. Finally, it incorporates rfAlign, a generative diffusion model to restore geometric details while ensuring semantic consistency. Comprehensive evaluation demonstrates that Stereo-Fi achieves quasi-vision reconstruction accuracy across diverse targets and environments, significantly outperforming state-of-the-art RF-based 3D reconstruction methods.
Xueqiang Han, Tianyue Zheng, Jun Luo 0001
SenSys2
2026 EarPCG: Recovering Heart Sounds from in-Ear Audio via Physics-Informed Neural Network
abstract
While earables present a promising avenue for cardiac sensing, whether they may replace the stethoscope to perform heart sound (a.k.a. PCG) monitoring remains questionable. The latest effort attempts to generate PCG-like waveform out of in-ear audio collected via earphones, yet its data-driven approach does not seem to be grounded in the underlying physics. To this end, this paper introduces EarPCG, a system for continuous PCG monitoring leveraging physics-informed neural models. As opposed to the debatable belief that bone-conducted PCG appears within ear canal, EarPCG generates PCG waveforms from the (actually existing) photoplethysmography (PPG) waveforms conveyed via blood vessels. Arising from pressure variations induced by heartbeats, PPG can be mathematically described by a Partial Differential Equation (PDE). Therefore, solving this PDE inversely may reconstruct cardiac dynamics and in turn enable the generation of PCG waveforms with another PDE characterizing the pressure oscillations propagating through soft tissues. Pipelining the two PDE-solving neural models, EarPCG achieves accurate PCG monitoring from in-ear audio, while requiring minimal training. Our extensive experiments leveraging a custom-built prototype demonstrate the efficacy of our proposed system. Furthermore, we have conducted clinical trials, with clinicians reporting no perceptible difference between authentic PCG and the sounds reconstructed by EarPCG.
Junyi Zhou 0004, Henglin Pu, Peng Guo 0001, Tianyue Zheng, Chao Cai 0001, Jun Luo 0001
SenSys5
2026 Weaponizing Reflectivity for Pointcloud Deception with Forged Invisible Geometries
Hengwei Chen, Menglan Hu, Tianyue Zheng
SP3
2026 Unified Error Correction Code Transformer With Low Complexity
abstract
Channel coding is vital for reliable sixth-generation (6G) data transmission, employing diverse error correction codes for various application scenarios. Traditional decoders require dedicated hardware for each code, leading to high hardware costs. Recently, artificial intelligence (AI)-driven approaches, such as the error correction code Transformer (ECCT) and its enhanced version, the foundation error correction code Transformer (FECCT), have been proposed to reduce the hardware cost by leveraging the Transformer to decode multiple codes. However, their excessively high computational complexity ofO(N2) due to the self-attention mechanism in the Transformer limits scalability, whereNrepresents the sequence length. To reduce computational complexity, we propose a unified Transformer-based decoder that handles multiple linear block codes within a single framework. Specifically, a standardized unit is employed to align code length and code rate across different code types, while a redesigned low-rank unified attention module, with computational complexity ofO(N), is shared across various heads in the Transformer. Additionally, a sparse mask, derived from the parity-check matrix’s sparsity, is introduced to enhance the decoder’s ability to capture inherent constraints between information and parity-check bits, improving decoding accuracy and further reducing computational complexity by 86%. Extensive experimental results demonstrate that the proposed unified Transformer-based decoder outperforms existing methods and provides a high-performance, low-complexity solution for next-generation wireless communication systems.
Yongli Yan, Jieao Zhu, Tianyue Zheng, Linglong Dai
IEEE Internet Things J.3
2026 Large Language Model Enabled Multi-Task Physical Layer Network
abstract
The advance of Artificial Intelligence (AI) is continuously reshaping the future 6G wireless communications. Particularly, the development of Large Language Models (LLMs) offers a promising approach to effectively improve the performance and generalization of AI in different physical-layer (PHY) tasks. However, most existing works finetune dedicated LLM networks for a single wireless communication task separately. Thus, performing diverse PHY tasks requires extremely high training resources, memory usage, and deployment costs. To solve the problem, we propose a LLM-enabled multi-task PHY network to unify multiple tasks with a single LLM, by exploiting the excellent semantic understanding and generation capabilities of LLMs. Specifically, we first propose a multi-task LLM framework, which finetunes LLM to perform multiple tasks including multi-user precoding, signal detection, and channel prediction. Besides, the multi-task instruction module, input encoders, as well as output decoders, are elaborately designed to distinguish different tasks and adapt LLM for different tasks in the wireless domain. Moreover, low-rank adaptation (LoRA) is utilized for LLM fine-tuning. To reduce the memory requirement during LLM fine-tuning, a LoRA fine-tuning-aware quantization method is introduced. Extensive numerical simulations are also displayed to verify the effectiveness of the proposed method.
Tianyue Zheng, Linglong Dai
IEEE Trans. Commun.1
2026 E-M2: Efficient Multimodal Sensing via Adaptive Sensor-Computation Activation
abstract
Multimodal sensing systems have gained widespread adoption in IoT and edge intelligence, due to their ability to collect more comprehensive information of the target, thus improving sensing accuracy. However, this improved accuracy often comes at the cost of increased power consumption and computational overhead. Specifically, sensors require continuous power to collect data, and data processing algorithms not only demand intensive computational resources, but also consume significant energy to analyze the data, hindering the deployment of such systems on the edge. To address this issue, we propose E-M$^{2}$, a framework for efficient multimodal sensing by adaptive sensor-computation activation. First, E-M$^{2}$selectively disables redundant modalities and corresponding data processing modules, effectively reducing unnecessary power consumption and computational overhead. Second, E-M$^{2}$employs an exploration mechanism to reactivate disabled modalities, thus preventing “dead” modalities and enhancing overall system utilization. Finally, E-M$^{2}$conditions the data processing algorithms on the on/standby states of the modalities, thus alleviating the negative impacts of sensor deactivation. Extensive evaluations demonstrate that E-M$^{2}$reduces average power consumption by 40.75% and computational overhead by 48.84% across various sensing tasks, all while maintaining the sensing performance.
Jinyi Cui, Tianyue Zheng
IEEE Trans. Mob. Comput.2
2026 Energy-Aware Service Mesh Deployment and Online Request Routing in Edge: A Hierarchical Deep Reinforcement Learning Approach
abstract
Service meshes built upon ubiquitous microservice architectures, as an emerging paradigm, promise to enhance the flexibility, scalability, and portability of energy-consuming and latency-sensitive applications in edge with limited resources. However, due to intricate microservice dependencies, service multiplexing, and parallel distributed instances, microservice deployment and request routing are highly interdependent. To reduce response latency and energy consumption, such collaborative optimization for efficient service mesh orchestration is necessary, but significantly challenging. Besides, strict service level objective (SLO) requirements and f ine-grained latency analysis with multi-nest routing further impose great difficulties to online orchestration. When considering multi instance modeling and multi-hop data communications for numerous microservices, the difficulty is extremely amplified. Nevertheless, most prevailing work failed to design sophisticate models and methods for addressing the above difficulties, and ignored the inherent transmission energy consumption for highly-concurrent multi-hop data interactions. Therefore, this paper investigates the energy aware service mesh deployment and online request routing in edge. First, we establish a multi-instance queuing network model to accurately analyze the end-to-end response latency with complicated dependencies and multi-hop communications, and optimize energy consumption in a fine-grained manner. Then, to boost the overall performance, we design an efficient multi-dimensional hierarchical deep reinforcement learning algorithm, which enables edges and service instances to cooperate with each other to handle massively concurrent requests. Besides, we propose an energy-aware proactive autoscaling algorithm to carefully adapt to exceedingly dynamic scenarios. Finally, extensive experiments are performed to show our superior performance compared to other baselines.
Junhui Hu, Menglan Hu, Kai Peng 0001, Tianyue Zheng, Chao Cai 0001, Zehui Xiong
IEEE Trans. Mob. Comput.6
2026 Rising From Pieces: Effective Inference at the Edge via Robust Split ML
abstract
The increasing processing demands of today's mobile deep learning applications impose stringent requirements on edge devices. Offloading these tasks to the cloud, while being a potential solution, often results in significant data transfer overhead, as well as privacy and connectivity concerns. To address these challenges, split machine learning (split ML) has emerged as an innovative paradigm, enabling task distribution among edge devices themselves. However, split ML systems inherently exhibit instability due to the hardware and communication limitations of mobile devices, which frequently result in failures and malfunctions of client nodes. In light of these challenges, we present Axolotl, a fault-tolerant edge split ML inference system for addressing node failure with minimal performance impact. Specifically, we first design a novel curriculum dropout mechanism to enhance the model's resilience by gradually exposing it to potential server node failures. We then design inverse-proximal weight consolidation to mitigate catastrophic forgetting caused by curriculum dropout. To further tackle potential node failures, we innovate in a resource-aware substitution module that offload the functions of a failed node to neighboring ones, ensuring efficient information flow. Extensive experiments demonstrate the effectiveness and robustness of Axolotl in various deep learning networks and tasks in edge environments.
Yuxuan Weng, Tianyue Zheng, Zhe Chen 0015, Menglan Hu, Jun Luo 0001
IEEE Trans. Mob. Comput.2
2026 FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition
abstract
Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution when low-light, obstructions, or privacy concerns render computer vision impractical. However, thescarcityof labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough offoundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi 2.0, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based, multi-person HAR systems. FM-Fi 2.0 first employs the intrinsic capabilities of FM and RF modality to associate both intra- and cross-modal features of each subject, while simultaneously filtering out irrelevant features to achieve better alignment between the two modalities. FM-Fi 2.0 also employs a cross-modalcontrastiveknowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi 2.0 rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi 2.0's generalizability across various environments.
Yuxuan Weng, Tianyue Zheng, Yanbing Yang 0001, Jun Luo 0001
IEEE Trans. Mob. Comput.2
2026 A Joint Game-Theoretic Approach for Multicast Routing and Load Balancing in LEO Satellite Networks
abstract
Low Earth Orbit (LEO) satellite networks, with their low latency, high bandwidth, and global coverage, are becoming key technologies for applications like real-time video transmission. As satellite networks expand, effectively managing multicast traffic and optimizing bandwidth utilization have become major challenges for efficient video distribution. Although Software-Defined Multicast (SDM) technology has made progress in bandwidth optimization, existing SDM methods are still focused on constructing Steiner trees, making it difficult to address the dynamic changes and high-load issues in LEO satellite networks. This paper frames the multicast tree construction problem as a Joint Path Optimization Game (JPOG). We propose a Cooperative Game-Theoretic Routing (CGMR) Algorithm based on game theory, which optimizes multicast path selection and achieves load balancing by introducing a link cost-sharing mechanism. Additionally, we propose a two-stage A* path generation algorithm to improve path search efficiency. Theoretically, this paper proves that JPOG is a potential game and can converge to a pure strategy Nash equilibrium (PSNE) within a finite number of iterations. The results showed that JPOG outperformed other algorithms, achieving lower link load, path cost, and superior load balancing, demonstrating its effectiveness in optimizing multicast routing and resource management in large-scale LEO satellite networks.
Yan Dong 0001, Menglan Hu, Chao Cai 0001, Tianyue Zheng, Kai Peng 0001
IEEE Trans. Netw. Serv. Manag.6
2026 Dimension-Independent Channel Estimation for RIS-Assisted Communications: From Cascaded to Separate
abstract
Channel estimation in reconfigurable intelligent surface (RIS) assisted communications requires high pilot overhead due to numerous RIS elements incapable of signal processing. Recently, research on sensing RIS has provided a dimension-independent channel estimation scheme with merely three pilots. Nevertheless, it assumes that the BS-RIS channel is perfectly known to the RIS and remains invariant over a prolonged period, while inducing high hardware and power consumption. To address these issues, this paper introduces a generalized approximate message passing (GAMP) based channel estimation framework to achieve dimension-independent estimation of separate channels, without the assumption of known BS-RIS channel. Specifically, we first formulate the channel estimation problem in RIS assisted communications as a compressive phase retrieval problem. Based on the phaseless power observations, we leverage the GAMP algorithm to retrieve original sparse signals, which inherently supports the sparse-sampling sensing RIS architecture. Furthermore, by exploiting the intrinsic mapping between the sparse representations of the channels and the power observations in the angular domain, we propose a learned GAMP network to enhance the convergence stability and estimation accuracy. Finally, simulation results demonstrate that the proposed approach can efficiently estimate both the BS-RIS and UE-RIS channels with four pilots, while eliminating the requirement for prior knowledge of the BS-RIS channel and significantly reducing hardware and power consumption.
Tianyue Zheng, Jieao Zhu, Shenheng Xu, Linglong Dai
IEEE Trans. Wirel. Commun.2
2026 MUSE-FM: Multi-Task Environment-Aware Foundation Model for Wireless Communications
abstract
Recent advancements in foundation models (FMs) have attracted increasing attention in the wireless communication domain. Leveraging the powerful multi-task learning capability, FMs hold the promise of unifying multiple tasks of wireless communication with a single framework. Nevertheless, existing wireless FMs face limitations in the uniformity to address multiple tasks with diverse inputs/outputs across different communication scenarios. In this paper, we propose a MUlti-taSk Environment-aware FM (MUSE-FM) with a unified architecture to handle multiple tasks in wireless communications, while effectively incorporating scenario information. Specifically, to achieve task uniformity, we propose a unified prompt-guided data encoder-decoder pair to handle data with heterogeneous formats and distributions across different tasks. Besides, we integrate the environmental context as a multi-modal input, which serves as prior knowledge of environment and channel distributions and facilitates cross-scenario feature extraction. Simulation results illustrate that the proposed MUSE-FM outperforms existing methods for various tasks, and its prompt-guided encoder-decoder pair facilitates few-shot adaptation to new task configurations. Moreover, the incorporation of environment information improves the ability to adapt to different scenarios.
Tianyue Zheng, Jiajia Guo 0001, Linglong Dai, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Wirel. Commun.1
2025 LLM4NF: LLM-Empowered Near-Field Communications in Low-Altitude Economy
abstract
The low-altitude economy (LAE) has recently received widespread attention from both academia and industry. To facilitate and support the successful implementation of the LAE, we fortunately find that the LAE and near-field communications in extremely large-scale MIMO (XL-MIMO) are a natural combination. Specifically, the LAE can utilize the near-field beamfocusing characteristic to accurately focus the beam energy to the positions of different UAVs, and utilize the new distance dimension to further enhance the entire spectrum efficiency. However, most existing works on near-field communications only consider the ideal scenario in a 2D horizontal plane and how to efficiently achieve near-field communications for LAE is still a blank in the literature and faces several challenges. To fill in this blank, inspired by the powerful large language models (LLM) which can act as a general wireless communications optimization solver, in this paper, we first apply LLM to solve the spectrum efficiency maximization problem of near-field communications for LAE. Specifically, our proposed LLM-based scheme can accurately distinguish far-field and near-field users and achieve joint optimization of precoding and power allocation through elaborately designing adapters and finetuning the pretrained GPT2. Simulation results substantiate the efficacy and excellence of our proposed scheme compared to the existing benchmark schemes.
Tianyue Zheng, Linglong Dai
GLOBECOM2
2025 Unified Physical Layer Network for Multiple Tasks based on Large Language Model
Tianyue Zheng, Linglong Dai
GLOBECOM1
2025 Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment
Feng-Qi Cui, Yu-Tong Guo, Tianyue Zheng, Jinyang Huang
ICPADS3
2025 Poster: HeteroRF: Heterogeneity-Aware Federated Radio-Frequency Human Activity Recognition
Yanru Cui, Tianyue Zheng
MobiSys2
2025 Poster: GlueRT: Bridging the Gap Between Ray-tracing and Real-World Wireless Channel
abstract
Despite ray-tracing's theoretical foundation in radio frequency propagation modeling, significant discrepancies persist between simulated and measured channel characteristics in real environments. This paper presents GlueRT, a hybrid framework leveraging ray-tracing's physical interpretability while using a targeted neural network to correct simulation-reality gaps. GlueRT applies neural network-generated corrections to both amplitude and phase components of ray-tracing channel estimates, maintaining explainability while requiring less training data than pure learning methods. Experiments across various environments demonstrate GlueRT achieves higher channel estimation accuracy than traditional ray-tracing or neural network approaches.
Tianyue Zheng
MobiSys2
2025 Global Microservice Autoscaling Over Heterogeneous Edge Environments for Internet Applications: A Reinforcement Learning Approach
abstract
The integration of microservice architecture and edge computing offers innovative solutions for highly interactive, low-latency Internet applications. To manage the dynamic nature of requests in edge computing, microservice autoscaling techniques are frequently employed. However, the resource limitation of individual edge servers and the heterogeneity among edge servers present significant challenges for autoscaling in edge computing. Meanwhile, few studies have considered the long-term optimization and the joint optimization of instance adjustment and request routing in edge computing. This paper aims to fill these gaps. First, we propose Global Horizontal Pod Autoscaler (GHPA), a novel framework that addresses microservice autoscaling from the perspective of edge server clusters. Second, we consider the joint optimization of instance adjustment and request routing, and formulate a long-term optimization problem. Third, we transform the long-term optimization problem into a Markov Decision Problem (MDP) and use reinforcement learning techniques to solve it. Finally, we conduct extensive experiments using both real and synthetic data. The experiment results demonstrate that our algorithm achieves at least a 10% performance improvement in various test environments compared to state-of-the-art algorithms.
Kai Peng 0001, Jie Rao, Tianyue Zheng, Menglan Hu
IEEE Internet Things J.6
2025 Coded Beam Training
abstract
In extremely large-scale multiple-input-multiple-output (XL-MIMO) systems for future sixth-generation (6G) communications, codebook-based beam training stands out as a promising technology to acquire channel state information (CSI). Despite their effectiveness, existing beam training methods suffer from significant achievable rate degradation for remote users with low signal-to-noise ratio (SNR). To tackle this challenge, leveraging the error-correcting capability of channel codes, we incorporate channel coding theory into beam training to enhance the training accuracy, thereby extending the coverage area. Specifically, we establish the duality between hierarchical beam training and channel coding, and build on it to propose a general coded beam training framework. Then, we present two specific implementations exemplified by coded beam training methods based on Hamming codes and convolutional codes, during which the beam encoding and decoding processes are refined respectively to better accommodate to the beam training problem. Simulation results have demonstrated that, the proposed coded beam training method can enable reliable beam training performance for remote users with low SNR, while keeping training overhead low.
Tianyue Zheng, Jieao Zhu, Qiumo Yu, Yongli Yan, Linglong Dai
IEEE J. Sel. Areas Commun.1
2025 LLM-Empowered Near-Field Communications for Low-Altitude Economy
abstract
The low-altitude economy (LAE) has recently received widespread attention from both academia and industry. To facilitate and support the successful implementation of the LAE, we fortunately find that the LAE and near-field communications in extremely large-scale MIMO (XL-MIMO) systems are a natural combination. Specifically, the LAE can utilize the near-field beamfocusing characteristic to accurately focus the beam energy to the positions of different unmanned aerial vehicles, and utilize the new distance dimension to further enhance the entire spectrum efficiency. However, most existing works on near-field communications only consider the ideal scenario in a horizontal plane and how to efficiently achieve near-field communications for LAE is still a blank in the literature and faces several challenges. To fill in this blank, inspired by the powerful large language models (LLM) which can act as a general wireless communications optimization solver, in this paper, we first apply LLM to solve the spectrum efficiency maximization problem of near-field communications for LAE. Specifically, our proposed LLM-based scheme can accurately distinguish far-field and near-field users and achieve joint optimization of precoding and power allocation through elaborately designing adapters and finetuning the pretrained GPT-2. Simulation results substantiate the efficacy and excellence of our proposed scheme compared to the existing benchmark schemes.
Tianyue Zheng, Linglong Dai
IEEE Trans. Commun.2
2025 Loki: Physical-World Adversarial Attacks on Wireless Indoor Localization via Differentiable Object Placement
abstract
As a cornerstone for numerous sensing applications, wireless indoor localization has been a pivotal area of research over the last two decades. While techniques such as jamming, spoofing, and adversarial perturbation have been exploited to compromise wireless indoor localization, existing attacks face challenges in accessibility to wireless systems and stealthiness. To address these limitations, we introduceLoki, a novel physical-world attack on wireless indoor localization via differentiable object placement. Specifically, we develop a differentiable wireless ray-tracing technique that allows us to optimize object placement in the scene. By repositioning an existing object in the scene by just a few centimeters,Lokifools existing wireless indoor localization systems into generating erroneous localization results. We also show via experiments that the object placement generated byLokialigns with wireless sensing theory (e.g., the forward scattering region and Fresnel zone), confirming its explainability. Additionally,Lokiproves effective across various localization models and scenarios, highlighting its generalizability.
Xueqiang Han, Jinyang Huang, Meng Li 0006, Chao Cai 0001, Tianyue Zheng
IEEE Trans. Inf. Forensics Secur.5
2025 Echoes of Fingertip: Unveiling POS Terminal Passwords Through Wi-Fi Beamforming Feedback
abstract
Recent years, point-of-sale (POS) terminals are no longer limited to wired connections, with many relying on Wi-Fi for data transmission. Although Wi-Fi offers the convenience of wireless connectivity, it introduces significant security vulnerabilities. This work presents a non-intrusive method for eavesdropping POS passwords via Wi-Fi sensing, named${\mathsf {BeamThief}}$. Instead of conventional Wi-Fi Channel State Information (CSI) readings, our approach employs Wi-Fi Beamforming Feedback Information (BFI) for an eavesdropping attack. Compared to CSI, which can only be extracted through intruding into the Access Point (AP) or from a limited selection of commercial Wi-Fi cards (e.g., Intel-5300), BFI readings can be more readily obtained from a broad array of commercial Wi-Fi devices. A key technological contribution of${\mathsf {BeamThief}}$is the development of an analysis model for predicting finger motion trajectories. This model is based on the physical relationship between BFI readings and finger motion, thus eliminating the need for extensive labeled training data. Furthermore, we employ Maximum Ratio Combining (MRC) to enhance the BFI series, ensuring performance across various scenarios. We implement${\mathsf {BeamThief}}$using everyday commercial Wi-Fi devices and conduct a series of experiments to assess the impact of this attack. Experimental results demonstrate that${\mathsf {BeamThief}}$achieves an accuracy rate 79$\%$in inferring 6-digit POS passwords within the top-100 attempts.
Siyu Chen 0017, Hongbo Jiang 0001, Jingyang Hu, Tianyue Zheng, Zhu Xiao, Daibo Liu, Jun Luo 0001
IEEE Trans. Mob. Comput.4
2025 t-READi: Transformer-Powered Robust and Efficient Multimodal Inference for Autonomous Driving
abstract
Given the wide adoption of multimodal sensors (e.g., camera, lidar, radar) byautonomous vehicles (AVs), deep analytics to fuse their outputs for a robust perception become imperative. However, existing fusion methods often make two assumptions rarely holding in practice: i) similar data distributions for all inputs and ii) constant availability for all sensors. Because, for example, lidars have various resolutions and failures of radars may occur, such variability often results in significant performance degradation in fusion. To this end, we present t-READi, an adaptive inference system that accommodates the variability of multimodal sensory data and thus enables robust and efficient perception. t-READi identifies variation-sensitive yetstructure-specificmodel parameters; it then adapts only these parameters while keeping the rest intact. t-READi also leverages a cross-modality contrastive learning method to compensate for the loss from missing modalities. Both functions are implemented to maintain compatibility with existing multimodal deep fusion methods. The extensive experiments evidently demonstrate that compared with the status quo approaches, t-READi not only improves the average inference accuracy by more than 6% but also reduces the inference latency by almost 15× with the cost of only 5% extra memory overhead in the worst case under realistic data and modal variations.
Pengfei Hu 0001, Yuhang Qian, Tianyue Zheng, Ang Li 0005, Zhe Chen 0015, Yue Gao 0001, Xiuzhen Cheng, Jun Luo 0001
IEEE Trans. Mob. Comput.3
2025 RF-Eye: Commodity RFID Can Know What You Write and Who You Are Wherever You Are
abstract
Handwriting recognition systems have greatly enhanced AIoT applications, especially in human-computer interaction. Wireless-based methods, favored for their non-invasive nature and ease of deployment, are becoming more common. However, existing works, which typically depend on the user’s position, often perform poorly in varied writing positions. Additionally, they do not incorporate user identity information, which could lead to security vulnerabilities by failing to reject unauthorized users. To address these issues, this article introduces RF-Eye , a system that enables contactless, position-independent handwriting recognition and user identification without prior training. Its innovative approach uses each Radio-frequency identification (RFID) tag as a unique viewpoint for observing hand movements and employs pairs of tags to track directional changes. Specifically, building upon the signal transmission model and the Fresnel Zone, we propose a novel feature, DCG , to capture changes in gesture direction and confirm its consistency across different positions. Based on DCG , we develop unique patterns for common handwriting symbols that enhance our recognition algorithm. Moreover, to strengthen the system security, we link these patterns with distinct handwriting styles through the extraction of finer-grained features, thus, preventing the misuse of the system by unauthorized users. Extensive experiments demonstrate RF-Eye ’s efficacy, which achieves recognition accuracies of 93.5%, 95.2%, and 95.8% for 26 lowercase letters, 10 digits, and 10 graphic symbols, respectively, and identifying unauthorized users with 98.6% accuracy.
Yuanhao Feng, Jinyang Huang, Xiang Zhang 0011, Meng Li 0006, Fusang Zhang, Tianyue Zheng, Anran Li 0001, Mianxiong Dong, Zhi Liu 0002
ACM Trans. Sens. Networks7
2025 Near-Field Wideband Beam Training Based on Distance-Dependent Beam Split
abstract
Near-field beam training is essential for acquiring channel state information in 6G extremely large-scale multiple input multiple output (XL-MIMO) systems. To achieve low-overhead beam training, existing method has been proposed to leverage the near-field beam split effect, which deploys true-time-delay arrays to simultaneously search multiple angles of the entire angular range in a distance ring with a single pilot. However, the method still requires exhaustive search in the distance domain, which limits its efficiency. To address the problem, we propose a distance-dependent beam-split-based beam training method to further reduce the training overheads. Specifically, we first reveal the new phenomenon of distance-dependent beam split, where by manipulating the configurations of time-delay and phase-shift, beams at different frequencies can simultaneously scan the angular domain in multiple distance rings. Leveraging the phenomenon, we propose a near-field beam training method where both different angles and distances can simultaneously be searched in one time slot. Thus, a few pilots are capable of covering the whole angle-distance space for wideband XL-MIMO. Theoretical analysis and numerical simulations are also displayed to verify the superiority of the proposed method on beamforming gain and training overhead.
Tianyue Zheng, Mingyao Cui, Zidong Wu, Linglong Dai
IEEE Trans. Wirel. Commun.1
2024 M2-Fi: Multi-person Respiration Monitoring via Handheld WiFi Devices
abstract
Wi-Fi signals are commonly used for conventional communication, yet they can also realize low-cost and non-invasive human sensing. However, Wi-Fi sensing in Multi-person scenarios is still a challenging problem. In this paper, we propose M2-Fi to achieve multi-person respiration monitoring using a handheld device. M2-Fi leverages Wi-Fi BFI (beamforming feedback information) performs respiration monitoring. As a compressed version of the uplink CSI (channel state information), BFI transmission is unencrypted, easily obtained using frame capture, and does not require specific firmware to obtain. M2-Fi is based on an interesting experiment phenomenon that when a Wi-Fi device is very close to a subject, near-field channel changes caused by the subject significantly cancel out changes from other subjects. We employed VMD (Variational Mode Decomposition) to eliminate the interference caused by hand movement in the BFI time series. Subsequently, we devised a deep learning architecture based on GAN (Generative Adversarial Networks) to recover fine-grained respiration waveforms from the respiration patterns extracted from the BFI time series. Our experiments on collected 50-hour data from 8 subjects show that M2-Fi can accurately recover the respiration waveforms of multiple persons with handheld devices.
Jingyang Hu, Hongbo Jiang 0001, Tianyue Zheng, Jingzhi Hu, Hangcheng Cao, Zhe Chen 0015, Jun Luo 0001
INFOCOM3
2024 Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition
abstract
Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution for applications unamenable to techniques requiring computer visions. However, the scarcity of labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough of foundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based HAR systems. FM-Fi involves a novel cross-modal contrastive knowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. It also employs the intrinsic capabilities of FM and RF to remove extraneous features for better alignment between the two modalities. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi's generalizability across various environments.
Yuxuan Weng, Guoquan Wu, Tianyue Zheng, Yanbing Yang 0001, Jun Luo 0001
SenSys3
2024 pi-Jack: Physical-World Adversarial Attack on Monocular Depth Estimation with Perspective Hijacking
Tianyue Zheng, Jingzhi Hu, Yinqian Zhang, Ying He 0001, Jun Luo 0001
USENIX Security Symposium1
2024 Distance-Dependent Beam Split Aided Low-Overhead Near-field Wideband Beam Training
abstract
In extremely large-scale multiple input multiple output (XL-MIMO) systems for future 6G communications, codebook-based beam training stands out as a promising technology to acquire channel state information. Due to sharp increase in the number of antennas, the transition of electromagnetic propagation from the far-field to the near-field introduces extremely high overhead when employing exhaustive search in both angle and distance domain. To reduce the overwhelming overheads, existing method has been proposed to utilize the dispersed directional beams produced by controllable beam split to simultaneously search multiple directions in a distance ring in wideband communications. However, the method still requires exhaustive search in the distance domain and thus cannot achieve low-overhead beam training for XL-MIMO. To address the problem, achieving a transition from distance-independent beam split to distance-dependent beam split, we propose a near-field beam training method where both different angles and distances can be searched in a pilot simultaneously, and thus a few pilots are capable of covering the whole two-dimensional space for wideband XL-MIMO. Numerical simulations are displayed to verify the performance of the proposed low-overhead method.
Tianyue Zheng, Linglong Dai
VTC Fall1
2024 Adv-4-Adv: Thwarting changing adversarial perturbations via adversarial domain adaptation
Tianyue Zheng, Zhe Chen 0015, Shuya Ding, Chao Cai 0001, Jun Luo 0001
Neurocomputing1
2024 HandKey: Knocking-Triggered Robust Vibration Signature for Keyless Unlocking
abstract
Door lock is regarded as a critical line of defending the privacy and security of personal areas. However, for inner doors in environments like factories, existing locking mechanisms can be poor in user-friendliness and high in cost. For instance, mechanical locks require carrying keys that inevitably compromise user experiences, while smart locks always require non-trivial sensors. Therefore, inner doors urgently require a lightweight unlocking scheme that can properly balance user-friendliness, cost, and security. To this end, we propose HandKey as a keyless unlocking scheme to supplement existing lock systems. HandKey relies on two principles: the simplicity of hand knocking doors and the uniqueness of vibration triggered by the knocking force. In other words, a door and a hand knocking it jointly form a unique physical system that generates hand-dependent and user-specific vibration signatures uniquely representing a user identity. In designing HandKey, we first analyze the vibration mechanism behind it and the impacts of gestures and door materials on vibration signatures. Then we innovatively construct a signal processing and deep learning-based pipeline to extract signatures robust to variable knocking behaviors for representing user identity. Finally, we implement a HandKey prototype and use extensive evaluation to demonstrate its security and effectiveness.
Hangcheng Cao, Daibo Liu, Hongbo Jiang 0001, Chao Cai 0001, Tianyue Zheng, John C. S. Lui, Jun Luo 0001
IEEE Trans. Mob. Comput.5
2024 MuKI-Fi: Multi-Person Keystroke Inference With BFI-Enabled Wi-Fi Sensing
abstract
The contact-free sensing nature of Wi-Fi has been leveraged to achieve privacy breaches such askeystroke inference(KI). However, the use ofchannel state information(CSI) in existing attacks is highly questionable due to its signal instability and hardness to acquire. Moreover, such Wi-Fi-based attacks are confined to only one victim because Wi-Fi sensing offers insufficient range resolution to physically differentiate multiple victims. To this end, we propose MuKI-Fi to enable, for the first time,multi-personKI, leveragingbeamforming feedback information(BFI), a new feature offered by latest Wi-Fi hardware, transmitted in clear-text by smartphones. BFI's characteristics, clear-text communication and signal stability, make it readily acquirable and usable by any other Wi-Fi devices switching to monitor mode without the need forlow-levelhacking on hardware. Moreover, to improve upon existing KI methods offering very limited generalizability across diversified application scenarios, MuKI-Fi innovates in an adversarial learning scheme to enable its inference generalizable towards unseen scenarios. Finally, we discover that, as a smartphone is in close proximity to a victim, the variations of BFI caused by that victim's keystrokes in suchnear-fieldsubstantially outweigh those caused by other distant victims; this phenomenon naturally allows for multi-person KI. Our extensive evaluations clearly demonstrate that MuKI-Fi can effectively eavesdrop on the keystrokes of multiple subjects, achieving 87.1% accuracy for individual keystrokes and up to 81% top-100 accuracy for stealing passwords from mobile applications(e.g., WeChat) on average.
Jingyang Hu, Tianyue Zheng, Jingzhi Hu, Zhe Chen 0015, Hongbo Jiang 0001, Yuanjin Zheng, Jun Luo 0001
IEEE Trans. Mob. Comput.3
2023 Password-Stealing without Hacking: Wi-Fi Enabled Practical Keystroke Eavesdropping
abstract
The contact-free sensing nature of Wi-Fi has been leveraged to achieve privacy breaches, yet existing attacks relying on Wi-Fi CSI (channel state information) demand hacking Wi-Fi hardware to obtain desired CSIs. Since such hacking has proven prohibitively hard due to compact hardware, its feasibility in keeping up with fast-developing Wi-Fi technology becomes very questionable. To this end, we propose WiKI-Eve to eavesdrop keystrokes on smartphones without the need for hacking. WiKI-Eve exploits a new feature, BFI (beamforming feedback information), offered by latest Wi-Fi hardware: since BFI is transmitted from a smartphone to an AP in clear-text, it can be overheard (hence eavesdropped) by any other Wi-Fi devices switching to monitor mode. As existing keystroke inference methods offer very limited generalizability, WiKI-Eve further innovates in an adversarial learning scheme to enable its inference generalizable towards unseen scenarios. We implement WiKI-Eve and conduct extensive evaluation on it; the results demonstrate that WiKI-Eve achieves 88.9% inference accuracy for individual keystrokes and up to 65.8% top-10 accuracy for stealing passwords of mobile applications (e.g., WeChat).
Jingyang Hu, Tianyue Zheng, Jingzhi Hu, Zhe Chen 0015, Hongbo Jiang 0001, Jun Luo 0001
CCS3
2023 OCHID-Fi: Occlusion-Robust Hand Pose Estimation in 3D via RF-Vision
abstract
Hand Pose Estimation (HPE) is crucial to many applications, but conventional cameras-based CM-HPE methods are completely subject to Line-of-Sight (LoS), as cameras cannot capture occluded objects. In this paper, we propose to exploit Radio-Frequency-Vision (RF-vision) capable of bypassing obstacles for achieving occluded HPE, and we introduce OCHID-Fi as the first RF-HPE method with 3D pose estimation capability. OCHID-Fi employs wideband RF sensors widely available on smart devices (e.g., iPhones) to probe 3D human hand pose and extract their skeletons behind obstacles. To overcome the challenge in labeling RF imaging given its human incomprehensible nature, OCHID-Fi employs a cross-modality and cross-domain training process. It uses a pre-trained CM-HPE network and a synchronized CM/RF dataset, to guide the training of its complex-valued RF-HPE network under LoS conditions. It further transfers knowledge learned from labeled LoS domain to unlabeled occluded domain via adversarial learning, enabling OCHID-Fi to generalize to unseen occluded scenarios. Experimental results demonstrate the superiority of OCHID-Fi: it achieves comparable accuracy to CM-HPE under normal conditions while maintaining such accuracy even in occluded scenarios, with empirical evidence for its generalizability to new domains.
Tianyue Zheng, Zhe Chen 0015, Jingzhi Hu, Abdelwahed Khamis, Jiajun Liu 0004, Jun Luo 0001
ICCV2
2023 Wider is Better? Contact-free Vibration Sensing via Different COTS-RF Technologies
Zhe Chen 0015, Tianyue Zheng, Chao Cai 0001, Yue Gao 0001, Pengfei Hu 0001, Jun Luo 0001
INFOCOM2
2023 MUSE-Fi: Contactless MUti-person SEnsing Exploiting Near-field Wi-Fi Channel Variation
abstract
Having been studied for more than a decade, Wi-Fi human sensing still faces a major challenge in the presence of multiple persons, simply because the limited bandwidth of Wi-Fi fails to provide a suficient range resolution to physically separate multiple subjects. Existing solutions mostly avoid this challenge by switching to radars with GHz bandwidth, at the cost of cumbersome deployments. Therefore, could Wi-Fi human sensing handle multiple subjects remains an open question. This paper presents MUSE-Fi, the first Wi-Fi multi-person sensing system with physical separability. The principle behind MUSE-Fi is that, given a Wi-Fi device (e.g., smartphone) very close to a subject, the near-field channel variation caused by the subject significantly overwhelms variations caused by other distant subjects. Consequently, focusing on the channel state information (CSI) carried by the trafic in and out of this device naturally allows for physically separating multiple subjects. Based on this principle, we propose three sensing strategies for MUSE-Fi: i) uplink CSI, ii) downlink CSI, and iii) downlink beamforming feedback, where we specifically tackle signal recovery from sparse (per-user) trafic under realistic multi-user communication scenarios. Our extensive evaluations clearly demonstrate that MUSE-Fi is able to successfully handle multi-person sensing with respect to three typical applications: respiration monitoring, gesture detection, and activity recognition.
Jingzhi Hu, Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001
MobiCom2
2023 AutoFed: Heterogeneity-Aware Federated Multimodal Learning for Robust Autonomous Driving
abstract
Object detection with on-board sensors (e.g., lidar, radar, and camera) is crucial to autonomous driving (AD), and these sensors complement each other in modalities. While crowdsensing may potentially exploit these sensors (of huge quantity) to derive more comprehensive knowledge, federated learning (FL) appears to be the necessary tool to reach this potential: it enables autonomous vehicles (AVs) to train machine learning models without explicitly sharing raw sensory data. However, the multimodal sensors introduce various data heterogeneity across distributed AVs (e.g., label quantity skews and varied modalities), posing critical challenges to effective FL. To this end, we present AutoFed as a heterogeneity-aware FL framework to fully exploit multimodal sensory data on AVs and thus enable robust AD. Specifically, we first propose a novel model leveraging pseudo labeling to avoid mistakenly treating unlabeled objects as the background. We also propose an autoencoder-based data imputation method to fill missing data modality (of certain AVs) with the available ones. To further reconcile the heterogeneity, we finally present a client selection mechanism based on client model similarities to improve training stability and convergence rate. Our experiments confirm that AutoFed substantially improves over status quo in both precision and recall, while demonstrating strong robustness to adverse weather.
Tianyue Zheng, Ang Li 0005, Zhe Chen 0015, Jun Luo 0001
MobiCom1
2023 Algorithmic Sensing: A Joint Sensing and Learning Perspective
abstract
Sensing hardware technologies and algorithms to interpret sensing data are two main pillars of high-level perception. Existing works either handle these two pillars independently, hence resulting in a loss of efficiency, or plainly adopt deep learning as a one-size-fits-all solution, attempting to directly fit sensor data to desired outputs (e.g., decisions or predictions). Consequently, there are urgent needs for improving sensing efficiency while improving its practicality and robustness in the face of diversified application scenarios that may severely interfere with the sensing data. To this end, we propose algorithmic sensing to integrate these two pillars, and to avoid blindly applying deep-learning algorithms to sensing. In particular, our joint sensing and learning scheme involves designing adaptable sensing platforms for various learning algorithms, as well as adapting algorithms to fit existing sensing infrastructure. We illustrate the two sides of algorithmic sensing by using past work as examples, thus providing a comprehensive framework for designing future algorithmic sensing systems and suggesting open research directions.
Tianyue Zheng
MobiSys1
2023 HoloFed: Environment-Adaptive Positioning via Multi-Band Reconfigurable Holographic Surfaces and Federated Learning
abstract
Positioning is an essential service for various applications and is expected to be integrated with existing communication infrastructures in 5G and 6G. Though current Wi-Fi and cellular base stations (BSs) can be used to support this integration, the resulting precision is unsatisfactory due to the lack of precise control of the wireless signals. Recently, BSs adopting reconfigurable holographic surfaces (RHSs) have been advocated for positioning as RHSs’ large number of antenna elements enable generation of arbitrary and highly-focused signal beam patterns. However, existing designs face two major challenges: i) RHSs only have limited operating bandwidth, and ii) the positioning methods cannot adapt to the diverse environments encountered in practice. To overcome these challenges, we present HoloFed, a system providing high-precision environment-adaptive user positioning services by exploitingmulti-band(MB)-RHS andfederated learning(FL). For improving the positioning performance, a lower bound on the error variance is obtained and utilized for guiding MB-RHS’s digital and analog beamforming design. For better adaptability while preserving privacy, an FL framework is proposed for users to collaboratively train a position estimator, where we exploit the transfer learning technique to handle the lack of position labels of the users. Moreover, a scheduling algorithm for the BS to select which users train the position estimator is designed, jointly considering the convergence and efficiency of FL. Our performance evaluation based on simulations confirms that HoloFed achieves a 57% lower positioning error variance compared to a beam-scanning baseline and can effectively adapt to diverse environments.
Jingzhi Hu, Zhe Chen 0015, Tianyue Zheng, Robert Schober, Jun Luo 0001
IEEE J. Sel. Areas Commun.3
2023 RF-Based Human Activity Recognition Using Signal Adapted Convolutional Neural Network
abstract
Human activity recognition (HAR) plays a critical role in a wide range of real-world applications, and it is traditionally achieved via wearable sensing. Recently, to avoid the burden and discomfort caused by wearable devices, device-free approaches exploiting radio-frequency (RF) signals arise as a promising alternative for HAR. Most of the latest device-free approaches require training a large deep neural network model in either time or frequency domain, entailing extensive storage to contain the model and intensive computations to infer human activities. Consequently, even with some major advances on device-free HAR, current device-free approaches are still far from practical in real-world scenarios where the computation and storage resources possessed by, for example, edge devices, are limited. To overcome these weaknesses, we introduce HAR-SAnet which is a novel RF-based HAR framework. It adopts an original signal adapted convolutional neural network architecture: instead of feeding the handcraft features of RF signals into a classifier, HAR-SAnet fuses them adaptively from both time and frequency domains to design an end-to-end neural network model. We apply point-wise grouped convolution and depth-wise separable convolutions to confine the model scale and to speed up the inference execution time. The experiment results show that the recognition accuracy of HAR-SAnet substantially outperforms the state-of-the-art algorithms and systems.
Zhe Chen 0015, Chao Cai 0001, Tianyue Zheng, Jun Luo 0001, Jie Xiong 0001, Xin Wang 0002
IEEE Trans. Mob. Comput.3
2023 Catch Your Breath: Simultaneous RF Tracking and Respiration Monitoring With Radar Pairs
abstract
Continuous respiration monitoring is significant for real-life healthcare applications, but realizing it is extremely hard as wearable sensors are cumbersome and contact-free sensors largely fail to tolerate user movements. Meanwhile, tracking users indoors mostly demands user-held devices, while device-free localization can barely tell what and who it tracks. Fortunately, as both contact-free respiration monitoring and device-free localization may rely on Radio-Frequency (RF) sensing, fusing them together creates a novel system capable of continuously tracking users while recovering their fine-grained respiratory waveforms. To this end, we propose BreathCatcher as a continuous human respiration tracking system for indoor applications. To build this system, we employ commercial-grade compact radar pairs to capture RF reflections containing respiratory signals. We then propose a hybrid human respiration and position tracking algorithm to locate and identify respiratory signals from complex RF reflection mixtures. Finally, we design an encoder-decoder deep neural network driven by variational inference to recover fine-grained respiratory waveforms. Essentially, BreathCatcher cannot only obtain respiratory waveforms from multiple walking users, but also identify each user according to the latent properties of the respiratory signals. We evidently demonstrate the accuracy of both tracking and respiration monitoring via experiments involving 12 subjects and 80 man-hour data.
Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001
IEEE Trans. Mob. Comput.1
2022 Can We Obtain Fine-grained Heartbeat Waveform via Contact-free RF-sensing?
abstract
Contact-free vital-signs monitoring enabled by radio frequency (RF) sensing is gaining increasing attention, thanks to its non-intrusiveness, noise-resistance, and low cost. Whereas most of these systems only perform respiration monitoring or retrieve heart rate, few can recover fine-grained heartbeat waveform. The major reason is that, though both respiration and heartbeat cause detectable micro-motions on human bodies, the former is so strong that it overwhelms the latter. In this paper, we aim to answer the question in the paper title, by demystifying how heartbeat waveform can be extracted from RF-sensing signal. Applying several mainstream methods to recover heartbeat waveform from raw RF signal, our results reveal that these methods may not achieve what they have claimed, mainly because they assume linear signal mixing whereas the composition between respiration and heartbeat can be highly nonlinear. To overcome the difficulty of decomposing nonlinear signal mixing, we leverage the power of a novel deep generative model termed variational encoder-decoder (VED). Exploiting the universal approximation ability of deep neural networks and the generative potential of variational inference, VED demonstrates a promising capability in recovering fine-grained heartbeat waveform from RF-sensing signal; this is firmly validated by our experiments with 12 subjects and 48-hour data.
Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001
INFOCOM2
2022 Sound of Motion: Real-time Wrist Tracking with A Smart Watch-Phone Pair
abstract
Proliferation of smart environments entails the need for real-time and ubiquitous human-machine interactions through, mostly likely, hand/arm motions. Though a few recent efforts attempt to track hand/arm motions in real-time with COTS devices, they either obtain a rather low accuracy or have to rely on a carefully designed infrastructure and some heavy signal processing. To this end, we propose SoM (Sound of Motion) as a lightweight system for wrist tracking. Requiring only a smart watch-phone pair, SoM entails very light computations that can operate in resource constrained smartwatches. SoM uses embedded IMU sensors to perform basic motion tracking in the smartwatch, and it depends on the fixed smartphone to act as an "acoustic anchor": regular beacons sent by the phone are received in an irregular manner due to the watch motion, and such variances provide useful hints to adjust the drifting of IMU tracking. Using extensive experiments on our SoM prototype, we demonstrate that the delicately engineered system achieves a satisfactory wrist tracking accuracy and strikes a good balance between complexity and performance.
Tianyue Zheng, Chao Cai 0001, Zhe Chen 0015, Jun Luo 0001
INFOCOM1
2022 CORE-lens: simultaneous communication and object recognition with disentangled-GAN cameras
abstract
Optical camera communication (OCC) enabled by LED and embedded cameras has attracted extensive attention, thanks to its rich spectrum availability and ready deployability. However, the close interactions between OCC and the indoor spaces have created two major challenges. On one hand, the stripe pattern incurred by OCC may greatly damage the accuracy of image-based object recognition. On the other hand, the patterns inherent to indoor spaces can significantly degrade the decoding performance of reflected OCC. To this end, we propose CORE-Lens as a pipeline to make the mutual interference transparent to existing OR and OCC algorithms. Essentially, CORE-Lens treats the two challenges as two sides of a signal mixture issue: the signals transmitted by OCC get mixed with background images so well that their features become entangled. Consequently, CORE-Lens exploits the idea of disentangled representation learning to separate the mixed signals in the feature space: while the GAN-reconstructed clean background images are used to perform object recognition, OCC decoding is conducted on the residual of the original image after subtracting the reconstructed background. Our extensive experiments on evaluating the real-life performance of CORE-Lens evidently demonstrate its superiority over conventional approaches.
Ziwei Liu 0002, Tianyue Zheng, Yanbing Yang 0001, Yimao Sun, Zhe Chen 0015, Liangyin Chen, Jun Luo 0001
MobiCom2
2022 Quantifying the Physical Separability of RF-Based Multi-Person Respiration Monitoring via SINR
abstract
Recent years have witnessed a growing interest in contact-free respiration monitoring leveraging radio-frequency (RF) technologies. However, the proposed solutions mostly consider single-person scenarios, whereas a few multi-person monitoring proposals simply apply blind source separation to handle inter-person interference, without drawing a clear line between physical and algorithmic separability. In this paper, we set out to answer: under what condition(s) one may physically separate multiple respiration signals sensed by diversified RF technologies? Drawing inspiration from conventional signal processing, we propose respiration-to-interference-plus-noise ratio (RINR) as a novel metric, taking into account the impact from both background noise and various interfering sources. Instead of attenuation in Euclidean distance, RINR has to be evaluated upon range/angle bins where physical separation actually take place. As signal attenuation has never been modeled in this manner, we rise to this challenge by levering a deep learning model to fit a spread function upon range/angle bins. The resulting RINR model allows us to concretely indicate the limit of physical separability of RF-based multi-person respiration monitoring. Our extensive experiments firmly validate the RINR model, thus evidently demonstrating the benefits of employing RINR model as a guideline for conducting respiration monitoring with different RF technologies.
Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001
SenSys2
2021 Octopus: a practical and versatile wideband MIMO sensing platform
abstract
Radio frequency (RF) technologies have achieved a great success in data communication. In recent years, pervasive RF signals are further exploited for sensing; RF sensing has since attracted attentions from both academia and industry. Existing developments mainly employ commodity Wi-Fi hardware or rely on sophisticated SDR platforms. While promising in many aspects, there still remains a gap between lab prototypes and real-life deployments. On one hand, due to its narrow bandwidth and communication-oriented design, Wi-Fi sensing offers a coarse sensing granularity and its performance is very unstable in harsh real-world environments. On the other hand, SDR-based designs may hardly be adopted in practice due to its large size and high cost. To this end, we propose, design, and implement Octopus, a compact and flexible wideband MIMO sensing platform, built using commercial-grade low-power impulse radio. Octopus provides a standalone and fully programmable RF sensing solution; it allows for quick algorithm design and application development, and it specifically leverages the wideband radio to achieve a competent and robust performance in practice. We evaluate the performance of Octopus via micro-benchmarking, and further demonstrate its applicability using representative RF sensing applications, including passive localization, vibration sensing, and human/object imaging.
Zhe Chen 0015, Tianyue Zheng, Jun Luo 0001
MobiCom2
2021 MoVi-Fi: motion-robust vital signs waveform recovery via deep interpreted RF sensing
abstract
Vital signs are crucial indicators for human health, and researchers are studying contact-free alternatives to existing wearable vital signs sensors. Unfortunately, most of these designs demand a subject human body to be relatively static, rendering them very inconvenient to adopt in practice where body movements occur frequently. In particular, radio-frequency (RF) based contact-free sensing can be severely affected by body movements that overwhelm vital signs. To this end, we introduce MoVi-Fi as a motion-robust vital signs monitoring system, capable of recovering fine-grained vital signs waveform in a contact-free manner. Being a pure software system, MoVi-Fi can be built on top of virtually any commercial-grade radars. What inspires our design is that RF reflections caused by vital signs, albeit weak, do not totally disappear but are composited with other motion-incurred reflections in a nonlinear manner. As nonlinear blind source separation is inherently hard, MoVi-Fi innovatively employs deep contrastive learning to tackle the problem; this self-supervised method requires no ground truth in training, and it exploits contrastive signal features to distinguish vital signs from body movements. Our experiments with 12 subjects and 80hour data demonstrate that MoVi-Fi accurately recovers vital signs waveform under severe body movements.
Zhe Chen 0015, Tianyue Zheng, Chao Cai 0001, Jun Luo 0001
MobiCom2
2021 SiWa: see into walls via deep UWB radar
abstract
Being able to see into walls is crucial for diagnostics of building health; it enables inspections of wall structure without undermining the structural integrity. However, existing sensing devices do not seem to offer a full capability in mapping the in-wall structure while identifying their status (e.g., seepage and corrosion). In this paper, we design and implement SiWa as a low-cost and portable system for wall inspections. Built upon a customized IR-UWB radar, SiWa scans a wall as a user swipes its probe along the wall surface; it then analyzes the reflected signals to synthesize an image and also to identify the material status. Although conventional schemes exist to handle these problems individually, they require troublesome calibrations that largely prevent them from practical adoptions. To this end, we equip SiWa with a deep learning pipeline to parse the rich sensory data. With innovative construction and training, the deep learning modules perform structural imaging and the subsequent analysis on material status, without the need for repetitive parameter tuning and calibrations. We build SiWa as a prototype and evaluate its performance via extensive experiments and field studies; results evidently confirm that SiWa accurately maps in-wall structures, identifies their materials, and detects possible defects, suggesting a promising solution for diagnosing building health with minimal effort and cost.
Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001, Lin Ke, Chaoyang Zhao, Yaowen Yang
MobiCom1
2021 MoRe-Fi: Motion-robust and Fine-grained Respiration Monitoring via Deep-Learning UWB Radar
abstract
Crucial for healthcare and biomedical applications, respiration monitoring often employs wearable sensors in practice, causing inconvenience due to their direct contact with human bodies. Therefore, researchers have been constantly searching for contact-free alternatives. Nonetheless, existing contact-free designs mostly require human subjects to remain static, largely confining their adoptions in everyday environments where body movements are inevitable. Fortunately, radio-frequency (RF) enabled contact-free sensing, though suffering motion interference inseparable by conventional filtering, may offer a potential to distill respiratory waveform with the help of deep learning. To realize this potential, we introduce MoRe-Fi to conduct fine-grained respiration monitoring under body movements. MoRe-Fi leverages an IR-UWB radar to achieve contact-free sensing, and it fully exploits the complex radar signal for data augmentation. The core of MoRe-Fi is a novel variational encoder-decoder network; it aims to single out the respiratory waveforms that are modulated by body movements in a non-linear manner. Our experiments with 12 subjects and 66-hour data demonstrate that MoRe-Fi accurately recovers respiratory waveform despite the interference caused by body movements. We also discuss potential applications of MoRe-Fi for pulmonary disease diagnoses.
Tianyue Zheng, Zhe Chen 0015, Chao Cai 0001, Jun Luo 0001
SenSys1
2020 RF-net: a unified meta-learning framework for RF-enabled one-shot human activity recognition
abstract
Radio-Frequency (RF) based device-free Human Activity Recognition (HAR) rises as a promising solution for many applications. However, device-free (or contactless) sensing is often more sensitive to environment changes than device-based (or wearable) sensing. Also, RF datasets strictly require on-line labeling during collection, starkly different from image and text data collections where human interpretations can be leveraged to perform off-line labeling. Therefore, existing solutions to RF-HAR entail a laborious data collection process for adapting to new environments. To this end, we propose RF-Net as a meta-learning based approach to one-shot RF-HAR; it reduces the labeling efforts for environment adaptation to the minimum level. In particular, we first examine three representative RF sensing techniques and two major meta-learning approaches. The results motivate us to innovate in two designs: i) a dual-path base HAR network, where both time and frequency domains are dedicated to learning powerful RF features including spatial and attention-based temporal ones, and ii) a metric-based meta-learning framework to enhance the fast adaption capability of the base network, including an RF-specific metric module along with a residual classification module. We conduct extensive experiments based on all three RF sensing techniques in multiple real-world indoor environments; all results strongly demonstrate the efficacy of RF-Net compared with state-of-the-art baselines.
Shuya Ding, Zhe Chen 0015, Tianyue Zheng, Jun Luo 0001
SenSys3
2017 Deep probabilities for age estimation
abstract
Human age can provide important demographic information. In this paper, we tackle the estimation of age in face images with probabilities. The design of the proposed method is based on the relative order of age labels in the database. The age estimation problem is transformed into a series of binary classifications achieved by convolution neural network. Each classifier is used to judge whether the age of input image is larger than a certain age and the estimated age is obtained by adding probability values of these classification problems. The proposed method: Deep Probabilities (DP) of facial age shows improvements over direct regression and multi-classification methods.
Tianyue Zheng, Weihong Deng, Jiani Hu
VCIP1