Wei Xi 0003

dblp:10/264-3 · DBLP profile ↗
← Back
143ranked-venue papers
7as first author
89since 2021 · last 2026
0000-0001-9348-2982ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 66 · 5 first-author · 30 since 2021Artificial intelligence and machine learning · 24 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 20 since 2021Systems, architecture and hardware · 16 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count Prior
Dachao Han, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003
INFOCOM7
2026 Zero-Effort Cross-Domain Wireless Respiration Monitoring Under Free Movements With Commercial UWB Devices
abstract
Respiratory monitoring using wireless technologies has garnered significant attention for its potential in healthcare, smart cockpits, and various applications. Though extensively studied, existing systems face practical challenges in adapting to new data domains without substantial customization efforts. Current solutions attempt to address this limitation through domain-independent feature extraction or cross-domain feature translation, employing either knowledge-based sensing models or data-driven neural networks. However, these approaches typically require additional data collection or model retraining for new domains, significantly hindering their practical deployment. This paper proposes RF-Carer, a fully zero-effort cross-domain respiration monitoring system. Our key innovation lies in building an explainable propagation model to transform any heterogeneous signals under unknown domains into a unified form in the signal processing layer. To further address accidental irrelevant factors, we propose to align the feature spaces while suppressing the noisy ones with contrastive learning. On this basis, we develop a one-fits-all model that requires only one-time training but can adapt to 12 domains with 57 cases like unconstrained movements, unknown users, untrained environments, etc.. To the best of our knowledge, RF-Carer is the first zero-effort cross-domain respiration monitoring work with wireless RF signals and would be a fundamental step toward real-world deployments.
Ge Wang 0003, Jiazheng Chen, Zhe Chen 0015, Fei Wang 0037, Cong Zhao 0006, Han Ding 0002, Cui Zhao, Wei Xi 0003, Jinsong Han
SenSys9
2026 DARL: Diffusion-augmented representation learning via disentangled contrastive pre-training for industrial anomaly detection
Hongliang Luo, Wei Xi 0003
Neurocomputing2
2026 FedPRS: A Privacy-preserving Representation Synthesis Framework for Federated Contribution Evaluation
abstract
Federated Learning (FL) enables the collaborative training of a global model while protecting participants’ privacy. Evaluating each participant’s contribution is essential to providing a high-quality model, ensuring fairness, and mitigating potential biases. Most existing contribution evaluation approaches for FL assume that the server has a public validation dataset. However, it is almost impossible to obtain a validation dataset due to privacy concerns. In this article, we propose a Federated Privacy-preserving Representation Synthesis (FedPRS) framework to synthesize a validation dataset for contribution evaluation. The proposed FedPRS framework first transforms each participant’s private validation dataset into its representation. Then, a random-region desensitization strategy is developed to further desensitize the dataset without compromising its utility. The desensitized representation dataset of each participant is collected by the server to evaluate federated contribution, which considers both equity and privacy protection. Moreover, we instantiate and integrate three specific contribution evaluation approaches in this framework. We perform experiments on various FL settings, including independently identically distributed (IID) and non-IID data distributions. Experimental results demonstrate that the contribution evaluation results obtained using the validation dataset synthesized by the FedPRS framework are closely aligned with those obtained using a real, private validation dataset.
Yuan Yao 0011, Wei Xi 0003, Zelei Liu, Lixin Fan, Qiang Yang 0001
ACM Trans. Intell. Syst. Technol.3
2026 High-Efficiency Cellular Backscatter With Ambient Traffic
abstract
We present HEScatter, a high-efficiency ambient backscatter system that simultaneously improves carrier, power, and transmission efficiency. To improve carrier efficiency, we choose cellular signal as the carrier due to its continuous transmission nature. Specifically, to ensure low power, we design low-power periodic template matching based on the periodicity of cellular signals to trade time for synchronization accuracy. Further, we calibrate the drift introduced by Sampling Frequency Offset (SFO) to increase carrier utilization. In addition, we exploit Reference Signal (RS)-based demodulation to demodulate tag and ambient data from backscattered signals alone in various traffic patterns for efficient transmission. We prototype HEScatter using off-the-shelf FPGAs and SDRs. Extensive experiments show that HEScatter performs well in carrier utilization, power consumption and data transmission. The carrier utilization rate of HEScatter is as high as 99.97%, which is 3.0x higher than the counterpart of SyncLTE. In end-to-end transmission, the energy efficiency of HEScatter is 1.6x and 19.2x higher than LScatter+ and SyncLTE, while LScatter suffers from transmission failures. We also demonstrate the high transmission efficiency of HEScatter, as its aggregate goodput is 1.5x and 3.8x better than LScatter+ and SyncLTE respectively.
Yunyun Feng, Xianjun Deng, Shuai Wang 0021, Wei Xi 0003, Wei Gong 0001
IEEE Trans. Mob. Comput.5
2026 Active Domain Adaptation for mmWave-Based HAR via R$\acute{e}$e'nyi Entropy-Based Uncertainty Estimation
abstract
Human Activity Recognition (HAR) using mmWave radar provides a non-invasive alternative to traditional sensor-based methods but suffers from domain shift, where model performance declines in new users, positions, or environments. To address this, we propose mmADA, an Active Domain Adaptation (ADA) framework that efficiently adapts mmWave-based HAR models with minimal labeled data. mmADA enhances adaptation by introducing Rényi Entropy-based uncertainty estimation to identify and label the most informative target samples. Additionally, it leverages contrastive learning and pseudo-labeling to refine feature alignment using unlabeled data. Evaluations with a TI IWR1443BOOST radar across multiple users, positions, and environments show that mmADA achieves over 90% accuracy in various cross-domain settings. Comparisons with five baselines confirm its superior adaptation performance, while further tests on unseen users, environments, and two additional open-source datasets validate its robustness and generalization.
Mingzhi Lin, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003
IEEE Trans. Mob. Comput.7
2026 InfSquad: SLO-Aware Serverless Machine Learning Inference With Wasm-Assisted Hybrid Functions
abstract
Though serverless computing offers transformative benefits for deploying machine learning (ML) services, it faces challenges in meeting strict real-time service-level objectives (SLOs) of ML inference while maintaining resource efficiency. Fortunately, a new binary instruction format, WebAssembly (or Wasm), offers a promising solution for serverless ML inferences thanks to its short startup time and efficient execution compared with traditional container-based solutions. Therefore, we introduce InfSquad, a serverless ML inference framework designed to balance SLO-aware execution with resource efficiency. The key design of InfSquad is a Wasm-assisted hybrid serverless function runtime to harness the complementary strengths of Wasm and traditional containerized function runtime. Based on the hybrid runtime, InfSquad advocates an SLO-aware runtime scheduling approach that delivers efficient serverless inference. InfSquad leverages a proactive runtime recycling mechanism to increase resource efficiency further. Experiments on real-world applications show that, compared with state-of-the-art serverless systems, InfSquad achieves 19.8% 85.1% SLO violation reduction with 32.5% 74.8% fewer resources.
Borui Li 0001, Hai Wang 0019, Wei Xi 0003, Shuai Wang 0008, Tian He 0001
IEEE Trans. Serv. Comput.4
2025 TEMPEST-LoRa: Cross-Technology Covert Communication
abstract
Electromagnetic (EM) covert channels pose significant threats to computer and communications security in air-gapped networks. Previous works exploit EM radiation from various components (e.g., video cables, memory buses, CPUs) to secretly send sensitive information. These approaches typically require the attacker to deploy highly specialized receivers near the victim, which limits their real-world impact. This paper reports a new EM covert channel, TEMPEST-LoRa, that builds on Cross-Technology Covert Communication (CTCC), which could allow attackers to covertly transmit EM-modulated secret data from air-gapped networks to widely deployed operational LoRa receivers from afar. We reveal the potential risk and demonstrate the feasibility of CTCC by tackling practical challenges involved in manipulating video cables to precisely generate the EM leakage that could readily be received by third-party commercial LoRa nodes/gateways. Experiment results show that attackers can reliably decode secret data modulated by the EM leakage from a video cable at a maximum distance of 87.5m or a rate of 21.6 kbps. We note that the secret data transmission can be performed with monitors turned off (therefore covertly).
Xieyang Sun, Yuanqing Zheng, Wei Xi 0003, Zuhao Chen, Zhizhen Chen, Zhiping Jiang, Sheng Zhong 0002
CCS3
2025 Chinese Speech Processing via Chinese Character Feature
abstract
This paper focuses on the basic structure of Chinese characters: semantic-phonetic compound characters. This paper takes advantage of this feature of Chinese characters and innovatively proposes a Chinese speech-processing method based on character shape. We use the association between the character shape and pronunciation of Chinese characters to construct a new character stroke-based dataset. We use two neural network structures, RNN and Transformer, to verify our proposed Chinese speech processing method.It is proved through experiments that the method improves the performance of Mandarin ASR(Automatic Speech Recognition) by about 2% and AEC(ASR Error Correction) by about 1%. Theoretically, this method applies to all Chinese speech-processing algorithms based on the attention mechanism.
Wei Xi 0003, Xiao Fu 0001, Jizhong Zhao
ICASSP3
2025 Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning
abstract
The utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existing multimodal models face limitations in handling the modality-missing inference. In this work, we propose a Latent Modality Interaction with mutual information Maximization (LMIM) model architecture for multimodal learning, which effectively integrates multimodal cues by learning essential modality information and reducing the redundant information. We employ a group of latent tokens as pivots to filter out noise and redundancy across different modalities. Simultaneously, mutual information maximization and distribution alignment are utilized to preserve task-related information through multimodal fusion. Furthermore, a random modality masking training strategy is employed to mitigate potential over-reliance on dominant modality. Extensive experiments demonstrate that our model achieves significant improvement over current competitive baselines on two datasets, including UCF51 and Kinetics-Sounds datasets.
Xiao Fu 0001, Wei Xi 0003, Jizhong Zhao
ICASSP4
2025 Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition
abstract
Audio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or fully finetuning a pre-trained model, both of which incur significant computational costs and are often impractical for large-scale speech foundation models. This gap highlights the need for more efficient methods to leverage visual and acoustic information in AVSR tasks. To address this challenge, we propose AVWhisper, a parameter-efficient model that integrates visual and acoustic representations by injecting visual features from the AV-HuBERT encoder into the pre-trained Whisper model. Our approach leverages the existing attention mechanisms in Whisper to facilitate cross-modal interaction and integrates auxiliary visual information through lightweight adapters based on Low-Rank Adaptation (LoRA) and prompt-based techniques. Furthermore, a two-phase training strategy is adopted to effectively handle cross-domain differences and visual information injection problems respectively. Extensive experiments on the LRS3-TED dataset demonstrate that AVWhisper consistently outperforms state-of-the-art methods across various noise conditions, offering a more efficient and scalable solution for audio-visual speech recognition.
Yue Heng Yeo, Xiao Fu 0001, Weiguang Chen, Wei Xi 0003, Jizhong Zhao
ICASSP6
2025 UniConvNet: Expanding Effective Receptive Field While Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
abstract
Convolutional neural networks (ConvNets) with large effective receptive field (ERF), still in their early stages, have demonstrated promising effectiveness while constrained by high parameters and FLOPs costs and disrupted asymptotically Gaussian distribution (AGD) of ERF. This paper proposes an alternative paradigm: rather than merely employing extremely large ERF, it is more effective and efficient to expand the ERF while maintaining AGD of ERF by proper combination of smaller kernels, such as $7\times{7}$, $9\times{9}$, $11\times{11}$. This paper introduces a Three-layer Receptive Field Aggregator and designs a Layer Operator as the fundamental operator from the perspective of receptive field. The ERF can be expanded to the level of existing large-kernel ConvNets through the stack of proposed modules while maintaining AGD of ERF. Using these designs, we propose a universal model for ConvNet of any scale, termed UniConvNet. Extensive experiments on ImageNet-1K, COCO2017, and ADE20K demonstrate that UniConvNet outperforms state-of-the-art CNNs and ViTs across various vision recognition tasks for both lightweight and large-scale models with comparable throughput. Surprisingly, UniConvNet-T achieves $84.2\%$ ImageNet top-1 accuracy with $30M$ parameters and $5.1G$ FLOPs. UniConvNet-XL also shows competitive scalability to big data and large models, acquiring $88.4\%$ top-1 accuracy on ImageNet. Code and models are publicly available at https://github.com/ai-paperwithcode/UniConvNet.
Wei Xi 0003
ICCV2
2025 Numerical Estimation of Spatial Distributions Under Differential Privacy
abstract
Estimating spatial distributions is important in data analysis, such as traffic flow forecasting and epidemic prevention. To achieve accurate spatial distribution estimation, the analysis needs to collect sufficient user data. However, collecting data directly from individuals could compromise their privacy. Most previous works focused on private distribution estimation for one-dimensional data, which does not consider spatial data relation and leads to poor accuracy for spatial distribution estimation. In this paper, we address the problem of private spatial distribution estimation, where we collect spatial data from individuals and aim to minimize the distance between the actual distribution and estimated one under Local Differential Privacy (LDP). To leverage the numerical nature of the domain, we project spatial data and its relationships onto a one-dimensional distribution. We then use this projection to estimate the overall spatial distribution. Specifically, we propose a reporting mechanism called Disk Area Mechanism (DAM), which projects the spatial domain onto a line and optimizes the estimation using the sliced Wasserstein distance. Through extensive experiments, we show the effectiveness of our DAM approach on both real and synthetic data sets, compared with the state-of-the-art methods, such as Multi-dimensional Square Wave Mechanism (MDSW) and Subset Exponential Mechanism with Geo-I (SEM-Geo-I). Our results show that our DAM always performs better than MDSW and is better than SEM-Geo-I when the data granularity is fine enough.
Leilei Du 0001, Peng Cheng 0003, Libin Zheng 0001, Xiang Lian 0001, Lei Chen 0002, Wei Xi 0003, Wangze Ni
ICDE6
2025 Fed3D: Enhancing Security in Federated Learning with Dataset Distillation
abstract
Dataset Distillation (DD) compresses large datasets into compact representations while preserving performance, offering substantial benefits for Federated Learning (FL). However, using distilled datasets introduces new security vulnerabilities, as adversaries can easily embed backdoors into the distilled data. In this paper, we extend existing backdoor attack strategies in DD to the Federated Learning context (DD-FL) and empirically demonstrate their effectiveness. To address these threats, we propose Fed3D, the first defense algorithm specifically designed for DD-FL. Fed3D incorporates a dual-layer defense mechanism, combining intra-client diversity detection with inter-client clustering based on reconstructed feature representations. Comprehensive experiments show that Fed3D effectively reduces attack success rates (<1.5%) while maintaining the performance of distilled datasets (<1.1%). These results establish Fed3D as a robust and promising solution for mitigating backdoor attacks in DD-FL systems.
Canhui Wu, Wei Xi 0003, Yuhao Shen 0001, Jizhong Zhao
ICME2
2025 Advanced Backdoor Threats and Countermeasures in Dataset Condensation
abstract
Dataset Condensation (DC) aims to distill a large original dataset into a compact synthetic counterpart while preserving its utility. Although most research on DC focuses on improving performance, its security aspects remain underexplored. DC is inherently vulnerable to backdoor attacks, where malicious triggers embedded in the original data set become imperceptible, yet remain highly effective after condensation. This paper introduces a novel backdoor attack framework tailored for DC, which iteratively refines triggers by leveraging DC-specific characteristics. Based on distinct alignment strategies, this paper proposes two attack variants, DCA-MIN and DCA-MAX, and incorporates regularization terms to enhance trigger stealthiness. To counter these threats, we propose DC-Judge, a detection defense mechanism that identifies backdoors by reconstructing sample-level features and analyzing intra-class dispersion. Experimental results highlight the superior effectiveness and stealth of our proposed attack compared to existing methods, as well as DC-Judge's robustness in detecting backdoor-contaminated condensed datasets.
Canhui Wu, Wei Xi 0003, Jizhong Zhao
ICME2
2025 Rethinking Uplink Multi-User Access for Dense WLANs with Heterogeneous User States
abstract
The advent of modern WLAN standards, beginning with IEEE 802.11ax, has established a new paradigm for uplink multi-user access through foundational technologies like MU-MIMO and OFDMA. While these technologies promise unprecedented concurrent transmission speeds for dense deployments, their practical performance is severely constrained in networks with a massive number of users exhibiting heterogeneous states. Specifically, a user with a poor channel can contaminate an entire MU-MIMO group via error propagation, while a highly mobile user exacerbates the accumulation of channel sounding overhead. This paper systematically investigates these performance bottlenecks and their root causes. We propose Nexus, a simple yet effective cross-layer solution that consistently improves the uplink throughput in various environments. At the core of Nexus are two synergistic designs: (1) channel qualityaware user grouping, which isolates users with poor channel conditions to mitigate error propagation; and (2) opportunistic CSI multiplexing, which dynamically adapts the reuse of channel information according to each user's mobility pattern, thereby drastically reducing sounding overhead. We have prototyped Nexus on a software-defined radio platform, and our experimental results demonstrate that it improves the overall uplink throughput by up to$2 \times$compared to the standard 802.11ax, significantly outperforming existing state-of-the-art baselines.
Wei Xi 0003
ICPADS2
2025 Enhancing Multimodal Model Robustness Under Missing Modalities via Memory-Driven Prompt Learning
abstract
Existing multimodal models typically assume the availability of all modalities, leading to significant performance degradation when certain modalities are missing. Recent methods have introduced prompt learning to adapt pretrained models to incomplete data, achieving remarkable performance when the missing cases are consistent during training and inference. However, these methods rely heavily on distribution consistency and fail to compensate for missing modalities, limiting their ability to generalize to unseen missing cases. To address this issue, we propose Memory-Driven Prompt Learning, a framework that adaptively compensates for missing modalities through prompt learning. The compensation strategies are achieved by two types of prompts: generative prompts and shared prompts. Generative prompts retrieve semantically similar samples from a predefined prompt memory that stores modality-specific semantic information, while shared prompts leverage available modalities to provide cross-modal compensation. Extensive experiments demonstrate the effectiveness of the proposed model, achieving significant improvements across diverse missing-modality scenarios, with average performance increasing from 34.76% to 40.40% on MM-IMDb, 62.71% to 77.06% on Food101, and 60.40% to 62.77% on Hateful Memes. The code is available at https://github.com/zhao-yh20/MemPrompt.
Yihan Zhao, Wei Xi 0003, Xiao Fu 0001, Jizhong Zhao
IJCAI2
2025 A Unified Multi-Class Anomaly Detection Framework Based on Vision Foundation Models
abstract
Despite significant advances in unsupervised anomaly detection, two critical challenges persist. First, current anomaly simulation methods generate unrealistic anomalies due to insufficient constraints on anomalous regions, compromising training data quality. Second, the necessity to train separate models for different object classes leads to prohibitive computational and storage costs, severely limiting practical deployment. To address these issues, we propose a novel unified multi-class anomaly detection framework leveraging vision foundation models: (1) a foreground-constrained anomaly simulation method that utilizes vision foundation models to restrict anomaly generation to semantically meaningful foreground regions, producing more realistic training samples; (2) a Unified Anomaly Detection Model (UADM) built upon CLIP that enables efficient multi-class detection with a single model. UADM incorporates three components: an adapter module that processes multi-scale features from CLIP’s intermediate layers to capture anomalies at various granularities, a memory bank that efficiently stores normal pattern representations for rapid comparison during inference, and a Semantic Cluster module that effectively aggregates and distinguishes anomaly-specific semantic information. Extensive evaluations demonstrate the superiority of our approach. In multi-class detection scenarios, our framework achieves competitive performance with 98.5% image-level AUROC and 97.1% pixel-level AUROC on MVTec AD, along with 95.3% image-level AUROC and 98.4% pixel-level AUROC on VisA. Remarkably, our method maintains exceptional performance even in few-shot settings, underscoring its data efficiency and practical utility. These results establish our framework as a significant advancement in developing efficient, accurate, and generalizable multi-class anomaly detection systems.
Wei Xi 0003
IJCNN2
2025 One Snapshot is All You Need: A Generalized Method for mmWave Signal Generation
Han Ding 0002, Wenxin Sun, Cui Zhao, Ge Wang 0003, Fei Wang 0037, Kun Zhao 0002, Zhi Wang 0002, Wei Xi 0003
INFOCOM9
2025 Visually-Adaptive Guided Robust Speech Recognition with Parameter-Efficient Adaptation
Yue Heng Yeo, Xiao Fu 0001, Wei Xi 0003, Jizhong Zhao
INTERSPEECH5
2025 LES-CLIP: A Lightweight Emotion-Sensitive Adaptation of CLIP for Precise Similar Emotion Discrimination
abstract
CLIP has been widely adopted in affective computing for its strong vision-language representation capabilities. However, it fails to accurately distinguish visually similar yet label-distinct facial expressions. This limitation is rooted in CLIP's encoding paradigm and large-scale contrastive pretraining, which bias the model toward focusing primarily on globally salient visual features and aligning them with broad semantic concepts. Such alignment overlooks subtle facial variations and induces representational shortcuts, where emotionally distinct categories are projected into overlapping regions of the shared semantic space. This semantic entanglement severely compromises the model's ability to preserve emotional separability. We propose LES-CLIP, a Lightweight and Emotion-Sensitive framework that adapts CLIP for precise discrimination of similar emotions. LES-CLIP achieves fine-grained emotional sensitivity using only simple text prompts and facial images. It introduces three novel components: 1) an Emotion-Sensitive Adaptive Mixture-of-Experts, which pre-adapts representations for subtle expression discrimination; 2) a Prompt-Guided Emotion Discrimination module that activates CLIP's visual sensitivity to fine-grained facial cues; and 3) a LES hybrid loss that guides contrastive learning toward accurate emotion-label alignment. Extensive experiments demonstrate that LES-CLIP achieves state-of-the-art performance, reaching 70.18% on the 8-class AffectNet dataset. Moreover, it converges faster and requires significantly fewer parameters.
Xiao Fu 0001, Wei Xi 0003, Kun Zhao 0002, Jiadong Feng, Jizhong Zhao
ACM Multimedia3
2025 Poster: Zero-effort Cross-domain Wireless Respiration Monitoring under Free Body Movement
abstract
Wireless respiratory monitoring has garnered significant attention for its potential in various applications. However, existing systems face practical challenges in adapting to new data domains without substantial customization efforts. Current solutions attempt to address this limitation through domain-independent feature extraction or cross-domain feature translation, employing either knowledge-based sensing models or data-driven neural networks. However, these approaches typically require additional data collection or model retraining for new domains, significantly hindering their practical deployment. This paper proposes RF-Carer, a fully zero-effort cross-domain respiration monitoring system. Our key innovation lies in building an explainable propagation model to transform any heterogeneous signals under unknown domains into a unified form in the signal processing layer. To further address accidental irrelevant factors, we propose to align the feature spaces while suppressing the noisy ones with contrastive learning. On this basis, we develop a one-fits-all model that requires only one-time training but can adapt to unknown scenarios with unconstrained user movements, postures, positions, etc. To the best of our knowledge, RF-Carer is the first zero-effort cross-domain respiration monitoring work with wireless RF signals and would be a fundamental step toward real-world deployments.Chen
Jiazheng Chen, Ge Wang 0003, Zhe Chen 0015, Fei Wang 0037, Wei Xi 0003, Jinsong Han
MobiCom5
2025 SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation
abstract
Dataset condensation aims to synthesize compact yet informative datasets that retain the training efficacy of full-scale data, offering substantial gains in efficiency. Recent studies reveal that the condensation process can be vulnerable to backdoor attacks, where malicious triggers are injected into the condensation dataset, manipulating model behavior during inference. While prior approaches have made progress in balancing attack success rate and clean test accuracy, they often fall short in preserving stealthiness, especially in concealing the visual artifacts of condensed data or the perturbations introduced during inference. To address this challenge, we introduce \textsc{Sneakdoor}, which enhances stealthiness without compromising attack effectiveness. \textsc{Sneakdoor} exploits the inherent vulnerability of class decision boundaries and incorporates a generative module that constructs input-aware triggers aligned with local feature geometry, thereby minimizing detectability. This joint design enables the attack to remain imperceptible to both human inspection and statistical detection. Extensive experiments across multiple datasets demonstrate that \textsc{Sneakdoor} achieves a compelling balance among attack success rate, clean test accuracy, and stealthiness, substantially improving the invisibility of both the synthetic data and triggered samples while maintaining high attack efficacy. The code is available at \url{https://github.com/XJTU-AI-Lab/SneakDoor}.
Dongyi Lv, Wei Xi 0003, Jizhong Zhao
NeurIPS4
2025 Multi-objective federated learning: Balancing global performance and individual fairness
Yuhao Shen 0001, Wei Xi 0003, Yunyun Cai, Jizhong Zhao
Future Gener. Comput. Syst.2
2025 mmYodar+: Robust Human Detection Using mmWave Signals
abstract
The detection of human objects can be crucial for various real-world applications, such as surveillance and autonomous driving. However, traditional vision-based approaches suffer from limitations such as low lighting conditions, occlusions, and privacy concerns. To address these challenges, we introduce mmYodar+, a novel mmWave-based automatic human detection system. Our system processes mmWave signals to generate a 3D point cloud, which is then transformed into a 2D radar image for easier visualization and analysis. To enhance human profiling, we filter the point cloud using biometric information and expand human-related points in the image based on radar angle resolution, incorporating color to improve the differentiation. Additionally, we employ a deep mutual learning (DML) framework, enabling efficient human detection using a lightweight DNN. Experimental results show that mmYodar+ achieves an average precision of 96.29% in various scenarios, including indoor and outdoor environments, various lighting conditions, and in the presence of occlusions. These results demonstrate the effectiveness of using mmWave radar signals for reliable and accurate human detection.
Yuance Chang, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Zhi Wang 0002, Wei Xi 0003
IEEE Internet Things J.8
2025 Graph-Empowered Multidimensional Target Full-Coverage Reliability for Internet of Everything
abstract
Wireless sensor network plays a crucial role in sensing everything in Internet of Everything (IoE) applications. Network reliability, which measures the ability of the network to satisfy specific requirements, is one of the core factors influencing the quality of service of the network and a vital support for ensuring the normal operation of IoE applications. Existing reliability evaluation methods are mainly based on minimum cutsets or paths, which are inefficient and not suitable for large-scale networks. Furthermore, most work either focuses on coverage functionality or connectivity functionality, lacking energy awareness. To address these limitations, this article proposes a multidimensional target full-coverage reliability (TFCR). TFCR comprehensively considers various factors affecting network reliability. To evaluate TFCR, a graph-empowered confident information coverage (CIC) and signal-to-interference and noise ratio (SINR)-based energy-aware reliability algorithm (CSERA) is proposed. This algorithm evaluates network coverage based on the CIC model. Additionally, graph neural networks and the SINR-based fade tail connectivity (FTC) model are used to evaluate network connectivity functionality. CSERA balances computational accuracy and efficiency, providing reliability evaluation values within an acceptable margin of error. Extensive simulations and comparative experiments from multiple perspectives demonstrate the superiority of the proposed method CSERA over existing approaches.
Chenlu Zhu, Wujie Zheng, Xiaoxuan Fan, Xianjun Deng, Shenghao Liu, Lingzhi Yi, Wei Xi 0003, Young-Sik Jeong
IEEE Internet Things J.7
2025 HeartIt: Low-Power Smoking Detection with a Smartwatch on Either Wrist
Jiao Ma, Tianzhang Xing, Wei Xi 0003, Kun Zhao 0002, Xiaojiang Chen
J. Comput. Sci. Technol.3
2025 Evolutionary cross-client network aggregation for personalized federated learning
Wei Xi 0003, Yuhao Shen 0001, Jizhong Zhao
Knowl. Based Syst.2
2025 Infinite Stream Estimation under Personalized w-Event Privacy
abstract
Streaming data collection is indispensable for stream data analysis, such as event monitoring. However, publishing these data directly leads to privacy leaks. w -event privacy is a valuable tool to protect individual privacy within a given time window while maintaining high accuracy in data collection. Most existing w -event privacy studies on infinite data stream only focus on homogeneous privacy requirements for all users. In this paper, we propose personalized w -event privacy protection that allows different users to have different privacy requirements in private data stream estimation. Specifically, we design a mechanism that allows users to maintain constant privacy requirements at each time slot, namely Personalized Window Size Mechanism (PWSM). Then, we propose two solutions to accurately estimate stream data statistics while achieving w -Event є -Personalized Differential Privacy (( w,є )-EPDP), namely Personalized Budget Distribution (PBD) and Personalized Budget Absorption (PBA). PBD always provides at least the same privacy budget for the next time step as the amount consumed in the previous release. PBA fully absorbs the privacy budget from the previous k time slots, while also borrowing from the privacy budget of the next k time slots, to increase the privacy budget for the current time slot. We prove that both PBD and PBA outperform the state-of-the-art private stream estimation methods while satisfying the privacy requirements of all users. We demonstrate the efficiency and effectiveness of our PBD and PBA on both real and synthetic datasets, compared with the recent uniformity w -event approaches, Budget Distribution (BD) and Budget Absorption (BA). Our PBD achieves 68% less error than BD on average on real datasets. Besides, our PBA achieves 24.9% less error than BA on average on synthetic datasets.
Leilei Du 0001, Peng Cheng 0003, Lei Chen 0002, Heng Tao Shen, Xuemin Lin 0001, Wei Xi 0003
Proc. VLDB Endow.6
2025 MUSE: A Trustworthy Vertical Federated Feature Selection Framework
abstract
Vertical federated feature selection can select effective features and avoid overfitting in vertical federated learning. However, existing privacy-preserving techniques for vertical federated feature selection are limited to selecting task-related features and cannot reduce redundant features among clients, resulting in performance loss. This article introduces a mutual information-based federated feature selection (MUSE) framework to address these issues. In the MUSE framework, the correlation of cross-device feature–feature and feature–class is estimated by our defined privacy-preserving mutual information, called federated mutual information (FMI). To compute FMI, we propose the anonymous bin matching (ABM) algorithm, which only uses the intersection size of bins rather than bin elements to avoidsample-IDsleakage. With FMI, MUSE can support the minimized dependency feature selection criteria for removing redundant features. Additionally, we propose the local feature preselection to reduce the computation cost of FMI. It is theoretically and experimentally proved as a close approximation of the global optimum under certain constraints. We evaluate the effectiveness of our MUSE framework on various datasets. The experimental results demonstrate that our methods consistently outperform the state-of-the-art federated feature selection methods across most datasets. Moreover, our method shows potential in multimodal data as well.
Xinyuan Ji, Olga Gadyatskaya, Zixiang Mao, Wei Xi 0003
IEEE Trans. Comput. Soc. Syst.6
2025 Ten Challenging Problems in Federated Foundation Models
abstract
Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large foundation models and the small local domain models at the remote clients to learn from each other in a teacher-student learning setting. This paper provides a comprehensive summary of the ten challenging problems inherent in FedFMs, encompassing foundational theory, utilization of private data, continual learning, unlearning, Non-IID and graph data, bidirectional knowledge transfer, incentive mechanism design, game mechanism design, model watermarking, and efficiency. The ten challenging problems manifest in five pivotal aspects: “Foundational Theory,” which aims to establish a coherent and unifying theoretical framework for FedFMs. “Data,” addressing the difficulties in leveraging domain-specific knowledge from private data while maintaining privacy; “Heterogeneity,” examining variations in data, model, and computational resources across clients; “Security and Privacy,” focusing on defenses against malicious attacks and model theft; and “Efficiency,” highlighting the need for improvements in training, communication, and parameter efficiency. For each problem, we offer a clear mathematical definition on the objective function, analyze existing methods, and discuss the key challenges and potential solutions. This in-depth exploration aims to advance the theoretical foundations of FedFMs, guide practical implementations, and inspire future research to overcome these obstacles, thereby enabling the robust, efficient, and privacy-preserving FedFMs in various real-world applications.
Tao Fan 0002, Hanlin Gu, Xuemei Cao 0001, Chee Seng Chan, Qian Chen 0023, Yiqiang Chen 0001, Yihui Feng, Yang Gu 0001, Jiaxiang Geng, Bing Luo 0002, Shuoling Liu, WinKent Ong, Chao Ren 0006, Jiaqi Shao, Xiaoli Tang 0001, Hong Xi Tae, Yongxin Tong, Shuyue Wei 0001, Fan Wu 0006, Wei Xi 0003, Mingcong Xu, Xin Yang 0012, Jiangpeng Yan, Hao Yu 0023, Han Yu 0001, Xiaojin Zhang 0002, Zhenzhe Zheng 0001, Lixin Fan, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.21
2025 HPST-GT: Full-Link Delivery Time Estimation Via Heterogeneous Periodic Spatial-Temporal Graph Transformer
abstract
A warehouse-distribution integration (WDI) e-commerce platform is an approach that combines warehousing and distribution processes, which is increasingly adopted in industry to enhance business efficiency. In the WDI e-commerce, one of the most important problems is to estimate the full-link delivery time for decision-making. Traditional methods designed for separate warehouse-distribution models struggle to address challenges in integrated systems. The difficulties stem from two main factors: (i) the contextual influence exerted by neighboring units within heterogeneous delivery networks, and (ii) the uncertainty in delivery times caused by dynamic and periodic temporal factors such as fluctuations in online sales volumes and the varying characteristics of different delivery units (e.g., warehouses and sorting centers). To address these challenges, we propose a novel full-link delivery time estimation framework calledHeterogeneousPeriodicSpatial-TemporalGraphTransformer (HPST-GT). First, we develop heterogeneous graph transformers to capture the hierarchical and diverse information of the warehouse-distribution network. Next, we design spatial-temporal transformers based on heterogeneous features to analyze the correlation between spatial and temporal information. Finally, we create a heterogeneous spatial-temporal graph prediction module to estimate full-link delivery time. Our method, evaluated on a one-month dataset from a leading e-commerce platform, surpasses current benchmarks across multiple performance metrics.
Shuai Wang 0008, Hai Wang 0019, Li Lin 0011, Xiaohui Zhao 0006, Tian He 0001, Dian Shen, Wei Xi 0003
IEEE Trans. Knowl. Data Eng.7
2025 Individualized Data Generation in Personalized Federated Learning
abstract
Most Personalized Federated Learning (PFL) algorithms merge the model parameters of each client with other (similar or generic) model parameters to optimize the personalized model (PM). However, the merged model parameters in these algorithms may fit low relevance data, thereby limiting the performance of PM. In this paper, we generate similar data for each client through the collaboration of a generic model (GM) on the server, rather than merging model parameters. To train a generator capable of generating data for all classes on the server without real data, we employ the GM as the discriminator in adversarial training with the generator. Additionally, we introduce a similarity assessment metric, which allows for the assessment of the similarity between local data and data from other classes. Nevertheless, the presence of non-IID data among clients can weaken the performance of the GM, consequently impacting the training of the generator and similarity assessment. To address this issue, we design a directive mechanism so that GM can be optimized during adversarial training without the need for additional training. The experimental results validate the superiority of our algorithm over state-of-the-art algorithms in terms of accuracy, loss, and convergence speed.
Yunyun Cai, Wei Xi 0003, Yuhao Shen 0001, Cerui Sun, Shuai Wang 0008, Wei Gong 0001, Jizhong Zhao
IEEE Trans. Mob. Comput.2
2025 Federated Multi-Source Domain Adaptation for mmWave-Based Human Activity Recognition
abstract
Contactless mmWave-based human activity recognition (HAR) is essential for various applications, yet most existing approaches often assume consistent environments. Integrating domain adaptation offers a promising solution to this challenge. This prevailing paradigm works well when the source and target data are centralized on a single server while learning to adapt. However, in more universal and practical situations, such as personal health records, users’ biometric information, and financial issues, the raw data is typically protected by different privacy-preserving policies and is stored by multiple parties. Additionally, labeling RF signals in the target domain is a non-trivial and labor-intensive task for most end-users. To address these problems, this paper introduces FMDA, a federated multi-source domain adaptation framework for mmWave-based HAR. FMDA assesses the contribution of each source and performs weighted parameter aggregation for knowledge transfer. This facilitates unsupervised training of the target HAR model without requiring access to any source domain data. Moreover, the model is optimized by minimizing the generalization gaps between the source and target models, benefiting all participants during the learning process and enhancing overall performance. Extensive experiments demonstrate the effectiveness of FMDA. The results indicate that in the target domain, FMDA achieves comparable performance to supervised learning approaches, while also enhancing the efficacy of source domain models to varying degrees.
Cui Zhao, Guotong Fang, Han Ding 0002, Fei Wang 0037, Ge Wang 0003, Kun Zhao 0002, Zhi Wang 0002, Wei Xi 0003
IEEE Trans. Mob. Comput.9
2025 mm-Fall: Practical and Robust Fall Detection via mmWave Signals
abstract
Falls pose a significant risk to the health and wellbeing of older adults, driving the development of various fall detection systems. Existing solutions have explored wearable and vision sensors, while non-invasive RF-based approaches have raised a growing interest due to their convenience and privacy considerations. Despite major advancements in RF-based passive estimation, current approaches still face challenges in handling complex real-world scenarios. They often lack the ability to generalize to new domains (i.e., people, position, environment), and struggle to accurately detect and localize a fallen person in the presence of unknown activities from nearby objects (e.g., pet animal and robot vacuum cleaner) or persons. To address these challenges, we present mm-Fall, a novel mmWave-based non-invasive fall detection system that utilizes Range-Angle (RA) energy maps to separate and localize multiple moving targets, and further accurately estimate their states. Unlike previous approaches, mm-Fall is capable of working with new domains and effectively distinguishing falls from non-fall motions that may appear similar. Additionally, it performs well in challenging conditions, such as poor lighting and occluded scenarios. Our design of mm-Fall is evaluated in 13 environments with over 16 individuals performing 24+ types of motions. The results demonstrate an impressive average recall of 0.969 and precision of 0.996 in detecting falls, whether involving single or multiple moving targets simultaneously. The code and dataset will be made publicly available.
Cui Zhao, Qiumin Luo, Han Ding 0002, Ge Wang 0003, Kun Zhao 0002, Zhi Wang 0002, Wei Xi 0003, Jizhong Zhao
IEEE Trans. Mob. Comput.7
2025 Robust and Rotation-Equivariant Contrastive Learning
abstract
Contrastive learning (CL) methods achieve great success by learning the invariant representation from various transformations. However, rotation transformations are considered harmful to CL and are rarely used, which results in failure when the objects show unseen orientations. This article proposes a representation focus shift network (RefosNet), which adds the rotation transformations to CL methods to improve the robustness of representation. First, the RefosNet constructs the rotation-equivariant mapping between the features of the original image and the rotated ones. Then, the RefosNet learns semantic-invariant representations (SIRs) based on explicitly decoupling the rotation-invariant features and the rotation-equivariant features. Moreover, an adaptive gradient passivation strategy is introduced to gradually shift the representation focus to invariant representations. This strategy can prevent catastrophic forgetting of the rotation equivariance, which is beneficial to the generalization of representations in both seen and unseen orientations. We adapt the baseline methods (i.e., "SimCLR" and "momentum contrast (MoCo) v2") to work with RefosNet to verify the performance. Extensive experimental results show that our method achieves significant improvements on the task of recognition. On ObjectNet-13 with unseen orientations, RefosNet gains 7.12% in terms of classification accuracy compared with SimCLR. On datasets in seen orientation, the performance improves by 5.5% on ImageNet-100, 7.29% on STL10, and 1.93% on CIFAR10. In addition, RefosNet has strong generalization on Place205, PASCAL VOC, and Caltech 101. Our method has also achieved satisfactory results in image retrieval tasks.
Gairui Bai, Wei Xi 0003, Xiaopeng Hong, Songwen Zhao
IEEE Trans. Neural Networks Learn. Syst.2
2025 EchScatter: Enriching Codeword Translation for High-Throughput Ambient ZigBee Backscatter
abstract
We present EchScatter, a novel backscatter system that takes productive ZigBee signals as excitations and enables high-throughput ZigBee backscatter communication. Compared with the existing ZigBee backscatter, EchScatter does not need to control the carrier, ensuring the universality of the transmitter. EchScatter has realized chip-level modulation of productive ambient ZigBee backscatter for the first time. It is capable of transmitting more tag bits simultaneously, thus increasing the throughput of the system. We first design 16 different 32-chip phase modulation sequences to realize the translation of any two ZigBee symbols. Therefore, our system can transmit four tag bits through one symbol. After that, we use an average energy detection-based synchronization method to make sure the synchronization error of the EchScatter meets the requirements of chip-level modulation. We prototyped EchScatter using an off-the-shelf FPGA and commodity ZigBee transceivers. Through extensive experiments and field studies, we show that EchScatter can work universally on commodity ZigBee transceivers. In line-of-sight and non-line-of-sight scenarios, EchScatter is able to achieve 247 kbps and 245 kbps throughput at a signal strength of around -60 dBm, respectively. The throughput of EchScatter is 32 times higher than that of FreeRider. We believe it will have more pervasive applications.
Jiuwei Li, Shixin Wang 0008, Zhaoyuan Xu, Wei Xi 0003, Shuai Wang 0008, Wei Gong 0001
ACM Trans. Sens. Networks4
2025 Towards Stable WiFi-based HAR from Imbalanced Data and Changing Circumstances
abstract
WiFi-based human activity recognition (WiFi-based HAR) has emerged as a technology in recent decades, offering convenient and privacy-friendly applications. However, existing frameworks designed for stable environments encounter challenges when faced with changing circumstances and imbalanced training datasets in realistic scenarios. In this article, we address both issues from a unified perspective by exploring a more generalized local minima. Initially, we revisit existing solutions and empirically observe the presence of sharp minima in trained long-tailed WiFi-based HAR models. Consequently, we propose a novel method called Class Region Flattening ( CRF ) to identify class-conditional flat minima. This approach effectively mitigates bias caused by the long-tailed distribution and enhances generalization capabilities in the face of changing circumstances. Furthermore, we introduce a selective flattening operation to prevent optimization conflicts among different activity categories and reduce computational overhead. We integrate CRF into mainstream WiFi-based HAR models and evaluate their performance using our collected WiFi-based HAR dataset. Through extensive experiments, we demonstrate that the incorporation of CRF leads to significant improvements in performance. These findings underscore the effectiveness of CRF in addressing the challenges posed by changing circumstances and imbalanced training datasets in WiFi-based HAR.
Youquan Wang, Shuai Wang 0008, Xianjun Deng, Wei Xi 0003, Wei Gong 0001
ACM Trans. Sens. Networks5
2025 Leveraging Time-Shifted Orthogonal Codes for Concurrent Backscatter Communication
abstract
Backscatter communication has attracted significant attention due to its low power consumption and energy efficiency. Enabling concurrent backscatter allows multiple tags to operate simultaneously, and their data can be recovered from collided signals. This capability is crucial for enhancing management efficiency in smart logistics and mitigating multi-tag collisions in Internet-of-Things (IoT) scenarios where multiple tags work collaboratively. However, existing concurrent backscatter schemes are vulnerable to noise and asynchronous signals, causing limited performance. To address these challenges, we introduce Ortho-CodeA, a backscatter scheme that enables reliable concurrent backscatter communication despite high noise levels and asynchronous signals. The underlying concept is to take advantage of coding mechanisms to combat noise and employ time-shifted orthogonal codes to mitigate the effects of asynchronous signals. Specifically, we design a set of time-shifted orthogonal codes that maintain code orthogonality despite asynchronous signals. Built upon the designed codes, we develop a multi-tag decoding scheme to recover data from each tag. We theoretically analyze the feasibility of our scheme and validate its performance through extensive experimental simulations. The results demonstrate that Ortho-CodeA achieves a BER of about 0.0036% in the case of 7 tags with an SNR of 10 dB and a maximum time delay of$1 \,\mu \text{s}$.
Weiqi Wu, Wei Xi 0003, Xianjun Deng, Shuai Wang 0021, Haoquan Zhou, Wei Gong 0001
IEEE Trans. Sustain. Comput.2
2024 FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning
abstract
Federated Learning (FL) heavily depends on label quality for its performance. However, the label distribution among individual clients is always both noisy and heterogeneous. The high loss incurred by client-specific samples in heterogeneous label noise poses challenges for distinguishing between client-specific and noisy label samples, impacting the effectiveness of existing label noise learning approaches. To tackle this issue, we propose FedFixer, where the personalized model is introduced to cooperate with the global model to effectively select clean client-specific samples. In the dual models, updating the personalized model solely at a local level can lead to overfitting on noisy data due to limited samples, consequently affecting both the local and global models’ performance. To mitigate overfitting, we address this concern from two perspectives. Firstly, we employ a confidence regularizer to alleviate the impact of unconfident predictions caused by label noise. Secondly, a distance regularizer is implemented to constrain the disparity between the personalized and global models. We validate the effectiveness of FedFixer through extensive experiments on benchmark datasets. The results demonstrate that FedFixer can perform well in filtering noisy label samples on different clients, especially in highly heterogeneous label noise scenarios.
Xinyuan Ji, Zhaowei Zhu, Wei Xi 0003, Olga Gadyatskaya, Zilong Song, Yang Liu 0018
AAAI3
2024 UFDA: Universal Federated Domain Adaptation with Practical Assumptions
abstract
Conventional Federated Domain Adaptation (FDA) approaches usually demand an abundance of assumptions, which makes them significantly less feasible for real-world situations and introduces security hazards. This paper relaxes the assumptions from previous FDAs and studies a more practical scenario named Universal Federated Domain Adaptation (UFDA). It only requires the black-box model and the label set information of each source domain, while the label sets of different source domains could be inconsistent, and the target-domain label set is totally blind. Towards a more effective solution for our newly proposed UFDA scenario, we propose a corresponding methodology called Hot-Learning with Contrastive Label Disambiguation (HCLD). It particularly tackles UFDA's domain shifts and category gaps problems by using one-hot outputs from the black-box models of various source domains. Moreover, to better distinguish the shared and unknown classes, we further present a cluster-level strategy named Mutual-Voting Decision (MVD) to extract robust consensus knowledge across peer classes from both source and target domains. Extensive experiments on three benchmark datasets demonstrate that our method achieves comparable performance for our UFDA scenario with much fewer assumptions, compared to previous methodologies with comprehensive additional assumptions.
Luping Zhou, Dong Xu 0001, Wei Xi 0003, Gairui Bai, Yihan Zhao, Jizhong Zhao
AAAI5
2024 Adapting Multi-view Correlation Learning for Enhanced Cervical Dysplasia Diagnosis
abstract
Analyzing relationships among pathological images under varying conditions during colposcopy is crucial for diagnosing cervical dysplasia. However, images from different conditions present challenges in accurate alignment and scaling, and the lesion features exhibit inconsistencies. These factors increase the difficulty of precisely capturing correlations between lesions across different views, especially without lesion annotations. To address these challenges, we propose the Dual Modalities for Multi-view Correlation Learning (D2MC) method. It enhances lesion representation by considering connections between text and visual data and relationships among multiple views. Specifically, the model first aligns lesion-specific text attributes with images at a fine-grained level, adaptively adjusting the visual model’s attention to focus on the lesion. It then incorporates a Temporal Correlation Adapter (TC-Adapter) into the encoder and introduces a Multi-View Fusion Module (MVFM) to strengthen multi-view correlation learning from a visual attention perspective. Finally, it aligns learnable text prompts with diagnostic outcomes and integrates a visual feature classifier to enhance diagnostic precision. Experimental results demonstrate that D2MC advances colposcopic image diagnosis by effectively capturing key correlations across different images.
Gairui Bai, Wei Xi 0003, Jialin Zhuo
BIBM2
2024 Towards Seamless Single Receiver Backscatter with Uncontrolled Ambient OFDM WiFi
abstract
OFDM WiFi backscatter with uncontrolled ambient signals is a promising approach for realizing passive Internet of Things (IoT) systems. However, existing OFDM backscatter systems are often limited by coarse modulation granularity, typically constrained to the packet or OFDM symbol levels, and commonly require dual receivers for data demodulation. To address these challenges, we propose DFTScatter, a novel sub-symbol level backscatter system that utilizes a single receiver for demodulation. DFTScatter leverages the Discrete Fourier Transform (DFT) shift theorem to achieve sub-symbol level modulation through frequency domain cyclic shifts. This method enhances data transmission efficiency and operates within a single-symbol bandwidth, thereby optimizing spectrum utilization. Additionally, we introduce a single-receiver decoding technique that exploits invariant frequency domain information of the reflected signals for accurate demodulation. Extensive experiments demonstrate that DFTScatter outperforms existing methods and effectively operates with various ambient OFDM WiFi signals, including WiFi 3/4/5/6, paving the way for scalable, low-cost IoT ecosystems in smart cities, healthcare, industrial automation, and beyond.
Chenhong Cao, Wei Xi 0003, Shuai Wang 0008, Wei Gong 0001
HPCC2
2024 O2O Logistics Customer Value Prediction with Periodic Asynchronous Vertical Federated Learning
abstract
Recent years have witnessed significant advancements in O2O logistics, which require predicting the volume of shipments generated by customers, commonly referred to as customer value. The essence of accurately predicting customer value in O2O logistics involves analyzing both online buying habits and offline logistics operations. Existing customer value prediction efforts focus solely on online or offline features, making them unsuitable for O2O scenarios. In this paper, we investigate the integration of both online and offline features for customer value prediction, which faces challenges including (i) data silos issues between logistics platforms and e-commerce platforms, and (ii) feature heterogeneity between online and offline data. To address these challenges, we propose a Periodic Asynchronous Vertical Federated Learning framework with adaptive Feature Selection (PAVFL-FS), enabling logistics platforms to efficiently and accurately predict customer value in collaboration with the e-commerce platforms. PAVFL-FS consists of two components: (i) a periodic asynchronous vertical federated learning algorithm to handle data silos problem and enable efficient model training; (ii) a local feature selection algorithm based on stochastic gates to address the cross-platform feature heterogeneity. We have conducted experimental validation on a large-scale real-world dataset collected from a major logistics company in China, including over 2.4 million waybill records and over 6 million online sales records from more than 3,500 merchants. The experimental results demonstrate that PAVFL-FS achieves a mean absolute error of 0.52 in O2O logistics customer value prediction, outperforming 24.6% to the baseline.
Ruize Li, Baoshen Guo, Shuai Wang 0008, Xiaolei Zhou 0001, Wei Xi 0003
HPCC6
2024 Byzantine Robust Aggregation in Federated Distillation with Adversaries
abstract
Federated learning empowers privacy-preserving, multi-party secure model training without the necessity of sharing raw data. In recent years, knowledge distillation has emerged as a promising solution to address the significant challenge of model heterogeneity within federated learning. However, current research often overlooks the potential threats posed by Byzantine attacks, which can significantly compromise the security of federated distillation. Previous work on Byzantine attacks has been primarily focused on manipulating local gradients to compromise global model, lacking attacks on logits in knowledge distillation scenarios. In this paper, we introduce two innovative attacks, shedding light on the inherent risks in federated distillation. The proposed attacks include a top-k attack, which perturbs the top k values of logits in each column, and an impersonation attack, which emulates knowledge significantly deviating from the norm. To counter such attacks, we propose a robust aggregation strategy-FedTGD (Federated Top Guard Distillation), designed to ensure robust distillation with heterogeneous models. Specifically, FedTGD incorporates Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and maximum cosine similarity on top-k values of logits to select benign knowledge. Experimental evaluations conducted on FEMNIST and CIFAR100 datasets, considering scenarios for both IID and Non-IID, reveal that top-k attack results in a substantial 27.16% accuracy reduction for FedMD. In contrast, our aggregation method shows a marginal 0.7% accuracy decrease under top-k attacks, outperforming state-of-the-art baselines.
Hanlin Gu, Sheng Wan, Zhirong Luan, Wei Xi 0003, Lixin Fan, Qiang Yang 0001, Badong Chen
ICDCS5
2024 MRFER: Multi-Channel Robust Feature Enhanced Fusion for Multi-Modal Emotion Recognition
abstract
In multi-modal emotion recognition, previous studies focus on obtaining more distinguishable unimodal features and expanding complementary information across modalities. However, a considerable amount of latent emotional information is neglected. It leads to insufficient intra-modal representations and a one-sided perspective on inter-modal relationship learning. To address these challenges, we propose a novel framework named MRFER, which explores strategies to reduce the loss of emotional information. It models robust unimodal features through a multi-path feature extractor and captures more comprehensive inter-modal relationships through a text-guided dual attention fusion module. Systematic evaluation covers generalization and overall performance, showcasing MRFER’s advancement beyond existing state-of-the-art approaches.
Xiao Fu 0001, Wei Xi 0003, Dianwen Ng, Jizhong Zhao
ICME2
2024 Attention Shifting to Pursue Optimal Representation for Adapting Multi-granularity Tasks
Gairui Bai, Wei Xi 0003, Yihan Zhao, Jizhong Zhao
IJCAI2
2024 CoPL: Parameter-Efficient Collaborative Prompt Learning for Audio-Visual Tasks
abstract
Parameter-Efficient Fine Tuning (PEFT) has been demonstrated to be effective and efficient for transferring foundation models to downstream tasks. Transferring pretrained uni-modal models to multi-modal downstream tasks helps alleviate substantial computational costs for retraining multi-modal models. However, existing approaches primarily focus on multi-modal fusion, while neglecting the modal-specific fine-tuning, which is also crucial for multi-modal tasks. To this end, we propose parameter-efficient Collaborative Prompt Learning (CoPL) to fine-tune both uni-modal and multi-modal features. Specifically, the collaborative prompts consist of modal-specific prompts and modal-interaction prompts. The modal-specific prompts are tailored for fine-tuning each modality, while the modal-interaction prompts are customized to explore inter-modality association. Furthermore, prompt bank-based mutual coupling is introduced to extract instance-level features, further enhancing the model's generalization ability. Extensive experimental results demonstrate that our approach achieves comparable or higher performance on various audio-visual downstream tasks while utilizing approximately 1% extra trainable parameters.
Yihan Zhao, Wei Xi 0003, Yuhang Cui, Gairui Bai, Jizhong Zhao
ACM Multimedia2
2024 Robust Contrastive Learning Against Audio-Visual Noisy Correspondence
Yihan Zhao, Wei Xi 0003, Gairui Bai, Jizhong Zhao
PRCV (5)2
2024 MiniPFL: Mini federations for hierarchical personalized federated learning
Wei Xi 0003, Hengyi Zhu, Jizhong Zhao
Future Gener. Comput. Syst.2
2024 Meta Generative Flow Networks with personalization for task-specific adaptation
Xinyuan Ji, Xu Zhang 0011, Wei Xi 0003, Haozhi Wang, Olga Gadyatskaya, Yinchuan Li
Inf. Sci.3
2024 Heterogeneous Interactive Graph Network for Audio-Visual Question Answering
Yihan Zhao, Wei Xi 0003, Gairui Bai, Jizhong Zhao
Knowl. Based Syst.2
2024 Genre Classification Empowered by Knowledge-Embedded Music Representation
abstract
This paper introduces a pioneering framework for music representation learning, which harnesses knowledge graph embeddings to enrich genre classification. Leveraging metadata from publicly available datasets like FMA and OpenMIC-2018, the constructed knowledge graph delineates intricate relationships among genres, artists, and instruments, offering valuable insights for genre representation. Within this framework, we propose two models tailored for distinct genre classification scenarios: fixed-set genre classification and open-set genre classification. These models exploit the knowledge graph to unveil correlations among different genres and integrate this knowledge into the audio representation. Notably, our approach is the first to merge audio data with high-level knowledge for music genre classification. Experimental results demonstrate that our proposed methods outperform state-of-the-art approaches, achieving an average genre classification accuracy of 68.07% on the FMA-medium dataset and 42.4% for open-set classification on the FMA-large dataset.
Han Ding 0002, Linwei Zhai, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003, Zhi Wang 0002, Jizhong Zhao
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Adaptive Client Clustering for Efficient Federated Learning Over Non-IID and Imbalanced Data
abstract
Federated learning (FL) is an emerging distributed and privacy-preserving machine learning framework. However, the performance of traditional FL methods is seriously impaired by the real-world data, which appear to be non-IID. The recent clustered federated learning (CFL) methods eliminate the impact of non-IID data by grouping clients with similar data distribution into the same cluster. Unfortunately, existing CFL methods heavily rely on the pre-setting of the cluster number, failing to achieve adaptive client clustering. We also experimentally observe that imbalanced data largely degrade their correctness of client clustering. In this paper, we present a novel CFL method without manual intervention, named AutoCFL, which can eliminate both effects of non-IID and imbalanced data simultaneously. To deal with imbalanced data, the local training adjustment strategy adaptively adjusts the number of local training epochs for each client. To further improve the clustering correctness and adaptability, the weighted voting-based client clustering strategy automatically groups each client into an appropriate cluster. Extensive experiments are conducted to evaluate the design of AutoCFL with three popular datasets under various data settings. Experimental results demonstrate that AutoCFL outperforms state-of-the-art methods, e.g., on average improving model accuracy by 9.24%, while reducing communication costs by 4.67 in an adaptive manner.
Biyao Gong, Tianzhang Xing, Zhidan Liu 0001, Wei Xi 0003, Xiaojiang Chen
IEEE Trans. Big Data4
2024 RoseAgg: Robust Defense Against Targeted Collusion Attacks in Federated Learning
abstract
Recent defense approaches against targeted model poisoning attacks aim to prevent specific prediction failures in federated learning (FL). However, these defenses remain susceptible to targeted collusion attacks, particularly under conditions of high proportions of malicious clients and attack density. To address these vulnerabilities, we propose RoseAgg, which dynamically identifies a plausible clean ingredient from local updates and leverages it to constrain the influence of poisoned updates. Firstly, RoseAgg recognizes and confines common characteristics found in poisoned updates, such as scaled-up magnitudes or similar directional contributions. Furthermore, RoseAgg dynamically extracts a plausible clean ingredient using a dimension-reduction method. This clean ingredient becomes the foundation for the server to bootstrap credit scores for each local update, ensuring the dominance of benign updates over poisoned ones. Ultimately, the server computes a weighted average of local updates based on credit scores, generating a global update for refining the global model. Comprehensive evaluations on four benchmark datasets showcase RoseAgg’s effectiveness against seven advanced attacks. The code is available athttps://github.com/SleepedCat/RoseAgg.
Wei Xi 0003, Yuhao Shen 0001, Canhui Wu, Jizhong Zhao
IEEE Trans. Inf. Forensics Secur.2
2024 Heartbeating With LTE Networks for Ambient Backscatter
abstract
Different from intermittent ISM signals like Bluetooth and WiFi, LTE signals are continuous in time and more pervasive in space, which makes them suitable carriers for ambient backscatter systems. However, due to the continuous LTE traffic and complex frame structures, existing ambient backscatter systems, such as HitchHike and LScatter, cannot reliably backscatter LTE signals in a standard-compatible way. We observe that the primary cause of their failures is that tags cannot accurately synchronize with LTE excitations. To address this issue, we propose SyncLTE, an LTE backscatter system that achieves high-accuracy synchronization and standard-compatible backscatter communication. The key novelty is a new tag design that uses the periodicity of LTE signals for synchronization and provides a customized single-symbol modulation scheme for LTE carriers. Our design is prototyped using FPGAs, SDR LTE eNBs and UEs. Comprehensive experiments have been done in various scenarios, including LoS, NLoS, indoor and outdoor. Results show that SyncLTE is 22.4x and 7.4x better than LScatter and Multiscatter in terms of the 80th percentiles of synchronization errors. Also, SyncLTE can deliver throughputs of up to 200 bps using BPSK and 400 bps using QPSK while other systems suffer from failures.
Yunyun Feng, Si Chen 0003, Wei Xi 0003, Shuai Wang 0008, Jia Zhao 0006, Wei Gong 0001
IEEE Trans. Mob. Comput.3
2024 Towards Hierarchical Clustered Federated Learning With Model Stability on Mobile Devices
abstract
Clustered federated learning (CFL) has proved to be an effective way to alleviate the non-IID (not independently and identically distributed) data challenge, which severely restricts the wider application of federated learning. However, existing approaches either lack adaptability,i.e., they require an additional number of clusters as a guide when clustering, or lack effectiveness in terms of communication. In this paper, we explore the differences in the ability of different layers in a model to represent non-IID data, and propose a hierarchical CFL approach, namedHiCFL, which considers both adaptivity and communication efficiency. The improvement of communication efficiency is due to our proposed novel concept of model stability, which characterizes the variation of model weights during training. Based on model stability,HiCFLcan find the proper time to bi-partition the clusters of mobile devices in a hierarchical manner more quickly. We conduct extensive experiments based on popular datasets with various non-IID data settings. The results show thatHiCFLachieves excellent performance effectiveness and efficiency. Compared to state-of-the-art approaches,HiCFLcan improve the model accuracy by$2.0\% \sim 9.0\%$, while reducing the communication overheads by$27.3\% \sim 80.6\%$.
Biyao Gong, Tianzhang Xing, Zhidan Liu 0001, Wei Xi 0003, Xiaojiang Chen
IEEE Trans. Mob. Comput.4
2024 Enabling Multi-Frequency and Wider-Band RFID Sensing Using COTS Device
abstract
RFID shows great potentials to build useful sensing applications. However, current RFID sensing can obtain mainly a single-dimensional sensing measurement from each reader-to-tag query, such as phase, RSS, etc. This is sufficient to fulfill the designs that are bound to the tag’s movement, e.g., the localization of tags. However, it imposes inevitable uncertainty on many sensing tasks relying on the features extracted from the RFID signals. These traditional sensing measurements limit the fidelity of RFID sensing fundamentally and prevent its broader usage in more sophisticated sensing scenarios. This paper presents RF-Wise to push the limit of RFID-based sensing, motivated by an insightful observation to customize RFID signals. RF-Wise can enrich the existing single-dimensional feature measure to a channel state information (CSI)-like measure with up to 150-dimensional samples across different frequencies concurrently. More importantly, RF-Wise is a software solution atop the standard EPC Gen2 protocol without using any extra hardware. It requires only one tag for sensing and works within the ISM band. RF-Wise, so far as we know, is the first system of such a kind. Extensive experiments show that RF-Wise does not impact underlying RFID communications, while by using the features extracted by RF-Wise, applications’ sensing performance can be improved remarkably. The source codes of RF-Wise are available at https://cui-zhao.github.io/RF-WISE/.
Cui Zhao, Zhenjiang Li 0001, Han Ding 0002, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao
IEEE/ACM Trans. Netw.5
2023 Knowledge-Graph Augmented Music Representation for Genre Classification
abstract
In this paper, we propose KGenre, a knowledge-embedded music representation learning framework for improved genre classification. We construct the knowledge graph from the metadata in the open-source FMA-medium and OpenMIC-2018 datasets, with no extra information/effort required. KGenre then mines the correlation between different genres from the knowledge graph and embeds such correlation in audio representation. To our knowledge, KGenre is the first method fusing the audio with high-level knowledge for music genre classification. Experimental results demonstrate the embedded knowledge can effectively enhance the audio feature representation, and the genre classification performance surpasses the state-of-the-art methods.
Han Ding 0002, Wenjing Song, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao
ICASSP6
2023 Dynamic Private Task Assignment under Differential Privacy
abstract
Data collection is indispensable for spatial crowd-sourcing services, such as resource allocation, policymaking, and scientific explorations. However, privacy issues make it challenging for users to share their information unless receiving sufficient compensation. Differential Privacy (DP) is a promising mechanism to release helpful information while protecting individuals’ privacy. However, most DP mechanisms only consider a fixed compensation for each user’s privacy loss. In this paper, we design a task assignment scheme that allows workers to dynamically improve their utility with dynamic distance privacy leakage. Specifically, we propose two solutions to improve the total utility of task assignment results, namely Private Utility Conflict-Elimination (PUCE) approach and Private Game Theory (PGT) approach, respectively. We prove that PUCE achieves higher utility than the state-of-the-art works. We demonstrate the efficiency and effectiveness of our PUCE and PGT approaches on both real and synthetic data sets compared with the recent distance-based approach, Private Distance Conflict-Elimination (PDCE). PUCE is always better than PDCE slightly. PGT is 50% to 63% faster than PDCE and can improve 16% utility on average when worker range is large enough.
Leilei Du 0001, Peng Cheng 0003, Libin Zheng 0001, Wei Xi 0003, Xuemin Lin 0001, Wenjie Zhang 0001
ICDE4
2023 Dual Acoustic Linguistic Self-supervised Representation Learning for Cross-Domain Speech Recognition
Dianwen Ng, Chong Zhang 0003, Xiao Fu 0001, Wei Xi 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001, Jizhong Zhao
INTERSPEECH6
2023 A Unified Recognition and Correction Model under Noisy and Accent Speech Conditions
Dianwen Ng, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH5
2023 Dual-Memory Multi-Modal Learning for Continual Spoken Keyword Spotting with Confidence Selection and Diversity Enhancement
Dianwen Ng, Xizhe Li, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH6
2023 WiHunter: Enabling Real-time Small Object Detection via Wireless Sensing
abstract
Rodent infestation is a great danger to human society, continuously threatening food safety and inducing disease spread. Existing methods to deal with rodent infestation are mainly based on passive bait traps and poisoning. These methods lack timeliness and effectiveness due to the missing of real-time detection. In this paper, we develop WiHunter, a new wireless sensing system to discover small objects (e.g., rat). Our idea is to exploit reflection signal effect of wireless channels induced by the movement of small objects around the receiver antenna. However, existing wireless sensing works usually employ customized or costly device-dependency Network Interface Cards(NIC), which are impractical to be densely deployed in reality. We implement WiHunter with several CSI-enabled standalone IoT nodes. We show how such devices enable moving small object detection via WiFi signal. The rationale behind this is 1) thanks to the widespread deployments of WiFi infrastructures and IoT devices, the WiFi signal covers almost every location of the corner, 2) the signal amplitude of each device is related to the small object near the receiver antenna. This ability gives us the opportunity to sense object as small as a rat. We implement WiHunter with ESP32 microcontroller on Espressif IoT Development Framework (both of them are cheap commodity off-the-shelf (COTS) devices) and design a practical small object intrusion detection system. Comprehensive and real-world experiments demonstrate that our system is effective in detecting the presence of small objects with an average accuracy of 92.1%.
Jianwei Liu 0008, Jinsong Han, Wei Xi 0003, Zhi Wang 0002
IWQoS4
2023 mmYodar: Lightweight and Robust Object Detection using mmWave Signals
abstract
The detection of human objects can be crucial for various real-world applications, such as surveillance and autonomous driving. However, traditional vision-based approaches suffer from limitations such as low lighting conditions, occlusions, and privacy concerns. To overcome these limitations, we propose a novel automatic object detection system, called mmYodar, which utilizes millimeter-wave (mmWave) radar signals. Our system collects mmWave signals and calculates a 3D point cloud, which is transformed into a radar image for easier visualization and analysis. To improve the system's human profiling capability, we expand the corresponding points in the image with color based on the radar angle resolution. Then, a designed deep mutual learning framework is employed to detect human objects from the expanded image. Experimental results show that mmYodar achieves nearly real-time detection with an average precision of 90.35% in various scenarios, including indoor and outdoor environments, various lighting conditions, and in the presence of occlusions. These results demonstrate the effectiveness of using mmWave radar signals for reliable and accurate human object detection. Our code and dataset are available at https:llgithub.comlbrave20005lmmYodar.
Yuance Chang, Han Ding 0002, Dachao Han, Ge Wang 0003, Cui Zhao, Fei Wang 0037, Wei Xi 0003, Jizhong Zhao
SECON8
2023 Concurrent Rate-Adaptive Reading With Passive RFIDs
abstract
Radio frequency identification (RFID)-assisted management systems have been widely applied in warehousing, logistics, retailing, etc. In these scenarios, RFID-aided applications, e.g., object tracking and human behavior sensing, rely on a high-efficiency tag reading to realize accurate analyses and timely responses. However, serious tag collisions in those large-scale RFID systems will inevitably lead to significant decreases in the tag reading rates. To meet the strict timeliness requirements of those practical applications, we aim to treat the individual reading rate for each item tag differently and focus more attention on those user-interactive ones. However, due to unpredictable user behaviors, it is impractical to infer the user-interactive tags in advance. In addition, keeping focusing on them for continuous monitoring despite user movements and multipath-prevalent environments is also challenging. To solve these problems, we propose Spotlight, the first concurrent rate-adaptive reading system in passive RFIDs. Spotlight screens the ID-agnostic user-interactive tags by proposing a multichannel feature for narrow-band RFID systems without any hardware or protocol modification and achieves rate-adaptive reading by implementing real-time MU-MIMO beamforming. Substantial experiments with 1000+ COTS RFID tags exhibit that Spotlight outperforms the commercial reader by$2.7\times $and the SDR-based reader by$6.12\times $. In addition, Spotlight first proposes the online parallel decoding method to realize concurrency among multiple users, which breaks the commercial protocol’s throughput ceiling (37%) and achieves up to 59% throughputs.
Ge Wang 0003, Shouqian Shi, Huazhe Wang, Yi Liu 0115, Chen Qian 0001, Cong Zhao 0006, Wei Xi 0003, Han Ding 0002, Zhiping Jiang, Jizhong Zhao
IEEE Internet Things J.7
2023 FedRich: Towards efficient federated learning for heterogeneous clients using heuristic scheduling
Wei Xi 0003, Yuhao Shen 0001, Xinyuan Ji, Cerui Sun, Jizhong Zhao
Inf. Sci.2
2023 Co-MDA: Federated Multisource Domain Adaptation on Black-Box Models
abstract
Federated domain adaptation (FDA) is an effective method for performing learning tasks over distributed networks, which well improves data privacy and portability in unsupervised multi-source domain adaptation (UMDA) tasks. Despite the impressive gains achieved, two common limitations exist in current FDA works. First, most previous studies require access to the model parameters or gradient details of each source party. However, the raw source data can be reconstructed from the model gradients or parameters, which may leak individual information. Second, these works assume that different parties share an identical network architecture, which is impractical and not desirable for low- or high-resource target users. To address these issues, in this work, we propose a more practical UMDA setting, called Federated Multi-source Domain Adaptation on Black-box Models (B2FDA), where all data are stored locally and only the input-output interface of the source model is available. To tackle B2FDA, we propose an effective method, termed Co2-Learning with Multi-Domain Attention (Co-MDA). Experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method. Notably, Co-MDA performs comparably with traditional UMDA methods where the source data or the trained model are fully available.
Wei Xi 0003, Wen Li 0001, Dong Xu 0001, Gairui Bai, Jizhong Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Generalized Method to Combat Multipaths for RFID Sensing
abstract
There have been increasing interests in exploring the sensing capabilities of RFID to enable numerous IoT applications, including object localization, trajectory tracking, and human behavior sensing. However, most existing methods rely on the signal measurement either in a low multipath environment, which is unlikely to exist in many practical situations, or with special devices, which increase the operating cost. This paper investigates the possibility of measuring ‘multi-path-free’ signal information in multipath-prevalent environments simply using a commodity RFID reader. The proposed solution, Clean Physical Information Extraction (CPIX), is universal, accurate, and compatible to standard protocols and devices. CPIX improves RFID sensing quality with near zero cost – it requires no extra device. We implement CPIX and study three major RFID sensing applications: tag localization, device calibration and human behavior sensing. CPIX reduces the localization error by 30% to 50% and achieves the MOST accurate localization by commodity readers compared to existing work. It also significantly improves the quality of device calibration and human behaviour sensing.
Ge Wang 0003, Haofan Cai, Chen Qian 0001, Han Ding 0002, Wei Xi 0003, Kun Zhao 0002, Jizhong Zhao, Jinsong Han
IEEE/ACM Trans. Netw.6
2022 UTIO: Universal, Targeted, Imperceptible and Over-the-air Audio Adversarial Example
abstract
The audio adversarial example has been demonstrated to be an effective attack which leads to prediction errors of the intelligent voice control system (e.g., deep neural network based speech recognition service), despite resembling a valid input to our human beings. An ideal adversarial example attack should have four major advantages, including 1) utilizing a universal adversarial perturbation against arbitrary voice commands, 2) tricking a model to get an incorrect and targeted result, 3) imperceptible to users even in a silent place and 4) validating in an over-the-air (OTA) scenario as well. However, existing studies mainly involve several but not all of these criteria. In this paper, we propose UTIO, a universal, targeted, imperceptible and OTA audio adversarial example design, which leverages one perturbation to fool a speech recognition model in OTA scenarios. Moreover, a variety of speeches can be misled to a targeted threat command imperceptibly. To harvest such benefits, we leverage two targeted loss functions to generate adversarial perturbations, and employ the psychoacoustic principle to further conceal the attack. Finally, we actively embed additional distortions, occurred during the physical propagation, in the process of perturbation generation to make UTIO still valid in an OTA scenario. Extensive experiments show that UTIO can perform 94.15% success attack rate locally, i.e., without physical propagation, while retaining 93.44% attack rate in an OTA scenario. In addition, three types of defensive strategies are also introduced to resist against our attack.
Cui Zhao, Zhenjiang Li 0001, Han Ding 0002, Wei Xi 0003
ICPADS4
2022 ACFNet: An Adaptive Context Fusion Network for Skin Lesion Segmentation
abstract
Skin lesion segmentation is a key step in computer-aided diagnosis. Fully convolution network-based approaches have achieved great results. However, inadequate feature extraction has lost many important features as well as rough fusion has introduced many adverse noises. All of these limit the accuracy of skin lesion segmentation. In this paper, we propose a new and effective adaptive context fusion network for skin lesion segmentation. The proposed network is based on the U-Net architecture. An adaptive context fusion module (ACF) is used to extract more adequate feature. Combining the advantages of atrous convolution and attention mechanism, it can not only reduce the loss of spatial information, but also further refine the extracted feature. The gated residual fusion module (GRF) is added to skip architecture to make the fusion process more refined. It can suppress invalid information in the fusion process with less spatial information loss. We evaluate the proposed method on two benchmark datasets: ISIC 2016 and ISIC 2017, and the experimental results show that the proposed method achieves significant accuracy improvement compared with the existing several networks. The ablation experiments also prove the effectiveness of the proposed two modules.
Jingtong Sun, Wei Xi 0003, Gairui Bai
IJCNN2
2022 Learning Multi-Level Consistency for Noisy Labels
abstract
Recent methods performing well on Learning with Noisy Label (LNL) problem generally are based on semi-supervised learning and consistency regularization. It usually consists of three stages: warm-up, noisy/clean data division, and semi-supervised learning. However, these methods trained purely with classification consistency suffer from the confirmation bias problem and tend to memorize the noisy labels, resulting in accumulated error and degraded performance. Leveraging the compositional and relational peculiarities of the noisy data, we propose a graph-based Multi-Level Consistency (MLC) framework that jointly exploits multi-level relation consistencies between graphs and classification consistency which can better correct wrong labels by continuing to learn the multi-level differences between clean data and noisy data. Moreover, we propose a Dynamic Filter Module (DFM) which effectively improves the reliability of divided data by re-filtering noisy data despite its simplicity. Our method achieves the state-of-the-art performance on multiple benchmark datasets. On Cifar-100 with 90% noisy labels, our method achieves a top-1 accuracy of 49.1%, outperforming DivideMix by 17.6%.
Ziye Tong, Wei Xi 0003
IJCNN2
2022 LRTD: A Low-rank Transformer with Dynamic Depth and Width for Speech Recognition
abstract
Though Transformer-based models have achieved great success in the automatic speech recognition (ASR) field, they are generally resource-hungry and computation-intensive which makes them difficult to deploy in resource-restricted devices. In this paper, we propose LRTD, a lightweight Transformer for end-to-end ASR. LRTD compresses the model size using matrix decomposition during training and further reduces the depth and width of the trained model dynamically leveraging two structured pruning strategies. The performance of the pruned model reply on two additional structured dropout methods and two search techniques that can recognize the important layers and attention heads. The experimental results show that our proposed model can achieve competitive performance on Aishell-1 even with about 2.27 × fewer model parameters compared to the baseline transformer model. Moreover, we experimentally demonstrate that the matrix decomposition technique achieves a higher compression rate while two pruning methods can adjust the model size flexibly.
Wei Xi 0003, Ziye Tong, Jingtong Sun
IJCNN2
2022 Unsupervised Multimodal Image-to-Image Translation: Generate What You Want
abstract
Unsupervised image-to-image translation is one of the most important research topics in the field of computer vision. Traditional unsupervised image-to-image translation methods can only learn one-to-one mapping relation between two domains. However, the target domain is multimodal in general. Therefore, some existing methods generate diverse and multimodal outputs for a given source image by introducing random style codes or noise vectors. But they can only generate images with random modal of target domain. In some scenarios, we expect the generated images have desired modal. To solve this problem, we propose a novel framework for unsupervised multimodal image-to-image translation, called MuGAN, which can learn one-to-many mapping relations from source domain to target domain. And by controlling the modal label of target domain, our approach can generate images with desired modal. We verify the effectiveness and superiority of our framework by qualitative and quantitative evaluations on several image datasets.
Wei Xi 0003, Gairui Bai, Jingtong Sun
IJCNN2
2022 RF-Wise: Pushing the Limit of RFID-based Sensing
abstract
RFID shows great potentials to build useful sensing applications. However, current RFID sensing can obtain mainly a single-dimensional sensing measurement from each reader-to-tag query, such as phase, RSS, etc. This is sufficient to fulfill the designs that are bounded to the tag’s own movement, e.g., the localization of tags. However, it imposes inevitable uncertainty to many sensing tasks relying on the features extracted from the RFID signals, which limits the fidelity of RFID sensing fundamentally and prevents its broader usage in more sophisticated sensing scenarios. This paper presents RF-Wise to push the limit of the RFID-based sensing, motivated by an insightful observation to customize RFID signals. RF-Wise can enrich the existing single-dimensional feature measure to a channel state information (CSI)-like measure with up to 150 dimensional samples across different frequencies concurrently. More importantly, RF-Wise is a software solution atop the standard EPC Gen2 protocol without using any extra hardware, requires only one tag for sensing and works within the ISM band. RF-Wise, so far as we know, is the first system of such a kind. Extensive experiments show that RF-Wise does not impact underlying RFID communications, while by using the features extracted by RF-Wise, applications’ sensing performance can be improved remarkably. The source codes of RF-Wise are available at https://cui-zhao.github.io/RF-WISE/.
Cui Zhao, Zhenjiang Li 0001, Han Ding 0002, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao
INFOCOM5
2022 ScreenInformer: Whispering Secret Information via an LCD Screen
abstract
In this paper, we observe an acoustic covert channel by modulating capacitor squeal on the monitor's power supply unit, and then present ScreenInformer to build a covert communication within a physically isolated network system. Unlike traditional electromagnetic side channels, capacitor squeal is usually not covered by safety shielding measures. To precisely modulate the capacitor squeal, we reveal the relationships among the displayed content on the screen, voltage variation on the power unit, and the frequency of acoustic leakage. In order to improve the demodulating ability for acoustic side-channel leakage with a low signal-to-noise ratio (SNR), we design a cross-correlation demodulation algorithm for the rich harmonics of leakage. An off-the-shelf mobile phone can support ScreenIn-former for exfiltrating sensitive information under ambient noise of up to 55 dB. Our various real-world experimental results show that ScreenInformer can achieve a communication distance of up to 130 cm and a maximum throughput of 170 bps. In addition, our observed capacitor squeal is widespread, which has the potential to enable common electronic devices to have communication capability.
Xieyang Sun, Wei Xi 0003, Zhiping Jiang, Zuhao Chen
SECON2
2022 High-efficient hierarchical federated learning on non-IID data with progressive collaboration
Yunyun Cai, Wei Xi 0003, Yuhao Shen 0001, Youcheng Peng, Shixuan Song, Jizhong Zhao
Future Gener. Comput. Syst.2
2022 M2N: Mutual constraint network for multi-level unsupervised domain adaptation
Wei Xi 0003, Gairui Bai, Zhilin Liu, Jizhong Zhao
Neurocomputing2
2022 Utilizing Tag Interference for Refined Localization of Passive RFID
abstract
We study a new problem, refined localization, in this article. Refined localization calculates the location of an object in high precision, given that the object is in a relatively small region such as the surface of a table. Refined localization is useful in many cyber–physical systems such as industrial autonomous robots. Existing vision-based approaches suffer from several disadvantages, including good lighting conditions, line of sight, prelearning process, and high computation overhead. Also, vision-based approaches cannot differentiate objects with similar colors and shapes. This article presents a new refined localization system, called Trio, which uses passive radio frequency identification (RFID) tags for low cost and easy deployment. Trio utilizes RF interference for tag localization by modeling the equivalent circuits of coupled tags. We implement our prototype using commercial off-the-shelf RFID reader and tags. Extensive experiment results demonstrate that Trio effectively achieves high accuracy of refined localization, i.e., < 1 cm errors for several types of main stream tags.
Han Ding 0002, Cui Zhao, Ge Wang 0003, Kun Zhao 0002, Wei Xi 0003, Jizhong Zhao
IEEE Internet Things J.5
2022 Eliminating the Barriers: Demystifying Wi-Fi Baseband Design and Introducing the PicoScenes Wi-Fi Sensing Platform
abstract
The research on Wi-Fi sensing has been thriving over the past decade but the process has not been smooth. Three barriers always hamper the research: 1) unknown baseband design and its influence; 2) inadequate hardware; and 3) the lack of versatile and flexible measurement software. This article tries to eliminate these barriers through the following work.First, we present an in-depth study of the baseband design of the Qualcomm Atheros AR9300 (QCA9300) NIC. We identify a missing item of the existing channel state information (CSI) model, namely, the CSI distortion, and identify the baseband filter as its origin. We also propose a distortion removal method.Second, we reintroduce both the QCA9300 and software-defined radio (SDR) as powerful hardware for research. For the QCA9300, we unlock the arbitrary tuning of both the carrier frequency and bandwidth. For SDR, we develop a high-performance software implementation of the 802.11a/g/n/ac/ax baseband, allowing users to fully control the baseband and access the complete physical-layer information.Third, we release the PicoScenes software, which supports concurrent CSI measurement from multiple QCA9300, Intel Wireless Link (IWL5300), and SDR hardware. PicoScenes features rich low-level controls, packet injection, and software baseband implementation. It also allows users to develop their own measurement plugins.Finally, we report state-of-the-art results in the extensive evaluations of the PicoScenes system, such as the >2-GHz available spectrum on the QCA9300, concurrent CSI measurement, and up to 40 and 1 kHz CSI measurement rates achieved by the QCA9300 and SDR. PicoScenes is available athttps://ps.zpj.io.
Zhiping Jiang, Tom H. Luan, Xincheng Ren, Dongtao Lv, Kun Zhao 0002, Wei Xi 0003, Yueshen Xu, Rui Li 0047
IEEE Internet Things J.8
2022 Wisual: Indoor Crowd Density Estimation and Distribution Visualization Using Wi-Fi
abstract
Driven by the Internet of Things (IoT), many device-free crowd density estimation techniques can roughly estimate the crowd density based on the relationship between the dynamic crowd and the variation of wireless signals. However, they cannot distinguish the path information of different persons in a fine-grained manner. In this article, we propose Wisual, a channel state information (CSI)-based device-free crowd density estimation framework and can visualize the distribution of people. The major challenge of Wisual is how to extract proper quantifiable indexes to distinguish the path information of multiple targets and maximize the resolution of crowd density estimation. To address this challenge, Wisual first presents a method for estimating the frequency of the CSI propagation path (FoC) for the moving persons and constructs a joint multifeature parameter (JMFP) spectrum matrix with the other two parameters. Then the multitarget spectrum matrix is put into a proposed deep-learning model called CSI stream 3-D convolutional neural networks (CS-3DCNNs) for implementing crowd density estimation and the target path information differentiation. Finally, Wi-Fi imaging is implemented based on the 2-D-MUSIC algorithm, which shows the approximate distribution situation of indoor persons through the spectrograms. The experimental results in typical real-world scenes demonstrate that Wisual can forecast the crowd density with 98% precision and accurately display the frequency spectra of moving persons. Besides, the results also prove the superior effectiveness, scalability, and generalizability of the proposed framework.
Wei Xi 0003, Zuhao Chen, Jizhong Zhao
IEEE Internet Things J.2
2022 Arbitrator2.0: Preventing Unauthorized Access on Passive Tags
abstract
As the ultra high frequency (UHF) passive radio frequency identification (RFID) technology becomes increasingly deployed, it faces an array of new security attacks. In this paper, we consider a type of attack in which a malicious RFID reader could arbitrarily access the tags, e.g., retrieve or modify IDs or other data in the memory, via standard commands. To deal with this type of attack, we propose a physical-layer tag protection framework, namely Arbitrator2.0, that involves two operating mode, i.e., one is to passively listen on RF channels and identify unauthorized readers, the other is working as normal reader to access tag information but resilient to one-antenna eavesdropper. Our solution does not need to modify RFID tags or the underlying communication standards. In this study, we have implemented a prototype Arbitrator2.0 over the universal software radio peripheral (USRP) platform, and conducted extensive experiments to evaluate its performance. The results show that Arbitrator2.0 can effectively diminish the unauthorized access attacks and prevent eavesdropping.
Han Ding 0002, Jinsong Han, Cui Zhao, Ge Wang 0003, Wei Xi 0003, Zhiping Jiang, Jizhong Zhao
IEEE Trans. Mob. Comput.5
2022 A Fingertip Profiled RF Identifier
abstract
This paper presents RF-Mehndi, a passive commercial RFID tag array formed identifier. The key RF-Mehndi novelty is that when the user’s fingertip touching on the tag array surface during the communication, the backscattered signals by the tag array become user-dependent and unique. Hence, if we enhance the communication modality of many personal cards nowadays by RF-Mehndi, in case that a card gets lost or stolen, it cannot be used illegally by the adversaries. To harvest such a benefit, we leverage two key observations in designing RF-Mehndi. The first one is when tags are nearby, their interrogated currents can change each other’s circuit characteristics, based on which unique phase features can be obtained from backscattered signals. The second observation is that when the user’s fingertip touches the tag array surface during communication, the phase feature can be further profiled by this user. Based on these observations, the card and its holder can be potentially authenticated at the same time. To transfer the RF-Mehndi idea to a practical system, we further address technical challenges. We implement a prototype system. Extensive evaluations show the effectiveness of RF-Mehndi, achieving excellent authentication performance.
Cui Zhao, Zhenjiang Li 0001, Han Ding 0002, Wei Xi 0003, Ruowei Gui, Jinsong Han
IEEE Trans. Mob. Comput.4
2021 Human Motion Recognition Based on Wi-Fi Imaging
Liangliang Lin, Kun Zhao 0002, Wei Xi 0003, Jizhong Zhao
CollaborateCom (1)4
2021 Measuring and Modeling Multipath of Wi-Fi to Locate People in Indoor Environments
abstract
With the rapid development of the Internet of Things (IoT) technology, the position information of indoor people has become an indispensable factor in most fields. Most existing indoor positioning schemes require people to keep moving to detect significant variance of the signal as the location feature. Hence, this paper proposes a passive indoor positioning system based on commodity Wi-Fi called Wisite, which can implement indoor multipath signal measurement and static person positioning modeling. The biggest challenge is how to detect the dynamic features in the reflection path of the static person to achieve target path matching. To address this issue, Wisite proposes a MUSIC expectation-maximization (MEM) joint parameter estimation algorithm to estimate and enhance the indoor multipath parameters. Then, a dynamic path matching model based on signal change enhancement (SCE) is proposed to enhance the signal changes caused by human activities, which can amplify the weak signal changes introduced by human respiration when a person is in a static state. Finally, the multipath geometric positioning model is used to calculate the person's position. We implement Wisite using commercial off-the-shelf (COTS) IEEE 802.11n devices and evaluate its performance via extensive experiments in typical real-world scenes. The results show that Wisite outperforms the comparison approaches in estimating accuracy and effectiveness with the average indoor positioning error is less than 0.65cm.
Wei Xi 0003, Zuhao Chen, Jizhong Zhao
ICPADS4
2021 Audio DistilBERT: A Distilled Audio BERT for Speech Representation Learning
abstract
Self-supervised speech representation learning has been considered as an outstanding manner to improve the performance of downstream tasks. However, those models are often too cumbersome, which sets a barrier to deploy them on the edge and improves the threshold of the pre-training process. In this paper, we propose Audio DistilBERT, a distilled BERT-style speech representation learning method. It learns dark knowledge from a larger teacher model through one new designed loss which combines soft and hard targets. By doing this, it can achieve competitive performance with fewer parameters and faster inference time. The experimental results among two downstream tasks show that the proposed method can retain above 98% performance of the large model with about 1.8× smaller model size and over 1.6× faster inference speed. In a low-resource environment with very few labeled data and pretraining steps, our model also exhibits similar or even better performance compared to the large model. Furthermore, we explore the knowledge transfer competence between the teacher and student model.
Wei Xi 0003
IJCNN3
2021 KEEP: Secure and Efficient Communication for Distributed IoT Devices
abstract
Security over mobile Internet-of-Things (IoT) devices is critical due to the open nature of distributed wireless communication. To efficiently establish a secure connection between two communication parties, a fast mobile key extraction protocol, KEEP, is proposed. KEEP fastly generates similar bit sequences from two communication parties’ measurements of channel-state information (CSI) of different subcarriers. Then, a distributed “verification-recombination” mechanism is introduced to generate the same encryption key from bit sequences without the public-key authentication, digital signature, or key distribution center of the other party. We implemented real-world experiments using commercial off-the-shelf 802.11n devices to evaluate the performance of KEEP in various scenarios. Theoretical analysis and experimental verification show that KEEP is more secure, effective, and reliable than the state-of-the-art methods.
Wei Xi 0003, Meichen Duan, Xiuxiu Bai, Kun Zhao 0002, Lufeng Mo, Jizhong Zhao
IEEE Internet Things J.1
2021 Indoor Geofencing Based on Sensorless Motion Sensing and Fingerprint Self-Updating
Kun Zhao 0002, Wei Xi 0003, Zhiping Jiang, Zhi Wang 0002, Jizhong Zhao
Mob. Networks Appl.2
2021 Corrections to "HMO: Ordering RFID Tags With Static Devices in Mobile Environments"
abstract
Presents corrections to the acknowledgement section for the above named article.
Ge Wang 0003, Chen Qian 0001, Longfei Shangguan, Han Ding 0002, Jinsong Han, Kaiyan Cui, Wei Xi 0003, Jizhong Zhao
IEEE Trans. Mob. Comput.7
2020 TAB: CSI Lossless Compression for MU-MIMO Network
Qigui Xu, Wei Xi 0003, Lubing Han, Kun Zhao 0002
CollaborateCom (1)2
2020 Speech2Stroke: Generate Chinese Character Strokes Directly from Speech
Yinhui Zhang, Wei Xi 0003, Sitao Men, Jizhong Zhao
CollaborateCom (1)2
2020 MufiNet: Multiscale Fusion Residual Networks for Medical Image Segmentation
abstract
U-Net has been considered as an outstanding deep learning neural network in medical image segmentation problems. The segmentation results of the U-Net based model, however, are always too conservative and smooth. MufiNet, a segmentation model using multiple U-Net chains (with multiple encoder-decoder branches), is proposed in this paper. It can fuse the receptive fields obtained from different scales. The convolution layer of 1 × 1 is introduced to add the residual connection to enhance the adaptability to the depth of the network. The multi-scale fusion module with residuals is combined with the U-Net chain architecture to retain more information flow paths, and the multi-scale context information is used to improve the performance and robustness of the segmented network. MufiNet model is extensively evaluated on three datasets in this paper, including two benchmark datasets (lung segmentation and skin cancer lesion segmentation) and cervical cancer dataset jointly constructed with a hospital. The experimental results show that MufiNet could yield better performance in medical image segmentation tasks than U-Net and LadderNet models.
Zhi Wang 0002, Wei Xi 0003, Gairui Bai, Ruimeng Wang, Meichen Duan
IJCNN3
2020 A Universal Method to Combat Multipaths for RFID Sensing
abstract
There have been increasing interests in exploring the sensing capabilities of RFID to enable numerous IoT applications, including object localization, trajectory tracking, and human behavior sensing. However, most existing methods rely on the signal measurement either in a low multipath environment, which is unlikely to exist in many practical situations, or with special devices, which increase the operating cost. This paper investigates the possibility of measuring `multi-path-free' signal information in multipath-prevalent environments simply using a commodity RFID reader. The proposed solution, Clean Physical Information Extraction (CPIX), is universal, accurate, and compatible to standard protocols and devices. CPIX improves RFID sensing quality with near zero cost - it requires no extra device. We implement CPIX and study two major RFID sensing applications: tag localization and human behavior sensing. CPIX reduces the localization error by 30% to 50% and achieves the MOST accurate localization by commodity readers compared to existing work. It also significantly improves the quality of human behaviour sensing.
Ge Wang 0003, Chen Qian 0001, Kaiyan Cui, Han Ding 0002, Wei Xi 0003, Jizhong Zhao, Jinsong Han
INFOCOM6
2020 RFnet: Automatic Gesture Recognition and Human Identification Using Time Series RFID Signals
Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Zhiping Jiang, Wei Xi 0003, Jizhong Zhao
Mob. Networks Appl.7
2020 HMO: Ordering RFID Tags with Static Devices in Mobile Environments
abstract
Passive Radio Frequency Identification (RFID) tags have been widely applied in many applications, such as logistics, retailing, and warehousing. In many situations, the order of objects is more important than their absolute locations. However, state-of-art ordering methods need a continuing movement of tags and readers, which limit the application domain and scalability. In this paper, we propose a 2-dimension ordering approach for passive tags that requires no device movement. Instead, our method utilizes signal changes caused by arbitrary movement of human beings around tags, who carry no device for horizontal dimension ordering. Hence, our method is called Human Movement based Ordering (HMO). The basic idea of HMO is that when people pass between the reader antenna and tags, the received signal strength will change. By observing the time-series RSS changes of tags, HMO can obtain the order of tags along with a specific horizontal direction. For vertical dimension, we employ a linear programming method that is tolerant of tiny errors in practice. We implement HMO with commodity off-the-shelf RFID devices. The experimental results show that HMO can achieve up to 88.71 and 90.86 percent average accuracies in the signal-and multi-person cases, respectively.
Ge Wang 0003, Chen Qian 0001, Longfei Shangguan, Han Ding 0002, Jinsong Han, Kaiyan Cui, Wei Xi 0003, Jizhong Zhao
IEEE Trans. Mob. Comput.7
2020 Hu-Fu: Replay-Resilient RFID Authentication
abstract
We provide the first solution to an important question, “how a physical-layer authentication method can defend against signal replay attacks”. It was believed that if an attacker can replay the exact same reply signal of a legitimate authentication object (such as an RFID tag), any physical-layer authentication method will fail. This paper presents Hu-Fu, the first physical layer RFID authentication protocol that is resilient to the major attacks including tag counterfeiting, signal replay, signal compensation, and brute-force feature reply. Hu-Fu is built on two fundamental ideas, namely inductive coupling of two tags and signal randomization. Hu-Fu does not require any hardware or protocol modification on COTS passive tags and can be implemented with COTS devices. We implement a prototype of Hu-Fu and demonstrate that it is accurate and robust to device diversity and environmental changes, including locations, distance, and temperature. Hu-Fu provides a new direction of battery-free/low-power device authentication that enables numerous IoT applications.
Ge Wang 0003, Haofan Cai, Chen Qian 0001, Jinsong Han, Shouqian Shi, Xin Li 0057, Han Ding 0002, Wei Xi 0003, Jizhong Zhao
IEEE/ACM Trans. Netw.8
2020 dWatch: A Reliable and Low-Power Drowsiness Detection System for Drivers Based on Mobile Devices
abstract
Drowsiness detection is critical to driver safety, considering thousands of deaths caused by drowsy driving annually. Professional equipment is capable of providing high detection accuracy, but the high cost limits their applications in practice. The use of mobile devices such as smart watches and smart phones holds the promise of providing a more convenient, practical, non-invasive method for drowsiness detection. In this article, we propose a real-time driver drowsiness detection system based on mobile devices, referred to as dWatch, which combines physiological measurements with motion states of a driver to achieve high detection accuracy and low power consumption. Specifically, based on heart rate measurements, we design different methods for calculating heart rate variability (HRV) and sensing yawn actions, respectively, which are combined with steering wheel motion features extracted from motion sensors for drowsiness detection. We also design a driving posture detection algorithm to control the operation of the heart rate sensor to reduce system power consumption. Extensive experimental results show that the proposed system achieves a detection accuracy up to 97.1% and reduces energy consumption by 33%.
Tianzhang Xing, Qing Wang 0024, Chase Qishi Wu, Wei Xi 0003, Xiaojiang Chen
ACM Trans. Sens. Networks4
2019 Wi-Fi Imaging Based Segmentation and Recognition of Continuous Activity
Yang Zi, Wei Xi 0003, Kun Zhao 0002, Zhi Wang 0002
CollaborateCom2
2019 Accurate CSI Estimation to Eliminate Unnecessary Transmission for MU-MIMO Networks
Wei Xi 0003, Qigui Xu, Kun Zhao 0002, Yuanhang Cai
EWSN2
2019 A (Near) Zero-cost and Universal Method to Combat Multipaths for RFID Sensing
abstract
There have been increasing interests in exploring the sensing capabilities of RFID to enable numerous IoT applications, including object localization, trajectory tracking, and human behavior sensing. However, most existing methods rely on the signal measurement either in a low multipath environment, which is unlikely to exist in many practical situations, or with special devices, which increase the operating cost. This paper investigates the possibility of measuring `multipath-free' signal information in multipath-prevalent environments simply using a commodity RFID reader. The proposed solution, Clean Physical Information Extraction (CPIX), is universal, accurate, and compatible to standard protocols and devices. CPIX improves RFID sensing quality with near zero cost - it requires no extra device. We implement CPIX and evaluate its effectiveness on improving the performance on tag localization. The results show that CPIX reduces the localization error by 30% to 50% and achieves the MOST accurate localization by commodity readers compared to existing work.
Ge Wang 0003, Chen Qian 0001, Kaiyan Cui, Han Ding 0002, Haofan Cai, Wei Xi 0003, Jinsong Han, Jizhong Zhao
ICNP6
2019 RF-Mehndi: A Fingertip Profiled RF Identifier
abstract
This paper presents RF-Mehndi, a passive commercial RFID tag array formed identifier. The key RF-Mehndi novelty is that when the user's fingertip touching on the tag array surface during the communication, the backscattered signals by the tag array become user-dependent and unique. Hence, if we enhance the communication modality of many personal cards nowadays by RF-Mehndi, in case that a card gets lost or stolen, it cannot be used illegally by the adversaries. To harvest such a benefit, we have two key observations in designing RF-Mehndi. The first observation is when tags are nearby, their interrogated currents can change each other's circuit characteristics, based on which unique phase features can be obtained from backscattered signals. The second observation is that when the user's fingertip touches the tag array surface during communication, the phase feature can be further profiled by this user. Based on these observations, the card and its holder can be potentially authenticated at the same time. To transfer the RF-Mehndi idea to a practical system, we further address technical challenges. We implement a prototype system. Extensive evaluations show the effectiveness of RF-Mehndi, achieving excellent authentication performance.
Cui Zhao, Zhenjiang Li 0001, Han Ding 0002, Jinsong Han, Wei Xi 0003, Ruowei Gui
INFOCOM6
2019 Close-Proximity Detection for Hand Approaching Using Backscatter Communication
abstract
Smart environments and security systems require automatic detection of human behaviors including approaching to or departing from an object. Existing human motion detection systems usually require human beings to carry special devices, which limits their applications. In this paper, we present a system called APID to detect hand approaching behaviors by analyzing backscatter communication signals from a passive RFID tag on the object. APID does not require human beings to carry any device. The idea is based on the influence of hand movements to the vibration of backscattered tag signals. APID is compatible with commodity off-the-shelf devices and the EPCglobal Class-1 Generation-2 protocol. In APID, a commercial RFID reader continuously queries tags through emitting RF signals and tags simply respond with their IDs. A USRP monitor passively analyzes the communication signals and reports the approach and departure behaviors. We have implemented the APID system for both single-object and multi-object scenarios. Extensive evaluations demonstrate that APID can achieve high detection accuracy in both scenarios.
Han Ding 0002, Chen Qian 0001, Jinsong Han, Jian Xiao 0002, Xingjun Zhang, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao
IEEE Trans. Mob. Comput.7
2019 Counting Human Objects Using Backscattered Radio Frequency Signals
abstract
In this paper, we propose a system called R# to estimate the number of human objects using passive RFID tags but without attaching anything to human objects. The idea is based on our observation that the more human objects are present, the higher the variation in the RSS values of the tag backscattered RF signals. Thus, based on the received RF signals, the reader can estimate the number of human objects. R# includes an RFID reader and some (say 20) passive tags, which are deployed in the region that we want to monitor the number of human objects, such as the region in front of a supermarket shelf. The RFID reader periodically emits RF signals to identify all tags and the tags simply respond with their IDs via EPCglobal Class 1 Generation 2 protocol. We implemented R# using commercial Impinj H47 passive RFID tags and Impinj reader model R420. We conducted experiments in a simulated picking aisle area of the supermarket environment. The experimental results show that R# can achieve high estimation accuracy (more than 90 percent with up to ten human objects).
Han Ding 0002, Jinsong Han, Alex X. Liu, Wei Xi 0003, Jizhong Zhao, Panlong Yang, Zhiping Jiang
IEEE Trans. Mob. Comput.4
2019 Verifiable Smart Packaging with Passive RFID
abstract
Smart packaging adds sensing abilities to traditional packages. This paper investigates the possibility of using RF signals to test the internal status of packages and detect abnormal internal changes. Towards this goal, we design and implement a nondestructive package testing and verification system using commodity passive RFID systems, called Echoscope. Echoscope extracts unique features from the backscatter signals penetrating the internal space of a package and compares them with the previously collected features during the check-in phase. The use of backscatter signals guarantees that there is no difference in RF sources and the features reflecting the internal status will not be affected. Compared to other nondestructive testing methods such as X-ray and ultrasound, Echoscope is much cheaper and provides ubiquitous usage. Our experiments in practical environments show that Echoscope can achieve very high accuracy and is very sensitive to various types abnormal changes.
Ge Wang 0003, Jinsong Han, Chen Qian 0001, Wei Xi 0003, Han Ding 0002, Zhiping Jiang, Jizhong Zhao
IEEE Trans. Mob. Comput.4
2018 Trio: Utilizing Tag Interference for Refined Localization of Passive RFID
abstract
We study a new problem, refined localization, in this paper. Refined localization calculates the location of an object in high precision, given that the object is in a relatively small region such as the surface of a table. Refined localization is useful in many cyber-physical systems such as industrial autonomous robots. Existing vision-based approaches suffer from several disadvantages, including good lighting conditions, line of sight, pre-learning process, and high computation overhead. Also vision-based approaches cannot differentiate objects with similar colors and shapes. This paper presents a new refined localization system, called Trio, which uses passive Radio Frequency Identification (RFID) tags for low cost and easy deployment. Trio provides a new angle to utilize RF interference for tag localization by modeling the equivalent circuits of coupled tags. We implement our prototype using commercial off-the-shelf RFID reader and tags. Extensive experiment results demonstrate that Trio effectively achieves high accuracy of refined localization, i.e., <; 1 cm errors for several types of main stream tags.
Han Ding 0002, Jinsong Han, Chen Qian 0001, Fu Xiao 0001, Ge Wang 0003, Wei Xi 0003, Jian Xiao 0002
INFOCOM7
2018 Preventing Unauthorized Access on Passive Tags
abstract
As the Ultra High Frequency (UHF) passive Radio Frequency IDentification (RFID) technology becomes increasingly deployed, it faces an array of new security attacks. In this paper, we consider a type of attack in which a malicious RFID reader could arbitrarily modify the tags via standard commands, e.g., IDs or other data in the memory. To deal with this type of attack, we propose a physical-layer RF signal based reader authentication solution, namely Arbitrator, that involves passively listening on RF channels, analyzing the communication signals, identifying unauthorized readers and jamming the commands from such readers. Our solution does not need to modify RFID devices or the underlying communication standards, hence fully compatible with the existing RFID infrastructure. In this study, we have implemented a prototype Arbitrator over the Universal Software Radio Peripheral (USRP) platform, and conducted extensive experiments to evaluate its performance. Our results show that Arbitrator can detect unauthorized RFID readers with high accuracy, and thus effectively diminish the unauthorized access attacks.
Han Ding 0002, Jinsong Han, Yanyong Zhang, Fu Xiao 0001, Wei Xi 0003, Ge Wang 0003, Zhiping Jiang
INFOCOM5
2017 RFIPad: Enabling Cost-Efficient and Device-Free In-air Handwriting Using Passive Tags
abstract
An important function of smart environments is the ubiquitous access of computing devices. In public areas such as hospitals, libraries, and airports, people may want to interact with nearby computing systems to get information, such as directions to a hospital room, locations of books, and flight departure/arrival information. Touch screen based displays and kiosks, which are commonly used today, may incur extra hardware cost or even possible germ and bacteria infection. This work provides a new solution: users can make queries and inputs by performing in-air handwriting to an array of passive RFID tags, named RFIPad. This input method does not require human hands to carry any device and hence is convenient for applications in public areas. Besides the mobile and contactless property, this system is a cost-efficient extension to current RFID systems: an existing reader can monitor multiple RFIPads while performing its regular applications such as identification and tracking. We implement a prototype of RFIPad using commercial off-the-shelf UHF RFID devices. Experimental results show that RFIPad achieves >91% accuracy in recognizing basic touch-screen operations and English letters.
Han Ding 0002, Chen Qian 0001, Jinsong Han, Ge Wang 0003, Wei Xi 0003, Kun Zhao 0002, Jizhong Zhao
ICDCS5
2017 HMRL: Relative Localization of RFID Tags with Static Devices
abstract
Passive Radio Frequency Identification (RFID) tags have been widely applied in many applications, such as logistics, retailing, and warehousing. In many situations the relative locations of objects are more important than their absolute locations. However, state-of-art relative localization methods need continuing movement of tags and readers, which limit the application domain and scalability. In this paper, we propose a relative localization approach for passive tags that requires no device movement. Instead, our method utilizes signal changes caused by arbitrary movement of human beings around tags, who carry no device. Hence our method is called Human Movement based Relative Localization (HMRL). The basic idea of HMRL is that when people pass between reader antenna and tags, the received signal strength will change. By observing the time-series RSS changes of tags, HMRL can obtain the order of tags along a specific horizontal direction. HMRL can also get the order of tags in a vertical direction using hyperbolic positioning. We implement HMRL with commodity off-the-shelf RFID devices. The experimental results show that HMRL achieves high accuracy for relative localization of passive tags.
Ge Wang 0003, Chen Qian 0001, Longfei Shangguan, Han Ding 0002, Jinsong Han, Wei Xi 0003, Jizhong Zhao
SECON7
2017 SALM: Smartphone-Based Identity Authentication Using Lip Motion Characteristics
abstract
With rapid development and popularity, smartphones have been of importance in our daily life. Despite of its convenience in communication and computing, smartphones also lead potential security threats to users. Existing methods on smartphones for protecting user's privacy mainly depend on password or fingerprint based authentication. Most smartphone passwords are very simple and easy to guess or crack, and fingerprinting requires extra hardware and hence increases the price of smartphones. In this paper, we present a smartphone-based identity authentication method based on user's lip motion characteristics, called SALM, which can be used as an additional authentication with password. SALM extracts the feature of lip movements as the authentication token, which is unique for each user. We implement SALM using off-the-shelf smartphones and evaluate its performance via extensive experiments. The results show that the overall accuracy of user authentication using SALM (without password) is higher than 96%.
Yaoxuan Yuan, Jizhong Zhao, Wei Xi 0003, Chen Qian 0001, Zhi Wang 0002
SMARTCOMP3
2017 A Platform for Free-Weight Exercise Monitoring with Passive Tags
abstract
Regular free-weight exercise helps to strengthen natural movements and stabilize muscles that are important to strength, balance, and posture of human beings. Prior works have exploited wearable sensors or RF signal changes for activity sensing, recognition, and counting, etc.. However, none of them have incorporated three key factors necessary for a practical free-weight exercise monitoring system: recognizing free-weight activities on site, assessing their qualities, and providing useful feedbacks to the bodybuilder promptly. Our FEMO system provides an integrated free-weight exercise monitoring service that incorporates all the essential functionalities mentioned above. FEMO achieves this by attaching passive RFID tags on the dumbbells and leveraging the Doppler shift profile of the reflected backscatter signals for on-site free-weight activity recognition and assessment. The rationale behind FEMO is 1) since each free-weight activity owns unique arm motions, the corresponding Doppler shift profile should be distinguishable to each other. 2) Doppler profile of each activity has a strong spatial-temporal correlation that implicitly reflects the quality of the activity. We implement FEMO with COTS RFID devices and conduct a two-week experiment. The preliminary result from 15 volunteers demonstrates that FEMO can be applied to a variety of free-weight activities, and provide valuable feedbacks for activity alignment.
Han Ding 0002, Jinsong Han, Longfei Shangguan, Wei Xi 0003, Zhiping Jiang, Zheng Yang 0002, Zimu Zhou, Panlong Yang, Jizhong Zhao
IEEE Trans. Mob. Comput.4
2016 Instant and Robust Authentication and Key Agreement among Mobile Devices
abstract
Device-to-device communication is important to emerging mobile applications such as Internet of Things and mobile social networks. Authentication and key agreement among multiple legitimate devices is the important first step to build a secure communication channel. Existing solutions put the devices into physical proximity and use the common radio environment as a proof of identities and the common secret to agree on a same key. However they experience very slow secret bit generation rate and high errors, requiring several minutes to build a 256-bit key. In this work, we design and implement an authentication and key agreement protocol for mobile devices, called The Dancing Signals (TDS), being extremely fast and error-free. TDS uses channel state information (CSI) as the common secret among legitimate devices. It guarantees that only devices in a close physical proximity can agree on a key and any device outside a certain distance gets nothing about the key. Compared with existing solutions, TDS is very fast and robust, supports group key agreement, and can effectively defend against predictable channel attacks. We implement TDS using commodity off-the-shelf 802.11n devices and evaluate its performance via extensive experiments. Results show that TDS only takes a couple of seconds to make devices agree on a 256-bit secret key with high entropy.
Wei Xi 0003, Chen Qian 0001, Jinsong Han, Kun Zhao 0002, Sheng Zhong 0002, Xiang-Yang Li 0001, Jizhong Zhao
CCS1
2016 Device-free detection of approach and departure behaviors using backscatter communication
abstract
Smart environments and security systems require automatic detection of human behaviors including approaching to or departing from an object. Existing human motion detection systems usually require human beings to carry special devices, which limits their applications. In this paper, we present a system called APID to detect arm reaching by analyzing backscatter communication signals from a passive RFID tag on the object. APID does not require human beings to carry any device. The idea is based on the influence of human movements to the vibration of backscattered tag signals. APID is compatible with commodity off-the-shelf devices and the EPCglobal Class-1 Generation-2 protocol. In APID an commercial RFID reader continuously queries tags through emitting RF signals and tags simply respond with their IDs. A USRP monitor passively analyzes the communication signals and reports the approach and departure behaviors. We have implemented the APID system for both single-object and multi-object scenarios in both horizontal and vertical deployment modes. The experimental results show that APID can achieve high detection accuracy.
Han Ding 0002, Chen Qian 0001, Jinsong Han, Ge Wang 0003, Zhiping Jiang, Jizhong Zhao, Wei Xi 0003
UbiComp7
2016 Verifiable smart packaging with passive RFID
abstract
Smart packaging adds sensing abilities to traditional packages. This paper investigates the possibility of using RF signals to test the internal status of packages and detect abnormal internal changes. Towards this goal, we design and implement a nondestructive package testing and verification system using commodity passive RFID systems, called Echoscope. Echoscope extracts unique features from the backscatter signals penetrating the internal space of a package and compares them with the previously collected features during the check-in phase. The use of backscatter signals guarantees that there is no difference in RF sources and the features reflecting the internal status will not be affected. Compared to other nondestructive testing methods such as X-ray and ultrasound, Echoscope is much cheaper and provides ubiquitous usage. Our experiments in practical environments show that Echoscope can achieve very high accuracy and is very sensitive to various types abnormal changes.
Ge Wang 0003, Chen Qian 0001, Jinsong Han, Wei Xi 0003, Han Ding 0002, Zhiping Jiang, Jizhong Zhao
UbiComp4
2016 VADS: Visual attention detection with a smartphone
abstract
Identifying the object that attracts human visual attention is an essential function for automatic services in smart environments. However, existing solutions can compute the gaze direction without providing the distance to the target. In addition, most of them rely on special devices or infrastructure support. This paper explores the possibility of using a smartphone to detect the visual attention of a user. By applying the proposed VADS system, acquiring the location of the intended object only requires one simple action: gazing at the intended object and holding up the smartphone so that the object as well as user's face can be simultaneously captured by the front and rear cameras. We extend the current advances of computer vision to develop efficient algorithms to obtain the distance between the camera and user, the user's gaze direction, and the object's direction from camera. The object's location can then be computed by solving a trigonometric problem. VADS has been prototyped on commercial off-the-shelf (COTS) devices. Extensive evaluation results show that VADS achieves low error (about 1.5° in angle and 0.15m in distance for objects within 12m) as well as short latency. We believe that VADS enables a large variety of applications in smart environments.
Zhiping Jiang, Jinsong Han, Chen Qian 0001, Wei Xi 0003, Kun Zhao 0002, Han Ding 0002, Shaojie Tang 0001, Jizhong Zhao, Panlong Yang
INFOCOM4
2016 CSI feedback reduction by checking its validity period: poster
abstract
Multi-user MIMO (MU-MIMO) is proposed in 802.11ac to achieve more than 3x faster than 802.11n. In the real world no-one gets close to theoretical speeds. The primary reason for this anomaly are the various overheads of channel access and channel state information (CSI) feedback. In order to achieve concurrent data transmission, (CSI) feedback from users is required. However, this overhead can easily overwhelm the actual channel time spent on data transmission in large-scale network. Moreover, due to spontaneous uplink traffic, which makes the problem even more challenging.
Yuanhang Cai, Wei Xi 0003, Zhi Wang 0002, Kun Zhao 0002, Jinsong Han, Chen Qian 0001, Han Ding 0002, Jizhong Zhao
MobiCom2
2016 Leveraging Topic Model for CSI Based Human Activity Recognition
abstract
Activity recognition plays an important role in human-computer interactions. Recently, Channel State Information (CSI), known as a fine-grained information capturing the properties of WiFi signal propagation, has been widely used for activity recognition in a device-free pattern. Since CSI is much sensitive to ambient changes, CSI can be used as fingerprints as human activities. However, existing approaches require tremendous overhead in the model training and suffer from failures due to environmental interferences. In this paper, we propose HAR, a CSI based human activity recognition system. HAR investigates the CSI intra-correlation structure (termed as topics) of different human activities. We leverage an unsupervised machine learning method, namely topic model, to extract action characters. Compared to prior works, HAR only requests minor manual intervention, significantly reducing manpower costs in the model training. We implement HAR using commodity WiFi devices to evaluate its performance under different environment settings. The results show that the extracted features are stable to different devices and volunteers, facilitating HAR to achieving an average matching accuracy, i.e., > 90%.
Kun Zhao 0002, Wei Xi 0003, Zhiping Jiang, Zhi Wang 0002, Hongliang Luo, Jizhong Zhao
MSN2
2016 CBID: A Customer Behavior Identification System Using Passive Tags
abstract
Different from online shopping, in-store shopping has few ways to collect the customer behaviors before purchase. In this paper, we present the design and implementation of an on-site Customer Behavior IDentification system based on passive RFID tags, named CBID. By collecting and analyzing wireless signal features, CBID can detect and track tag movements and further infer corresponding customer behaviors. We model three main objectives of behavior identification by concrete problems and solve them using novel protocols and algorithms. The design innovations of this work include a Doppler effect based protocol to detect tag movements, an accurate Doppler frequency estimation algorithm, an image-based human count estimation protocol and a tag clustering algorithm using cosine similarity. We have implemented a prototype of CBID in which all components are built by off-the-shelf devices. We have deployed CBID in real environments and conducted extensive experiments to demonstrate the accuracy and efficiency of CBID in customer behavior identification.
Jinsong Han, Han Ding 0002, Chen Qian 0001, Wei Xi 0003, Zhi Wang 0002, Zhiping Jiang, Longfei Shangguan, Jizhong Zhao
IEEE/ACM Trans. Netw.4
2016 Twins: Device-Free Object Tracking Using Passive Tags
abstract
Device-free object tracking provides a promising solution for many localization and tracking systems to monitor non-cooperative objects, such as intruders, which do not carry any transceiver. However, existing device-free solutions mainly use special sensors or active RFID tags, which are much more expensive compared to passive tags. In this paper, we propose a novel motion detection and tracking method using passive RFID tags, named Twins. The method leverages a newly observed phenomenon called critical state caused by interference among passive tags. We contribute to both theory and practice of this phenomenon by presenting a new interference model that precisely explains it and using extensive experiments to validate it. We design a practical Twins based intrusion detection system and implement a real prototype by commercial off-the-shelf RFID reader and tags. Experimental results show that Twins is effective in detecting the moving object, with very low location errors of 0.75 m in average (with a deployment spacing of 0.6 m).
Jinsong Han, Chen Qian 0001, Dan Ma 0006, Jizhong Zhao, Wei Xi 0003, Zhiping Jiang, Zhi Wang 0002
IEEE/ACM Trans. Netw.6
2016 GenePrint: Generic and Accurate Physical-Layer Identification for UHF RFID Tags
abstract
Physical-layer identification utilizes unique features of wireless devices as their fingerprints, providing authenticity and security guarantee. Prior physical-layer identification techniques on radio frequency identification (RFID) tags require nongeneric equipments and are not fully compatible with existing standards. In this paper, we propose a novel physical-layer identification system, GenePrint, for UHF passive tags. The GenePrint prototype system is implemented by a commercial reader, a USRP-based monitor, and off-the-shelf UHF passive tags. Our solution is generic and completely compatible with the existing standard, EPCglobal C1G2 specification. GenePrint leverages the internal similarity among pulses of tags' RN16 preamble signals to extract a hardware feature as the fingerprint. We conduct extensive experiments on over 10 000 RN16 preamble signals from 150 off-the-shelf RFID tags. The results show that GenePrint achieves a high identification accuracy of 99.68% +. The feature extraction of GenePrint is resilient to various malicious attacks, such as the feature replay attack.
Jinsong Han, Chen Qian 0001, Panlong Yang, Dan Ma 0006, Zhiping Jiang, Wei Xi 0003, Jizhong Zhao
IEEE/ACM Trans. Netw.6
2015 NFV: Near Field Vibration Based Group Device Pairing
Zhiping Jiang, Jinsong Han, Wei Xi 0003, Jizhong Zhao
CollaborateCom3
2015 EMoD: Efficient Motion Detection of Device-Free Objects Using Passive RFID Tags
abstract
Efficient and accurate tracking of device-free objects is critical for anti-intrusion systems. Prior solutions for device-free object tracking are mainly based on costly sensing infrastructures, resulting in barriers to practical applications. In this paper, we propose an accurate and efficient motion detection system, named EMoD, to track device-free objects based on cheap passive RFID tags. EMoD is the first RFID system that can estimate the moving direction as well as the current location of a device-free object by measuring critical power variation sequences of passive tags. Compared with previous solutions, the unique advantage of EMoD, i.e., the capability to estimate moving directions, enables object tracking using a much sparser tag deployment. We contribute to both theory and practice of this phenomenon by presenting the interference model that precisely explains it and using extensive experiments to validate it. We design a practical EMoD based intrusion detection system and implement a prototype by commercial off-the-shelf (COTS) RFID reader and tags. The real-world experiments results show that EMoD is effective in tracking the trajectory of moving object in various environments.
Kun Zhao 0002, Chen Qian 0001, Wei Xi 0003, Jinsong Han, Xue (Steve) Liu, Zhiping Jiang, Jizhong Zhao
ICNP3
2015 Accelerating Crowdsourcing Based Indoor Localization Using CSI
abstract
Indoor localization is of importance for many applications. Crowdsourcing individual users' measurements can provide accurate localization without costly site-survey. However, crowdsourcing based approaches suffer from the cold start problem, in which at the beginning of system deployment, there are insufficient users to contribute their measurements, resulting in inaccurate and time-inefficient localization. In this paper, we propose a hybrid indoor localization method to solve such problem, called ACIL. We first employ the inertial navigation technique to localize some core positions or paths. To tackle the inaccuracy problem, we propose an effective method that utilizes the channel state information (CSI) of wireless signals for accurate distance estimation. This method is based on a new observation: there is a ripple-like fading pattern in wireless signals upon moving objects. Leveraging this observation, our system is capable of calculating the distance of human's movement and his/her direction. We also propose a graph-matching algorithm to setup the correlation between the trajectory and floor map. With those extra obtained location information, the impact of cold start issue will be significantly mitigated, while the LBS can be guaranteed with high localization accuracy. Extensive experiments show that the effectiveness in the human localization and movement detection. Extensive experiments validate the great performance of our protocol in case of various human locations and diverse channel conditions.
Hai-Jiang Xie, Li Lin 0011, Zhiping Jiang, Wei Xi 0003, Kun Zhao 0002, Meiyong Ding, Jizhong Zhao
ICPADS4
2015 Human object estimation via backscattered radio frequency signal
abstract
In this paper, we propose a system called R# to estimate the number of human objects using passive RFID tags but without attaching anything to human objects. The idea is based on our observation that the more human objects are present, the higher the variance in the RSS values of the tag backscattered RF signal. Thus, based on the received RF signal, the reader can estimate the number of human objects. R# includes an RFID reader and some (say 20) passive tags, which are deployed in the region that we want to monitor the number of human objects, such as the region in front of a painting. The RFID reader periodically emits RF signal to identify all tags and the tags simply respond with their IDs via C1G2 standard protocols. We implemented R# using commercial Impinj H47 passive RFID tags and Impinj reader model R420. We conducted experiments in a simulated picking aisle area of the supermarket environment. The experimental results show that R# can achieve high estimation accuracy (more than 90%).
Han Ding 0002, Jinsong Han, Alex X. Liu, Jizhong Zhao, Panlong Yang, Wei Xi 0003, Zhiping Jiang
INFOCOM6
2015 FEMO: A Platform for Free-weight Exercise Monitoring with RFIDs
abstract
Regular free-weight exercise helps to strengthen the body's natural movements and stabilize muscles that are important to strength, balance, and posture of human beings. Prior works have exploited wearable sensors or RF signal changes (e.g., WiFi and Blue tooth) for activity sensing, recognition and countingetc.. However, none of them have incorporate three key factors necessary for a practical free-weight exercise monitoring system: recognizing free-weight activities on site, assessing their qualities, and providing useful feedbacks to the bodybuilder promptly. Our FEMO system responds to these demands, providing an integrated free-weight exercise monitoring service that incorporates all the essential functionalities mentioned above. FEMO achieves this by attaching passive RFID tags on the dumbbells and leveraging the Doppler shift profile of the reflected backscatter signals for on-site free-weight activity recognition and assessment. The rationale behind FEMO is 1): since each free-weight activity owns unique arm motions, the corresponding Doppler shift profile should be distinguishable to each other and serves as a reliable signature for each activity. 2): the Doppler profile of each activity has a strong spatial-temporal correlation that implicitly reflects the quality of each performed activity. We implement FEMO with COTS RFID devices and conduct a two-week experiment. The preliminary result from 15 volunteers demonstrates that FEMO can be applied to a variety of free-weight activities and users, and provide valuable feedbacks for activity alignment.
Han Ding 0002, Longfei Shangguan, Zheng Yang 0002, Jinsong Han, Zimu Zhou, Panlong Yang, Wei Xi 0003, Jizhong Zhao
SenSys7
2014 CBID: A Customer Behavior Identification System Using Passive Tags
abstract
Different from online shopping, in-store shopping has few ways to collect the customer behaviors before purchase. In this paper, we present the design and implementation of an on-site Customer Behavior Identification system based on passive RFID tags, named CBID. By collecting and analyzing wireless signal features, CBID can detect and track tag movements and further infer corresponding customer behaviors. We model three main objectives of behavior identification by concrete problems and solve them using novel protocols and algorithms. The design innovations of this work include a Doppler effect based protocol to detect tag movements, an accurate Doppler frequency estimation algorithm, a multi-RSS based tag localization protocol, and a tag clustering algorithm using cosine similarity. We have implemented a prototype of CBID in which all components are built by off-the-shelf devices. We have deployed CBID in real environments and conducted extensive experiments to demonstrate the accuracy and efficiency of CBID in customer behavior identification.
Jinsong Han, Han Ding 0002, Chen Qian 0001, Dan Ma 0006, Wei Xi 0003, Zhi Wang 0002, Zhiping Jiang, Longfei Shangguan
ICNP5
2014 A fine-grained indoor localization using multidimensional Wi-Fi fingerprinting
abstract
Although fingerprint based localization is promising for indoor applications, its accuracy still remains a huge challenge. Most of existing approaches rely on the Radio Signal Strength (RSS) to generate fingerprints. However, merely using RSS is unable to accurately localize objects since such an one-dimensional fingerprint will be seriously influenced by the interference and multi-path effect in the indoor environment. In this paper, we propose a new localization approach based on multidimensional Wi-Fi fingerprint. Instead of only using RSS to construct fingerprint, we employ RSS, transmitted power, and channel information to construct an integrated fingerprint. The extended fingerprint enables fine-grained localization and tracking services. We also deign a cosine similarity based matching algorithm and enhanced particle filter mechanism to achieve accurate localization and tracking. Extensive experiment and implementation results show that the new fingerprint and proposed algorithms can achieve an accuracy within two meters in 90% of testing points, while demonstrating a good adaptability to complex indoor environments.
Deng Chen, Zhiping Jiang, Wei Xi 0003, Jinsong Han, Kun Zhao 0002, Jizhong Zhao, Zhi Wang 0002, Rui Li 0047
ICPADS4
2014 Twins: Device-free object tracking using passive tags
abstract
Device-free based object tracking provides a promising solution for many localization and tracking systems to monitor non-cooperative objects which do not carry any transceiver such as intruders. However, existing device-free solutions mainly use sensors and active RFID tags, which are much more expensive compared to passive tags. In this paper, we propose a novel motion detection and tracking method using passive RFID tags, named Twins. The method leverages a phenomenon called critical state caused by interference among passive tags. We theoretically explain this phenomenon via an interference model and conduct extensive experiment to validate it. We design a practical Twins based intrusion detection system and implement a real prototype with commercial off-the-shelf reader and tags. Experimental results show that Twins is effective in detecting the moving object, with low location errors of 0.75m in average.
Jinsong Han, Chen Qian 0001, Dan Ma 0006, Jizhong Zhao, Pengfeng Zhang, Wei Xi 0003, Zhiping Jiang
INFOCOM7
2014 Electronic frog eye: Counting crowd using WiFi
abstract
Crowd counting, which count or accurately estimate the number of human beings within a region, is critical in many applications, such as guided tour, crowd control and marketing research and analysis. A crowd counting solution should be scalable and be minimally intrusive (i.e., device-free) to users. Image-based solutions are device-free, but cannot work well in a dim or dark environment. Non-image based solutions usually require every human being carrying device, and are inaccurate and unreliable in practice. In this paper, we present FCC, a device-Free Crowd Counting approach based on Channel State Information (CSI). Our design is motivated by our observation that CSI is highly sensitive to environment variation, like a frog eye. We theoretically discuss the relationship between the number of moving people and the variation of wireless channel state. A major challenge in our design of FCC is to find a stable monotonic function to characterize the relationship between the crowd number and various features of CSI. To this end, we propose a metric, the Percentage of nonzero Elements (PEM), in the dilated CSI Matrix. The monotonic relationship can be explicitly formulated by the Grey Verhulst Model, which is used for crowd counting without a labor-intensive site survey. We implement FCC using off-the-shelf IEEE 802.11n devices and evaluate its performance via extensive experiments in typical real-world scenarios. Our results demonstrate that FCC outperforms the state-of-art approaches with much better accuracy, scalability and reliability.
Wei Xi 0003, Jizhong Zhao, Xiang-Yang Li 0001, Kun Zhao 0002, Shaojie Tang 0001, Xue (Steve) Liu, Zhiping Jiang
INFOCOM1
2014 KEEP: Fast secret key extraction protocol for D2D communication
abstract
Device to device (D2D) communication is expected to become a promising technology of the next-generation wireless communication systems. Security issues have become technical barriers of D2D communication due to its “open-air” nature and lack of centralized control. Generating symmetric keys individually on different communication parties without key exchange or distribution is desirable but challenging. Recent work has proposed to extract keys from the measurement of physical layer random variations of a wireless channel, e.g., the channel state information (CSI) from orthogonal frequency-division multiplexing (OFDM). Existing CSI-based key extraction methods usually use the measurement results of individual subcarriers. However, our real world experiment results show that CSI measurements from near-by subcarriers have strong correlations and a generated key may have a large proportion of repeated bit segments. Hence attackers may crack the key in a relatively short time and hence reduce the security level of the generated keys. In this work, we propose a fast secret key extraction protocol, called KEEP. KEEP uses a validation-recombination mechanism to obtain consistent secret keys from CSI measurements of all subcarriers. It achieves high security level of the keys and fast key-generation rate. We implement KEEP using off-the-shelf 802.11n devices and evaluate its performance via extensive experiments. Both theoretical analysis and experimental results demonstrate that KEEP is safer and more effective than the state-of-the-art approaches.
Wei Xi 0003, Xiang-Yang Li 0001, Chen Qian 0001, Jinsong Han, Shaojie Tang 0001, Jizhong Zhao, Kun Zhao 0002
IWQoS1
2014 Poster: locating RFID tags by rotation
abstract
Locating objects labeled with RFID tags is an important issue which should be addressed in many applications, such as warehouse management, goods management in supermarket and finding of lost objects. Some existing works use large numbers of reference tags which involve lots of manpower to deploy them. Others achieve high accuracy, but rely on sophisticated equipments which are hardly available in large scale to the industry. This work exploits the radiation pattern of existing directional panel antenna which is steerable and derives angle-of-arrival (AoA) information from the energy reflected by the target tag when the antenna is rotating. We use Commercial Off-The-Shelf (COTS) equipments and get median position accuracy of 29cm in our preliminary experiment.
Wei Xi 0003, Shaojie Tang 0001, Jinsong Han, Jizhong Zhao, Xiang-Yang Li 0001, Zhi Wang 0002, Zhiping Jiang
MobiCom2
2014 Communicating Is Crowdsourcing: Wi-Fi Indoor Localization with CSI-Based Speed Estimation
Zhiping Jiang, Wei Xi 0003, Xiang-Yang Li 0001, Shaojie Tang 0001, Jizhong Zhao, Jinsong Han, Kun Zhao 0002, Zhi Wang 0002
J. Comput. Sci. Technol.2
2014 Assessing Diagnosis Approaches for Wireless Sensor Networks: Concepts and Analysis
Rui Li 0047, Kebin Liu 0001, Xiang-Yang Li 0001, Yuan He 0004, Wei Xi 0003, Zhi Wang 0002, Jizhong Zhao, Meng Wan
J. Comput. Sci. Technol.5
2014 Efficient and secure key extraction using channel state information
Zhi Wang 0002, Jinsong Han, Wei Xi 0003, Jizhong Zhao
J. Supercomput.3
2013 Rejecting the attack: Source authentication for Wi-Fi management frames using CSI Information
abstract
Comparing to well protected data frames, Wi-Fi management frames (MFs) are extremely vulnerable to various attacks. Since MFs are transmitted without encryption or authentication, attackers can easily launch various attacks by forging the MFs. In a collaborative environment with many Wi-Fi sniffers, such attacks can be easily detected by sensing the anomaly RSS changes. However, it is quite difficult to identify these spoofing attacks without assistance from other nodes. By exploiting some unique characteristics (e.g., rapid spatial decorrelation, independence of Txpower, and much richer dimensions) of 802.11n Channel State Information (CSI), we design and implement CSITE, a prototype system to authenticate the Wi-Fi management frames on PHY layer merely by one station. Our system CSITE, built upon off-the-shelf hardware, achieves precise spoofing detection without collaboration and in-advance fingerprint. Several novel techniques are designed to address the challenges caused by user mobility and channel dynamics. To verify the performances of our solution, we conduct extensive evaluations in various scenarios. Our test results show that our design significantly outperforms the RSS-based method. We observe about 8 times improvement by CSITE over RSS-based method on the falsely accepted attacking frames.
Zhiping Jiang, Jizhong Zhao, Xiang-Yang Li 0001, Jinsong Han, Wei Xi 0003
INFOCOM5
2013 Wi-Fi Fingerprint Based Indoor Localization without Indoor Space Measurement
abstract
Numerous indoor localization techniques have been proposed recently to meet the intensive demand for location-based service. Fingerprint-based approach is one of most popular and inexpensive solution. In terms of constructing the fingerprint database, there have to be a synchronized measurement for both indoor space(eg by labor-intensive site-survey or sensor-based crowd sensing) and fingerprint space, by this means the fingerprints database is established. It is the indoor space measurement hinders the usability of fingerprint-based localization system. In this work, we propose a sensor-free crowds ensing indoor localization scheme, protocol. The main contribution of our protocol is that we don't need indoor space measurement. Floor plan and RSS samples temporal sequence is the only requirement. The core of our method is a graph matching based manifold alignment process, which automatically finds the best correspondence between floor plan and wireless fingerprint transition structure. With no more need of indoor space measurement, the system deployment complexity and cost are significantly reduced. We implement our protocol at AP-end and deploy it in a 2000m^2 office environment. The evaluation has shown that our protocol can handle complex environment mapping and achieve high localization & tracking accuracy.
Zhiping Jiang, Jizhong Zhao, Jinsong Han, Shaojie Tang 0001, Wei Xi 0003
MASS6
2013 Localization of Wireless Sensor Networks in the Wild: Pursuit of Ranging Quality
abstract
Localization is a fundamental issue of wireless sensor networks that has been extensively studied in the literature. Our real-world experience from GreenOrbs, a sensor network system deployed in a forest, shows that localization in the wild remains very challenging due to various interfering factors. In this paper, we propose CDL, a Combined and Differentiated Localization approach for localization that exploits the strength of range-free approaches and range-based approaches using received signal strength indicator (RSSI). A critical observation is that ranging quality greatly impacts the overall localization accuracy. To achieve a better ranging quality, our method CDL incorporates virtual-hop localization, local filtration, and ranging-quality aware calibration. We have implemented and evaluated CDL by extensive real-world experiments in GreenOrbs and large-scale simulations. Our experimental and simulation results demonstrate that CDL outperforms current state-of-art localization approaches with a more accurate and consistent performance. For example, the average location error using CDL in GreenOrbs system is 2.9 m, while the previous best method SISR has an average error of 4.6 m.
Jizhong Zhao, Wei Xi 0003, Yuan He 0004, Yunhao Liu 0001, Xiang-Yang Li 0001, Lufeng Mo, Zheng Yang 0002
IEEE/ACM Trans. Netw.2
2013 WILL: Wireless Indoor Localization without Site Survey
abstract
Indoor localization is of great importance for a range of pervasive applications, attracting many research efforts in the past two decades. Most radio-based solutions require a process of site survey, in which radio signatures are collected and stored for further comparison and matching. Site survey involves intensive costs on manpower and time. In this work, we study unexploited RF signal characteristics and leverage user motions to construct radio floor plan that is previously obtained by site survey. On this basis, we design WILL, an indoor localization approach based on off-the-shelf WiFi infrastructure and mobile phones. WILL is deployed in a real building covering over 1600 m2, and its deployment is easy and rapid since site survey is no longer needed. The experiment results show that WILL achieves competitive performance comparing with traditional approaches.
Chenshu Wu, Zheng Yang 0002, Yunhao Liu 0001, Wei Xi 0003
IEEE Trans. Parallel Distributed Syst.4
2012 WILL: Wireless indoor localization without site survey
abstract
Indoor localization is of great importance for a range of pervasive applications, attracting many research efforts in the past two decades. Most radio-based solutions require a process of site survey, in which radio signatures are collected and stored for further comparison and matching. Site survey involves intensive costs on manpower and time. In this work, we study unexploited RF signal characteristics and leverage user motions to construct radio floor plan that is previously obtained by site survey. On this basis, we design WILL, an indoor localization approach based on off-the-shelf WiFi infrastructure and mobile phones. WILL is deployed in a real building covering over 1600m2, and its deployment is easy and rapid since site survey is no longer needed. The experiment results show that WILL achieves competitive performance comparing with traditional approaches.
Chenshu Wu, Zheng Yang 0002, Yunhao Liu 0001, Wei Xi 0003
INFOCOM4
2011 Exploiting the Associated Information to Locate Mobile Users in Ubiquitous Computing Environment
abstract
Although GPS is deemed as ubiquitous outdoor localization technology, we are still far from a similar technology for indoor environments. Though a number of techniques are proposed for indoor localization, they are separated efforts that are way from a real ubiquitous localization system. Our real-world experience from InSpace, a pervasive computing system with wireless devices to provide intelligent services to users, shows that locating mobile users remains very challenging due to various interfering factors. We analyze real traces of mobile phones carried by users and find that mobile users exhibit temporal-spatial stability and neighborhood relativity. Motivated by this observation, we develop a Mobile Boundary Localization approach, MBL, to exploit the associated information to locate mobile users. This localization approach uses different treatment in different conditions and lets each mobile phone try to estimate its possible location range. We have implemented and evaluated MBL by extensive real-world experiments in InSpace and simulations. The results demonstrate that MBL significantly outperforms state-of-the-art localization approaches with more accurate, efficient, and consistent performance.
Wei Xi 0003, Jizhong Zhao, Yuan He 0004, Zhi Wang 0002, Lufeng Mo
MASS1
2011 Lazy Schema: An Optimal Sampling Frequency Assignment for Real-Time Sensor Systems
abstract
How to reasonably allocate and schedule resources of wireless sensor system to maximum its potential capability has been an important area for research. In this paper, we focus on the Optimal Sampling Frequency Assignment (OSFA) in a real-time wireless sensor networks (RTWSN). An appropriate OSFA should both guarantee a good quality of real-time service and efficiently utilize the limited network resources as well. We propose a distributed optimization algorithm, called Lazy Schema (LySa), to obtain the optimal sampling rates of source nodes with low cost. The central idea is that redundancy reporting and constant adjust step size result in the excessive communicate overhead in iteration process. LySa adopts self-adaptive reporting rates and dynamically adjusts step size to reduce the traffic cost and accelerate the convergence of optimal solution. We have evaluated LySa together with related mainstream algorithms. The results demonstrate that LySa outperforms current state-of-art approaches in terms of low cost, high efficiency and scalability in RTWSN.
Jizhong Zhao, Yong Qi 0001, Shuo Lian, Wei Xi 0003
MSN5
2011 Crowd Density Estimation Using Wireless Sensor Networks
abstract
Estimation of crowd distribution is critical to various applications. Although most researches have provided solutions based on images and videos technologies, the high costs for deploying and an over-dependence on the bright light restrict its scope of application. In this paper, we use wireless sensor networks (WSNs) originally to make up for the lack of camera. Our approach is an iterative process which contains two phases in each time slot. In detection step, we divide the crowd density into different levels according to the RSSI data obtained by WSNs using K-means algorithm. In calibration step, we eliminate the noises and other deviations estimation based on the spatial-temporal correlation of crowd distribution. In addition, we have implemented and evaluated our algorithm by extensive real-world experiments using 16 sensor nodes and large-scale simulations. The results show that our algorithm has an accurate, efficient, and consistent performance.
Yaoxuan Yuan, Wei Xi 0003, Jizhong Zhao
MSN3
2010 Locating sensors in the wild: pursuit of ranging quality
abstract
Localization is a fundamental issue of wireless sensor networks that has been extensively studied in the literature. The real-world experience from GreenOrbs, a sensor network system in the forest, shows that localization in the wild remains very challenging due to various interfering factors. In this paper we propose CDL, a Combined and Differentiated Localization approach. The central idea is that ranging quality is the key that determines the overall localization accuracy. In its unremitting pursuit of better ranging quality, CDL incorporates virtual-hop localization, local filtration, and ranging-quality aware calibration. We have implemented CDL and evaluated it by extensive experiments and simulations. The results demonstrate that CDL outperforms current state-of-art approaches with better accuracy, efficiency and consistent performance.
Wei Xi 0003, Yuan He 0004, Yunhao Liu 0001, Jizhong Zhao, Lufeng Mo, Zheng Yang 0002, Jiliang Wang, Xiang-Yang Li 0001
SenSys1
2009 EUL: An Efficient and Universal Localization Method for Wireless Sensor Network
abstract
Localization is a crucial service for various applications in wireless sensor networks (WSNs). Although most researches assume stationary nodes, sensor mobility can enrich the application scenarios. Existing dynamic localization approaches require high seed density or incur a large communication overhead. In order to address these problems, we propose an efficient rang-free localization algorithm, EUL, which utilizes the relationship between neighboring nodes to estimate their possible location boundaries. Our algorithm not only allows all the nodes to remain static or move freely but also reduces the dependence on seeds, which achieves a uniform energy distribution to address the excessive energy drain around seeds and lengthen the network lifetime. We have evaluated EUL together with other major dynamic localization approaches. Simulation results show that EUL outperforms existing approaches in terms of accuracy under many different mobility conditions.
Wei Xi 0003, Jizhong Zhao, Xue (Steve) Liu, Xiang-Yang Li 0001, Yong Qi 0001
ICDCS1