Xiang Zhang 0011

dblp:91/4353-11 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0003-0413-6135ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition
abstract
The challenge in WiFi-based cross-domain Behavior Recognition lies in the significant interference of domain-specific signals on gesture variation. However, previous methods alleviate this interference by mapping the phase from multiple domains into a common feature space. If the Doppler Frequency Shift (DFS) signal is used to dynamically supplement the phase features to achieve better generalization, it enables the model to not only explore a wider feature space but also to avoid potential degradation of gesture semantic information. Specifically, we propose a novel Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition (Wi-CBR), which constructs a dual-branch self-attention module that captures temporal features from phase information reflecting dynamic path length variations while extracting kinematic features from DFS correlated with motion velocity. Moreover, we design a Saliency Guidance Module that employs group attention mechanisms to mine critical activity features and utilizes gating mechanisms to optimize information entropy, facilitating feature fusion and enabling effective interaction between salient and non-salient behavioral characteristics. Extensive experiments on two large-scale public datasets (Widar3.0 and XRF55) demonstrate the superior performance of our method in both in-domain and cross-domain scenarios.
Ruobei Zhang, Shengeng Tang, Xiang Zhang 0011, Jiabao Guo
AAAI4
2026 DiffLoc+: Toward Robust Wi-Fi Hidden Camera Localization Based on Electromagnetic Diffraction
abstract
The proliferation of hidden WiFi cameras has raised serious privacy concerns, making their accurate detection and localization essential for the secure development of future intelligent wireless networks. However, existing solutions often require substantial user involvement, large movement spaces, predefined system parameters, or pre-collected training data, limiting their practicality and scalability. In this paper, we present DiffLoc+, a novel and low-cost system that localizes hidden WiFi cameras by harnessing the fundamental physical principle of electromagnetic diffraction. When an obstacle crosses the line-of-sight path between a transmitter and a receiver, it causes a distinctive signal attenuation pattern. We theoretically analyze the feasibility of exploiting this phenomenon for localization and identify two key conditions for building an unbiased diffraction-based model: symmetry and observability. To satisfy these conditions, DiffLoc+ introduces a controllable diffraction generation mechanism that precisely rotates a small metal plate around a WiFi receiver (e.g. a Raspberry Pi), producing a stable and predictable diffraction “shadowing” effect. We then construct an unbiased localization model that maps this effect to the azimuth of the camera. To ensure the robustness of the theoretical model in real-world applications, DiffLoc+ further introduces two robustness-enhancing mechanisms: (1) an attenuation-region difference-driven subcarrier selection method, which filters subcarriers that reliably reflect the diffraction attenuation pattern by quantifying the signal contrast between diffraction- and reflection-dominated regions; and (2) an uncertainty evaluation framework that integrates result consistency and diffraction signal quality to eliminate unreliable estimates. Implemented entirely with commodity off-the-shelf (COTS) hardware, DiffLoc+ achieves an average angular error of 11.92° across six diverse indoor environments and eleven commercial camera models, demonstrating its effectiveness and robustness.
Huan Yan 0004, Jian Liu 0055, Xiang Zhang 0011, Zhi Liu 0002, Bin Liu 0016, Meng Li 0006, Ming Gao 0023, Fusang Zhang
IEEE J. Sel. Areas Commun.3
2026 Identifying Who You Are No Matter What You Write Through Abstracting Handwriting Style
abstract
With the increasing use of electronic devices, online handwriting verification has become crucial for biometricsbased identity authentication. Traditional methods, which rely on content-dependent verification of the writer's name, are vulnerable to forgery. This paper introduces a content-independent handwriting authentication system, Ph-Wri, designed for commodity smartphones. The core innovation is a multi-path attention feature fusion network that combines both static features (image of the handwritten text) and dynamic features (time-dependent properties during writing), to abstract the handwriting style instead of specific content for recognition, enabling robust user authentication. To extract handwriting style from dynamic writing features, we propose a polarity-aware attention strategy during training. This strategy incorporates Style Channel Attention (SCA) to capture direction-sensitive stylistic features, and Trajectory Spatial Attention (TSA) to highlight key handwriting trajectory regions. In the fine-tuning stage, the Correlation-Aware Attention (CAA) module models inter-channel structural correlations, mitigating the influence of content and enhancing style-consistent representations. By linking content-independent handwriting style to user identity, the system achieves accurate authentication. Extensive experiments on both the self-built CIEHD dataset and the public BiosecurID dataset demonstrate exceptional performance, achieving a 99% Verification Accuracy on CIEHD. Compared to state-of-theart methods that utilize only static or dynamic data, Ph-Wri significantly reduces the Equal Error Rate, showcasing the effectiveness and practicality of the proposed approach.
Jinyang Huang, Yuanhao Feng, Feng-Qi Cui, Xiang Zhang 0011, Zhi Liu 0002, Xin Liu 0104, Jianchun Liu, Fusang Zhang, Meng Li 0006
IEEE Trans. Dependable Secur. Comput.4
2025 Temporal Features for IoT Devices: Out-of-Distribution Detection without Upper-Layer Dependencies
abstract
The large-scale deployment of IoT devices accelerates intelligent applications but also brings significant security risks. Device detection helps mitigate these risks by identifying unauthorized or rogue devices and improving visibility into network activity. However, existing device detection methods based on network and transport layer protocols face two key challenges: encrypted traffic conceals protocol information, and most approaches fail to detect out-of-distribution (OOD) devices, limiting their effectiveness in real-world scenarios. To address these issues, this paper proposes an OOD detection method based on the 802.11 protocol. Specifically, we first extract intrinsic packet attributes from the 802.11 protocol headers, including transmission timing patterns and packet structure characteristics, without relying on any network or transport layer information. Then, these features are input into a bidirectional long short-term memory (LSTM) model to learn sequential dependencies, and the extracted feature embeddings are evaluated through k-nearest neighbor (KNN) distance calculation to detect both in-distribution (ID) and OOD samples. Experiments conducted on 12 commercial IoT devices spanning 8 categories demonstrate that the proposed method achieves effective device identification and OOD detection performance.
Jian Liu 0055, Huan Yan 0004, Jinyang Huang, Xiang Zhang 0011
GLOBECOM4
2025 Source-Free Domain Adaptation via Perceptual Semantic Decoupling for WiFi Gesture Recognition
abstract
Generalizable WiFi gesture recognition has gained increasing attention for its contactless operation, ubiquitous infrastructure and enhanced robustness. Among existing methods, source-free domain adaptation (SFDA) stands out by preserving privacy and reducing computational demands without relying on source data. Current methods typically process low-level WiFi signals and their high-level semantic representations from a unified perspective, making temporal semantic learning highly susceptible to low-level signal noise and lacking consistent semantic guidance for cross domain alignment, thereby limiting the effectiveness. In this paper, we propose ViFi, a novel SFDA framework specifically designed for cross-domain WiFi gesture recognition. Unlike prior work, ViFi introduces a viewpoint-hierarchical strategy that explicitly processes cross-domain sensing from two perspectives: the perceptual (signal-level) and the semantic (gesture-level). This separation mitigates the impact of signal noise on high-level semantics while preventing semantic space drift during domain alignment. ViFi operates in two key stages. First, it anchors the perceptual encoder and employs masked signal semantic reconstruction to learn robust high-level temporal semantics. Then, it freezes the semantic encoder and aligns the perceptual encoder across domains, again leveraging masked reconstruction to ensure alignment under a unified and meaningful semantic space. We evaluate ViFi on a public dataset, and experimental results show that our viewpoint-hierarchical method achieves over 15% improvement compared to the baseline and significantly outperforms state-of-the-art approaches.
Yelin Wei, Xiang Zhang 0011, Bin Liu 0016, Songming Jia, Jinyang Huang, Zhi Liu 0002, Huan Yan 0004
GLOBECOM2
2025 A Watermark Updating Framework for Multi-stage Image Content Distribution
abstract
Deep image watermarking embeds identification data into images to facilitate source tracking. However, existing schemes are primarily designed for single-stage transmission scenarios, and in practical multi-stage distribution requirements, current methods degrade image quality and reduce watermark extraction accuracy. In this paper, we introduces WaterUp, a deep watermark updating framework. WaterUp automatically updates watermark information as the image is transmitted, preserving image quality while accurately recording the transmission path for traceability. The core of WaterUp is a flow-based encoder-decoder (FED), which utilizes a forward and backward network to enable efficient watermark updating with minimal computational and storage demands. Experimental results show that WaterUp outperforms state-of-the-art methods, maintaining high visual quality with a PSNR exceeding 38 dB across multiple transmissions.
Bin Liu 0016, Jie Zhang 0073, Xiang Zhang 0011, Zehua Ma, Nenghai Yu
ICME4
2025 Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing
abstract
Peptide sequencing—the process of identifying amino acid sequences from mass spectrometry data—is a fundamental task in proteomics. Non-Autoregressive Transformers (NATs) have proven highly effective for this task, outperforming traditional methods. Unlike autoregressive models, which generate tokens sequentially, NATs predict all positions simultaneously, leveraging bidirectional context through unmasked self-attention. However, existing NAT approaches often rely on Connectionist Temporal Classification (CTC) loss, which presents significant optimization challenges due to CTC’s complexity and increases the risk of training failures. To address these issues, we propose an improved non-autoregressive peptide sequencing model that incorporates a structured protein sequence curriculum learning strategy. This approach adjusts protein’s learning difficulty based on the model’s estimated protein generational capabilities through a sampling process, progressively learning peptide generation from simple to complex sequences. Additionally, we introduce a self-refining inference-time module that iteratively enhances predictions using learned NAT token embeddings, improving sequence accuracy at a fine-grained level. Our curriculum learning strategy reduces NAT training failures frequency by more than 90% based on sampled training over various data distributions. Evaluations on nine benchmark species demonstrate that our approach outperforms all previous methods across multiple metrics and species. Model and source code are available at https://github.com/BEAM-Labs/denovo.
Xiang Zhang 0011, Zijie Qiu, Nanqing Dong
ICML1
2025 Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing
abstract
De novo peptide sequencing is a critical task in proteomics. However, the performance of current deep learning-based methods is limited by the inherent complexity of mass spectrometry data and the heterogeneous distribution of noise signals, leading to data-specific biases. We present RankNovo, the first deep reranking framework that enhances de novo peptide sequencing by leveraging the complementary strengths of multiple sequencing models. RankNovo employs a list-wise reranking approach, modeling candidate peptides as multiple sequence alignments and utilizing axial attention to extract informative features across candidates. Additionally, we introduce two new metrics, PMD (Peptide Mass Deviation) and RMD (ResidualMass Deviation), which offer delicate supervision by quantifying mass differences between peptides at both the sequence and residue levels. Extensive experiments demonstrate that RankNovo not only surpasses its base models used to generate training candidates for reranking pre-training, but also sets a new state-of-the-art benchmark. Moreover, RankNovo exhibits strong zero-shot generalization to unseen models—those whose generations were not exposed during training, highlighting its robustness and potential as a universal reranking framework for peptide sequencing. Our work presents a novel reranking strategy that fundamentally challenges existing single-model paradigms and advances the frontier of accurate de novo sequencing. Our source code is provided on GitHub.
Zijie Qiu, Xiang Zhang 0011, Nanqing Dong
ICML3
2025 STAR: A Benchmark for Astronomical Star Fields Super-Resolution
abstract
Super-resolution (SR) advances astronomical imaging by enabling cost-effective high-resolution capture, crucial for detecting faraway celestial objects and precise structural analysis. However, existing datasets for astronomical SR (ASR) exhibit three critical limitations: flux inconsistency, object-crop setting, and insufficient data diversity, significantly impeding ASR development. We propose STAR, a large-scale astronomical SR dataset containing 54,738 flux-consistent star field image pairs covering wide celestial regions. These pairs combine Hubble Space Telescope high-resolution observations with physically faithful low-resolution counterparts generated through a flux-preserving data generation pipeline, enabling systematic development of field-level ASR models. To further empower the ASR community, STAR provides a novel Flux Error (FE) to evaluate SR models in physical view. Leveraging this benchmark, we propose a Flux-Invariant Super Resolution (FISR) model that could accurately infer the flux-consistent high-resolution images from input photometry, suppressing several SR state-of-the-art methods by 24.84% on a novel designed flux consistency metric, showing the priority of our method for astrophysics. Extensive experiments demonstrate the effectiveness of our proposed method and the value of our dataset. Code and models are available at https://github.com/GuoCheng12/STAR.
Kuo-Cheng Wu, Guohang Zhuang, Jinyang Huang, Xiang Zhang 0011, Wanli Ouyang, Yan Lu 0001
NeurIPS4
2025 CamLopa: A Hidden Wireless Camera Localization Framework via Signal Propagation Path Analysis
abstract
Hidden wireless cameras pose significant privacy threats, necessitating effective detection and localization methods. However, existing localization solutions often require impractical activity spaces, expensive specialized devices, or pre-collected training data, limiting their practical deployment. To address these limitations, we introduce CamLopa, a training-free wireless camera localization framework that operates with minimal activity space constraints using low-cost, commercial-off-the-shelf (COTS) devices. CamLopa can achieve detection and localization in just 45 seconds of user activities with a Raspberry Pi board. During this short period, it analyzes the causal relationship between wireless traffic and user movement to detect the presence of a hidden camera. Upon detection, CamLopa utilizes a novel azimuth localization model based on wireless signal propagation path analysis for localization. This model leverages the time ratio of user paths crossing the First Fresnel Zone (FFZ) to determine the camera's azimuth angle. Subsequently, CamLopa refines the localization by identifying the camera's quadrant. We evaluate CamLopa across various devices and environments, demonstrating its effectiveness with a 95.37% detection accuracy for snooping cameras and an average localization error of 17.23°, under the significantly reduced activity space requirements and without the need for training. Our code and demo are available at https://github.com/CamLoPA/CamLoPA-Code.
Xiang Zhang 0011, Jie Zhang 0073, Zehua Ma, Jinyang Huang, Meng Li 0006, Huan Yan 0004, Peng Zhao 0024, Zijian Zhang 0001, Bin Liu 0016, Qing Guo 0005, Tianwei Zhang 0004, Nenghai Yu
SP1
2025 DiffLoc: WiFi Hidden Camera Localization Based on Electromagnetic Diffraction
Xiang Zhang 0011, Jie Zhang 0073, Huan Yan 0004, Jinyang Huang, Zehua Ma, Bin Liu 0016, Meng Li 0006, Kejiang Chen, Qing Guo 0005, Tianwei Zhang 0004, Zhi Liu 0002
USENIX Security Symposium1
2025 Wi-SFDAGR: WiFi-Based Cross-Domain Gesture Recognition via Source-Free Domain Adaptation
abstract
WiFi channel state information (CSI)-based gesture recognition offers unique advantages, including cost-effectiveness and enhanced privacy protection, and has garnered significant attention in recent years. However, existing WiFi-based gesture recognition solutions exhibit poor generalization ability when deployed in new environment, orientation, or location. Although some methods combine labeled source domain and unlabeled target domain to learn domain-independent features, factors, such as data privacy protection, hinder access to source data during practical environment adaptation. Consequently, we consider realistic scenario where source data is unavailable during adaptation of unlabeled test data, and instead, a trained source domain model is used. In this article, we propose Wi-SFDAGR, a WiFi-based source-free domain adaptation gesture recognition framework. Specifically, we treat cross-domain as an unsupervised clustering problem, aiming to ensure that features within local neighborhoods exhibit similar prediction results while those farther apart display different prediction outcomes in the feature space. We theoretically analyze the effect of enhanced prediction consistency between neighbor points extracted from gestures on generalization error. Furthermore, we employ an attraction-dispersion network to strengthen prediction consistency among closely located features in the feature space while reducing it for distantly located features. To mitigate noise introduced during nearest neighbor sample selection in the feature space (where predictions may not align with the input sample’s prediction), we progressively improve nearby sample feature aggregation by estimating uncertainty to reweight local neighborhood predictions. Finally, extensive experiments are conducted on the Widar 3.0 and XRF55 datasets and the results show our proposed framework outperforms most cross-domain methods.
Huan Yan 0004, Xiang Zhang 0011, Jinyang Huang, Yuanhao Feng, Meng Li 0006, Anzhi Wang, Weihua Ou, Zhi Liu 0002
IEEE Internet Things J.2
2025 Wi-Pulmo: Commodity WiFi Can Capture Your Pulmonary Function Without Mouth Clinging
abstract
Pulmonary function testing is a crucial examination for respiratory diseases. Current medical spirometers are bulky and inconvenient, while available portable spirometers are extremely expensive and often lack accuracy. Furthermore, both devices require direct contact, inevitably increasing the cross-infection risk. To tackle these challenges, we propose Wi-Pulmo, an end-to-end deep learning-based Wireless System that utilizes WiFi channel state information (CSI) to provide contact-free, convenient, cost-effective, and precise pulmonary function testing outside the clinical setting. Based on the analysis of thoracic and abdominal movement patterns, Wi-Pulmo first validates the feasibility of using WiFi to estimate pulmonary function. Then, Wi-Pulmo designs an efficient fine-grained sensing quality-based algorithm for complete exhalation segmentation. Additionally, a relevant interference-tolerant learning algorithm based on variational inference is proposed to accurately map the CSI of WiFi signals to pulmonary function. Extensive experiments achieved average monitoring error rates of 2.59% for normal subjects in daily scenarios and 5.87% for real patients in tertiary hospitals over a two-month period. These satisfactory results demonstrate the strong effectiveness and robustness of Wi-Pulmo. Furthermore, our findings in clinical reveal a close correlation between chronic diseases and pulmonary function.
Peng Zhao 0024, Jinyang Huang, Xiang Zhang 0011, Zhi Liu 0002, Huan Yan 0004, Meng Wang 0037, Guohang Zhuang, Yutong Guo, Xiao Sun 0003, Meng Li 0006
IEEE Internet Things J.3
2025 ReSup: Reliable Label Noise Suppression for Facial Expression Recognition
abstract
Because of the ambiguous and subjective property of the facial expression, the label noise is widely existing in the FER dataset. For this problem, in the training phase, current methods often directly predict whether the label is noised or not, aiming to reduce the contribution of the noised data. However, we argue that this kind of method suffers from the low reliability of such noise data decision operation. It makes that some mistakenly abounded clean data are not utilized sufficiently and some mistakenly kept noised data disturbing the model learning. In this paper, we propose a more reliable noise-label suppression method called ReSup. First, instead of directly predicting noised or not, ReSup makes the noise data decision by modeling the distribution of noise and clean labels simultaneously according to the disagreement between the prediction and the target. Specifically, to achieve optimal distribution modeling, ReSup models the similarity distribution of all samples. To further enhance the reliability of our noise decision results, ReSup uses two networks to jointly achieve noise suppression. Specifically, ReSup utilize the property that two networks are less likely to make the same mistakes, making two networks swap decisions and tending to trust decisions with high agreement. Extensive experiments on popular datasets shows the effectiveness of ReSup.
Xiang Zhang 0011, Yan Lu 0001, Huan Yan 0005, Jinyang Huang, Yu Gu 0003, Yusheng Ji, Zhi Liu 0002, Bin Liu 0016
IEEE Trans. Affect. Comput.1
2025 WiOpen: A Robust Wi-Fi-Based Open-Set Gesture Recognition Framework
abstract
Recent years have witnessed a growing interest in Wi-Fi-based gesture recognition. However, existing works have predominantly focused on closed-set paradigms, where all testing gestures are predefined during training. This poses a significant challenge in real-world applications, as unseen gestures might be misclassified as known class during testing. To address this issue, we propose WiOpen, a robust Wi-Fi-based open-set gesture recognition (OSGR) framework. Implementing OSGR requires addressing challenges caused by the unique uncertainty in Wi-Fi sensing. This uncertainty, resulting from noise and domains, leads to widely scattered and irregular data distributions in collected Wi-Fi sensing data. Consequently, data ambiguity between classes and challenges in defining appropriate decision boundaries to identify unknowns arise. To tackle these challenges, WiOpen adopts a twofold approach to eliminate uncertainty and define precise decision boundaries. Initially, it addresses uncertainty induced by noise during data preprocessing by utilizing the channel state information (CSI) ratio. Next, it designs the OSGR network based on an uncertainty quantification method. Throughout the learning process, this network effectively mitigates uncertainty stemming from domains. Ultimately, the network leverages relationships among samples' neighbors to dynamically define open-set decision boundaries, successfully realizing OSGR. Comprehensive experiments on publicly accessible datasets confirm WiOpen's effectiveness.
Xiang Zhang 0011, Jinyang Huang, Huan Yan 0004, Yuanhao Feng, Peng Zhao 0024, Guohang Zhuang, Zhi Liu 0002, Bin Liu 0016
IEEE Trans. Hum. Mach. Syst.1
2025 RF-Eye: Commodity RFID Can Know What You Write and Who You Are Wherever You Are
abstract
Handwriting recognition systems have greatly enhanced AIoT applications, especially in human-computer interaction. Wireless-based methods, favored for their non-invasive nature and ease of deployment, are becoming more common. However, existing works, which typically depend on the user’s position, often perform poorly in varied writing positions. Additionally, they do not incorporate user identity information, which could lead to security vulnerabilities by failing to reject unauthorized users. To address these issues, this article introduces RF-Eye , a system that enables contactless, position-independent handwriting recognition and user identification without prior training. Its innovative approach uses each Radio-frequency identification (RFID) tag as a unique viewpoint for observing hand movements and employs pairs of tags to track directional changes. Specifically, building upon the signal transmission model and the Fresnel Zone, we propose a novel feature, DCG , to capture changes in gesture direction and confirm its consistency across different positions. Based on DCG , we develop unique patterns for common handwriting symbols that enhance our recognition algorithm. Moreover, to strengthen the system security, we link these patterns with distinct handwriting styles through the extraction of finer-grained features, thus, preventing the misuse of the system by unauthorized users. Extensive experiments demonstrate RF-Eye ’s efficacy, which achieves recognition accuracies of 93.5%, 95.2%, and 95.8% for 26 lowercase letters, 10 digits, and 10 graphic symbols, respectively, and identifying unauthorized users with 98.6% accuracy.
Yuanhao Feng, Jinyang Huang, Xiang Zhang 0011, Meng Li 0006, Fusang Zhang, Tianyue Zheng, Anran Li 0001, Mianxiong Dong, Zhi Liu 0002
ACM Trans. Sens. Networks4
2024 DM-NAI: Dynamic Information Diffusion Model Incorporating Non-Adjacent Node Interaction
abstract
Describing the dynamics of information diffusion within social networks poses a formidable challenge. Despite multiple endeavors aimed at addressing this issue, only a limited number of studies have effectively replicated and forecasted the evolving course of information diffusion. In this paper, we propose a novel model, DM-NAI, which not only considers the information transfer between adjacent users but also takes into account the information transfer between non-adjacent users to comprehensively depict the information diffusion process. Extensive experiments are conducted on six datasets to predict the information diffusion range and the diffusion trend of the social network. The experimental results demonstrate an average prediction accuracy range of 94.62% to 96.71%, respectively, significantly outperforming state-of-the-art solutions. This finding illustrates that considering information transmission between non-adjacent users helps DM-NAI achieve more accurate information diffusion predictions.
Jinyang Huang, Xiang Zhang 0011, Peng Zhao 0024, Guohang Zhuang, Huan Yan 0004, Xiao Sun 0003, Meng Wang 0037
ICC3
2024 UAPE: Information Propagation Model Based on User Attitude and Public Opinion Environment
abstract
Modeling the information propagation process in social networks is a challenging problem. Despite numerous attempts to address this issue, existing studies often assume that user attitudes have only one opportunity to alter during the information propagation process. Additionally, these studies tend to consider the transformation of user attitudes as solely influenced by a single user, overlooking the dynamic and evolving nature of user attitudes and the impact of the public opinion environment. In this paper, we propose a novel model, UAPE, which considers the influence of the aforementioned factors on the information propagation process. Specifically, UAPE regards the user's attitude towards the topic as dynamically changing, with the change jointly affected by multiple users simultaneously. Furthermore, the joint influence of multiple users can be considered as the impact of the public opinion environment. Extensive experimental results demonstrate that the model achieves an accuracy range of 91.62% to 94.01 %, surpassing the performance of existing research.
Jinyang Huang, Xiang Zhang 0011, Peng Zhao 0024, Guohang Zhuang, Huan Yan 0004, Xiao Sun 0003, Meng Wang 0037
ICC3
2024 FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial Landmarks
abstract
Depression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment and management of depression. Recently, there are many end-to-end deep learning methods leveraging the facial expression features for automatic depression detection. However, most current methods overlook the temporal dynamics of facial expressions. Although very recent 3DCNN methods remedy this gap, they introduce more computational cost due to the selection of CNN-based backbones and redundant facial features. To address the above limitations, by considering the timing correlation of facial expressions, we propose a novel framework called FacialPulse, which recognizes depression with high accuracy and speed. By harnessing the bidirectional nature and proficiently addressing long-term dependencies, the Facial Motion Modeling Module (FMMM) is designed in FacialPulse to fully capture temporal features. Since the proposed FMMM has parallel processing capabilities and has the gate mechanism to mitigate gradient vanishing, this module can also significantly boost the training speed. Besides, to effectively use facial landmarks to replace original images to decrease information redundancy, a Facial Landmark Calibration Module (FLCM) is designed to eliminate facial landmark errors to further improve recognition accuracy. Extensive experiments on the AVEC2014 dataset and MMDA dataset (a depression dataset) demonstrate the superiority of FacialPulse on recognition accuracy and speed, with the average MAE (Mean Absolute Error) decreased by 21% compared to baselines, and the recognition speed increased by 100% compared to state-of-the-art methods. Codes are released at https://github.com/volatileee/FacialPulse.
Jinyang Huang, Jie Zhang 0042, Xin Liu 0104, Xiang Zhang 0011, Zhi Liu 0002, Peng Zhao 0024, Sigui Chen, Xiao Sun 0003
ACM Multimedia5
2024 Hidden WiFi Camera Localization via Signal Propagation Path Analysis
abstract
Hidden WiFi cameras pose significant privacy threats, necessitating effective localization methods. In this work, we introduce CamLoPA, a system designed for the detection and localization of WiFi cameras. CamLoPA achieves this in just 45 seconds of user walking. It begins by analyzing the causal relationship between WiFi traffic and user movement to identify the presence of a snooping camera. Upon detection, CamLoPA utilizes a novel azimuth location model based on WiFi signal propagation path analysis to localize the hidden camera. Comprehensive evaluations demonstrate that CamLoPA can accurately and swiftly detect and localize snooping WiFi cameras with minimal constraints.
Xiang Zhang 0011, Zehua Ma, Jinyang Huang, Huan Yan 0004, Meng Li 0006, Zhi Liu 0002, Bin Liu 0016
MobiCom1
2024 KeystrokeSniffer: An Off-the-Shelf Smartphone Can Eavesdrop on Your Privacy From Anywhere
abstract
With mobile phones becoming increasingly prevalent and embedding high-quality microphones, attackers have the ability to employ these microphones to eavesdrop user’s keyboard input. However, existing work usually assumes that keystroke eavesdropping is performed against known environments and victims, which inevitably makes attack systems lack generalization. To reveal the real threat of the acoustic signal-based attack strategy, this paper proposes a keystroke eavesdropping algorithm called KeystrokeSniffer, which is robust to unknown input environments and unknown victims. In particular, to mimic the real input environment of victims, an environment estimation algorithm is first designed by extracting the timbre-related characteristics to predict the keyboard type and identifying large-size key data from collected unlabeled samples to estimate the 3D microphone coordinates. Then, by imitating unknown environments and victim data, this algorithm achieves effective keystroke eavesdropping with a small training set. By further considering the commonalities of different keystroke habits, a robust feature extraction method that reflects the keystroke location is adopted to reduce the impact of individual input habits. Extensive experimental results using various commodity smartphones indicate that the scheme is capable of predicting keyboard input accurately under different unknown scenarios. Specifically, even when both the victims and keyboards are unknown, KeystrokeSniffer can still achieve high Top-5 accuracy, reaching 79.5% in predicting keystrokes and 96.7% in predicting meaningful words, which demonstrates KeystrokeSniffer has excellent generalization capabilities. By setting different parameter values of various impact factors, e.g., noise and hand length factors, the strong robustness of the system is demonstrated, which proves that KeystrokeSniffer can violate privacy in real situations.
Jinyang Huang, Jia-Xuan Bai, Xiang Zhang 0011, Zhi Liu 0002, Yuanhao Feng, Jianchun Liu, Xiao Sun 0003, Mianxiong Dong, Meng Li 0006
IEEE Trans. Inf. Forensics Secur.3
2024 PhyFinAtt: An Undetectable Attack Framework Against PHY Layer Fingerprint-Based WiFi Authentication
abstract
WiFi connection has been suffering from MAC forgery attacks due to the loose authentication mechanism between access points (APs) and clients. To address this problem, the physical (PHY) layer information-based fingerprint has been adopted for safe WiFi authentication. Since such a fingerprint is constant and unique for each specific network interface card (NIC), it can effectively prevent MAC forgery attacks. However, the PHY layer information-based fingerprint is still vulnerable to malicious attacks as it is extracted from Channel State Information (CSI), and its stability can be affected by the wireless environment. In this paper, we propose a novel undetectable attack framework, called PhyFinAtt, base on which the attacker can undermine the stability of the PHY layer-based authentication fingerprints through human movement and further attack the WiFi authentication protocols. Specifically, we first demonstrate that human movement at a designated location can affect the PHY fingerprint. We then illustrate the impact of human movement on the PHY fingerprint and the relationship between the movement and the channel quality to ensure that the PHY fingerprint is destroyed by the movement in an undetected way without affecting normal communication. Extensive experiments in real-world scenarios show that our proposed attack can effectively disrupt the stability of the PHY fingerprints and significantly degrade the performance of the authentication protocols based on such fingerprints. To the best of our knowledge, this is the first study on effective attacks against the PHY information-based WiFi authentication protocols. Furthermore, we also present a practical defense mechanism without involving any additional equipment to mitigate attacks similar to PhyFinAtt.
Jinyang Huang, Bin Liu 0016, Chenglin Miao, Xiang Zhang 0011, Jianchun Liu, Lu Su 0001, Zhi Liu 0002, Yu Gu 0003
IEEE Trans. Mob. Comput.4
2023 Bridging the Gap Between BabelNet and HowNet: Unsupervised Sense Alignment and Sememe Prediction
abstract
As the minimum semantic units of natural languages, sememes can provide interpretable representations of concepts.Despite the widespread utilization of lexical resources for semantic tasks, the use of sememes is limited by a lack of available sememe knowledge bases.Recent efforts have been made to connect Ba-belNet with HowNet by automating sememe prediction.However, these methods depend on large manually annotated datasets.Instead, we propose to use sense alignment via a novel unsupervised and explainable method.Our method consists of four stages, each relaxing predefined constraints until a complete alignment of BabelNet synsets to HowNet senses is achieved.Experimental results demonstrate the superiority of our unsupervised method over previous supervised ones by an improvement of 12% overall F1 score, setting a new state of the art.Our work is grounded in an interpretable propagation of sememe information between lexical resources, and may benefit downstream applications which can incorporate sememe information.
Xiang Zhang 0011, Ning Shi, Bradley Hauer, Grzegorz Kondrak
EACL1
2023 Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMs
abstract
Large language models (LLMs) have demonstrated exceptional natural language understanding abilities, and have excelled in a variety of natural language processing (NLP) tasks.Despite the fact that most LLMs are trained predominantly on English, multiple studies have demonstrated their capabilities in a variety of languages.However, fundamental questions persist regarding how LLMs acquire their multilingual abilities and how performance varies across different languages.These inquiries are crucial for the study of LLMs since users and researchers often come from diverse language backgrounds, potentially influencing how they use LLMs and interpret their output.In this work, we propose a systematic way of qualitatively and quantitatively evaluating the multilingual capabilities of LLMs.We investigate the phenomenon of cross-language generalization in LLMs, wherein limited multilingual training data leads to advanced multilingual capabilities.To accomplish this, we employ a novel prompt back-translation method.The results demonstrate that LLMs, such as GPT, can effectively transfer learned knowledge across different languages, yielding relatively consistent results in translation-equivariant tasks, in which the correct output does not depend on the language of the input.However, LLMs struggle to provide accurate results in translation-variant tasks, which lack this property, requiring careful user judgment to evaluate the answers.
Xiang Zhang 0011, Senyu Li, Bradley Hauer, Ning Shi, Grzegorz Kondrak
EMNLP1
2023 Conditional Convolution Residual Network for Efficient Super-Resolution
Yunsheng Guo, Jinyang Huang, Xiang Zhang 0011, Xiao Sun 0003, Yu Gu 0003
ICANN (10)3
2023 Dynamic Memory-Based Continual Learning with Generating and Screening
Siying Tao, Jinyang Huang, Xiang Zhang 0011, Xiao Sun 0003, Yu Gu 0003
ICANN (3)3
2023 WiFE: WiFi and Vision Based Unobtrusive Emotion Recognition via Gesture and Facial Expression
abstract
Emotion plays a critical role in making the computer more human-like. As the first and most essential step, emotion recognition emerges recently as a hot but relatively nascent topic, i.e., current research mainly focuses on single modality (e.g., facial expression) while human emotion expressions are multi-modal in nature. To this end, we propose an unobtrusive emotion recognition system leveraging two emotion-rich and tightly-coupled modalities, i.e., gesture and facial expression. The system design faces two major challenges, namely, how to capture the emotional expression in both modalities without disturbing the subject and how to leverage the relationship between modalities for recognizing the emotion. For the former, we explore WiFi and vision for unobtrusive and contactless gesture and facial expression sensing, respectively. For the latter, we propose a novel deep learning framework named Multi-Source Learning (MSL) to efficiently exploit both self-correlation in the modality and cross-correlation between modalities for fine-grained emotion recognition. To evaluate the proposed method, we prototype the system on low-cost commodity WiFi and vision devices, build a first-of-its-kind WiFi-Vision emotion dataset, and conduct extensive experiments. Empirical results not only verify the effectiveness of WiFE in emotion recognition, but also confirm the superiority of multi-modality over single-modality.
Yu Gu 0003, Xiang Zhang 0011, Huan Yan 0005, Jingyang Huang, Zhi Liu 0002, Mianxiong Dong, Fuji Ren
IEEE Trans. Affect. Comput.2
2023 Toward Facial Expression Recognition in the Wild via Noise-Tolerant Network
abstract
Facial Expression Recognition (FER) has recently emerged as a crucial area in Human-Computer Interaction (HCI) system for understanding the user’s inner state and intention. However, feature- and label-noise constitute the major challenge for FER in the wild due to the ambiguity of facial expressions worsened by low-quality images. To deal with this problem, in this paper, we propose a simple but effective Facial Expression Noise-tolerant Network (FENN) which explores the inter-class correlations for mitigating ambiguity that usually happens between morphologically similar classes. Specifically, FENN leverages a multivariate normal distribution to model such correlations at the final hidden layer of the neural network to suppress the heteroscedastic uncertainty caused by inter-class label noise. Furthermore, the discriminative ability of deep features is weakened by the subtle differences between expressions and the presence of feature noise. FENN utilizes a feature-noise mitigation module to extract compact intra-class feature representations under feature noise while preserving the intrinsic inter-class relationships. We conduct extensive experiments to evaluate the effectiveness of FENN on both original annotated images and synthetic noisy annotated images from RAF-DB, AffectNet, and FERPlus in-the-wild facial expression datasets. The results show that FENN significantly outperforms state-of-the-art FER methods.
Yu Gu 0003, Huan Yan 0005, Xiang Zhang 0011, Yantong Wang, Yusheng Ji, Fuji Ren
IEEE Trans. Circuits Syst. Video Technol.3
2023 Wital: A COTS WiFi Devices Based Vital Signs Monitoring System Using NLOS Sensing Model
abstract
Vital sign (breathing and heartbeat) monitoring is essential for patient care and sleep disease prevention. Most current solutions are based on wearable sensors or cameras; however, the former could affect sleep quality, while the latter often present privacy concerns. To address these shortcomings, we propose Wital, a contactless vital sign monitoring system based on low-cost and widespread commercial off-the-shelf (COTS) Wi-Fi devices. There are two challenges that need to be overcome. First, the torso deformations caused by breathing/heartbeats are weak. How can such deformations be effectively captured? Second, movements such as turning over affect the accuracy of vital sign monitoring. How can such detrimental effects be avoided? For the former, we propose a non-line-of-sight (NLOS) sensing model for modeling the relationship between the energy ratio of line-of-sight (LOS) to NLOS signals and the vital sign monitoring capability using Ricean K theory and use this model to guide the system construction to better capture the deformations caused by breathing/heartbeats. For the latter, we propose a motion segmentation method based on motion regularity detection that accurately distinguishes respiration from other motions, and we remove periods that include movements such as turning over to eliminate detrimental effects. We have implemented and validated Wital on low-cost COTS devices. The experimental results demonstrate the effectiveness of Wital in monitoring vital signs.
Xiang Zhang 0011, Yu Gu 0003, Huan Yan 0005, Yantong Wang, Mianxiong Dong, Kaoru Ota, Fuji Ren, Yusheng Ji
IEEE Trans. Hum. Mach. Syst.1
2022 WiFi and Vision enabled Multimodal Emotion Recognition
abstract
Emotion recognition plays a vital role in current research on human-computer interaction, and human emotion expressions are multi-modal. In this paper, we propose a passive multi-modal emotion recognition system based on facial expression and gesture. To achieve the system design, two major challenges must be addressed, namely, how to capture facial expression and gesture without disturbing the subject, and how to use the correlation between the two modalities to better recognize emotions. For the former, we use WiFi and vision for the passive gesture and facial expression capture, respectively. For the latter, we design a Multi-Source Learning method inspired by Multi-Task Learning to efficiently exploit the correlation between modalities for better emotion recognition. Finally, to evaluate the effectiveness of our system, we use low-cost vision and WiFi devices to prototype the system and build a WiFi-Vision emotion dataset for related research, and we verify the effectiveness of our system in emotion recognition and the superiority of multi-modality over single-modality through conduct extensive experiments.
Yuanwei Hou, Xiang Zhang 0011, Yu Gu 0003, Weiping Li 0002
ICC2
2022 Mitigating Label-Noise for Facial Expression Recognition in the Wild
abstract
Label-noise constitutes a major challenge for facial expression recognition in the wild due to the ambiguity of facial expressions worsened by low-quality images. To deal with this problem, we propose a simple but effective Label-noise Robust Network (LRN) which explores the inter-class correlations for mitigating ambiguity that usually happens between morphologically similar classes. Specifically, LRN leverages a multivariate normal distribution to model such correlations at the final hidden layer of the neural network to suppress the heteroscedastic uncertainty caused by inter-class label noise. Furthermore, LRN utilizes a confidence-based label-free loss to extract compact intra-class feature representations under label noise while preserving the intrinsic inter-class relationships. Experiments on three in-the-wild facial expression datasets demonstrates the superiority of our method.
Huan Yan 0005, Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Yusheng Ji, Fuji Ren
ICME3
2022 WiGRUNT: WiFi-Enabled Gesture Recognition Using Dual-Attention Network
abstract
Gestures constitute an important form of nonverbal communication where bodily actions are used for delivering messages alone or in parallel with spoken words. Recently, there exists an emerging trend of WiFi sensing-enabled gesture recognition due to its inherent merits like remote sensing, non-line-of-sight covering, and privacy-friendly. However, current WiFi-based approaches mainly reply on domain-specific training since they don’t know “where to look” and “when to look.” To this end, we propose WiGRUNT, a WiFi-enabled gesture recognition system using dual-attention network, to mimic how a keen human being intercepting a gesture regardless of the environment variations. The key insight is to train the network to dynamically focus on the domain-independent features of a gesture on the WiFi channel state information via a spatial-temporal dual-attention mechanism. WiGRUNT roots in a deep residual network (ResNet) backbone to evaluate the importance of spatial-temporal clues and exploit their inbuilt sequential correlations for fine-grained gesture recognition. We evaluate WiGRUNT on the open Widar3 dataset and show that it significantly outperforms its state-of-the-art rivals by achieving the best-ever performance in-domain or cross-domain.
Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Meng Wang 0001, Huan Yan 0005, Yusheng Ji, Zhi Liu 0002, Jianhua Li 0003, Mianxiong Dong
IEEE Trans. Hum. Mach. Syst.2
2021 Real-time Vital Signs Monitoring Based on COTS WiFi Devices
abstract
Real-time vital signs (breathing and heartbeat) monitoring is essential for patient care and sleep disease prevention. Current solutions are mostly based on wearable sensors or cameras, the former affects the quality of sleep, while the latter is not conducive to privacy protection, and the cost of these methods is usually expensive. In this paper, we propose Wital, a real-time vital signs monitoring system based on the low-cost and widespread COTS WiFi device. Most of the existing WiFi-based vital signs monitoring solutions utilize the line of sight (LOS) WiFi signals to achieve powerful performance. However, in our daily environments, NLOS sensing is more common. In this article, we first model the relationship between the energy ratio of LOS/NLOS signals and the ability to monitor vital signs based on the Ricean-K theory and theoretically prove that blocking LOS signals in NLOS sensing is more beneficial. We have also established a real-time vital signs monitoring system to verify our method, and the experimental results prove the effectiveness of our method.
Yu Gu 0003, Xiang Zhang 0011, Huan Yan 0005, Zhi Liu 0002, Yusheng Ji
BIBM2
2021 WiONE: One-Shot Learning for Environment-Robust Device-Free User Authentication via Commodity Wi-Fi in Man-Machine System
abstract
User authentication is the first and most critical step in protecting a man-machine system from a malicious spoofer. However, security and privacy are just like the two sides of one coin, hard to see both at the same time, especially by the current mainstream credential- and biometric-based approaches. To this end, we propose WiONE, a safe and privacy-preserving user authentication system leveraging the ubiquitous Wi-Fi infrastructure by exploring “how you behave” rather than “who you are”. The key idea is to apply deep learning to user physical behavior captured by Wi-Fi channel state information (CSI) to identify legitimate users while rejecting spoofers. The design of WiONE faces two challenges, namely, how to capture the subtle behavior, such as a keystroke on CSI, and how to mitigate the heavy environment-specific training required by deep learning. For the former, we design a behavior enhancement model based on the Rician fading to highlight the behavior-induced information by suppressing the behavior-unrelated information on channel response. For the latter, we develop a behavior characterization method tailored for the prototypical networks to facilitate the extraction of the domain-independent behavioral features and enable one-shot recognition of a new user in a new environment. Numerous experiments are conducted in several real-world environments, and the results show that WiONE outperforms its state-of-the-art rivals in authentication performance with much less training effort.
Yu Gu 0003, Huan Yan 0005, Mianxiong Dong, Meng Wang 0037, Xiang Zhang 0011, Zhi Liu 0002, Fuji Ren
IEEE Trans. Comput. Soc. Syst.5
2019 WiFi-Based Real-Time Breathing and Heart Rate Monitoring during Sleep
abstract
Good quality sleep is essential for good health and sleep monitoring becomes a vital research topic. This paper provides a low cost, continuous and contactless WiFi-based vital signs (breathing and heart rate) monitoring method. In particular, we set up the antennas based on Fresnel diffraction model and signal propagation theory, which enhances the detection of weak breathing/heartbeat motion. We implement a prototype system using the off-shelf devices and a real-time processing system to monitor vital signs in real time. The experimental results indicate the accurate breathing rate and heart rate detection performance. To the best of our knowledge, this is the first work to use a pair of WiFi devices and omnidirectional antennas to achieve real-time individual breathing rate and heart rate monitoring in different sleeping postures.
Yu Gu 0003, Xiang Zhang 0011, Zhi Liu 0002, Fuji Ren
GLOBECOM2
2018 Your WiFi Knows How You Behave: Leveraging WiFi Channel Data for Behavior Analysis
abstract
In this paper, we present WoSense, a device-free and real-time behavior analysis system leveraging only WiFi infrastructures. WoSense aims to remotely recognize various human behaviors like surfing, gaming and working around computers, which are considered to be an essential part of our daily lives both at work and at home. The key of WoSense is to exploit the signal distortions on channel data caused by gestures like finger and hand movements, and then identify possible behaviors via the composite of gestures. Therefore, two critical challenges need to be tackled: how to enhance such insignificant distortions led by micro gestures, how to segment the continuous signals according to different gestures in a real-time manner? For the former, instead of relying on empirical studies like our rivals, WoSense offers a Fresnel zone based model with theoretic understandings between the gestures and signal distortions. For the latter, WoSense employs a light-weight automatic segmentation algorithm exploring the variance feature of channel data. We prototype WoSense on the commodity low-cost WiFi devices and evaluate its performance in extensive real- world experiments. WoSense achieves an average 96.77% accuracy for distinguishing the typing and mousing gestures, and 92.5% accuracy for recognizing four different behaviors, i.e., stationary, surfing, gaming and working.
Yu Gu 0003, Xiang Zhang 0011, Chao Li 0009, Fuji Ren, Jie Li 0002, Zhi Liu 0002
GLOBECOM2