EDBT 2026 Demo / reviewers in the wild / expert
Fei Wang 0037
dblp:52/3194-37
· DBLP profile ↗
27ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-0750-6990ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count Prior
Dachao Han, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003 |
INFOCOM | 5 |
| 2026 | Zero-Effort Cross-Domain Wireless Respiration Monitoring Under Free Movements With Commercial UWB DevicesabstractRespiratory monitoring using wireless technologies has garnered significant attention for its potential in healthcare, smart cockpits, and various applications. Though extensively studied, existing systems face practical challenges in adapting to new data domains without substantial customization efforts. Current solutions attempt to address this limitation through domain-independent feature extraction or cross-domain feature translation, employing either knowledge-based sensing models or data-driven neural networks. However, these approaches typically require additional data collection or model retraining for new domains, significantly hindering their practical deployment. This paper proposes RF-Carer, a fully zero-effort cross-domain respiration monitoring system. Our key innovation lies in building an explainable propagation model to transform any heterogeneous signals under unknown domains into a unified form in the signal processing layer. To further address accidental irrelevant factors, we propose to align the feature spaces while suppressing the noisy ones with contrastive learning. On this basis, we develop a one-fits-all model that requires only one-time training but can adapt to 12 domains with 57 cases like unconstrained movements, unknown users, untrained environments, etc.. To the best of our knowledge, RF-Carer is the first zero-effort cross-domain respiration monitoring work with wireless RF signals and would be a fundamental step toward real-world deployments. Ge Wang 0003, Jiazheng Chen, Zhe Chen 0015, Fei Wang 0037, Cong Zhao 0006, Han Ding 0002, Cui Zhao, Wei Xi 0003, Jinsong Han |
SenSys | 4 |
| 2026 | Active Domain Adaptation for mmWave-Based HAR via R$\acute{e}$e'nyi Entropy-Based Uncertainty EstimationabstractHuman Activity Recognition (HAR) using mmWave radar provides a non-invasive alternative to traditional sensor-based methods but suffers from domain shift, where model performance declines in new users, positions, or environments. To address this, we propose mmADA, an Active Domain Adaptation (ADA) framework that efficiently adapts mmWave-based HAR models with minimal labeled data. mmADA enhances adaptation by introducing Rényi Entropy-based uncertainty estimation to identify and label the most informative target samples. Additionally, it leverages contrastive learning and pseudo-labeling to refine feature alignment using unlabeled data. Evaluations with a TI IWR1443BOOST radar across multiple users, positions, and environments show that mmADA achieves over 90% accuracy in various cross-domain settings. Comparisons with five baselines confirm its superior adaptation performance, while further tests on unseen users, environments, and two additional open-source datasets validate its robustness and generalization. Mingzhi Lin, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous DrivingabstractLiDAR point cloud semantic segmentation plays a crucial role in autonomous driving. In recent years, semi-supervised methods have gained popularity due to their significant reduction in annotation labor and time costs. Current semi-supervised methods typically focus on point cloud spatial distribution or consider short-term temporal representations, e.g., only two adjacent frames, often overlooking the rich long-term temporal properties inherent in autonomous driving scenarios. In driving experience, we observe that nearby objects, such as roads and vehicles, remain stable while driving, whereas distant objects exhibit greater variability in category and shape. This natural phenomenon is also captured by Li-DAR, which reflects lower temporal sensitivity for nearby objects and higher sensitivity for distant ones. To lever-age these characteristics, we propose HiLoTs, which learns high-temporal sensitivity and low-temporal sensitivity representations from continuous LiDAR frames. These representations are further enhanced and fused using a cross-attention mechanism. Additionally, we employ a teacher-student framework to align the representations learned by the labeled and unlabeled branches, effectively utilizing the large amounts of unlabeled data. Experimental results on the SemanticKITTI and nuScenes datasets demonstrate that our proposed HiLoTs outperforms state-of-the-art semi-supervised methods, and achieves performance close to Li-DAR+Camera multimodal approaches. R. D. Lin, Pengcheng Weng, Yinqiao Wang, Han Ding 0002, Jinsong Han, Fei Wang 0037 |
CVPR | 6 |
| 2025 | One Snapshot is All You Need: A Generalized Method for mmWave Signal Generation
Han Ding 0002, Wenxin Sun, Cui Zhao, Ge Wang 0003, Fei Wang 0037, Kun Zhao 0002, Zhi Wang 0002, Wei Xi 0003 |
INFOCOM | 6 |
| 2025 | Poster: Zero-effort Cross-domain Wireless Respiration Monitoring under Free Body MovementabstractWireless respiratory monitoring has garnered significant attention for its potential in various applications. However, existing systems face practical challenges in adapting to new data domains without substantial customization efforts. Current solutions attempt to address this limitation through domain-independent feature extraction or cross-domain feature translation, employing either knowledge-based sensing models or data-driven neural networks. However, these approaches typically require additional data collection or model retraining for new domains, significantly hindering their practical deployment. This paper proposes RF-Carer, a fully zero-effort cross-domain respiration monitoring system. Our key innovation lies in building an explainable propagation model to transform any heterogeneous signals under unknown domains into a unified form in the signal processing layer. To further address accidental irrelevant factors, we propose to align the feature spaces while suppressing the noisy ones with contrastive learning. On this basis, we develop a one-fits-all model that requires only one-time training but can adapt to unknown scenarios with unconstrained user movements, postures, positions, etc. To the best of our knowledge, RF-Carer is the first zero-effort cross-domain respiration monitoring work with wireless RF signals and would be a fundamental step toward real-world deployments.Chen Jiazheng Chen, Ge Wang 0003, Zhe Chen 0015, Fei Wang 0037, Wei Xi 0003, Jinsong Han |
MobiCom | 4 |
| 2025 | CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph DatabasesabstractXiangyan Liu, Bo Lan, Zhiyuan Hu, Yang Liu, Zhicheng Zhang, Fei Wang, Michael Qizhe Shieh, Wenmeng Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xiangyan Liu, Fei Wang 0037, Michael Shieh, Wenmeng Zhou |
NAACL (Long Papers) | 6 |
| 2025 | mmYodar+: Robust Human Detection Using mmWave SignalsabstractThe detection of human objects can be crucial for various real-world applications, such as surveillance and autonomous driving. However, traditional vision-based approaches suffer from limitations such as low lighting conditions, occlusions, and privacy concerns. To address these challenges, we introduce mmYodar+, a novel mmWave-based automatic human detection system. Our system processes mmWave signals to generate a 3D point cloud, which is then transformed into a 2D radar image for easier visualization and analysis. To enhance human profiling, we filter the point cloud using biometric information and expand human-related points in the image based on radar angle resolution, incorporating color to improve the differentiation. Additionally, we employ a deep mutual learning (DML) framework, enabling efficient human detection using a lightweight DNN. Experimental results show that mmYodar+ achieves an average precision of 96.29% in various scenarios, including indoor and outdoor environments, various lighting conditions, and in the presence of occlusions. These results demonstrate the effectiveness of using mmWave radar signals for reliable and accurate human detection. Yuance Chang, Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Zhi Wang 0002, Wei Xi 0003 |
IEEE Internet Things J. | 5 |
| 2025 | BullyDetect: Detecting School Physical Bullying With Wi-Fi and Deep Wavelet TransformerabstractMore than 246 million children and adolescents suffer from school violence and bullying, e.g., verbal harassment, social harassment, and physical bullying, every year, according to a report from the United Nations Educational, Scientific and Cultural Organization. School violence and bullying severely harm the physical and emotional well-being of the victims, increasing the risks of depression, anxiety, sleep difficulties, lower academic achievement, dropping out of school, and even suicide attempts. Since school physical bullying always happens in the low-visibility areas spots of surveillance cameras, in this article, we propose to utilize Wi-Fi, a widely deployed infrastructure, to detect school physical bullying. We design residual wavelet transformer networks to conduct noise removal and action feature learning in an end-to-end manner. Besides, we propose two data augmentation methods in the temporal domain of Wi-Fi signals to simulate the different speeds and extents of bullying actions performed. Extensive evaluation of 20-paired volunteers demonstrates that 1) Wi-Fi can effectively detect physical school bullying; 2) the proposed approaches outperform long-short-time-memory networks, ResNet-1D, vision transformer, etc.; and 3) the proposed data augmentation methods can work as plug-and-play modules to improve the detection accuracy of all the above-mentioned approaches. Fei Wang 0037, Lekun Xia, Fan Nai, Shiqiang Nie, Han Ding 0002, Jinsong Han |
IEEE Internet Things J. | 2 |
| 2025 | You Can Wash Hands Better: Accurate Daily Handwashing Assessment With a SmartwatchabstractHand hygiene is among the most effective daily practices for preventing infectious diseases such as influenza, malaria, and skin infections. While professional guidelines emphasize proper handwashing to reduce the risk of viral infections, surveys reveal that adherence to these recommendations remains low. To address this gap, we propose UWash, a wearable solution leveraging smartwatches to evaluate handwashing procedures, aiming to raise awareness and cultivate high-quality handwashing habits. We frame the task of handwashing assessment as an action segmentation problem, similar to those in computer vision, and introduce a simple yet efficient two-stream UNet-like network to achieve this goal. Experiments involving 51 subjects demonstrate that UWash achieves 92.27% accuracy in handwashing gesture recognition, an error of$\lt $0.5 seconds in onset/offset detection, and an error of$\lt $5 points in gesture scoring under user-dependent settings. The system also performs robustly in user-independent and user-independent-location-independent evaluations. Remarkably, UWash maintains high performance in real-world tests, including evaluations with 10 random passersby at a hospital 9 months later and 10 passersby in an in-the-wild test conducted 2 years later. UWash is the first system to score handwashing quality based on gesture sequences, offering actionable guidance for improving daily hand hygiene. The code and dataset are publicly available athttps://github.com/aiotgroup/UWash. Fei Wang 0037, Xilei Wu, Xin Wang 0195, Han Ding 0002, Jingang Shi, Jinsong Han |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Federated Multi-Source Domain Adaptation for mmWave-Based Human Activity RecognitionabstractContactless mmWave-based human activity recognition (HAR) is essential for various applications, yet most existing approaches often assume consistent environments. Integrating domain adaptation offers a promising solution to this challenge. This prevailing paradigm works well when the source and target data are centralized on a single server while learning to adapt. However, in more universal and practical situations, such as personal health records, users’ biometric information, and financial issues, the raw data is typically protected by different privacy-preserving policies and is stored by multiple parties. Additionally, labeling RF signals in the target domain is a non-trivial and labor-intensive task for most end-users. To address these problems, this paper introduces FMDA, a federated multi-source domain adaptation framework for mmWave-based HAR. FMDA assesses the contribution of each source and performs weighted parameter aggregation for knowledge transfer. This facilitates unsupervised training of the target HAR model without requiring access to any source domain data. Moreover, the model is optimized by minimizing the generalization gaps between the source and target models, benefiting all participants during the learning process and enhancing overall performance. Extensive experiments demonstrate the effectiveness of FMDA. The results indicate that in the target domain, FMDA achieves comparable performance to supervised learning approaches, while also enhancing the efficacy of source domain models to varying degrees. Cui Zhao, Guotong Fang, Han Ding 0002, Fei Wang 0037, Ge Wang 0003, Kun Zhao 0002, Zhi Wang 0002, Wei Xi 0003 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-FiabstractWi-Fi signals, in contrast to cameras, offer privacy protection and occlusion resilience for some practical scenarios such as smart homes, elderly care, and virtual reality. Recent years have seen remarkable progress in the estimation of single-person 2D pose, single-person 3D pose, and multi-person 2D pose. This paper takes a step forward by introducing Person-in- WiFi 3D, a pioneering Wi-Fi system that accomplishes multi-person 3D pose estimation. Person-in- WiFi 3D has two main updates. Firstly, it has a greater number of Wi-Fi devices to enhance the capability for capturing spatial reflections from multiple individuals. Secondly, it leverages the Transformer for end-to-end estimation. Compared to its predecessor, Person-in- WiFi 3D is storage-efficient and fast. We deployed a proof-of-concept system in$4m\times 3.5m$areas and collected a dataset of over 97K frames with seven volunteers. Person-in- WiFi 3D at-tains 3D joint localization errors of 9I.7mm (I-person), I08.Imm (2-person), and I25.3mm (3-person), comparable to cameras and millimeter-wave radars. The project page is at https:/laiotgroup.github.ioIPerson-in-WiFi-3D. Kangwei Yan, Fei Wang 0037, Han Ding 0002, Jinsong Han |
CVPR | 2 |
| 2024 | U-Shape Networks Are Unified Backbones for Human Action Understanding From Wi-Fi SignalsabstractWi-Fi is well-known in communication and deployment for indoor localization. Recently, Wi-Fi has been exploited for human action understanding, e.g., elder fall detection, smoking detection, and hand-gesture recognition. Many researchers have made great efforts to associate their own expertise on Wi-Fi signals with deep networks like convolutional neural networks, recurrent neural networks, and Transformers for representation learning, and demonstrate that expertise can promote action understanding accuracy. However, expert knowledge always relies on the existing personal understanding and assumptions of Wi-Fi signals, which limits the scalability of associated models and may introduce subjective bias to the models. Besides, requiring expertise raises an inescapable barrier and cost to the model design. Recent years have witnessed the great value of backbone networks, such as ResNet, in advancing the research progress in computer vision. We believe a backbone network for Wi-Fi signals will also play a crucial role. In this article, instead of proposing novel algorithms with expertise, we present that simple U-shape deep networks, such as FCN, U-Net, and U-Net++, are efficient and unified backbones for Wi-Fi-based human action understanding tasks, i.e., action recognition, action detection, and action segmentation. Results on three public data sets, Wi-Fi activity recognition, action recognition and indoor localization, and human-to-human interaction, show that all these U-shape deep networks have superior or competitive performance compared with original papers as well as state-of-the-art approaches, e.g., >97% recognition accuracy,90% segmentation accuracy. We envision this work breaks the barrier of network design and facilitates human action understanding from Wi-Fi signals. Fei Wang 0037, Yiao Gao, Han Ding 0002, Jingang Shi, Jinsong Han |
IEEE Internet Things J. | 1 |
| 2024 | Genre Classification Empowered by Knowledge-Embedded Music RepresentationabstractThis paper introduces a pioneering framework for music representation learning, which harnesses knowledge graph embeddings to enrich genre classification. Leveraging metadata from publicly available datasets like FMA and OpenMIC-2018, the constructed knowledge graph delineates intricate relationships among genres, artists, and instruments, offering valuable insights for genre representation. Within this framework, we propose two models tailored for distinct genre classification scenarios: fixed-set genre classification and open-set genre classification. These models exploit the knowledge graph to unveil correlations among different genres and integrate this knowledge into the audio representation. Notably, our approach is the first to merge audio data with high-level knowledge for music genre classification. Experimental results demonstrate that our proposed methods outperform state-of-the-art approaches, achieving an average genre classification accuracy of 68.07% on the FMA-medium dataset and 42.4% for open-set classification on the FMA-large dataset. Han Ding 0002, Linwei Zhai, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003, Zhi Wang 0002, Jizhong Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Exploiting Multi-Scale Parallel Self-Attention and Local Variation via Dual-Branch Transformer-CNN Structure for Face Super-ResolutionabstractRecently, deep learning technique has been widely employed to deal with face super-resolution (FSR) problem. It aims to predict the nonlinear relationship between the low-resolution (LR) face images and corresponding high-resolution (HR) ones, which could recover the high-frequency details from the LR degraded textures. However, either CNN-based or Transformer-based approaches mostly enhance the details by exploiting the relationship of local pixels or patches on LR features, the nonlocal features are not fully taken into account for producing high-frequency textures. To improve the above problem, we design a novel dual-branch module which consists of Transformer and CNN respectively. The Transformer branch extracts multiple scale feature embeddings and explores local and nonlocal self-attention simultaneously. Thus, the parallel self-attention mechanism has superior capabilities to capture the local and nonlocal dependencies on face image in the face reconstruction. Furthermore, the traditional CNNs usually extract features by combining pixels in a local convolutional kernel, it may be not effective to recover lost high-frequency details since the variations of local pixels are not well measured, which is important in recovering vivid edges and contours. To this end, we propose the local variation based attention block on the CNN branch, which could enhance the capabilities by directly extracting features from the variation of neighboring pixels. Finally, the Transformer-branch and CNN-branch are combined together by the modulation block to fuse both nonlocal and local advantages from two branches. Experimental results demonstrate the effectiveness of the proposed method when compared with state-of-the-art approaches. Jingang Shi, Yusi Wang, Zitong Yu, Guanxin Li, Xiaopeng Hong, Fei Wang 0037, Yihong Gong |
IEEE Trans. Multim. | 6 |
| 2023 | Knowledge-Graph Augmented Music Representation for Genre ClassificationabstractIn this paper, we propose KGenre, a knowledge-embedded music representation learning framework for improved genre classification. We construct the knowledge graph from the metadata in the open-source FMA-medium and OpenMIC-2018 datasets, with no extra information/effort required. KGenre then mines the correlation between different genres from the knowledge graph and embeds such correlation in audio representation. To our knowledge, KGenre is the first method fusing the audio with high-level knowledge for music genre classification. Experimental results demonstrate the embedded knowledge can effectively enhance the audio feature representation, and the genre classification performance surpasses the state-of-the-art methods. Han Ding 0002, Wenjing Song, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao |
ICASSP | 4 |
| 2023 | Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-ResolutionabstractRecently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information and induce blurry effect by the irrelevant textures. Some improved methods simply constrain self-attention in a local window to suppress the useless information. But it also limits the capability of recovering high-frequency details when flat areas dominate the local searching window. To improve the above issues, we propose a novel self-refinement mechanism which could adaptively achieve texture-aware reconstruction in a coarse-to-fine procedure. Generally, the primary self-attention is first conducted to reconstruct the coarse-grained textures and detect the fine-grained regions required further compensation. Then, region selection attention is performed to refine the textures on these key regions. Since self-attention considers the channel information on tokens equally, we employ a dual-branch feature integration module to privilege the important channels in feature extraction. Furthermore, we design the wavelet fusion module which integrate shallow-layer structure and deep-layer detailed feature to recover realistic face images in frequency domain. Extensive experiments demonstrate the effectiveness on a variety of datasets. Guanxin Li, Jingang Shi, Yuan Zong, Fei Wang 0037, Tian Wang 0002, Yihong Gong |
IJCAI | 4 |
| 2023 | mmYodar: Lightweight and Robust Object Detection using mmWave SignalsabstractThe detection of human objects can be crucial for various real-world applications, such as surveillance and autonomous driving. However, traditional vision-based approaches suffer from limitations such as low lighting conditions, occlusions, and privacy concerns. To overcome these limitations, we propose a novel automatic object detection system, called mmYodar, which utilizes millimeter-wave (mmWave) radar signals. Our system collects mmWave signals and calculates a 3D point cloud, which is transformed into a radar image for easier visualization and analysis. To improve the system's human profiling capability, we expand the corresponding points in the image with color based on the radar angle resolution. Then, a designed deep mutual learning framework is employed to detect human objects from the expanded image. Experimental results show that mmYodar achieves nearly real-time detection with an average precision of 90.35% in various scenarios, including indoor and outdoor environments, various lighting conditions, and in the presence of occlusions. These results demonstrate the effectiveness of using mmWave radar signals for reliable and accurate human object detection. Our code and dataset are available at https:llgithub.comlbrave20005lmmYodar. Yuance Chang, Han Ding 0002, Dachao Han, Ge Wang 0003, Cui Zhao, Fei Wang 0037, Wei Xi 0003, Jizhong Zhao |
SECON | 7 |
| 2022 | IDPT: Interconnected Dual Pyramid Transformer for Face Super-ResolutionabstractFace Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstructing detailed content from degraded images, FSR technology is still far away from being solved. In this paper, we propose a novel and effective face super-resolution framework based on Transformer, namely Interconnected Dual Pyramid Transformer (IDPT). Instead of straightly stacking cascaded feature reconstruction blocks, the proposed IDPT designs the pyramid encoder/decoder Transformer architecture to extract coarse and detailed facial textures respectively, while the relationship between the dual pyramid Transformers is further explored by a bottom pyramid feature extractor. The pyramid encoder/decoder structure is devised to adapt various characteristics of textures in different spatial spaces hierarchically. A novel fusing modulation module is inserted in each spatial layer to guide the refinement of detailed texture by the corresponding coarse texture, while fusing the shallow-layer coarse feature and corresponding deep-layer detailed feature simultaneously. Extensive experiments and visualizations on various datasets demonstrate the superiority of the proposed method for face super-resolution tasks. Jingang Shi, Yusi Wang, Songlin Dong, Xiaopeng Hong, Zitong Yu, Fei Wang 0037, Changxin Wang, Yihong Gong |
IJCAI | 6 |
| 2021 | Cascade Convolutional Neural Network With Progressive Optimization for Motor Fault Diagnosis Under Nonstationary ConditionsabstractRecently, convolutional neural networks (CNNs) have been successfully used for motor fault diagnosis because of its powerful feature extraction ability. However, there are still some barriers of traditional CNNs. Due to the fact of the hierarchical structure, feature resolution of CNNs will be reduced with layer growth, which can lead to the information loss. In addition, the fixed kernel size makes traditional CNNs not suitable for fault diagnosis of motors, which are widely used in nonstationary conditions. Therefore, starting from the physical characteristics of nonstationary vibration signals, a cascade CNN (C-CNN) with progressive optimization is proposed in this article. First, a cascade structure is built to avoid the information loss caused by consecutive convolution striding or pooling. Then, dilated convolution operations are implemented, which can extract the feature maps from different scales and extend the applications of CNN to nonstationary conditions. Furthermore, taking the advantage of the cascade structure, a progressive optimization algorithm is proposed for divide-and-conquer parameters optimization, which enables the C-CNN to converge to a more optimum state and improve the diagnosis performance. The proposed method is verified by two motor fault diagnosis experiments, which are conducted under constant speed and variable speed, respectively. The results show that the proposed method can achieve better performance when rotating speed is either constant or changing than exiting methods. Fei Wang 0037, Qinghua Hu, Xuefeng Chen 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Hierarchical and Interactive Refinement Network for Edge-Preserving Salient Object DetectionabstractSalient object detection has undergone a very rapid development with the blooming of Deep Neural Network (DNN), which is usually taken as an important preprocessing procedure in various computer vision tasks. However, the down-sampling operations, such as pooling and striding, always make the final predictions blurred at edges, which has seriously degenerated the performance of salient object detection. In this paper, we propose a simple yet effective approach, i.e., Hierarchical and Interactive Refinement Network (HIRN), to preserve the edge structures in detecting salient objects. In particular, a novel multi-stage and dual-path network structure is designed to estimate the salient edges and regions from the low-level and high-level feature maps, respectively. As a result, the predicted regions will become more accurate by enhancing the weak responses at edges, while the predicted edges will become more semantic by suppressing the false positives in background. Once the salient maps of edges and regions are obtained at the output layers, a novel edge-guided inference algorithm is introduced to further filter the resulting regions along the predicted edges. Extensive experiments on several benchmark datasets have been conducted, in which the results show that our method significantly outperforms a variety of state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Jimuyang Zhang, Fei Wang 0037, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | RFnet: Automatic Gesture Recognition and Human Identification Using Time Series RFID Signals
Han Ding 0002, Cui Zhao, Fei Wang 0037, Ge Wang 0003, Zhiping Jiang, Wei Xi 0003, Jizhong Zhao |
Mob. Networks Appl. | 4 |
| 2020 | Multiscale Kernel Based Residual Convolutional Neural Network for Motor Fault Diagnosis Under Nonstationary ConditionsabstractMotor fault diagnosis is imperative to enhance the reliability and security of industrial systems. However, since motors are often operated under nonstationary conditions, the high complexity of vibration signals raises notable difficulties for fault diagnosis. Therefore, considering the special physical characteristics of motor signals under nonstationary conditions, in this article, we propose a multiscale kernel based residual convolutional neural network (CNN) for motor fault diagnosis. Our contributions mainly fall into two aspects. First, we notice that each motor fault category has various patterns in vibration signals due to the changing operational conditions of the motor. To capture these patterns, a multiscale kernel algorithm is applied in the CNN architecture. Second, since the motor vibration signals are made up of many different components from different transfer paths, they are very complex and variable. To enable the architecture to extract fault features from deep and hierarchical representation spaces, sufficient depth of the network is needed, which will lead to the degradation problem. In the proposed method, residual learning is embedded into the multiscale kernel CNN to avoid performance degradation and build a deeper network. To validate the effectiveness of the proposed networks, a normal motor and five motors with different failures are tested. The results and comparisons with state-of-the-art methods highlight the superiority of the proposed method. Fei Wang 0037, Boyuan Yang 0002, S. Joe Qin |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | WiPIN: Operation-Free Passive Person Identification Using Wi-Fi SignalsabstractWi-Fi signals-based person identification attracts increasing attention in the booming Internet-of-Things era mainly due to its pervasiveness and passiveness. Most previous work applies gaits extracted from WiFi distortions caused by person walking to achieve the identification. However, to extract useful gait, a person must walk along a pre-defined path for several meters, which requires user high collaboration and increases identification time overhead, thus limiting use scenarios. Moreover, gait based work has severe shortcoming in identification performance, especially when the user volume is large. In order to eliminate above limitations, in this paper, we present an operation-free person identification system, namely WiPIN, that requires least user collaboration and achieves good performance. WiPIN is based on an entirely new insight that Wi-Fi signals would carry person body information when propagating through the body, which is potentially discriminated for person identification. Then we demonstrate the feasibility on commodity off-the-shelf Wi-Fi devices by well- designed signal pre-processing, feature extraction, and identity matching algorithms. Results show that WiPIN achieves 92% identification accuracy over 30 users, high robustness to various experimental settings, and low identifying time overhead, i.e., less than 300ms. Fei Wang 0037, Jinsong Han, Feng Lin 0004, Kui Ren 0001 |
GLOBECOM | 1 |
| 2019 | Person-in-WiFi: Fine-Grained Person Perception Using WiFiabstractFine-grained person perception such as body segmentation and pose estimation has been achieved with many 2D and 3D sensors such as RGB/depth cameras, radars (e.g. RF-Pose), and LiDARs. These solutions require 2D images, depth maps or 3D point clouds of person bodies as input. In this paper, we take one step forward to show that fine-grained person perception is possible even with 1D sensors: WiFi antennas. Specifically, we used two sets of WiFi antennas to acquire signals, i.e., one transmitter set and one receiver set. Each set contains three antennas horizontally lined-up as a regular household WiFi router. The WiFi signal generated by a transmitter antenna, penetrates through and reflects on human bodies, furniture, and walls, and then superposes at a receiver antenna as 1D signal samples. We developed a deep learning approach that uses annotations on 2D images, takes the received 1D WiFi signals as input, and performs body segmentation and pose estimation in an end-to-end manner. To our knowledge, our solution is the first work based on off-the-shelf WiFi antennas and standard IEEE 802.11n WiFi signals. Demonstrating comparable results to image-based solutions, our WiFi-based person perception solution is cheaper and more ubiquitous than radars and LiDARs, while invariant to illumination and has little privacy concern comparing to cameras. Fei Wang 0037, Sanping Zhou, Stanislav Panev, Jinsong Han |
ICCV | 1 |
| 2019 | Discriminative Feature Learning With Consistent Attention Regularization for Person Re-IdentificationabstractPerson re-identification (Re-ID) has undergone a rapid development with the blooming of deep neural network. Most methods are very easily affected by target misalignment and background clutter in the training process. In this paper, we propose a simple yet effective feedforward attention network to address the two mentioned problems, in which a novel consistent attention regularizer and an improved triplet loss are designed to learn foreground attentive features for person Re-ID. Specifically, the consistent attention regularizer aims to keep the deduced foreground masks similar from the low-level, mid-level and high-level feature maps. As a result, the network will focus on the foreground regions at the lower layers, which is benefit to learn discriminative features from the foreground regions at the higher layers. Last but not least, the improved triplet loss is introduced to enhance the feature learning capability, which can jointly minimize the intra-class distance and maximize the inter-class distance in each triplet unit. Experimental results on the Market1501, DukeMTMC-reID and CUHK03 datasets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Fei Wang 0037, Zeyi Huang, Jinjun Wang |
ICCV | 2 |
| 2019 | Continuous User Authentication by Contactless Wireless SensingabstractThis paper presents BodyPIN, which is a continuous user authentication system by contactless wireless sensing using commodity Wi-Fi. BodyPIN can track the current user's legal identity throughout a computer system's execution. In case the authentication fails, the consequent accesses will be denied to protect the system. The recent rich wireless-based user identification designs cannot be applied to BodyPIN directly, because they identify a user's various activities, rather than the user herself. The enforced to be performed activities can thus interrupt the user's operations on the system, highly inconvenient and not user-friendly. In this paper, we leverage the bio-electromagnetics domain human model for quantifying the impact of human body on the bypassing Wi-Fi signals and deriving the component that indicates a user's identity. Then, we extract suitable Wi-Fi signal features to fully represent such an identity component, based on which we fulfill the continuous user authentication design. We implement a BodyPIN prototype by commodity Wi-Fi NICs without any extra or dedicated wireless hardware. We show that BodyPIN achieves promising authentication performances, which is also lightweight and robust under various practical settings. Fei Wang 0037, Zhenjiang Li 0001, Jinsong Han |
IEEE Internet Things J. | 1 |