VLDB 2026 Research / reviewers in the wild / expert
Wei Cui 0002
dblp:42/3805-2
· DBLP profile ↗
26ranked-venue papers
2as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataabstractSensory Temporal Action Detection (STAD) aims to localize and classify human actions within long, untrimmed sequences captured by non-visual sensors such as WiFi or inertial measurement units (IMUs). Unlike video-based TAD, STAD poses unique challenges due to the low-dimensional, noisy, and heterogeneous nature of sensory data, as well as the real-time and resource constraints on edge devices. While recent STAD models have improved detection performance, their high computational cost hampers practical deployment. In this paper, we propose SlimSTAD, a simple yet effective framework that achieves both high accuracy and low latency for STAD. SlimSTAD features a novel Decoupled Channel Modeling (DCM) encoder, which preserves modality-specific temporal features and enables efficient inter-channel aggregation via lightweight graph attention. An anchor-free cascade predictor then refines action boundaries and class predictions in a two-stage design without dense proposals. Experiments on two real-world datasets demonstrate that SlimSTAD outperforms strong video-derived and sensory baselines by an average of 2.1 mAP, while significantly reducing GFLOPs, parameters, and latency, validating its effectiveness for real-world, edge-aware STAD deployment. Wei Cui 0002, Lukai Fan, Zhenghua Chen, Min Wu 0008, Shili Xiang, Haixia Wang 0003, Bing Li 0002 |
AAAI | 1 |
| 2026 | RF-Nav: A Robust Fusion-Based GNSS-Visual-Inertial Navigation SystemabstractAccurate vehicle navigation plays a critical role in vehicle-to-everything (V2X) applications, including connected transportation systems, intelligent traffic management, and autonomous driving. To address the stringent demands of these scenarios, the integration of the global navigation satellite system (GNSS) with the visual-inertial navigation system (VINS) has emerged as a pivotal advancement. Despite these strides, navigation systems remain susceptible to abnormal data. This data, originating from unpredictable external environments and internal device fallibility, poses a threat of substantial errors and system drift. In this paper, we present RF-Nav, a robust fusion-based GNSS-VINS navigation system with enhanced data processing and dynamic factor correction. The framework innovates with a dual-pronged approach: it first applies adaptive gamma correction with bilateral filtering and contrast-limited adaptive histogram equalization (AGCBF-CLAHE) to refine raw images; then, it deploys a long short-term memory (LSTM) denoising network enhanced with an advanced wavelet threshold for IMU data refinement. This dual enhancement of visual and IMU data integrity is further bolstered by a dynamic factor confidence correction mechanism, rooted in factor graph optimization (FGO), designed to counteract the adverse effects of abnormal data. Extensive experiments on large-scale public and real-field dataset demonstrate that RF-Nav exhibits superior robustness and accuracy in various environments. Pengju Si, Shenzhi Yang, Yongzhe Shi, Huan Wang 0019, Zhumu Fu, Jun Wang 0064, Wei Cui 0002 |
IEEE Internet Things J. | 7 |
| 2025 | SimCast: Enhancing Precipitation Nowcasting with Short-to-Long Term Knowledge DistillationabstractPrecipitation nowcasting predicts future radar sequences based on current observations, which is a highly challenging task driven by the inherent complexity of the Earth system. Accurate nowcasting is of utmost importance for addressing various societal needs, including disaster management, agriculture, transportation, and energy optimization. As a complementary to existing non-autoregressive nowcasting approaches, we investigate the impact of prediction horizons on nowcasting models and propose SimCast, a novel training pipeline featuring a short-to-long term knowledge distillation technique coupled with a weighted MSE loss to prioritize heavy rainfall regions. Improved nowcasting predictions can be obtained without introducing additional overhead during inference. As SimCast generates deterministic predictions, we further integrate it into a diffusion-based framework named CasCast, leveraging the strengths from probabilistic models to overcome limitations such as blurriness and distribution shift in deterministic outputs. Extensive experimental results on three benchmark datasets validate the effectiveness of the proposed framework, achieving mean CSI scores of 0.452 on SEVIR, 0.474 on HKO-7, and 0.361 on MeteoNet, which outperforms existing approaches by a significant margin. Yifang Yin, Shengkai Chen, Yiyao Li, Lu Wang 0003, Ruibing Jin, Wei Cui 0002, Shili Xiang |
ICME | 6 |
| 2025 | STADe: Sensory Temporal Action Detection via Temporal-Spectral Representation LearningabstractTemporal action detection (TAD) is a vital challenge in computer vision and the Internet of Things, aiming to detect and identify actions within temporal sequences. While TAD has primarily been associated with video data, its applications can also be extended to sensor data, opening up opportunities for various real-world applications. However, applying existing TAD models to sensory signals presents distinct challenges such as varying sampling rates, intricate pattern structures, and subtle, noise-prone patterns. In response to these challenges, we propose a Sensory Temporal Action Detection (STADe) model. STADe leverages Fourier kernels and adaptive frequency filtering to adaptively capture the nuanced interplay of temporal and frequency features underlying complex patterns. Moreover, STADe embraces adaptability by employing deep fusion at varying resolutions and scales, making it versatile enough to accommodate diverse data characteristics, such as the wide spectrum of sampling rates and action durations encountered in sensory signals. Unlike conventional models with unidirectional category-to-proposal dependencies, STADe adopts a cross-cascade predictor to introduce bidirectional and temporal dependencies within categories. To extensively evaluate STADe and promote future research in sensory TAD, we establish three diverse datasets using various sensors, featuring diverse sensor types, action categories, and sampling rates. Experiments across one public and our three new datasets demonstrate STADe's superior performance over state-of-the-art TAD models in sensory TAD tasks. Bing Li 0002, Haotian Duan, Yun Liu 0011, Le Zhang 0001, Wei Cui 0002, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | BiASAM: Bidirectional-Attention Guided Segment Anything Model for Very Few-Shot Medical Image SegmentationabstractThe Segment Anything Model (SAM) excels in general segmentation but encounters difficulties in medical imaging due to few-shot learning challenges, particularly with extremely limited annotated data. Existing approaches often suffer from insufficient feature extraction and inadequate loss function balancing, resulting in decreased accuracy and poor generalization. To address these issues, we propose BiASAM, which uniquely incorporates two bidirectional attention mechanisms into SAM for medical image segmentation. Firstly, BiASAM integrates a spatial-frequency attention module to improve feature extraction, enhancing the model's ability to capture both fine and coarse details. Secondly, we employ an attention-based gradient update mechanism that dynamically adjusts loss weights, boosting the model's learning efficiency and adaptability in data-scarce scenarios. Additionally, BiASAM utilizes the point and box fusion prompt to enhance segmentation precision at both global and local levels. Experiments across various medical datasets show BiASAM achieves performance comparable to fully supervised methods with just two labeled samples. Wei Zhou 0003, Guilin Guan, Wei Cui 0002, Yugen Yi |
IEEE Signal Process. Lett. | 3 |
| 2024 | Diffusion-driven Dual-flow Source-Free Domain Adaptation for Medical Image SegmentationabstractSource-Free Domain Adaptation (SFDA) aims to adapt a pre-trained model to unlabeled target domain data without access to source domain data, presenting a significant challenge for medical image segmentation. Most current approaches address this challenge through self-training, employing manually augmented target domain images and pseudo-labels to enforce consistency regularization. However, these approaches still encounter two primary issues. Firstly, manually augmented consistency self-training results in performance degradation due to the semantic mismatch between the target domain images and noisy pseudo-labels. Secondly, they fail to fully exploit the informative content present in the target domain, exhibiting inadequate adaptability, particularly in significant domain gaps. To address these, we introduce the Diffusion-driven Dual-flow SFDA (D2SFDA), the pioneering framework to integrate a diffusion model into SFDA for medical image segmentation. Our D2SFDA framework comprises two novel components: the Diffusion Perturbation Flow (DPF) and the Twin-Knowledge Investigation Flow (TKIF). DPF utilizes pseudo-labels to generate diverse and semantically consistent diffusion views, providing more realistic supervision, potentially enhancing model stability. Surprisingly, DPF using only diffusion images outperforms self-training using real images, as evidenced by the superior average Dice score on the BASE1 target domain of the RIGA+ dataset (90.31% vs. 85.64%). Additionally, TKIF rigorously analyzes the target domain with dual-focus consistency regularization on domain-invariant and target domain-specific knowledge, effectively reducing domain gaps, resulting in an improvement from 90.31% to 91.79%. Extensive experiments on two cross-domain datasets confirm that our D2SFDA surpasses state-of-the-art SFDA approaches in effectively addressing domain shift issues. The code is available at https://github.com/M4cheal/D2SFDA. Wei Zhou 0003, Jianhang Ji, Wei Cui 0002, Yugen Yi |
BIBM | 3 |
| 2024 | Unsupervised Domain Adaptation Fundus Image Segmentation via Multi-Scale Adaptive Adversarial LearningabstractSegmentation of the Optic Disc (OD) and Optic Cup (OC) is crucial for the early detection and treatment of glaucoma. Despite the strides made in deep neural networks, incorporating trained segmentation models for clinical application remains challenging due to domain shifts arising from disparities in fundus images across different healthcare institutions. To tackle this challenge, this study introduces an innovative unsupervised domain adaptation technique called Multi-scale Adaptive Adversarial Learning (MAAL), which consists of three key components. The Multi-scale Wasserstein Patch Discriminator (MWPD) module is designed to extract domain-specific features at multiple scales, enhancing domain classification performance and offering valuable guidance for the segmentation network. To further enhance model generalizability and explore domain-invariant features, we introduce the Adaptive Weighted Domain Constraint (AWDC) module. During training, this module dynamically assigns varying weights to different scales, allowing the model to adaptively focus on informative features. Furthermore, the Pixel-level Feature Enhancement (PFE) module enhances low-level features extracted at shallow network layers by incorporating refined high-level features. This integration ensures the preservation of domain-invariant information, effectively addressing domain variation and mitigating the loss of global features. Two publicly accessible fundus image databases are employed to demonstrate the effectiveness of our MAAL method in mitigating model degradation and improving segmentation performance. The achieved results outperform current state-of-the-art (SOTA) methods in both OD and OC segmentation. Wei Zhou 0003, Jianhang Ji, Wei Cui 0002, Yingyuan Wang, Yugen Yi |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Democratizing Federated WiFi-Based Human Activity Recognition Using Hypothesis TransferabstractHuman activity recognition (HAR) is a crucial task in IoT systems with applications ranging from surveillance and intruder detection to home automation and more. Recently, non-invasive HAR utilizing WiFi signals has gained considerable attention due to advancements in ubiquitous WiFi technologies. However, recent studies have revealed significant privacy risks associated with WiFi signals, raising concerns about bio-information leakage. To address these concerns, the decentralized paradigm, particularly federated learning (FL), has emerged as a promising approach for training HAR models while preserving data privacy. Nevertheless, FL models may struggle in end-user environments due to substantial domain discrepancies between the source training data and the target end-user environment. This discrepancy arises from the sensitivity of WiFi signals to environmental changes, resulting in notable domain shifts. As a consequence, FL-based HAR approaches often face challenges when deployed in real-world WiFi environments. Albeit there are pioneer attempts on federated domain adaptation, they typically require non-trivial communication and computation cost, which is prohibitively expensive especially considering edge-based hardware equipment of end-user environment. In this paper, we propose a model to democratize the WiFi-based HAR system by enhancing recognition accuracy in unannotated end-user environments while prioritizing data privacy. Our model leverages the hypothesis transfer and a lightweight hypothesis ensemble to mitigate negative transfer. We prove a tighter theoretical upper bound compared to existing multi-source federated domain adaptation models. Extensive experiments shows our model improves the average accuracy by approximately 10 absolute percentage points in both cross-person and cross-environment settings comparing several state-of-the-art baselines. Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Min Wu 0008, Joey Tianyi Zhou |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | AirFi: Empowering WiFi-Based Passive Human Gesture Recognition to Unseen Environment via Domain GeneralizationabstractWiFi-based smart human sensing technology enabled by Channel State Information (CSI) has received great attention in recent years. However, CSI-based sensing systems suffer from performance degradation when deployed in different environments. Existing works solve this problem by domain adaptation using massive unlabeled high-quality data from the new environment, which is usually unavailable in practice. In this paper, we propose a novel augmented environment-invariant robust WiFi gesture recognition system named AirFi that deals with the issue of environment dependency from a new perspective. The AirFi is a novel domain generalization framework that learns the critical part of CSI regardless of different environments and generalizes the model to unseen scenarios, which does not require collecting any data for adaptation to the new environment. AirFi extracts the common features from several training environment settings and minimizes the distribution differences among them. The feature is further augmented to be more robust to environments. Moreover, the system can be further improved by few-shot learning techniques. Compared to state-of-the-art methods, AirFi is able to work in different environment settings without acquiring any CSI data from the new environment. The experimental results demonstrate that our system remains robust in the new environment and outperforms the compared systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | GrapHAR: A Lightweight Human Activity Recognition Model by Exploring the Sub-Carrier CorrelationsabstractHuman activity recognition (HAR) is an important task due to its far-reaching applications, such as surveillance, healthcare systems, and human-computer interaction. Recently, Channel State Information (CSI)-based HAR has attracted increasing attention in the research community due to its ubiquitous availability, good user privacy, and fewer constraints on working conditions. Most of the existing methods for CSI-based HAR use various deep learning models, such as Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM), and Transformers, to distinguish activities based on their temporal patterns. Despite their remarkable effectiveness, these methods solely focus on temporal patterns while ignoring the correlations among sub-carriers. This limitation prevents them from achieving further performance improvement. Moreover, recent works often involve advanced yet massive and inefficient neural architectures, like Transformers, to obtain satisfactory recognition accuracy. The performance gain is traded off with a steep increase in model complexity, which leads to low efficacy and high training/inference costs outsides the small time window. To address these issues, we propose a lightweight CSI-based HAR model. Our model makes the first effort to explore the graphical correlations of CSI sub-carriers, working in conjunction with a temporal causal convolution module. The high efficacy design enables our model to be highly effective without requiring excessive model complexity. Extensive experiments conducted on four real-world datasets demonstrate that our model outperforms state-of-the-art methods, including a strong Transformer-based baseline. It achieves an average improvement of 8 percentage points in recognition accuracy, with only 10% of the parameters compared to the Transformer-based method (4.95M vs. 49.24M). Additionally, our model is significantly faster, with empirical training and execution times at least 2.07 times faster than the baseline. Wei Meng 0002, Zhicong Liu, Bing Li 0002, Wei Cui 0002, Joey Tianyi Zhou, Le Zhang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Pseudo-Label Clustering-Driven Dual-Level Contrast Learning Based Source-Free Domain Adaptation for Fundus Image Segmentation
Wei Zhou 0003, Jianhang Ji, Wei Cui 0002, Yugen Yi |
PRCV (5) | 3 |
| 2023 | CeHAR: CSI-Based Channel-Exchanging Human Activity RecognitionabstractDespite the intense effort from the research community, state-of-the-art WiFi-based human activity recognition (HAR) performance remains unsatisfactory. Current approaches usually use individual characteristics of CSI, i.e., amplitude or phase measurements, to model the relationship between the changes of channel state information (CSI) and human activities, which lead to their failure to achieve satisfactory accuracy due to information loss. To deal with this issue, this article proposes CeHAR, a CSI-based HAR using a channel-exchanging fusion network to deep fuse the CSI amplitude and phase features to obtain the informative features for HAR. The proposed CeHAR is a parameter-free dual-characteristic fusion framework that dynamically exchanges channels between subnetworks of two kinds of characteristics to comprehensively learn informative features from both. Specifically, the proposed approach employs two subnetworks using convolutional neural networks to learn features from each characteristic of CSI. The magnitude of the batch-normalization (BN) scaling factor is used to determine the channel importance of each characteristic, and then guides the exchange process. The proposed CeHAR also shares convolutional filters, but keeps private BNs layers in different characteristics, which, as an added benefit, allows our characteristic fusion network to be nearly as compact as a single-characteristic network. Extensive real-world experiments have been conducted to evaluate the performance of our proposed CeHAR, and the experimental results illustrate that our proposed approach outperforms baselines. Xiao Lu 0003, Yuli Li, Wei Cui 0002, Haixia Wang 0003 |
IEEE Internet Things J. | 3 |
| 2023 | DifFormer: Multi-Resolutional Differencing Transformer With Dynamic Ranging for Time Series AnalysisabstractTime series analysis is essential to many far-reaching applications of data science and statistics including economic and financial forecasting, surveillance, and automated business processing. Though being greatly successful of Transformer in computer vision and natural language processing, the potential of employing it as the general backbone in analyzing the ubiquitous times series data has not been fully released yet. Prior Transformer variants on time series highly rely on task-dependent designs and pre-assumed "pattern biases", revealing its insufficiency in representing nuanced seasonal, cyclic, and outlier patterns which are highly prevalent in time series. As a consequence, they can not generalize well to different time series analysis tasks. To tackle the challenges, we propose DifFormer, an effective and efficient Transformer architecture that can serve as a workhorse for a variety of time-series analysis tasks. DifFormer incorporates a novel multi-resolutional differencing mechanism, which is able to progressively and adaptively make nuanced yet meaningful changes prominent, meanwhile, the periodic or cyclic patterns can be dynamically captured with flexible lagging and dynamic ranging operations. Extensive experiments demonstrate DifFormer significantly outperforms state-of-the-art models on three essential time-series analysis tasks, including classification, regression, and forecasting. In addition to its superior performances, DifFormer also excels in efficiency - a linear time/memory complexity with empirically lower time consumption. Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Ce Zhu, Wei Wang 0011, Ivor W. Tsang, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Reinforced Adaptation Network for Partial Domain AdaptationabstractDomain adaptation enables generalized learning in new environments by transferring knowledge from label-rich source domains to label-scarce target domains. As a more realistic extension, partial domain adaptation (PDA) relaxes the assumption of fully shared label space, and instead deals with the scenario where the target label space is a subset of the source label space. In this paper, we propose a Reinforced Adaptation Network (RAN) to address the challenging PDA problem. Specifically, a deep reinforcement learning model is proposed to learn source data selection policies. Meanwhile, a domain adaptation model is presented to simultaneously determine rewards and learn domain-invariant feature representations. By combining reinforcement learning and domain adaptation techniques, the proposed network alleviates negative transfer by automatically filtering out less relevant source data and promotes positive transfer by minimizing the distribution discrepancy across domains. Experiments on three benchmark datasets demonstrate that RAN consistently outperforms seventeen existing state-of-the-art methods by a large margin. Keyu Wu 0002, Min Wu 0008, Zhenghua Chen, Ruibing Jin, Wei Cui 0002, Zhiguang Cao, Xiaoli Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Privacy-Preserving Cross-Environment Human Activity RecognitionabstractRecent studies have demonstrated the success of using the channel state information (CSI) from the WiFi signal to analyze human activities in a fixed and well-controlled environment. Those systems usually degrade when being deployed in new environments. A straightforward solution to solve this limitation is to collect and annotate data samples from different environments with advanced learning strategies. Although workable as reported, those methods are often privacy sensitive because the training algorithms need to access the data from different environments, which may be owned by different organizations. We present a practical method for the WiFi-based privacy-preserving cross-environment human activity recognition (HAR). It collects and shares information from different environments, while maintaining the privacy of individual person being involved. At the core of our approach is the utilization of the Johnson-Lindenstrauss transform, which is theoretically shown to be differentially private. Based on that, we further design an adversarial learning strategy to generate environment-invariant representations for HAR. We demonstrate the effectiveness of the proposed method with different data modalities from two real-life environments. More specifically, on the raw CSI dataset, it shows 2.18% and 1.24% improvements over challenging baselines for two environments, respectively. Moreover, with the discrete wavelet transform features, it further yields 5.71% and 1.55% improvements, respectively. Le Zhang 0001, Wei Cui 0002, Bing Li 0002, Zhenghua Chen, Min Wu 0008, Sin G. Teo |
IEEE Trans. Cybern. | 2 |
| 2022 | WiHGR: A Robust WiFi-Based Human Gesture Recognition System via Sparse Recovery and Modified Attention-Based BGRUabstractGesture recognition is an essential part in the field of human–computer interaction (HCI) and Internet of Things system. Compared with the existing technologies based on wearable sensors and dedicated devices, approaches using WiFi channel state information (CSI) signals are more desirable for passive and fine-grained gesture recognition. However, the existing CSI-based gesture recognition systems usually suffer from high model complexity and low accuracy caused by environmental dynamics. To address these issues, we propose a robust gesture recognition system (WiHGR) in this article. The WiHGR starts with a sparse recovery method to find the dominant paths from the multipath effect introduced by the orthogonal frequency division multiplexing (OFDM) technology, i.e., the main propagation paths disturbed by a human gesture. Then, the phase difference matrix is constructed according to the phase differences between two adjacent receiving antennas from the dominant paths. We propose a modified attention-based bi-directional gate recurrent unit (ABGRU) network to learn and extract discriminative features automatically from the phase difference matrix. The proposed attention mechanism assigns higher weights to the more important features, thus achieving a better recognition performance. The experimental results show that the WiHGR not only has a high accuracy for gesture recognition in the training environment, but also has a remarkable performance in new environment settings without retraining. Wei Meng 0002, Xingcan Chen, Wei Cui 0002, Jing Guo 0007 |
IEEE Internet Things J. | 3 |
| 2022 | CAUTION: A Robust WiFi-Based Human Authentication System via Few-Shot Open-Set RecognitionabstractExisting channel-state information (CSI)-based human authentication systems in the literature require a large amount of CSI data to train deep neural network (DNN) models and are ineffective for unknown intruder detection. To address this issue, we propose a CSI-based human authentication system (CAUTION) which is able to learn distinctive gait features of different users through CSI data to perform human authentication in this article. By taking advantage of few-shot learning, CAUTION is able to construct an accurate user identification model with a very limited number of CSI training data. By converting the CSI samples into low-dimensional representations on the feature plane, it computes central points for different users as their CSI profiles and introduces an intruder threshold to measure whether the CSI data matches one of the user classes by a margin. The intruder threshold is able to be optimized without any intruders’ data. CAUTION does not require a large number of training data and provides an effective way to train the system for unknown intruder detection. We have tested CAUTION at different places and compared it with state-of-the-art CSI-based authentication systems. The experimental results demonstrate that CAUTION is able to perform accurate human authentication with a limited amount of CSI training data (one-fifth of data needed by compared systems) and outperforms the compared human authentication systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Internet Things J. | 3 |
| 2021 | Two-Stream Convolution Augmented Transformer for Human Activity RecognitionabstractRecognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e.g., RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the time-over-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency. Bing Li 0002, Wei Cui 0002, Wei Wang 0011, Le Zhang 0001, Zhenghua Chen, Min Wu 0008 |
AAAI | 2 |
| 2021 | Multimodal CSI-Based Human Activity Recognition Using GANsabstractChannel state information (CSI)-based human activity recognition (HAR) has received great attention in recent years due to its advantages in privacy protection, insensitivity to illumination, and no requirement for wearable devices. In this article, we propose a multimodal channel state information-based activity recognition (MCBAR) system that leverages existing WiFi infrastructures and monitors human activities from CSI measurements. MCBAR aims to address the performances degradation of WiFi-based human recognition systems due to environmental dynamics. Specifically, we address the issue of nonuniformly distributed unlabeled data with rarely performed activities by taking advantages of the generative adversarial network (GAN) and semisupervised learning. We apply a multimodal generator to approximate the CSI data distribution in different environment settings with limited measured CSI data. The generated CSI data using the multimodal generator can provide better diversity for knowledge transfer. This multimodal generator improves the ability of MCBAR to recognize specific activities with various CSI patterns caused by environmental dynamics. Compared to state-of-the-art CSI-based recognition systems, MCBAR is more robust as it is able to handle the nonuniformly distributed CSI data collected from a new environment setting. In addition, diverse generated data from the multimodal generator improves the stability of the system. We have tested MCBAR under multiple experimental settings at different places. The experimental results demonstrate that our algorithm overcomes environmental dynamics and outperforms existing HAR systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Internet Things J. | 3 |
| 2021 | An Attention Based CNN-LSTM Approach for Sleep-Wake Detection With Heterogeneous SensorsabstractIn this article, we propose an attention based convolutional neural network long short-term memory (CNN-LSTM) approach for sleep-wake detection with heterogeneous sensor data, i.e., acceleration and heart rate variability (HRV). Since the three-dimensional acceleration data was sampled with a high frequency, we firstly design a CNN-LSTM structure to effectively learn latent features from the acceleration. Meanwhile, considering the unique format of the HRV data, some effective features are extracted based on domain knowledge. Next, we design a unified architecture to efficiently merge the features learned by CNN-LSTM approach from the acceleration and the extracted features from the HRV, which enables us to make full use of all the available information from these two heterogeneous sources. Taking into consideration that these two heterogeneous sources may have distinct contributions for the sleep and wake states, we propose an attention network to dynamically adjust the importance of features from the two sources. Real-world experiments have been conducted to verify the effectiveness of the proposed approach for sleep-wake detection. The results demonstrate that the proposed method outperforms all existing approaches for sleep-wake classification. In the evaluation of leave-one-subject-out (LOSO) cross-validation which is more challenging and practical, the proposed method achieves remarkable improvements ranging from 5% to 46% over the benchmark approaches. Zhenghua Chen, Min Wu 0008, Wei Cui 0002, Chengyu Liu 0001, Xiaoli Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Robust CSI-based Human Activity Recognition using Roaming GeneratorabstractChannel State Information (CSI) based human activity recognition has received great attention in recent years due to its advantages in privacy protection, insensitive to illumination and no requirement for wearable devices. However, for practical deployment, it needs to greatly enhance the performance robustness against dynamic changes of the surrounding environment. To address this problem, we propose a novel CSI based activity recognition using Roaming Generator (CSIRoG) system for human activity detection. CSIRoG leverages existing WiFi infrastructures and monitors human behaviours from CSI measurements. It utilizes the generative adversarial network (GAN) to transfer the CSI information from one environment to another with dynamic changes such as people passing by, furniture layout changes, etc. The proposed method aims to approximate the CSI distribution in the new environment setting which has very limited CSI data. Therefore, the system can learn to handle multiple environment dynamics. Compared to the existing works, CSIRoG leverages a multimodal system model for better diversity of the generated CSI data for knowledge transfer. This improves the ability of CSIRoG to recognize various kinds of CSI information for one specific user activity caused by various dynamic conditions, thus enhancing system robustness. We have tested CSIRoG under multiple environment settings at different places. The experimental results demonstrate that our algorithm overcomes environmental dynamics and outperforms existing human activity recognition systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
ICARCV | 3 |
| 2020 | WiFi-Based Indoor Robot Positioning Using Deep Fuzzy ForestsabstractAddressing the positioning problem of a mobile robot remains challenging to date despite many years of research. Indoor robot positioning strategies developed in the literature either rely on sophisticated computer vision techniques to handle visual inputs or require strong domain knowledge for nonvisual sensors. Although some systems have been deployed, the former may be lacking due to the intrinsic limitation of cameras (such as calibration, data association, system initialization, etc.) and the latter usually only works under certain environment layouts and additional equipment. To cope with those issues, we design a lightweight indoor robot positioning system which operates on cost-effective WiFi-based received signal strength (RSS) and could be readily pluggable into any existing WiFi network infrastructures. Moreover, a novel deep fuzzy forest is proposed to inherit the merits of decision trees and deep neural networks within an end-to-end trainable architecture. Real-world indoor localization experiments are conducted and results demonstrate the superiority of the proposed method over the existing approaches. Le Zhang 0001, Zhenghua Chen, Wei Cui 0002, Bing Li 0002, Cen Chen 0002, Zhiguang Cao, Kai-Zhou Gao |
IEEE Internet Things J. | 3 |
| 2019 | WiFi CSI Based Passive Human Activity Recognition Using Attention Based BLSTMabstractHuman activity recognition can benefit various applications including healthcare services and context awareness. Since human actions will influence WiFi signals, which can be captured by the channel state information (CSI) of WiFi, WiFi CSI based human activity recognition has gained more and more attention. Due to the complex relationship between human activities and WiFi CSI measurements, the accuracies of current recognition systems are far from satisfactory. In this paper, we propose a new deep learning based approach, i.e., attention based bi-directional long short-term memory (ABLSTM), for passive human activity recognition using WiFi CSI signals. The BLSTM is employed to learn representative features in two directions from raw sequential CSI measurements. Since the learned features may have different contributions for final activity recognition, we leverage on an attention mechanism to assign different weights for all the learned features. Real experiments have been carried out to evaluate the performance of the proposed ABLSTM for human activity recognition. The experimental results show that our proposed ABLSTM is able to achieve the best recognition performance for all activities when compared with some benchmark approaches. Zhenghua Chen, Le Zhang 0001, Chaoyang Jiang, Zhiguang Cao, Wei Cui 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2018 | An Adaptive Hierarchical Compositional Model for Phrase EmbeddingabstractPhrase embedding aims at representing phrases in a vector space and it is important for the performance of many NLP tasks. Existing models only regard a phrase as either full-compositional or non-compositional, while ignoring the hybrid-compositionality that widely exists, especially in long phrases. This drawback prevents them from having a deeper insight into the semantic structure for long phrases and as a consequence, weakens the accuracy of the embeddings. In this paper, we present a novel method for jointly learning compositionality and phrase embedding by adaptively weighting different compositions using an implicit hierarchical structure. Our model has the ability of adaptively adjusting among different compositions without entailing too much model complexity and time cost. To the best of our knowledge, our work is the first effort that considers hybrid-compositionality in phrase embedding. The experimental evaluation demonstrates that our model outperforms state-of-the-art methods in both similarity tasks and analogy tasks. Bing Li 0002, Xiaochun Yang 0001, Bin Wang 0015, Wei Wang 0011, Wei Cui 0002, Xianchao Zhang 0001 |
IJCAI | 5 |
| 2018 | Received Signal Strength Based Indoor Positioning Using a Random Vector Functional Link NetworkabstractFingerprinting based indoor positioning system is gaining more research interest under the umbrella of location-based services. However, existing works have certain limitations in addressing issues such as noisy measurements, high computational complexity, and poor generalization ability. In this work, a random vector functional link network based approach is introduced to address these issues. In the proposed system, a subset of informative features from many randomized noisy features is selected to both reduce the computational complexity and boost the generalization ability. Moreover, the feature selector and predictor are jointly learned iteratively in a single framework based on an augmented Lagrangian method. The proposed system is appealing as it can be naturally fit into parallel or distributed computing environment. Extensive real-world indoor localization experiments are conducted on users with smartphone devices and results demonstrate the superiority of the proposed method over the existing approaches. Wei Cui 0002, Le Zhang 0001, Bing Li 0002, Jing Guo 0007, Wei Meng 0002, Haixia Wang 0003, Lihua Xie 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Efficiently Mining High Quality Phrases from TextsabstractPhrase mining is a key research problem for semantic analysis and text-based information retrieval. The existing approaches based on NLP, frequency, and statistics cannot extract high quality phrases and the processing is also time consuming, which are not suitable for dynamic on-line applications. In this paper, we propose an efficient high-quality phrase mining approach (EQPM). To the best of our knowledge, our work is the first effort that considers both intra-cohesion and inter-isolation in mining phrases, which is able to guarantee appropriateness. We also propose a strategy to eliminate order sensitiveness, and ensure the completeness of phrases. We further design efficient algorithms to make the proposed model and strategy feasible. The empirical evaluations on four real data sets demonstrate that our approach achieved a considerable quality improvement and the processing time was 2.3X - 29X faster than the state-of-the-art works. Bing Li 0002, Xiaochun Yang 0001, Bin Wang 0015, Wei Cui 0002 |
AAAI | 4 |