VLDB 2026 Research / reviewers in the wild / expert
Jianfei Yang 0001
dblp:06/5852-1
· DBLP profile ↗
88ranked-venue papers
16as first author
61since 2021 · last 2026
0000-0002-8075-0439ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 7 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-author · 21 since 2021Computer networks · 25 · 7 first-author · 15 since 2021Systems, architecture and hardware · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mmPred: Radar-based Human Motion Prediction in the DarkabstractExisting Human Motion Prediction (HMP) methods based on RGB(D) cameras are sensitive to lighting conditions and raise privacy concerns, limiting their real-world applications such as firefighting and elderly care. Motivated by the robustness and privacy-preserving nature of millimeter-wave (mmWave) radar, this work introduces radar as a novel sensing modality for HMP for the first time. Nevertheless, radar signals often suffer from specular reflections and multipath effects, resulting in noisy and temporally inconsistent measurements, such as body-part miss-detection. To address these radar-specific artifacts, we propose mmPred, the first diffusion-based framework tailored for radar-based HMP. mmPred introduces a dual-domain historical motion representation to guide the generation process, combining a Time-domain Pose Refinement (TPR) branch for fine-grained details and a Frequency-domain Dominant Motion (FDM) branch for capturing global motion trends and suppressing frame-level inconsistency. Furthermore, we design a Global Skeleton-relational Transformer (GST) as the diffusion backbone to model global inter-joint cooperation, enabling corrupted joints to dynamically aggregate information from others. Extensive experiments show that mmPred achieves state-of-the-art performance, outperforming existing methods by 8.6% on mmBody and 22% on mm-Fi. Junqiao Fan, Haocong Rao, Jianfei Yang 0001, Lihua Xie 0001 |
AAAI | 4 |
| 2026 | Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesabstractGenerating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often struggle to model such complex and dynamic interactions. While recent studies show that masks can serve as effective control signals and enhance generation quality, obtaining dense and precise mask annotations remains a major challenge for real-world use. To overcome this limitation, we introduce Mask2IV, a novel framework specifically designed for interaction-centric video generation. It adopts a decoupled two-stage pipeline that first predicts plausible motion trajectories for both actor and object, then generates a video conditioned on these trajectories. This design eliminates the need for dense mask inputs from users while preserving the flexibility to manipulate the interaction process. Furthermore, Mask2IV supports versatile and intuitive control, allowing users to specify the target object of interaction and guide the motion trajectory through action descriptions or spatial position cues. To support systematic training and evaluation, we curate two benchmarks covering diverse action and object categories across both human-object interaction and robotic manipulation scenarios. Extensive experiments demonstrate that our method achieves superior visual realism and controllability compared to existing baselines. Gen Li 0008, Jianfei Yang 0001, Laura Sevilla-Lara |
AAAI | 3 |
| 2026 | Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time TrainingabstractUnderstanding complex human activities demands the ability to decompose motion into fine-grained, semantic-aligned sub-actions. This motion grounding process is crucial for behavior analysis, embodied AI and virtual reality. Yet, most existing methods rely on dense supervision with predefined action classes, which are infeasible in open-vocabulary, real-world settings. In this paper, we propose ZOMG, a zero-shot, open-vocabulary framework that segments motion sequences into semantically meaningful sub-actions without requiring any annotations or fine-tuning. Technically, ZOMG integrates (1) language semantic partition, which leverages large language models to decompose instructions into ordered sub-action units, and (2) soft masking optimization, which learns instance-specific temporal masks to focus on frames critical to sub-actions, while maintaining intra-segment continuity and enforcing inter-segment separation, all without altering the pretrained encoder. Experiments on three motion-language datasets demonstrate state-of-the-art effectiveness and efficiency of motion grounding performance, outperforming prior methods by 8.7% mAP on HumanML3D benchmark. Meanwhile, significant improvements also exist in downstream retrieval, establishing a new paradigm for annotation-free motion understanding. Yunjiao Zhou, Xinyan Chen 0002, Junlang Qian, Lihua Xie 0001, Jianfei Yang 0001 |
AAAI | 5 |
| 2026 | Online predictive generation: Safety-critical trajectory planning for robotic tower crane systems
Beiyu You, Boyu Ma, Xiaokai Zhou, Xinyu Zhou 0006, Jianfei Yang 0001 |
Expert Syst. Appl. | 6 |
| 2026 | SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action RecognitionabstractSkeleton-based action recognition leverages human pose keypoints to categorize human actions, which shows superior generalization and interoperability compared to regular end-to-end action recognition. Existing solutions use RGB cameras to annotate skeletal keypoints, but their performance declines in dark environments and raises privacy concerns, limiting their use in smart homes and hospitals. This paper explores non-invasive wireless sensors, i.e., LiDAR and mmWave, to mitigate these challenges as a feasible alternative. Two problems are addressed: (1) insufficient data on wireless sensor modality to train an accurate skeleton estimation model, and (2) skeletal keypoints derived from wireless sensors are noisier than RGB, causing great difficulties for subsequent action recognition models. Our work, SkeFi, overcomes these gaps through a novel cross-modal knowledge transfer method acquired from the data-rich RGB modality. We propose the enhanced Temporal Correlation Adaptive Graph Convolution (TC-AGC) with frame interactive enhancement to overcome the noise from missing or inconsecutive frames. Additionally, our research underscores the effectiveness of enhancing multiscale temporal modeling through dual temporal convolution. By integrating TC-AGC with temporal modeling for cross-modal transfer, our framework can extract accurate poses and actions from noisy wireless sensors. Experiments demonstrate that SkeFi realizes state-of-the-art performances on mmWave and LiDAR. The code is available at https://github.com/Huang0035/Skefi. Shunyu Huang, Yunjiao Zhou, Jianfei Yang 0001 |
IEEE Internet Things J. | 3 |
| 2026 | CiUAV: Scalable Device-Free Indoor UAV Localization via Multiobjective Optimized Network Using Channel State InformationabstractAccurate and scalable indoor localization for unmanned aerial vehicles (UAVs) is essential for Internet of Things (IoT) applications such as autonomous logistics, infrastructure inspection, and emergency response in GPS-denied environments. However, traditional methods often struggle with cost, deployment complexity, and sensitivity to environmental dynamics, limiting their practicality for large-scale IoT scenarios. This paper presents a method in which Channel State Information (CSI) from low-cost IoT sensors enables robust, device-free 3D UAV localization while optimizing accuracy, sensor adaptability, and data efficiency. We propose CiUAV, leveraging CSI captured by ESP32-S3 sensors, with a Robust CSI Signal Enhancement (RCSE) framework integrating Dynamic AGC Compensation (DAC) and Adaptive Noise Suppression and Outlier Removal (ANSOR), alongside a Sensor-in-Sample (SiS) multi-objective optimization model for adaptive multi-sensor fusion. Experimental evaluations in realistic indoor settings achieve a 3D root mean squared error (RMSE) of 0.2659 meters, outperforming baselines by up to 35% in accuracy and 50% in data efficiency. CiUAV offers a lightweight, scalable, and infrastructure-compatible solution for future IoT-enabled UAV systems. Cunyi Yin, Zhaoke Huang, Hao Jiang 0008, Jing Chen 0022, Xiren Miao, Shaocong Zheng, Jianfei Yang 0001, Zhiwen Chen 0001, Zhenghua Chen, Hong Yan 0001 |
IEEE Internet Things J. | 8 |
| 2026 | TENT: Connect Language Models With IoT Sensors for Zero-Shot Activity RecognitionabstractThe rapid expansion of the Internet of Things (IoT) has introduced new challenges in Human Activity Recognition (HAR), particularly in dynamic environments where new and unforeseen activities emerge. Traditional HAR models, relying on predefined labels, struggle to adapt to these scenarios, highlighting the need for zero-shot learning (ZSL) approaches that can generalize beyond fixed training categories. Recent advances in large language models (LLMs) have demonstrated remarkable zero-shot capability in textual and visual domains. However, extending this ability to IoT sensors is substantially more challenging due to their heterogeneous modalities, diverse data structures, and limited semantic annotations. In this paper, we propose TENT (IoT-sEnsorslanguage alignmEnt pre-Training), a novel framework that constructs a unified sensor-language semantic space for zero-shot HAR. Instead of aligning each sensor individually to text, TENT jointly aligns multiple heterogeneous modalities with language, treating them as peers rather than anchors. This balanced multi-modal alignment allows sensors to mutually regularize one another while being grounded in linguistic semantics, transforming heterogeneity from a barrier into a strength. To further enrich the semantic space, TENT incorporates detailed activity descriptions and learnable prompts, enhancing adaptability to unseen activities. Extensive experiments across datasets and evaluation protocols demonstrate that TENT not only achieves robust recognition of both seen and unseen activities but also significantly outperforms existing vision-language and sensor-language baselines, surpassing them by over 20% on zero-shot HAR tasks. These results establish TENT as a new paradigm for generalizable IoT representation learning. Yunjiao Zhou, Jianfei Yang 0001, Han Zou, Lihua Xie 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-LabelingabstractAs small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes audio array for 3D UAV trajectory estimation. Our approach incorporates a self-supervised learning model, starting with the conversion of audio data into mel-spectrograms, which are analyzed through an encoder to extract crucial temporal and spectral information. Simultaneously, UAV trajectories are estimated using LiDAR point clouds via unsupervised methods. These LiDAR-based estimations act as pseudo labels, enabling the training of an Audio Perception Network without requiring labeled data. In this architecture, the LiDAR-based system operates as the Teacher Network, guiding the Audio Perception Network, which serves as the Student Network. Once trained, the model can independently predict 3D trajectories using only audio signals, with no need for LiDAR data or external ground truth during deployment. To further enhance precision, we apply Gaussian Process modeling for improved spatiotemporal tracking. Our method delivers toptier performance on the MMAUD dataset, establishing a new benchmark in trajectory estimation using self-supervised learning techniques without reliance on ground truth annotations. Allen Lei, Tianchen Deng, Han Wang 0001, Jianfei Yang 0001, Shenghai Yuan 0001 |
ICASSP | 4 |
| 2025 | Unsupervised UAV 3D Trajectories Estimation with Sparse Point CloudsabstractCompact UAV systems, while advancing delivery and surveillance, pose significant security challenges due to their small size, which hinders detection by traditional methods. This paper presents a cost-effective, unsupervised UAV detection method using spatial-temporal sequence processing to fuse multiple LiDAR scans for accurate UAV tracking in real-world scenarios. Our approach segments point clouds into foreground and background, analyzes spatial-temporal data, and employs a scoring mechanism to enhance detection accuracy. Tested on a public dataset, our solution placed 4th in the CVPR 2024 UG2+ Challenge, demonstrating its practical effectiveness. We plan to open-source all designs, code and sample data for the research community @ github.com/lianghanfang/UnLiDAR-UAV-Est. Hanfang Liang, Yizhuo Yang 0001, Jinming Hu, Jianfei Yang 0001, Shenghai Yuan 0001 |
ICASSP | 4 |
| 2025 | GERA: Geometric Embedding for Efficient Point Registration AnalysisabstractPoint cloud registration aims to provide estimated transformations to align point clouds, which plays a crucial role in pose estimation of various navigation systems, such as surgical guidance systems and autonomous vehicles. Despite the impressive performance of recent models on benchmark datasets, many rely on complex modules like KPConv and Transformers, which impose significant computational and memory demands. These requirements hinder their practical application, particularly in resource-constrained environments such as mobile robotics. In this paper, we propose a novel point cloud registration network that leverages a pure MLP architecture, constructing geometric information offline. This approach eliminates the computational and memory burdens associated with traditional complex feature extractors and significantly reduces inference time and resource consumption. Our method is the first to replace 3D coordinate inputs with offline-constructed geometric encoding, improving generalization and stability, as demonstrated by Maximum Mean Discrepancy (MMD) comparisons. This efficient and accurate geometric representation marks a significant advancement in point cloud analysis, particularly for applications requiring fast and reliability. Haozhi Cao, Shenghai Yuan 0001, Jianfei Yang 0001 |
ICRA | 5 |
| 2025 | CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAMabstractRecent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes. However, these approaches are built on a tremendous number of redundant 3D Gaussian ellipsoids, leading to high memory and storage costs and slow training speed. To address this limitation, we propose a compact 3D Gaussian Splatting SLAM system that reduces the number and the parameter size of Gaussian ellipsoids. A sliding window-based masking strategy is first proposed to reduce the redundant ellipsoids. Then, a novel geometry codebook-based quantization method is proposed to further compress 3D Gaussian geometric attributes. Robust and accurate pose estimation is achieved by a local-to-global bundle adjustment method with reprojection loss. Extensive experiments demonstrate that our method achieves faster training, rendering speed, and low memory usage while maintaining the state-of-the-art (SOTA) quality of the scene representation. Tianchen Deng, Yaohui Chen 0003, Jianfei Yang 0001, Shenghai Yuan 0001, Jiuming Liu, Danwei Wang, Weidong Chen 0001 |
IROS | 3 |
| 2025 | QLIO: Quantized LiDAR-Inertial OdometryabstractLiDAR-Inertial Odometry (LIO) is widely used for autonomous navigation, but its deployment on Size, Weight, and Power (SWaP)-constrained platforms remains challenging due to the computational cost of processing dense point clouds. Conventional LIO frameworks rely on a single onboard processor, leading to computational bottlenecks and high memory demands, making real-time execution difficult on embedded systems. To address this, we propose QLIO, a multi-processor distributed quantized LIO framework that reduces computational load and bandwidth consumption while maintaining localization accuracy. QLIO introduces a quantized state estimation pipeline, where a co-processor pre-processes LiDAR measurements, compressing point-to-plane residuals before transmitting only essential features to the host processor. Additionally, an rQ-vector-based adaptive resampling strategy intelligently selects and compresses key observations, further reducing computational redundancy. Real-World evaluations demonstrate that QLIO achieves a 14.1× reduction in perobservation residual data while preserving localization accuracy. Furthermore, we release an open-source implementation to facilitate further research and real-world deployment. These results establish QLIO as an efficient and scalable solution for real-time autonomous systems operating under computational and bandwidth constraints. Boyang Lou, Shenghai Yuan 0001, Jianfei Yang 0001, Wenju Su, Yingjian Zhang, Enwen Hu |
IROS | 3 |
| 2025 | Cross-center Model Adaptive Tooth segmentation
Ruizhe Chen, Jianfei Yang 0001, Huimin Xiong, Ruiling Xu, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
Medical Image Anal. | 2 |
| 2025 | T3DNet: Compressing Point Cloud Models for Lightweight 3-D RecognitionabstractThe 3-D point cloud has been widely used in many mobile application scenarios, including autonomous driving and 3-D sensing on mobile devices. However, existing 3-D point cloud models tend to be large and cumbersome, making them hard to deploy on edged devices due to their high memory requirements and nonreal-time latency. There has been a lack of research on how to compress 3-D point cloud models into lightweight models. In this article, we propose a method called T3DNet (tiny 3-D network with augmentation and distillation) to address this issue. We find that the tiny model after network augmentation is much easier for a teacher to distill. Instead of gradually reducing the parameters through techniques, such as pruning or quantization, we predefine a tiny model and improve its performance through auxiliary supervision from augmented networks and the original model. We evaluate our method on several public datasets, including ModelNet40, ShapeNet, and ScanObjectNN. Our method can achieve high compression rates without significant accuracy sacrifice, achieving state-of-the-art performances on three datasets against existing methods. Amazingly, our T3DNet is 58 smaller and 54 faster than the original model yet with only 1.4 accuracy descent on the ModelNet40 dataset. Our code is available at https://github.com/Zhiyuan002/T3DNet. Yunjiao Zhou, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | Fully-Connected Spatial-Temporal Graph for Multivariate Time-Series DataabstractMultivariate Time-Series (MTS) data is crucial in various application fields. With its sequential and multi-source (multiple sensors) properties, MTS data inherently exhibits Spatial-Temporal (ST) dependencies, involving temporal correlations between timestamps and spatial correlations between sensors in each timestamp. To effectively leverage this information, Graph Neural Network-based methods (GNNs) have been widely adopted. However, existing approaches separately capture spatial dependency and temporal dependency and fail to capture the correlations between Different sEnsors at Different Timestamps (DEDT). Overlooking such correlations hinders the comprehensive modelling of ST dependencies within MTS data, thus restricting existing GNNs from learning effective representations. To address this limitation, we propose a novel method called Fully-Connected Spatial-Temporal Graph Neural Network (FC-STGNN), including two key components namely FC graph construction and FC graph convolution. For graph construction, we design a decay graph to connect sensors across all timestamps based on their temporal distances, enabling us to fully model the ST dependencies by considering the correlations between DEDT. Further, we devise FC graph convolution with a moving-pooling GNN layer to effectively capture the ST dependencies for learning effective representations. Extensive experiments show the effectiveness of FC-STGNN on multiple MTS datasets compared to SOTA methods. The code is available at https://github.com/Frank-Wang-oss/FCSTGNN. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
AAAI | 3 |
| 2024 | Graph-Aware Contrasting for Multivariate Time-Series ClassificationabstractContrastive learning, as a self-supervised learning paradigm, becomes popular for Multivariate Time-Series (MTS) classification. It ensures the consistency across different views of unlabeled samples and then learns effective representations for these samples. Existing contrastive learning methods mainly focus on achieving temporal consistency with temporal augmentation and contrasting techniques, aiming to preserve temporal patterns against perturbations for MTS data. However, they overlook spatial consistency that requires the stability of individual sensors and their correlations. As MTS data typically originate from multiple sensors, ensuring spatial consistency becomes essential for the overall performance of contrastive learning on MTS data. Thus, we propose Graph-Aware Contrasting for spatial consistency across MTS data. Specifically, we propose graph augmentations including node and edge augmentations to preserve the stability of sensors and their correlations, followed by graph contrasting with both node- and graph-level contrasting to extract robust sensor- and global-level features. We further introduce multi-window temporal contrasting to ensure temporal consistency in the data for each sensor. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on various MTS classification tasks. The code is available at https://github.com/Frank-Wang-oss/TS-GAC. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
AAAI | 3 |
| 2024 | Reliable Spatial-Temporal Voxels For Multi-modal Test-Time Adaptation
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Xingyu Ji, Shenghai Yuan 0001, Lihua Xie 0001 |
ECCV (28) | 3 |
| 2024 | Diffusion Model Is a Good Pose Estimator from 3D RF-Vision
Junqiao Fan, Jianfei Yang 0001, Yuecong Xu, Lihua Xie 0001 |
ECCV (16) | 2 |
| 2024 | Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?abstractUnsupervised domain adaptation (UDA) involves adapting a model trained on a label-rich source domain to an unlabeled target domain. However, in real-world scenarios, the absence of target-domain labels makes it challenging to evaluate the performance of UDA models. Furthermore, prevailing UDA methods relying on adversarial training and self-training could lead to model degeneration and negative transfer, further exacerbating the evaluation problem. In this paper, we propose a novel metric called the Transfer Score to address these issues. The proposed metric enables the unsupervised evaluation of UDA models by assessing the spatial uniformity of the classifier via model parameters, as well as the transferability and discriminability of deep representations. Based on the metric, we achieve three novel objectives without target-domain labels: (1) selecting the best UDA method from a range of available options, (2) optimizing hyperparameters of UDA models to prevent model degeneration, and (3) identifying which checkpoint of UDA model performs optimally. Our work bridges the gap between data-level UDA research and practical UDA scenarios, enabling a realistic assessment of UDA model performance. We validate the effectiveness of our metric through extensive empirical studies on UDA datasets of different scales and imbalanced distributions. The results demonstrate that our metric robustly achieves the aforementioned goals. Jianfei Yang 0001, Hanjie Qian, Yuecong Xu, Kai Wang 0036, Lihua Xie 0001 |
ICLR | 1 |
| 2024 | MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic SegmentationabstractMulti-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-imbalanced performance, restricting their adoption in real applications. This imbalanced performance is mainly caused by: 1) self-training with imbalanced data and 2) the lack of pixel-wise 2D supervision signals. In this work, we propose Multi-modal Prior Aided (MoPA) domain adaptation to improve the performance of rare objects. Specifically, we develop Valid Ground-based Insertion (VGI) to rectify the imbalance supervision signals by inserting prior rare objects collected from the wild while avoiding introducing artificial artifacts that lead to trivial solutions. Meanwhile, our SAM consistency loss leverages the 2D prior semantic masks from SAM as pixel-wise supervision signals to encourage consistent predictions for each object in the semantic mask. The knowledge learned from modal-specific prior is then shared across modalities to achieve better rare object segmentation. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MM-UDA benchmark. Code will be available at https://github.com/AronCao49/MoPA. Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001 |
ICRA | 3 |
| 2024 | MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone ThreatsabstractIn response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat detection methodologies by focusing on drone detection, UAV-type classification, and trajectory estimation. MMAUD stands out by combining diverse sensory inputs, including stereo vision, various Lidars, Radars, and audio arrays. It offers a unique overhead aerial detection vital for addressing real-world scenarios with higher fidelity than datasets captured on specific vantage points using thermal and RGB. Additionally, MMAUD provides accurate Leica-generated ground truth data, enhancing credibility and enabling confident refinement of algorithms and models, which has never been seen in other datasets. Most existing works do not disclose their datasets, making MMAUD an invaluable resource for developing accurate and efficient solutions. Our proposed modalities are cost-effective and highly adaptable, allowing users to experiment and implement new UAV threat detection tools. Our dataset closely simulates real-world scenarios by incorporating ambient heavy machinery sounds. This approach enhances the dataset’s applicability, capturing the exact challenges faced during proximate vehicular operations. It is expected that MMAUD can play a pivotal role in advancing UAV threat detection, classification, trajectory estimation capabilities, and beyond. Our dataset, codes, and designs will be available in https://ntu-aris.github.io/MMAUD. Shenghai Yuan 0001, Yizhuo Yang 0001, Thien Hoang Nguyen, Thien-Minh Nguyen, Jianfei Yang 0001, Jianping Li 0004, Han Wang 0001, Lihua Xie 0001 |
ICRA | 5 |
| 2024 | scPanel: a tool for automatic identification of sparse gene panels for generalizable patient classification using scRNA-seq datasetsabstractSingle-cell RNA sequencing (scRNA-seq) technologies can generate transcriptomic profiles at a single-cell resolution in large patient cohorts, facilitating discovery of gene and cellular biomarkers for disease. Yet, when the number of biomarker genes is large, the translation to clinical applications is challenging due to prohibitive sequencing costs. Here, we introduce scPanel, a computational framework designed to bridge the gap between biomarker discovery and clinical application by identifying a sparse gene panel for patient classification from the cell population(s) most responsive to perturbations (e.g. diseases/drugs). scPanel incorporates a data-driven way to automatically determine a minimal number of informative biomarker genes. Patient-level classification is achieved by aggregating the prediction probabilities of cells associated with a patient using the area under the curve score. Application of scPanel to scleroderma, colorectal cancer, and COVID-19 datasets resulted in high patient classification accuracy using only a small number of genes (<20), automatically selected from the entire transcriptome. In the COVID-19 case study, we demonstrated cross-dataset generalizability in predicting disease state in an external patient cohort. scPanel outperforms other state-of-the-art gene selection methods for patient classification and can be used to identify parsimonious sets of reliable biomarker candidates for clinical translation. Jianfei Yang 0001, John F. Ouyang, Enrico Petretto |
Briefings Bioinform. | 2 |
| 2024 | Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study
Yuecong Xu, Haozhi Cao, Jianxiong Yin, Zhenghua Chen, Xiaoli Li 0001, Zhengguo Li, Qianwen Xu 0001, Jianfei Yang 0001 |
Int. J. Comput. Vis. | 8 |
| 2024 | PowerSkel: A Device-Free Framework Using CSI Signal for Human Skeleton Estimation in Power StationabstractSafety monitoring of power operations in power stations is crucial for preventing accidents and ensuring stable power supply. However, conventional methods such as wearable devices and video surveillance have limitations such as high cost, dependence on light, and visual blind spots. WiFi-based human pose estimation is a suitable method for monitoring power operations due to its low cost, device-free, and robustness to various illumination conditions. In this paper, a novel Channel State Information (CSI)-based pose estimation framework, namely PowerSkel, is developed to address these challenges. PowerSkel utilizes self-developed CSI sensors to form a mutual sensing network and constructs a CSI acquisition scheme specialized for power scenarios. It significantly reduces the deployment cost and complexity compared to the existing solutions. To reduce interference with CSI in the electricity scenario, a sparse adaptive filtering algorithm is designed to preprocess the CSI. CKDformer, a knowledge distillation network based on collaborative learning and self-attention, is proposed to extract the features from CSI and establish the mapping relationship between CSI and keypoints. The experiments are conducted in a real-world power station, and the results show that the PowerSkel achieves high performance with a PCK@50 of 96.27%, and realizes a significant visualization on pose estimation, even in dark environments. Our work provides a novel low-cost and high-precision pose estimation solution for power operation. Cunyi Yin, Xiren Miao, Jing Chen 0022, Hao Jiang 0008, Jianfei Yang 0001, Yunjiao Zhou, Min Wu 0008, Zhenghua Chen |
IEEE Internet Things J. | 5 |
| 2024 | AdaPose: Toward Cross-Site Device-Free Human Pose Estimation With Commodity WiFiabstractWiFi-based pose estimation is a technology with great potential for the development of smart homes and metaverse avatar generation. However, current WiFi-based pose estimation methods are predominantly evaluated under controlled laboratory conditions with sophisticated vision models to acquire accurately labeled data. Furthermore, WiFi channel state information (CSI) is highly sensitive to environmental variables, and direct application of a pretrained model to a new environment may yield suboptimal results due to domain shift. In this article, we propose a domain adaptation algorithm, AdaPose, designed specifically for WiFi-based pose estimation. The proposed method aims to identify consistent human poses that are highly resistant to environmental dynamics and WiFi signal noises. To achieve this goal, we introduce instance-wise consistency alignment loss that aligns domain shifts considering instance-wise pose distribution variance, and cross-environment channel enhancement module that enhances WiFi CSI feature representation by emphasizing channel-wise similarity between source and target domains. We conduct extensive experiments on both our self-collected pose estimation data set and a large public MM-Fi data set. The results demonstrate the effectiveness and robustness of AdaPose in eliminating domain shift, thereby facilitating the widespread application of WiFi-based pose estimation in smart cities. Yunjiao Zhou, Jianfei Yang 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 2 |
| 2024 | SEA++: Multi-Graph-Based Higher-Order Sensor Alignment for Multivariate Time-Series Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) methods have been successful in reducing label dependency by minimizing the domain discrepancy between labeled source domains and unlabeled target domains. However, these methods face challenges when dealing with Multivariate Time-Series (MTS) data. MTS data typically originates from multiple sensors, each with its unique distribution. This property poses difficulties in adapting existing UDA techniques, which mainly focus on aligning global features while overlooking the distribution discrepancies at the sensor level, thus limiting their effectiveness for MTS data. To address this issue, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA, aiming to address domain discrepancy at both local and global sensor levels. At the local sensor level, we design endo-feature alignment, which aligns sensor features and their correlations across domains. To reduce domain discrepancy at the global sensor level, we design exo-feature alignment that enforces restrictions on global sensor features. We further extend SEA to SEA++ by enhancing the endo-feature alignment. Particularly, we incorporate multi-graph-based higher-order alignment for both sensor features and their correlations. Extensive empirical results have demonstrated the state-of-the-art performance of our SEA and SEA++ on six public MTS datasets for MTS-UDA. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Self-Supervised Video Representation Learning by Video Incoherence DetectionabstractThis article introduces a novel self-supervised method that leverages incoherence detection for video representation learning. It stems from the observation that the visual system of human beings can easily identify video incoherence based on their comprehensive understanding of videos. Specifically, we construct the incoherent clip by multiple subclips hierarchically sampled from the same raw video with various lengths of incoherence. The network is trained to learn the high-level representation by predicting the location and length of incoherence given the incoherent clip as input. Additionally, we introduce intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video. We evaluate our proposed method through extensive experiments on action recognition and video retrieval using various backbone networks. Experiments show that our proposed method achieves remarkable performance across different backbone networks and different datasets compared to previous coherence-based methods. Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie 0001, Jianxiong Yin, Simon See, Qianwen Xu 0001, Jianfei Yang 0001 |
IEEE Trans. Cybern. | 8 |
| 2024 | AirFi: Empowering WiFi-Based Passive Human Gesture Recognition to Unseen Environment via Domain GeneralizationabstractWiFi-based smart human sensing technology enabled by Channel State Information (CSI) has received great attention in recent years. However, CSI-based sensing systems suffer from performance degradation when deployed in different environments. Existing works solve this problem by domain adaptation using massive unlabeled high-quality data from the new environment, which is usually unavailable in practice. In this paper, we propose a novel augmented environment-invariant robust WiFi gesture recognition system named AirFi that deals with the issue of environment dependency from a new perspective. The AirFi is a novel domain generalization framework that learns the critical part of CSI regardless of different environments and generalizes the model to unseen scenarios, which does not require collecting any data for adaptation to the new environment. AirFi extracts the common features from several training environment settings and minimizes the distribution differences among them. The feature is further augmented to be more robust to environments. Moreover, the system can be further improved by few-shot learning techniques. Compared to state-of-the-art methods, AirFi is able to work in different environment settings without acquiring any CSI data from the new environment. The experimental results demonstrate that our system remains robust in the new environment and outperforms the compared systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | SecureSense: Defending Adversarial Attack for Secure Device-Free Human Activity RecognitionabstractDeep neural networks have empowered accurate device-free human activity recognition, which has wide applications. Deep models can extract robust features from various sensors and generalize well even in challenging situations such as data-insufficient cases. However, these systems could be vulnerable to input perturbations, i.e., adversarial attacks. We empirically demonstrate that both black-box Gaussian attacks and modern adversarial white-box attacks can render their accuracies to plummet. In this paper, we first point out that such phenomenon can bring severe safety hazards to device-free sensing systems, and then propose a novel learning framework, SecureSense, to defend common attacks. SecureSense aims to achieve consistent predictions regardless of whether there exists an attack on its input or not, alleviating the negative effect of distribution perturbation caused by adversarial attacks. Extensive experiments demonstrate that our proposed method can significantly enhance the model robustness of existing deep models, overcoming possible attacks. The results validate that our method works well on wireless human activity recognition and person identification systems. To the best of our knowledge, this is the first work to investigate adversarial attacks and further develop a novel defense framework for wireless human activity recognition in mobile computing research. Jianfei Yang 0001, Han Zou, Lihua Xie 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Aligning Correlation Information for Domain Adaptation in Action RecognitionabstractDomain adaptation (DA) approaches address domain shift and enable networks to be applied to different scenarios. Although various image DA approaches have been proposed in recent years, there is limited research toward video DA. This is partly due to the complexity in adapting the different modalities of features in videos, which includes the correlation features extracted as long-range dependencies of pixels across spatiotemporal dimensions. The correlation features are highly associated with action classes and proven their effectiveness in accurate video feature extraction through the supervised action recognition task. Yet correlation features of the same action would differ across domains due to domain shift. Therefore, we propose a novel adversarial correlation adaptation network (ACAN) to align action videos by aligning pixel correlations. ACAN aims to minimize the distribution of correlation information, termed as pixel correlation discrepancy (PCD). Additionally, video DA research is also limited by the lack of cross-domain video datasets with larger domain shifts. We, therefore, introduce a novel HMDB-ARID dataset with a larger domain shift caused by a larger statistical difference between domains. This dataset is built in an effort to leverage current datasets for dark video classification. Empirical results demonstrate the state-of-the-art performance of our proposed ACAN for both existing and the new video DA datasets. Yuecong Xu, Haozhi Cao, Kezhi Mao, Zhenghua Chen, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | SEnsor Alignment for Multivariate Time-Series Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) methods can reduce label dependency by mitigating the feature discrepancy between labeled samples in a source domain and unlabeled samples in a similar yet shifted target domain. Though achieving good performance, these methods are inapplicable for Multivariate Time-Series (MTS) data. MTS data are collected from multiple sensors, each of which follows various distributions. However, most UDA methods solely focus on aligning global features but cannot consider the distinct distributions of each sensor. To cope with such concerns, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA to reduce the domain discrepancy at both the local and global sensor levels. At the local sensor level, we design the endo-feature alignment to align sensor features and their correlations across domains, whose information represents the features of each sensor and the interactions between sensors. Further, to reduce domain discrepancy at the global sensor level, we design the exo-feature alignment to enforce restrictions on the global sensor features. Meanwhile, MTS also incorporates the essential spatial-temporal dependencies information between sensors, which cannot be transferred by existing UDA methods. Therefore, we model the spatial-temporal information of MTS with a multi-branch self-attention mechanism for simple and effective transfer across domains. Empirical results demonstrate the state-of-the-art performance of our proposed SEA on two public MTS datasets for MTS-UDA. The code is available at https://github.com/Frank-Wang-oss/SEA Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001 |
AAAI | 3 |
| 2023 | Multi-Modal Continual Test-Time Adaptation for 3D Semantic SegmentationabstractContinual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmentation. The key to MMCTTA is to adaptively attend to the reliable modality while avoiding catastrophic forgetting during continual domain shifts, which is out of the capability of previous TTA or CTTA methods. To fulfill this gap, we propose an MM-CTTA method called Continual Cross-Modal Adaptive Clustering (CoMAC) that addresses this task from two perspectives. On one hand, we propose an adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space. On the other hand, to perform test-time adaptation without catastrophic forgetting, we design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge. We further introduce two new benchmarks to facilitate the exploration of MM-CTTA in the future. Our experimental results show that our method achieves state-of-the-art performance on both benchmarks. Visit our project website at https://sites.google.com/view/mmcotta. Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001 |
ICCV | 3 |
| 2023 | Augmenting and Aligning Snippets for Few-Shot Video Domain AdaptationabstractFor video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video models. However, current VUDA methods rely on a vast amount of high-quality unlabeled target data, which may not be available in real-world cases. We thus consider a more realistic Few-Shot Video-based Domain Adaptation (FSVDA) scenario where we adapt video models with only a few target video samples. While a few methods have touched upon Few-Shot Domain Adaptation (FSDA) in images and in FSVDA, they rely primarily on spatial augmentation for target domain expansion with alignment performed statistically at the instance level. However, videos contain more knowledge in terms of rich temporal and semantic information, which should be fully considered while augmenting target domains and performing alignment in FSVDA. We propose a novel SSA2lign to address FSVDA at the snippet level, where the target domain is expanded through a simple snippet-level augmentation followed by the attentive alignment of snippets both semantically and statistically, where semantic alignment of snippets is conducted through multiple perspectives. Empirical results demonstrate state-of-the-art performance of SSA2lign across multiple cross-domain action recognition benchmarks. Code will be provided at: https://github.com/xuyu0010/SSA2lign. Yuecong Xu, Jianfei Yang 0001, Yunjiao Zhou, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001 |
ICCV | 2 |
| 2023 | Divide to Adapt: Mitigating Confirmation Bias for Domain Adaptation of Black-Box Predictors
Jianfei Yang 0001, Kai Wang 0036, Jiashi Feng, Lihua Xie 0001, Yang You 0001 |
ICLR | 1 |
| 2023 | AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian AwarenessabstractIn this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LIDARs to cover multiple views can be expensive and susceptible to issues such as changes in illumination, occlusion, and weather conditions. Our proposed solution replicates human perception for 3D pedestrian detection using low-cost audio and visual fusion. This study represents the first attempt to employ audio-visual fusion to monitor footstep sounds for the purpose of predicting the movements of pedestrians in the vicinity. The system is trained through self-supervised learning based on LIDAR-generated labels, making it a cost-effective alternative to LIDAR-based pedestrian awareness. AV-PedAware achieves comparable results to LIDAR-based systems at a fraction of the cost. By utilizing an attention mechanism, it can handle dynamic lighting and occlusions, overcoming the limitations of traditional LIDAR and camera-based systems. To evaluate our approach's effectiveness, we collected a new multimodal pedestrian detection dataset and conducted experiments that demonstrate the system's ability to provide reliable 3D detection results using only audio and visual data, even in extreme visual conditions. We will make our collected dataset and source code available online for the community to encourage further development in the field of robotics perception systems. Yizhuo Yang 0001, Shenghai Yuan 0001, Muqing Cao, Jianfei Yang 0001, Lihua Xie 0001 |
IROS | 4 |
| 2023 | MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensingabstract4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research. Jianfei Yang 0001, Yunjiao Zhou, Xinyan Chen 0002, Yuecong Xu, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001 |
NeurIPS | 1 |
| 2023 | GaitFi: Robust Device-Free Human Identification via WiFi and Vision Multimodal LearningabstractAs an important biomarker for human identification, human gait can be collected at a distance by passive sensors without subject cooperation, which plays an essential role in crime prevention, security detection, and other human identification applications. Presently, most research works are based on cameras and computer vision techniques to perform gait recognition. However, vision-based methods are not reliable when confronting poor illuminations, leading to degrading performances. In this article, we propose a novel multimodal gait recognition method, namely, GaitFi, which leverages WiFi signals and videos for human identification. In GaitFi, channel state information (CSI) that reflects the multipath propagation of WiFi is collected to capture human gaits, while videos are captured by cameras. To learn robust gait information, we propose a lightweight residual convolution network (LRCN) as the backbone network and further propose the two-stream GaitFi by integrating WiFi and vision features for the gait retrieval task. The GaitFi is trained by the triplet loss and classification loss on different levels of features. Extensive experiments are conducted in the real world, which demonstrates that the GaitFi outperforms state-of-the-art gait recognition methods based on single WiFi or camera, achieving 94.2% for human identification tasks of 12 subjects. Lang Deng, Jianfei Yang 0001, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 2 |
| 2023 | VariFi: Variational Inference for Indoor Pedestrian Localization and Tracking Using IMU and WiFi RSSabstractAccurate indoor pedestrian localization and tracking are crucial in many practical applications. One efficient yet low-cost sensing scheme is the integration of inertial measurement unit and WiFi received signal strength (RSS) due to the popularity of smart devices and WiFi networks. Many approaches have been proposed to enhance the localization performance. However, they heavily rely on prerequisites, including prior knowledge (e.g., map information) and beacon corrections, which degrades the generalization of the approaches and their accuracy in complex environments. To address this issue, in this article, we propose a novel localization approach named VariFi, which incorporates variational inference techniques to estimate the location of pedestrian. Variational inference is applied in this work, whose inference network can produce accurate estimates as its parameters are optimized in terms of the reconstruction loss and regularization loss in real time. A signal map is constructed to provide a conditional RSS distribution at any given location, which is further applied to generate the reconstruction loss based on the real measurements. Also, a filtering mechanism is designed to reduce local optimum cases in optimization by utilizing the prior estimate and RSS fingerprinting estimate. In addition, VariFi can be further applied to conduct online optimization following the existing localization approaches. We conduct experiments, including static localization and trajectory estimation scenarios to validate the performance of our approach. The trajectory estimation results show that our approach outperforms the mainstream approaches in terms of both localization accuracy and robustness, respectively. Furthermore, the combination of existing approaches and VariFi has also been validated effectively in the experiments of two environments, where VariFi has the ability to bring enhanced localization accuracy. Jianfei Yang 0001, Xu Fang 0001, Hao Jiang 0008, Lihua Xie 0001 |
IEEE Internet Things J. | 2 |
| 2023 | AutoFi: Toward Automatic Wi-Fi Human Sensing via Geometric Self-Supervised LearningabstractWi-Fi sensing technology has shown superiority in smart homes among various sensors for its cost-effective and privacy-preserving merits. It is empowered by channel state information (CSI) extracted from Wi-Fi signals and advanced machine learning models to analyze motion patterns in CSI. Many learning-based models have been proposed for kinds of applications, but they severely suffer from environmental dependency. Though domain adaptation methods have been proposed to tackle this issue, it is not practical to collect high-quality, well-segmented, and balanced CSI samples in a new environment for adaptation algorithms, but randomly captured CSI samples can be easily collected. In this article, we first explore how to learn a robust model from these low-quality CSI samples, and propose AutoFi, an annotation-efficient Wi-Fi sensing model based on a novel geometric self-supervised learning algorithm. The AutoFi fully utilizes unlabeled low-quality CSI samples that are captured randomly, and then transfers the knowledge to specific tasks defined by users, which is the first work to achieve cross-task transfer in Wi-Fi sensing. The AutoFi is implemented on a pair of Atheros Wi-Fi APs for evaluation. The AutoFi transfers knowledge from randomly collected CSI samples into human gait recognition and achieves state-of-the-art performance. Furthermore, we simulate cross-task transfer using public data sets to further demonstrate its capacity for cross-task learning. For the UT-HAR and Widar data sets, the AutoFi achieves satisfactory results on activity recognition and gesture recognition without any prior training. We believe that AutoFi takes a huge step toward automatic Wi-Fi sensing without any developer engagement. Our codes have been included inhttps://github.com/xyanchen/Wi-Fi-CSI-Sensing-Benchmark. Jianfei Yang 0001, Xinyan Chen 0002, Han Zou, Dazhuo Wang, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2023 | MetaFi++: WiFi-Enabled Transformer-Based Human Pose Estimation for Metaverse Avatar SimulationabstractIn the metaverse, digital avatar plays an important role in representing human beings for various interaction with virtual objects and environments, which puts a high demand on effective pose estimation. Though camera-based solutions yield remarkable performance, they encounter privacy issues and degraded performance caused by varying illumination, especially in the smart home. In this article, we propose a WiFi-based Internet of Things-enabled human pose estimation scheme for metaverse avatar simulation, namely, MetaFi++. Specifically, WPFormer is designed with a shared convolutional module and a Transformer block to map the channel state information of WiFi signals to human pose landmarks, effectively exploring spatial information of human pose through self-attention. It is enforced to learn the annotations from the accurate computer vision model, thus achieving cross-modal supervision. Due to the ubiquitous existence of WiFi and robustness to various illumination conditions, WiFi-based human poses are suitable to instruct the movement of digital avatars in the metaverse, promoting avatar applications in smart homes. The experiments are conducted in the real world, and the results show that the MetaFi++ achieves very high performance with a PCK@50 of 97.30%. Our codes are available inhttps://github.com/pridy999/metafi_pose_estimation. Yunjiao Zhou, Shenghai Yuan 0001, Han Zou, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Internet Things J. | 6 |
| 2023 | Cross-platform privacy-preserving CT image COVID-19 diagnosis based on source-free domain adaptation
Yuanyi Feng, Yuemei Luo, Jianfei Yang 0001 |
Knowl. Based Syst. | 3 |
| 2023 | Multi-Source Video Domain Adaptation With Temporal Attentive Moment Alignment NetworkabstractMulti-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios, which relaxes the assumption in conventional Unsupervised Domain Adaptation (UDA) that source data are sampled from a single domain and match a uniform data distribution. The MSDA is more challenging due to the existence of different domain shifts between distinct domain pairs. When considering videos, the negative transfer would be provoked by spatial-temporal features and can be formulated into a more challenging Multi-Source Video Domain Adaptation (MSVDA) problem. In this paper, we address the MSVDA problem by proposing a novel Temporal Attentive Moment Alignment Network (TAMAN) which aims for effective feature transfer by dynamically aligning both spatial and temporal feature moments. The TAMAN further constructs robust global temporal features by attending to dominant domain-invariant local temporal features with high local classification confidence and low disparity between global and local feature discrepancies. To facilitate future research on the MSVDA problem, we introduce comprehensive benchmarks, covering extensive MSVDA scenarios. Empirical results demonstrate a superior performance of the proposed TAMAN across multiple MSVDA benchmarks. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhengguo Li, Zhenghua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Advancing Imbalanced Domain Adaptation: Cluster-Level Discrepancy Minimization With a Comprehensive BenchmarkabstractUnsupervised domain adaptation methods have been proposed to tackle the problem of covariate shift by minimizing the distribution discrepancy between the feature embeddings of source domain and target domain. However, the standard evaluation protocols assume that the conditional label distributions of the two domains are invariant, which is usually not consistent with the real-world scenarios such as long-tailed distribution of visual categories. In this article, the imbalanced domain adaptation (IDA) is formulated for a more realistic scenario where both label shift and covariate shift occur between the two domains. Theoretically, when label shift exists, aligning the marginal distributions may result in negative transfer. Therefore, a novel cluster-level discrepancy minimization (CDM) is developed. CDM proposes cross-domain similarity learning to learn tight and discriminative clusters, which are utilized for both feature-level and distribution-level discrepancy minimization, palliating the negative effect of label shift during domain transfer. Theoretical justifications further demonstrate that CDM minimizes the target risk in a progressive manner. To corroborate the effectiveness of CDM, we propose two evaluation protocols according to the real-world situation and benchmark existing domain adaptation approaches. Extensive experiments demonstrate that negative transfer does occur due to label shift, while our approach achieves significant improvement on imbalanced datasets, including Office-31, Image-CLEF, and Office-Home. Jianfei Yang 0001, Jiangang Yang, Shizheng Wang, Shuxin Cao, Han Zou, Lihua Xie 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Global Context with Discrete Diffusion in Vector Quantised Modelling for Image GenerationabstractThe integration of Vector Quantised Variational AutoEncoder (VQ-VAE) with autoregressive models as generation part has yielded high-quality results on image generation. However, the autoregressive models will strictly follow the progressive scanning order during the sampling phase. This leads the existing VQ series models to hardly escape the trap of lacking global information. Denoising Diffusion Probabilistic Models (DDPM) in the continuous domain have shown a capability to capture the global context, while generating high-quality images. In the discrete state space, some works have demonstrated the potential to perform text generation and low resolution image generation. We show that with the help of a content-rich discrete visual codebook from VQ-VAE, the discrete diffusion model can also generate high fidelity images with global context, which compensates for the deficiency of the classical autoregressive model along pixel space. Meanwhile, the integration of the discrete VAE with the diffusion model resolves the drawback of conventional autoregressive models being oversized, and the diffusion model which demands excessive time in the sampling process when generating images. It is found that the quality of the generated images is heavily dependent on the discrete visual codebook. Extensive experiments demonstrate that the proposed Vector Quantised Discrete Diffusion Model (VQ-DDM) is able to achieve comparable performance to top-tier methods with low complexity. It also demonstrates outstanding advantages over other vectors quantised with autoregressive models in terms of image inpainting tasks without additional training. Minghui Hu 0001, Tat-Jen Cham, Jianfei Yang 0001, Ponnuthurai N. Suganthan |
CVPR | 4 |
| 2022 | Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition
Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhenghua Chen |
ECCV (34) | 2 |
| 2022 | Adversarial Cross-modal Domain Adaptation for Multi-modal Semantic Segmentation in Autonomous Drivingabstract3D semantic segmentation is a vital problem in autonomous driving. Vehicles rely on semantic segmentation to sense the surrounding environment and identify pedestrians, roads, and other vehicles. Though many datasets are publicly available, there exists a gap between public data and real-world scenarios due to the different weathers and environments, which is formulated as the domain shift. These days, the research for Unsupervised Domain Adaptation (UDA) rises for solving the problem of domain shift and the lack of annotated datasets. This paper aims to introduce adversarial learning and cross-modal networks (2D and 3D) to boost the performance of UDA for semantic segmentation across different datasets. With this goal, we design an adversarial training scheme with a domain discriminator and render the domain-invariant feature learning. Furthermore, we demonstrate that introducing 2D modalities can contribute to the improvement of 3D modalities by our method. Experimental results show that the proposed approach improves the mIoU by 7.53% compared to the baseline and has an improvement of 3.68% for the multi-modal performance. Mengqi Shi, Haozhi Cao, Lihua Xie 0001, Jianfei Yang 0001 |
ICARCV | 4 |
| 2022 | Improving Hazy Image Recognition by Unsupervised Domain AdaptationabstractDeep learning has achieved excellent performance in computer vision tasks, like image recognition, natural language processing, etc. However, in real-world applications, special circumstances brought about by the external world may create domain bias caused by distribution discrepancy between training and testing data, leading to degrading model performance. For example, when auto-driving meets hazy weather, the model performance will drop significantly. In this paper, we explore to solve this problem by utilizing modern Domain Adaptation (DA) methods, which generalizes from the source domain to the target domain by minimizing the distribution difference caused by dataset bias. We firstly propose the cross-domain haze image datasets and benchmark the five classic DA methods. The experiments show that DA methods can mitigate the negative effect of haze and significantly improves the model performance for visual recognition. Zhiyu Yuan, Jianfei Yang 0001 |
ICARCV | 3 |
| 2022 | Calibrating Class Weights with Multi-Modal Information for Partial Video Domain AdaptationabstractAssuming the source label space subsumes the target one, Partial Video Domain Adaptation (PVDA) is a more general and practical scenario for cross-domain video classification problems. The key challenge of PVDA is to mitigate the negative transfer caused by the source-only outlier classes. To tackle this challenge, a crucial step is to aggregate target predictions to assign class weights by up-weighing target classes and down-weighing outlier classes. However, the incorrect predictions of class weights can mislead the network and lead to negative transfer. Previous works improve the class weight accuracy by utilizing temporal features and attention mechanisms, but these methods may fall short when trying to generate accurate class weight when domain shifts are significant, as in most real-world scenarios. To deal with these challenges, we first propose the Multi-modality partial Adversarial Network (MAN), which utilizes multi-scale and multi-modal information to enhance PVDA performance. Based on MAN, we then propose Multi-modality Cluster-calibrated partial Adversarial Network (MCAN). It utilizes a novel class weight calibration method to alleviate the negative transfer caused by incorrect class weights. Specifically, the calibration method tries to identify and weigh correct and incorrect predictions using distributional information implied by unsupervised clustering. Extensive experiments are conducted on prevailing PVDA benchmarks, and the proposed MCAN achieves significant improvements when compared to state-of-the-art PVDA methods. Yuecong Xu, Jianfei Yang 0001, Kezhi Mao |
ACM Multimedia | 3 |
| 2022 | CAUTION: A Robust WiFi-Based Human Authentication System via Few-Shot Open-Set RecognitionabstractExisting channel-state information (CSI)-based human authentication systems in the literature require a large amount of CSI data to train deep neural network (DNN) models and are ineffective for unknown intruder detection. To address this issue, we propose a CSI-based human authentication system (CAUTION) which is able to learn distinctive gait features of different users through CSI data to perform human authentication in this article. By taking advantage of few-shot learning, CAUTION is able to construct an accurate user identification model with a very limited number of CSI training data. By converting the CSI samples into low-dimensional representations on the feature plane, it computes central points for different users as their CSI profiles and introduces an intruder threshold to measure whether the CSI data matches one of the user classes by a margin. The intruder threshold is able to be optimized without any intruders’ data. CAUTION does not require a large number of training data and provides an effective way to train the system for unknown intruder detection. We have tested CAUTION at different places and compared it with state-of-the-art CSI-based authentication systems. The experimental results demonstrate that CAUTION is able to perform accurate human authentication with a limited amount of CSI training data (one-fifth of data needed by compared systems) and outperforms the compared human authentication systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Internet Things J. | 2 |
| 2022 | EfficientFi: Toward Large-Scale Lightweight WiFi Sensing via CSI CompressionabstractWiFi technology has been applied to various places due to the increasing requirement of high-speed Internet access. Recently, besides network services, WiFi sensing is appealing in smart homes since it is device free, cost effective and privacy preserving. Though numerous WiFi sensing methods have been developed, most of them only consider single smart home scenario. Without the connection of powerful cloud server and massive users, large-scale WiFi sensing is still difficult. In this article, we first analyze and summarize these obstacles, and propose an efficient large-scale WiFi sensing framework, namely, EfficientFi. The EfficientFi works with edge computing at WiFi access points and cloud computing at center servers. It consists of a novel deep neural network that can compress fine-grained WiFi channel state information (CSI) at edge, restore CSI at cloud, and perform sensing tasks simultaneously. A quantized autoencoder and a joint classifier are designed to achieve these goals in an end-to-end fashion. To the best of our knowledge, the EfficientFi is the first Internet of Things-cloud-enabled WiFi sensing framework that significantly reduces communication overhead while realizing sensing tasks accurately. We utilized human activity recognition (HAR) and identification via WiFi sensing as two case studies, and conduct extensive experiments to evaluate the EfficientFi. The results show that it compresses CSI data from 1.368 Mb/s to 0.768 kb/s with extremely low error of data reconstruction and achieves over 98% accuracy for HAR. Jianfei Yang 0001, Xinyan Chen 0002, Han Zou, Dazhuo Wang, Qianwen Xu 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2021 | Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive NetworkabstractPartial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos, such negative transfer could be triggered by both spatial and temporal features, which leads to a more challenging Partial Video Domain Adaptation (PVDA) problem. In this paper, we propose a novel Partial Adversarial Temporal Attentive Network (PATAN) to address the PVDA problem by utilizing both spatial and temporal features for filtering source-only classes. Besides, PATAN constructs effective overall temporal features by attending to local temporal features that contribute more toward the class filtration process. We further introduce new benchmarks to facilitate research on PVDA problems, covering a wide range of PVDA scenarios. Empirical results demonstrate the state-of-the-art performance of our proposed PATAN across the multiple PVDA benchmarks. Code will be provided at: https://github.com/xuyu0010/PATAN. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Zhenghua Chen, Kezhi Mao |
ICCV | 2 |
| 2021 | Deep Reinforcement Learning Boosted Partial Domain AdaptationabstractDomain adaptation is critical for learning transferable features that effectively reduce the distribution difference among domains. In the era of big data, the availability of large-scale labeled datasets motivates partial domain adaptation (PDA) which deals with adaptation from large source domains to small target domains with less number of classes. In the PDA setting, it is crucial to transfer relevant source samples and eliminate irrelevant ones to mitigate negative transfer. In this paper, we propose a deep reinforcement learning based source data selector for PDA, which is capable of eliminating less relevant source samples automatically to boost existing adaptation methods. It determines to either keep or discard the source instances based on their feature representations so that more effective knowledge transfer across domains can be achieved via filtering out irrelevant samples. As a general module, the proposed DRL-based data selector can be integrated into any existing domain adaptation or partial domain adaptation models. Extensive experiments on several benchmark datasets demonstrate the superiority of the proposed DRL-based data selector which leads to state-of-the-art performance for various PDA tasks. Keyu Wu 0002, Min Wu 0008, Jianfei Yang 0001, Zhenghua Chen, Zhengguo Li, Xiaoli Li 0001 |
IJCAI | 3 |
| 2021 | Bi-Adversarial Discrepancy Minimization for Unsupervised Domain Adaptation on 3D Point CloudabstractDomain Adaptation (DA) has brought significant progress to a wide range of computer vision tasks. However, there are few methods that explores domain adaptation on 3D point cloud data, which plays crucial role in 3D recognition tasks (i.e., autonomous driving and robotics). In a nutshell, the unique challenges of point cloud data consist of abundant spatial global and local geometric information, complicated data distribution, and more ambiguous margins. In this paper, we introduce a novel Bi-Adversarial Discrepancy Minimization network (BADM), which aligns the cross-domain point cloud data by adopting two adversarial learning units with respect to global representations. Based on PointNet, we also utilize local alignment to reduce the discrepancy of local structures. In addition, our BADM method helps define more accurate margin and decision bounds. For this 3D point cloud DA scenario, we use PointDA-10 as the benchmark, which is extracted from ModelNet, ShapeNet, and ScanNet for cross-domain 3D object classification (10 mutual classes). Extensive experiments demonstrate that our approach achieves state-of-the-art performance, outperforming other 3D DA approaches. Changwei Xu, Jianfei Yang 0001 |
IJCNN | 3 |
| 2021 | Towards Realistic Visual Dubbing with Heterogeneous SourcesabstractThe task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art. Tianyi Xie, Liucheng Liao, Benlai Tang, Xiang Yin 0006, Jianfei Yang 0001, Jiali Yao, Yang Zhang 0088, Zejun Ma 0001 |
ACM Multimedia | 6 |
| 2021 | Improving WiFi-based Human Activity Recognition with Adaptive Initial State via One-shot LearningabstractWiFi-based human activity recognition technology has attracted widespread attention for its prominent application value and theoretical significance. Existing approaches have made great achievements in the same domain sensing, which means the activity samples applied for training the model have a similar distribution with the testing data. However, in practical application, we hope that the same activity of different people with various states and habits in different locations can be accurately recognized and produce the same reaction. Therefore, cross-domain sensing technology is pretty important. Some studies explore the location-independent and environment-independent methods, but few attempts consider the influence of the initial states of the users, such as standing and sitting, which actually have very different effects on the transmission of the wireless signal. This paper presents a human activity recognition method adapted to different initial states. Meanwhile, we solve the accompanying issue of the small sample size sensing, obviating the need for the cumbersome wok resulting from the massive data collection. We take advantage of the idea of metric learning and few-shot learning to realize cross-domain sensing with very few samples. The experiments demonstrate the feasibility and excellent performance of our method, which could recognize human activities with different initial states as the training data. Xue Ding 0001, Ting Jiang 0008, Yi Zhong 0002, Sheng Wu 0001, Jianfei Yang 0001, Wenling Xue |
WCNC | 5 |
| 2021 | Exploiting inter-frame regional correlation for efficient action recognition
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Expert Syst. Appl. | 2 |
| 2021 | PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition
Yuecong Xu, Haozhi Cao, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Neurocomputing | 3 |
| 2021 | Robust adversarial discriminative domain adaptation for real-world cross-domain visual recognition
Jianfei Yang 0001, Han Zou, Yuxun Zhou, Lihua Xie 0001 |
Neurocomputing | 1 |
| 2021 | Multimodal CSI-Based Human Activity Recognition Using GANsabstractChannel state information (CSI)-based human activity recognition (HAR) has received great attention in recent years due to its advantages in privacy protection, insensitivity to illumination, and no requirement for wearable devices. In this article, we propose a multimodal channel state information-based activity recognition (MCBAR) system that leverages existing WiFi infrastructures and monitors human activities from CSI measurements. MCBAR aims to address the performances degradation of WiFi-based human recognition systems due to environmental dynamics. Specifically, we address the issue of nonuniformly distributed unlabeled data with rarely performed activities by taking advantages of the generative adversarial network (GAN) and semisupervised learning. We apply a multimodal generator to approximate the CSI data distribution in different environment settings with limited measured CSI data. The generated CSI data using the multimodal generator can provide better diversity for knowledge transfer. This multimodal generator improves the ability of MCBAR to recognize specific activities with various CSI patterns caused by environmental dynamics. Compared to state-of-the-art CSI-based recognition systems, MCBAR is more robust as it is able to handle the nonuniformly distributed CSI data collected from a new environment setting. In addition, diverse generated data from the multimodal generator improves the stability of the system. We have tested MCBAR under multiple experimental settings at different places. The experimental results demonstrate that our algorithm overcomes environmental dynamics and outperforms existing HAR systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
IEEE Internet Things J. | 2 |
| 2021 | Learning decomposed hierarchical feature for better transferability of deep modelsabstractDeep models have achieved prominent results in pattern recognition tasks, especially computer vision and natural language processing. However, the dataset bias caused by the distribution discrepancy between the training and testing data hinders the generalization ability of deep models. Though many domain adaptation approaches have been proposed to mitigate such negative effect, most of them improve the transferability of features by aligning global distributions of deep models. Few researchers pay attention to the versatility of deep features which can play a vital role in cross-domain recognition. In this paper, we propose to enrich the classic deep learning models by capturing high-low-frequency information and multi-scale features, which deal with the domain shift that cannot be easily addressed by merely feature-level alignment. The Hierarchical Transfer Network (HTN) leverages octave convolution, pyramid features, and self-attention mechanism for revamping the classic models, which can be further integrated with any domain alignment approaches by replacing the feature extractor with the proposed HTN. Extensive experiments have been conducted on three public domain adaptation benchmarks. The results show that the proposed HTN can effectively improve adversarial-based, statistics-based, and norm-based domain adaptation approaches, achieving competitive performance without involving model complexity. Jianfei Yang 0001, Hanjie Qian, Han Zou, Lihua Xie 0001 |
Inf. Sci. | 1 |
| 2021 | Effective action recognition with embedded key point shifts
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Pattern Recognit. | 3 |
| 2020 | Suppressing Uncertainties for Large-Scale Facial Expression RecognitionabstractAnnotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognition (FER) in data-driven deep learning era. To address this problelm, this paper proposes to suppress the uncertainties by a simple yet efficient Self-Cure Network (SCN). Specifically, SCN suppresses the uncertainty from two different aspects: 1) a self-attention mechanism over FER dataset to weight each sample in training with a ranking regularization, and 2) a careful relabeling mechanism to modify the labels of these samples in the lowest-ranked group. Experiments on synthetic FER datasets and our collected WebEmotion dataset validate the effectiveness of our method. Results on public benchmarks demonstrate that our SCN outperforms current state-of-the-art methods with \textbf{88.14}\% on RAF-DB, \textbf{60.23}\% on AffectNet, and \textbf{89.35}\% on FERPlus. Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Shijian Lu, Yu Qiao 0001 |
CVPR | 3 |
| 2020 | Suppressing Mislabeled Data via Grouping and Self-attention
Xiaojiang Peng, Kai Wang 0036, Zhaoyang Zeng, Qing Li 0058, Jianfei Yang 0001, Yu Qiao 0001 |
ECCV (16) | 5 |
| 2020 | Mind the Discriminability: Asymmetric Adversarial Domain Adaptation
Jianfei Yang 0001, Han Zou, Yuxun Zhou, Zhaoyang Zeng, Lihua Xie 0001 |
ECCV (24) | 1 |
| 2020 | Robust CSI-based Human Activity Recognition using Roaming GeneratorabstractChannel State Information (CSI) based human activity recognition has received great attention in recent years due to its advantages in privacy protection, insensitive to illumination and no requirement for wearable devices. However, for practical deployment, it needs to greatly enhance the performance robustness against dynamic changes of the surrounding environment. To address this problem, we propose a novel CSI based activity recognition using Roaming Generator (CSIRoG) system for human activity detection. CSIRoG leverages existing WiFi infrastructures and monitors human behaviours from CSI measurements. It utilizes the generative adversarial network (GAN) to transfer the CSI information from one environment to another with dynamic changes such as people passing by, furniture layout changes, etc. The proposed method aims to approximate the CSI distribution in the new environment setting which has very limited CSI data. Therefore, the system can learn to handle multiple environment dynamics. Compared to the existing works, CSIRoG leverages a multimodal system model for better diversity of the generated CSI data for knowledge transfer. This improves the ability of CSIRoG to recognize various kinds of CSI information for one specific user activity caused by various dynamic conditions, thus enhancing system robustness. We have tested CSIRoG under multiple environment settings at different places. The experimental results demonstrate that our algorithm overcomes environmental dynamics and outperforms existing human activity recognition systems. Dazhuo Wang, Jianfei Yang 0001, Wei Cui 0002, Lihua Xie 0001, Sumei Sun |
ICARCV | 2 |
| 2020 | Domain Adaptation for Degraded Remote Scene ClassificationabstractRemote scene classification serves a vital role in many applications. However, satellite images are often blurred and degraded due to aerosol scattering under fog, haze, and other weather conditions, reducing the image contrast and color fidelity. State-of-the-art remote sensing classification models building upon convolutional neural networks (CNNs) are mostly trained on annotated datasets of clear satellite images. When applied to blurred images, they will suffer a great degradation in performance. To address this problem, we adopt the domain adaptation algorithm TADA and propose Transferable Attention enhanced Adversarial Adaptation Network (TA3N), which utilizes annotated data in clear images by applying knowledge transferring from clear image domain to blurred image domain. Our TA3N first integrates spatial attention to focus on salient areas which are discriminative and transferable. In addition, domain discriminator and adversarial training via gradient reversal layer are used to minimize the discrepancies in extracted features from clear and degraded domains. We synthesize degraded remote scene classification dataset SSI based on FoHIS model. Experiments on degraded SSI showed that TA3N significantly outperforms baseline and other state-of-the-art domain adaptation methods. Jianfei Yang 0001, Hailin Chen, Yuecong Xu, Ziji Shi, Ruikang Luo, Lihua Xie 0001, Rong Su 0001 |
ICARCV | 1 |
| 2020 | MobileDA: Toward Edge-Domain AdaptationabstractDeep neural networks (DNNs) have made significant advances in computer vision and sensor-based smart sensing. DNNs achieve prominent results based on standard data sets and powerful servers, whereas, in real applications with domain-shift data and resource-constrained environments such as Internet-of-Things (IoT) devices in the edge computing, DNNs are likely to have degraded performance in terms of accuracy and efficiency. To this end, we develop the MobileDA framework that learns transferable features while keeping the simple structure of the deep model. Our method allows a novel teacher network trained in the server to distill the knowledge for a student network running in the edge device, which is achieved by a cross-domain distillation. Leveraging unlabeled data in the new environment, our student model amends the feature learning to be domain invariant, then being our objective model running in the edge device. Our approach is evaluated on a challenging IoT-based WiFi gesture recognition scenario, and three classic visual adaptation benchmarks. The empirical studies corroborate the effectiveness of distillation for domain transfer, and the overall results show that our model achieves state-of-the-art performance merely using a simple network. Jianfei Yang 0001, Han Zou, Shuxin Cao, Zhenghua Chen, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2020 | Adversarial Learning-Enabled Automatic WiFi Indoor Radio Map Construction and Adaptation With Mobile RobotabstractLocation-based service (LBS) has become an indispensable part of our daily lives. Realizing accurate LBS in indoor environments is still a challenging task. WiFi fingerprinting-based indoor positioning system (IPS) has achieved encouraging results recently, but the time and labor overhead of constructing a dense WiFi radio map remains the key bottleneck that hinders it for real-world large-scale implementation. In this article, we propose WiGAN an automatic fine-grained indoor ratio map construction and the adaptation scheme empowered by the Gaussian process regression conditioned least-squares generative adversarial networks (GPR-GANs) with a mobile robot. First, we develop a mobile robotic platform that constructs the spatial map and radio map simultaneously in the easily accessed free space. GPR-GAN first establishes a Gaussian process regression (GPR) model using the real received signal strength (RSS) measurements collected by our robotic platform via LiDAR SLAM in the free space. Then, the outputs of the GPR are adopted as the input of GAN's generator. The learning objective of GAN is to synthesize realistic RSS data in a constrained space where it has not been covered and model the irregular RSS distributions in complex indoor environments. Real-world experiments were conducted in a real-world indoor environment, which confirms the feasibility, high accuracy, and superiority of WiGAN over existing solutions in terms of both RSS estimation accuracy and localization accuracy. Han Zou, Maoxun Li, Jianfei Yang 0001, Yuxun Zhou, Lihua Xie 0001, Costas J. Spanos |
IEEE Internet Things J. | 4 |
| 2020 | Region Attention Networks for Pose and Occlusion Robust Facial Expression RecognitionabstractOcclusion and pose variations, which can change facial appearance significantly, are two major obstacles for automatic Facial Expression Recognition (FER). Though automatic FER has made substantial progresses in the past few decades, occlusion-robust and pose-invariant issues of FER have received relatively less attention, especially in real-world scenarios. This paper addresses the real-world pose and occlusion robust FER problem in the following aspects. First, to stimulate the research of FER under real-world occlusions and variant poses, we annotate several in-the-wild FER datasets with pose and occlusion attributes for the community. Second, we propose a novel Region Attention Network (RAN), to adaptively capture the importance of facial regions for occlusion and pose variant FER. The RAN aggregates and embeds varied number of region features produced by a backbone convolutional neural network into a compact fixed-length representation. Last, inspired by the fact that facial expressions are mainly defined by facial action units, we propose a region biased loss to encourage high attention weights for the most important regions. We validate our RAN and region biased loss on both our built test datasets and four popular datasets: FERPlus, AffectNet, RAF-DB, and SFEW. Extensive experiments show that our RAN and region biased loss largely improve the performance of FER with occlusion and variant pose. Our method also achieves state-of-the-art results on FERPlus, AffectNet, RAF-DB, and SFEW. Code and the collected test data will be publicly available. Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Debin Meng, Yu Qiao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 63 |
| 2019 | Consensus Adversarial Domain AdaptationabstractWe propose a novel domain adaptation framework, namely Consensus Adversarial Domain Adaptation (CADA), that gives freedom to both target encoder and source encoder to embed data from both domains into a common domaininvariant feature space until they achieve consensus during adversarial learning. In this manner, the domain discrepancy can be further minimized in the embedded space, yielding more generalizable representations. The framework is also extended to establish a new few-shot domain adaptation scheme (F-CADA), that remarkably enhances the ADA performance by efficiently propagating a few labeled data once available in the target domain. Extensive experiments are conducted on the task of digit recognition across multiple benchmark datasets and a real-world problem involving WiFi-enabled device-free gesture recognition under spatial dynamics. The results show the compelling performance of CADA versus the state-of-the-art unsupervised domain adaptation (UDA) and supervised domain adaptation (SDA) methods. Numerical experiments also demonstrate that F-CADA can significantly improve the adaptation performance even with sparsely labeled data in the target domain. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Huihan Liu, Hari Prasanna Das, Costas J. Spanos |
AAAI | 3 |
| 2019 | Kervolutional Neural NetworksabstractConvolutional neural networks (CNNs) have enabled the state-of-the-art performance in many computer vision tasks. However, little effort has been devoted to establishing convolution in non-linear space. Existing works mainly leverage on the activation layers, which can only provide point-wise non-linearity. To solve this problem, a new operation, kervolution (kernel convolution), is introduced to approximate complex behaviors of human perception systems leveraging on the kernel trick. It generalizes convolution, enhances the model capacity, and captures higher order interactions of features, via patch-wise kernel functions, but without introducing additional parameters. Extensive experiments show that kervolutional neural networks (KNN) achieve higher accuracy and faster convergence than baseline CNN. Chen Wang 0033, Jianfei Yang 0001, Lihua Xie 0001, Junsong Yuan 0001 |
CVPR | 2 |
| 2019 | Exploring Regularizations with Face, Body and Image Cues for Group Cohesion PredictionabstractThis paper presents our approach for the group cohesion prediction sub-challenge in the EmotiW 2019. The task is to predict group cohesiveness in images. We mainly explore several regularizations with three types of visual cues, namely face, body ,and global image. Our main contribution is two-fold. First, we jointly train the group cohesion prediction task and group emotion recognition task using multi-task learning strategy with all visual cues. Second, we elaborately design two regularizations, namely a rank loss and a hourglass loss, where the former aims to give a margin between the distance of distant categories and near categories and the later to avoid centralization predictions with only MSE loss. With careful evaluations, we finally achieve the second place in this sub-challenge with MSE of 0.43821 on the testing set. https://github.com/DaleAG/Group_Cohesion_Prediction Da Guo, Kai Wang 0036, Jianfei Yang 0001, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 3 |
| 2019 | Bootstrap Model Ensemble and Rank Loss for Engagement Intensity RegressionabstractThis paper presents our approach for the engagement intensity regression task of EmotiW 2019. The task is to predict the engagement intensity value of a student when he or she is watching an online MOOCs video in various conditions. Based on our winner solution last year, we mainly explore head features and body features with a bootstrap strategy and two novel loss functions in this paper. We maintain the framework of multi-instance learning with long short-term memory (LSTM) network, and make three contributions. First, besides of the gaze and head pose features, we explore facial landmark features in our framework. Second, inspired by the fact that engagement intensity can be ranked in values, we design a rank loss as a regularization which enforces a distance margin between the features of distant category pairs and adjacent category pairs. Third, we use the classical bootstrap aggregation method to perform model ensemble which randomly samples a certain training data by several times and then averages the model predictions. We evaluate the performance of our method and discuss the influence of each part on the validation dataset. Our methods finally win 3rd place with MSE of 0.0626 on the testing set. https://github.com/kaiwang960112/EmotiW_2019_ engagement_regression Kai Wang 0036, Jianfei Yang 0001, Da Guo, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 2 |
| 2019 | Semantic-filtered Soft-Split-Aware video captioning with audio-augmented feature
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao |
Neurocomputing | 2 |
| 2019 | Learning Gestures From WiFi: A Siamese Recurrent Convolutional ArchitectureabstractWe propose a gesture recognition system that leverages existing WiFi infrastructures and learns gestures from channel state information (CSI) measurements. Having developed an innovative OpenWrt-based platform for commercial WiFi devices to extract CSI data, we propose a novel deep Siamese representation learning architecture for one-shot gesture recognition. Technically, our model extends the capacity of spatio-temporal patterns learning for the standard Siamese structure by incorporating convolutional and bidirectional recurrent neural networks. More importantly, the representation learning is ameliorated by our Siamese framework and transferable pairwise loss which helps to remove structured noise, such as individual heterogeneity and various measurement conditions during domain-different training. Meanwhile, our Siamese model also enables one-shot learning for higher availability in reality. We prototype our system on commercial WiFi routers. The experiments demonstrate that our model outperforms state-of-the-art solutions for temporal-spatial representation learning and achieves satisfactory results under one-shot conditions. Jianfei Yang 0001, Han Zou, Yuxun Zhou, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2019 | Unsupervised WiFi-Enabled IoT Device-User Association for Personalized Location-Based ServiceabstractA fundamental building block toward personalized location-based service and context-aware service in smart buildings is the knowledge about the identity and mobility of users in indoor environments. Conventional user identification systems require the deployment of dedicated infrastructure or the active user involvement. Motivated by the widespread usage of the WiFi-enabled mobile device (MD), e.g., people usually carry at least one MD in their daily lives, in this paper, we propose WinDUA, a WiFi-enabled nonintrusive device and user association scheme to infer user identity and mobility via a novel unsupervised association learning algorithm. First, we utilize our WiFi-based indoor positioning system to obtain the historical location data of each MD using only existing WiFi infrastructure in a nonintrusive manner. Then, we classify all the MDs into two categories: 1) static device (SD) and 2) mobile phone (MP), according to their location variations and overnight presences. Subsequently, we estimate the correct mapping between each SD and its user through hierarchical clustering and location similarity matching between its location and user's personal space. Finally, we make possible pairs of MP and SD according to their duration of coexistence as well as the historical location similarity to associate the owner of each MP. Real-world experiments are conducted in an office, verifying that WinDUA is able to associate the MD to the correct users in a nonintrusive and unsupervised manner. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Costas J. Spanos |
IEEE Internet Things J. | 3 |
| 2018 | WiFi-Based Human Identification via Convex Tensor Shapelet LearningabstractWe propose AutoID, a human identification system that leverages the measurements from existing WiFi-enabled Internet of Things (IoT) devices and produces the identity estimation via a novel sparse representation learning technique. The key idea is to use the unique fine-grained gait patterns of each person revealed from the WiFi Channel State Information (CSI) measurements, technically referred to as shapelet signatures, as the "fingerprint" for human identification. For this purpose, a novel OpenWrt-based IoT platform is designed to collect CSI data from commercial IoT devices. More importantly, we propose a new optimization-based shapelet learning framework for tensors, namely Convex Clustered Concurrent Shapelet Learning (C3SL), which formulates the learning problem as a convex optimization. The global solution of C3SL can be obtained efficiently with a generalized gradient-based algorithm, and the three concurrent regularization terms reveal the inter-dependence and the clustering effect of the CSI tensor data. Extensive experiments are conducted in multiple real-world indoor environments, showing that AutoID achieves an average human identification accuracy of 91% from a group of 20 people. As a combination of novel sensing and learning platform, AutoID attains substantial progress towards a more accurate, cost-effective and sustainable human identification system for pervasive implementations. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Weixi Gu, Lihua Xie 0001, Costas J. Spanos |
AAAI | 3 |
| 2018 | DeepSense: Device-Free Human Activity Recognition via Autoencoder Long-Term Recurrent Convolutional NetworkabstractIn the era of Internet of Things (IoT), human activity recognition is becoming the vital underpinning for a myriad of emerging applications in smart home and smart buildings. Existing activity recognition approaches require either the deployment of extra infrastructure or the cooperation of occupants to carry dedicated devices, which are expensive, intrusive and inconvenient for pervasive implementation. In this paper, we propose DeepSense, a device-free human activity recognition scheme that can automatically identify common activities via deep learning using only commodity WiFi-enabled IoT devices. We design a novel OpenWrt-based IoT platform to collect Channel State Information (CSI) measurements from commercial IoT devices. Moreover, an innovative deep learning framework, Autoencoder Long-term Recurrent Convolutional Network (AE-LRCN), is proposed. It consists of an autoencoder module, a convolutional neural network (CNN) module and a long short-term memory (LSTM) module, which aims to sanitize the noise in raw CSI data, extract high-level representative features and reveal the inherent temporal dependencies among data for accurate human activity recognition, respectively. All the hyperparameters in AE-LRCN are fine-tuned end-to-end automatically. Extensive experiments are conducted in typical indoor environments and the experimental results demonstrate that DeepSense outperforms existing methods and achieves a 97.6% activity recognition accuracy without human intervention. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Hao Jiang 0008, Lihua Xie 0001, Costas J. Spanos |
ICC | 3 |
| 2018 | Robust WiFi-Enabled Device-Free Gesture Recognition via Unsupervised Adversarial Domain AdaptationabstractAccurate human gesture recognition is becoming a cornerstone for myriad emerging applications in human-computer interaction. Existing gesture recognition systems either require dedicated extra infrastructure or user's active cooperation. Although some WiFi-enabled gesture recognition systems have been proposed, they are vulnerable to environmental dynamics and rely on the tedious data re-labeling and expert knowledge each time being implemented in a new environment. In this paper, we propose a WiFi- enabled device-free adaptive gesture recognition scheme, WiADG, that is able to identify human gestures accurately and consistently under environmental dynamics via adversarial domain adaptation. Firstly, a novel OpenWrt-based IoT platform is developed, enabling the direct collection of Channel State Information (CSI) measurements from commercial IoT devices. After constructing an accurate source classifier with labeled source CSI data via the proposed convolutional neural network in the source domain (original environment), we design an unsupervised domain adaptation scheme to reduce the domain discrepancy between the source and the target domain (new environment) and thus improve the generalization performance of the source classifier. The domain- adversarial objective is to train a generator (target encoder) to map the unlabeled target data to a domain invariant latent feature space so that a domain discriminator cannot distinguish the domain labels of the data. In the phase of implementation, we utilize the trained target encoder to map the target CSI frame to the latent feature space and use the source classifier to identify various gestures performed by the user. We implement WiADG on commercial WiFi routers and conduct experiments in multiple indoor environments. The results validate that WiADG achieves 98% gesture recognition accuracy in the original environment. Furthermore, the proposed unsupervised adversarial domain adaptation is able to enhance the recognition accuracy of WiADG by 25% on average without the needs of labeled data collection and new classifier generation when implements it in new environments. Han Zou, Jianfei Yang 0001, Yuxun Zhou, Lihua Xie 0001, Costas J. Spanos |
ICCCN | 2 |
| 2018 | Cascade Attention Networks For Group Emotion Recognition with Face, Body and Image CuesabstractThis paper presents our approach for group-level emotion recognition sub-challenge in the EmotiW 2018. The task is to classify an image into one of the group emotions such as positive, negative, and neutral. Our approach mainly explores three cues, namely face, body and global image with recent deep networks. Our main contribution is two-fold. First, we introduce body based Convolutional Neural Networks (CNNs) into this task based on our previous winner method [18]. For body based CNNs, we crop all bodies in an image with the state-of-the-art human pose estimation method and train CNNs with the image-level label to capture. The body cue captures a full view of an individual. Second, we propose a cascade attention network for the face cue in images. This network exploits the importance of each face in an image to generates a global representation based on all faces. The cascade attention network is not only complementary with other models but also improves the naive average pooling method by about 2%. We finally achieve the second place in this sub-challenge with classification accuracies of 86.9% and 67.48% on the validation set and testing set, respectively. Kai Wang 0036, Xiaoxing Zeng, Jianfei Yang 0001, Debin Meng, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 3 |
| 2018 | Deep Recurrent Multi-instance Learning with Spatio-temporal Features for Engagement Intensity PredictionabstractThis paper elaborates the winner approach for engagement intensity prediction in the EmotiW Challenge 2018. The task is to predict the engagement level of a subject when he or she is watching an educational video in diverse conditions and different environments. Our approach formulates the prediction task as a multi-instance regression problem. We divide an input video sequence into segments and calculate the temporal and spatial features of each segment for regressing the intensity. Subject engagement, that is intuitively related with body and face changes in time domain, can be characterized by long short-term memory (LSTM) network. Hence, we build a multi-modal regression model based on multi-instance mechanism as well as LSTM. To make full use of training and validation data, we train different models for different data split and conduct model ensemble finally. Experimental results show that our method achieves mean squared error (MSE) of 0.0717 in the validation set, which improves the baseline results by 28%. Our methods finally win the challenge with MSE of 0.0626 on the testing set. Jianfei Yang 0001, Kai Wang 0036, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 1 |
| 2018 | Joint Adversarial Domain Adaptation for Resilient WiFi-Enabled Device-Free Gesture RecognitionabstractHuman gesture recognition plays a critical role in numerous applications of human-computer interaction. By analyzing how gesture alters the WiFi propagation among WiFi-enabled IoT devices to identify the gestures in a device-free manner could be a promising solution. However, existing methods require tedious data collection and labeling process each time being implemented in a new environment. The classifier constructed by SVM or random forest is vulnerable to spatial dynamics. In this paper, we proposed JADA, a novel unsupervised Joint adversarial domain adaptation (JADA) scheme that realizes accurate and resilient WiFi-enabled device-free gesture recognition without collecting and labeling training data in new environments. After constructing a source encoder and a source classifier in the source domain by convolutional neural network, JADA trains a target encoder and also fine-tunes the source encoder through adversarial learning to map both unlabeled target data and labeled source data to a domain-invariant feature space such that a domain discriminator cannot distinguish the domain labels of the data. After training a shared classifier with the labeled source data while fixing the parameters of the source encoder, we employ the trained target encoder to embed the test target samples into the domain-invariant feature space and infer its class using the shared classifier. We develop a novel Channel State Information (CSI) enabled IoT platform that could obtain fine-grained CSI time series data directly from IoT devices and transform them into CSI frames. Real-world experiments with COTS WiFi routers were conducted in 2 indoor environments. The experimental results demonstrate that JADA achieves 98.75% gesture recognition accuracy in the original environment. Moreover, when the environmental scenario is altered, it is able to reduce the domain discrepancy across domains without collecting any labeled data in the new context. Han Zou, Jianfei Yang 0001, Yuxun Zhou, Costas J. Spanos |
ICMLA | 2 |
| 2018 | Fine-grained adaptive location-independent activity recognition using commodity WiFiabstractDevice-free activity recognition is appealing in smart home applications. It not only is convenient, but also causes no privacy concern, as compared to other activity recognition techniques such as the vision based technique. Existing WiFi-based methods have achieved high accuracy in static circumstances but have limitations in adapting changes in environment and activities locations. In this paper, we propose a fine-grained adaptive location-independent activity recognition system (FALAR) which leverages WiFi signals to characterize and recognize common activities regardless of inconsistency of mutative surroundings. FALAR applies fine-grained channel state information (CSI) to achieve accurate recognitions. To address the issue of environmental changes, we present a Kernel Density Estimation (KDE) based motion extraction method and a coarse-to-fine search strategy for speedy processing. After a denoising scheme, we introduce Class Estimated Basis Space Singular Value Decomposition (CSVD) to efface the static path in the background, and use nonnegative matrix factorization to distinguish various activities by looking into the signal profiles. We evaluate FALAR using two commodity WiFi routers in a typical office environment. Our results show that it achieves remarkable performance. Jianfei Yang 0001, Han Zou, Hao Jiang 0008, Lihua Xie 0001 |
WCNC | 1 |
| 2018 | Device-Free Occupant Activity Sensing Using WiFi-Enabled IoT Devices for Smart HomesabstractIntelligent occupancy sensing is becoming a vital underpinning for various emerging applications in smart homes, such as security surveillance and human behavior analysis. However, prevailing approaches mainly rely on video camera, ambient sensors, or wearable devices, which either requires arduous deployment or arouses privacy concerns. In this paper, we present a novel real-time, device-free, and privacy-preserving WiFi-enabled Internet of Things platform for occupancy sensing, which can promote a myriad of emerging applications. It is designed to achieve an optimal tradeoff between performance and scalability. Our system empowers commercial off-the-shelf WiFi routers to collect channel state information (CSI) measurements and provides an efficient cloud server for computing via a lightweight communication protocol. To demonstrate the usefulness of our platform, an occupancy detection system is developed by exploiting the CSI curve of human presence. Furthermore, we also design an innovative activity recognition system based on our platform and machine learning techniques with high availability and extensibility. In the evaluation, the experimental results show that our platform enables these applications efficiently, with the accuracy of 96.8% and 90.6% in terms of occupancy detection and recognition, respectively. Jianfei Yang 0001, Han Zou, Hao Jiang 0008, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2017 | FreeCount: Device-Free Crowd Counting with Commodity WiFiabstractIn the era of Internet of Things, crowd counting, which estimates the number of people within a region, becomes the underpinning for many emerging applications, such as occupancy estimation in smart building and queuing management and product placement in shopping center. Existing vision based crowd counting schemes require favorable lighting conditions and also raise privacy concerns. RF based approaches rely on specialized sensors and require users to carry RF devices. Thus, an accurate, reliable and non-intrusive crowd counting scheme is still desired. In this paper, we propose FreeCount, a device-free crowd counting scheme that is able to precisely estimate the number of people within a region using only commodity WiFi routers. To this end, the channel state information (CSI) data in PHY layer is obtained directly by upgrading the router's software. We propose an information theory based feature selection scheme to select the most representative features that are sensitive to human motion. To build a classifier that is robust to temporal and environmental disparities, we adopt transfer kernel learning, which minimizes the difference between the source and target distributions in the reproducing kernel Hilbert space, is adopted to process the real-time CSI feature data. Experiments were conducted in moderate sized rooms and the results demonstrated that FreeCount is able to accurately estimate the number of people with 96% crowd counting accuracy consistently over temporal and environmental variation. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Weixi Gu, Lihua Xie 0001, Costas J. Spanos |
GLOBECOM | 3 |
| 2017 | Multiple Kernel Representation Learning for WiFi-Based Human Activity RecognitionabstractHuman activity recognition is becoming the vital underpinning for a myriad of emerging applications in the field of human-computer interaction, mobile computing, and smart grid. Besides the utilization of up-to-date sensing techniques, modern activity recognition systems also require a machine learning (ML) algorithm that leverages the sensory data for identification purposes. In view of the unique characteristics of the measurement data and the ML challenges thereof, we propose a non-intrusive human activity recognition system that only uses existing commodity WiFi routers. The core of our system is a novel multiple kernel representation learning (MKRL) framework that automatically extracts and combines informative patterns from the Channel State Information (CSI) measurements. The MKRL firstly learns a kernel string representation from time, frequency, wavelet, and shape domains with an efficient greedy algorithm. Then it performs information fusion from diverse perspectives based on multi-view kernel learning. Moreover, different stages of MKRL can be seamlessly integrated into a multiple kernel learning framework to build up a robust and comprehensive activity classifier. Extensive experiments are conducted in typical indoor environments and the experimental results demonstrate that the proposed system outperforms existing methods and achieves a 98\% activity recognition accuracy. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Weixi Gu, Lihua Xie 0001, Costas J. Spanos |
ICMLA | 3 |
| 2017 | Poster: WiFi-based Device-Free Human Activity Recognition via Automatic Representation LearningabstractExisting human activity recognition approaches require either the deployment of extra infrastructure or the cooperation of occupants to carry dedicated devices, which are expensive, intrusive and inconvenient for pervasive implementation. In this paper, we propose SmartSense, a device-free human activity recognition system based on a novel machine learning algorithm with existing commercial off-the-shelf (COTS) WiFi routers. By exploiting the prevalence of existing WiFi infrastructure in buildings, we developed a novel OpenWrt based firmware for COTS WiFi routers to collect the CSI measurements from regular data frames. To identify different human activities, an automatic kernel representation learning method, namely auto-HSRL, is established to selection informative Hilbert space patterns from time, frequency, wavelet, and shape domains. A new information fusion tool based on multi-view kernel learning is proposed to combine the representations extracted from diverse perspectives and build up a robust and comprehensive activity classifier. Extensive experiments were conducted in an office and the experimental results demonstrate that SmartSense outperforms existing methods and achieves a 98% activity recognition accuracy. Han Zou, Yuxun Zhou, Jianfei Yang 0001, Weixi Gu, Lihua Xie 0001, Costas J. Spanos |
MobiCom | 3 |