VLDB 2026 Research / reviewers in the wild / expert
Yunjiao Zhou
dblp:327/3251
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0009-7515-1739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time TrainingabstractUnderstanding complex human activities demands the ability to decompose motion into fine-grained, semantic-aligned sub-actions. This motion grounding process is crucial for behavior analysis, embodied AI and virtual reality. Yet, most existing methods rely on dense supervision with predefined action classes, which are infeasible in open-vocabulary, real-world settings. In this paper, we propose ZOMG, a zero-shot, open-vocabulary framework that segments motion sequences into semantically meaningful sub-actions without requiring any annotations or fine-tuning. Technically, ZOMG integrates (1) language semantic partition, which leverages large language models to decompose instructions into ordered sub-action units, and (2) soft masking optimization, which learns instance-specific temporal masks to focus on frames critical to sub-actions, while maintaining intra-segment continuity and enforcing inter-segment separation, all without altering the pretrained encoder. Experiments on three motion-language datasets demonstrate state-of-the-art effectiveness and efficiency of motion grounding performance, outperforming prior methods by 8.7% mAP on HumanML3D benchmark. Meanwhile, significant improvements also exist in downstream retrieval, establishing a new paradigm for annotation-free motion understanding. Yunjiao Zhou, Xinyan Chen 0002, Junlang Qian, Lihua Xie 0001, Jianfei Yang 0001 |
AAAI | 1 |
| 2026 | SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action RecognitionabstractSkeleton-based action recognition leverages human pose keypoints to categorize human actions, which shows superior generalization and interoperability compared to regular end-to-end action recognition. Existing solutions use RGB cameras to annotate skeletal keypoints, but their performance declines in dark environments and raises privacy concerns, limiting their use in smart homes and hospitals. This paper explores non-invasive wireless sensors, i.e., LiDAR and mmWave, to mitigate these challenges as a feasible alternative. Two problems are addressed: (1) insufficient data on wireless sensor modality to train an accurate skeleton estimation model, and (2) skeletal keypoints derived from wireless sensors are noisier than RGB, causing great difficulties for subsequent action recognition models. Our work, SkeFi, overcomes these gaps through a novel cross-modal knowledge transfer method acquired from the data-rich RGB modality. We propose the enhanced Temporal Correlation Adaptive Graph Convolution (TC-AGC) with frame interactive enhancement to overcome the noise from missing or inconsecutive frames. Additionally, our research underscores the effectiveness of enhancing multiscale temporal modeling through dual temporal convolution. By integrating TC-AGC with temporal modeling for cross-modal transfer, our framework can extract accurate poses and actions from noisy wireless sensors. Experiments demonstrate that SkeFi realizes state-of-the-art performances on mmWave and LiDAR. The code is available at https://github.com/Huang0035/Skefi. Shunyu Huang, Yunjiao Zhou, Jianfei Yang 0001 |
IEEE Internet Things J. | 2 |
| 2026 | TENT: Connect Language Models With IoT Sensors for Zero-Shot Activity RecognitionabstractThe rapid expansion of the Internet of Things (IoT) has introduced new challenges in Human Activity Recognition (HAR), particularly in dynamic environments where new and unforeseen activities emerge. Traditional HAR models, relying on predefined labels, struggle to adapt to these scenarios, highlighting the need for zero-shot learning (ZSL) approaches that can generalize beyond fixed training categories. Recent advances in large language models (LLMs) have demonstrated remarkable zero-shot capability in textual and visual domains. However, extending this ability to IoT sensors is substantially more challenging due to their heterogeneous modalities, diverse data structures, and limited semantic annotations. In this paper, we propose TENT (IoT-sEnsorslanguage alignmEnt pre-Training), a novel framework that constructs a unified sensor-language semantic space for zero-shot HAR. Instead of aligning each sensor individually to text, TENT jointly aligns multiple heterogeneous modalities with language, treating them as peers rather than anchors. This balanced multi-modal alignment allows sensors to mutually regularize one another while being grounded in linguistic semantics, transforming heterogeneity from a barrier into a strength. To further enrich the semantic space, TENT incorporates detailed activity descriptions and learnable prompts, enhancing adaptability to unseen activities. Extensive experiments across datasets and evaluation protocols demonstrate that TENT not only achieves robust recognition of both seen and unseen activities but also significantly outperforms existing vision-language and sensor-language baselines, surpassing them by over 20% on zero-shot HAR tasks. These results establish TENT as a new paradigm for generalizable IoT representation learning. Yunjiao Zhou, Jianfei Yang 0001, Han Zou, Lihua Xie 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | T3DNet: Compressing Point Cloud Models for Lightweight 3-D RecognitionabstractThe 3-D point cloud has been widely used in many mobile application scenarios, including autonomous driving and 3-D sensing on mobile devices. However, existing 3-D point cloud models tend to be large and cumbersome, making them hard to deploy on edged devices due to their high memory requirements and nonreal-time latency. There has been a lack of research on how to compress 3-D point cloud models into lightweight models. In this article, we propose a method called T3DNet (tiny 3-D network with augmentation and distillation) to address this issue. We find that the tiny model after network augmentation is much easier for a teacher to distill. Instead of gradually reducing the parameters through techniques, such as pruning or quantization, we predefine a tiny model and improve its performance through auxiliary supervision from augmented networks and the original model. We evaluate our method on several public datasets, including ModelNet40, ShapeNet, and ScanObjectNN. Our method can achieve high compression rates without significant accuracy sacrifice, achieving state-of-the-art performances on three datasets against existing methods. Amazingly, our T3DNet is 58 smaller and 54 faster than the original model yet with only 1.4 accuracy descent on the ModelNet40 dataset. Our code is available at https://github.com/Zhiyuan002/T3DNet. Yunjiao Zhou, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | PowerSkel: A Device-Free Framework Using CSI Signal for Human Skeleton Estimation in Power StationabstractSafety monitoring of power operations in power stations is crucial for preventing accidents and ensuring stable power supply. However, conventional methods such as wearable devices and video surveillance have limitations such as high cost, dependence on light, and visual blind spots. WiFi-based human pose estimation is a suitable method for monitoring power operations due to its low cost, device-free, and robustness to various illumination conditions. In this paper, a novel Channel State Information (CSI)-based pose estimation framework, namely PowerSkel, is developed to address these challenges. PowerSkel utilizes self-developed CSI sensors to form a mutual sensing network and constructs a CSI acquisition scheme specialized for power scenarios. It significantly reduces the deployment cost and complexity compared to the existing solutions. To reduce interference with CSI in the electricity scenario, a sparse adaptive filtering algorithm is designed to preprocess the CSI. CKDformer, a knowledge distillation network based on collaborative learning and self-attention, is proposed to extract the features from CSI and establish the mapping relationship between CSI and keypoints. The experiments are conducted in a real-world power station, and the results show that the PowerSkel achieves high performance with a PCK@50 of 96.27%, and realizes a significant visualization on pose estimation, even in dark environments. Our work provides a novel low-cost and high-precision pose estimation solution for power operation. Cunyi Yin, Xiren Miao, Jing Chen 0022, Hao Jiang 0008, Jianfei Yang 0001, Yunjiao Zhou, Min Wu 0008, Zhenghua Chen |
IEEE Internet Things J. | 6 |
| 2024 | AdaPose: Toward Cross-Site Device-Free Human Pose Estimation With Commodity WiFiabstractWiFi-based pose estimation is a technology with great potential for the development of smart homes and metaverse avatar generation. However, current WiFi-based pose estimation methods are predominantly evaluated under controlled laboratory conditions with sophisticated vision models to acquire accurately labeled data. Furthermore, WiFi channel state information (CSI) is highly sensitive to environmental variables, and direct application of a pretrained model to a new environment may yield suboptimal results due to domain shift. In this article, we propose a domain adaptation algorithm, AdaPose, designed specifically for WiFi-based pose estimation. The proposed method aims to identify consistent human poses that are highly resistant to environmental dynamics and WiFi signal noises. To achieve this goal, we introduce instance-wise consistency alignment loss that aligns domain shifts considering instance-wise pose distribution variance, and cross-environment channel enhancement module that enhances WiFi CSI feature representation by emphasizing channel-wise similarity between source and target domains. We conduct extensive experiments on both our self-collected pose estimation data set and a large public MM-Fi data set. The results demonstrate the effectiveness and robustness of AdaPose in eliminating domain shift, thereby facilitating the widespread application of WiFi-based pose estimation in smart cities. Yunjiao Zhou, Jianfei Yang 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Augmenting and Aligning Snippets for Few-Shot Video Domain AdaptationabstractFor video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video models. However, current VUDA methods rely on a vast amount of high-quality unlabeled target data, which may not be available in real-world cases. We thus consider a more realistic Few-Shot Video-based Domain Adaptation (FSVDA) scenario where we adapt video models with only a few target video samples. While a few methods have touched upon Few-Shot Domain Adaptation (FSDA) in images and in FSVDA, they rely primarily on spatial augmentation for target domain expansion with alignment performed statistically at the instance level. However, videos contain more knowledge in terms of rich temporal and semantic information, which should be fully considered while augmenting target domains and performing alignment in FSVDA. We propose a novel SSA2lign to address FSVDA at the snippet level, where the target domain is expanded through a simple snippet-level augmentation followed by the attentive alignment of snippets both semantically and statistically, where semantic alignment of snippets is conducted through multiple perspectives. Empirical results demonstrate state-of-the-art performance of SSA2lign across multiple cross-domain action recognition benchmarks. Code will be provided at: https://github.com/xuyu0010/SSA2lign. Yuecong Xu, Jianfei Yang 0001, Yunjiao Zhou, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001 |
ICCV | 3 |
| 2023 | MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensingabstract4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research. Jianfei Yang 0001, Yunjiao Zhou, Xinyan Chen 0002, Yuecong Xu, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001 |
NeurIPS | 3 |
| 2023 | MetaFi++: WiFi-Enabled Transformer-Based Human Pose Estimation for Metaverse Avatar SimulationabstractIn the metaverse, digital avatar plays an important role in representing human beings for various interaction with virtual objects and environments, which puts a high demand on effective pose estimation. Though camera-based solutions yield remarkable performance, they encounter privacy issues and degraded performance caused by varying illumination, especially in the smart home. In this article, we propose a WiFi-based Internet of Things-enabled human pose estimation scheme for metaverse avatar simulation, namely, MetaFi++. Specifically, WPFormer is designed with a shared convolutional module and a Transformer block to map the channel state information of WiFi signals to human pose landmarks, effectively exploring spatial information of human pose through self-attention. It is enforced to learn the annotations from the accurate computer vision model, thus achieving cross-modal supervision. Due to the ubiquitous existence of WiFi and robustness to various illumination conditions, WiFi-based human poses are suitable to instruct the movement of digital avatars in the metaverse, promoting avatar applications in smart homes. The experiments are conducted in the real world, and the results show that the MetaFi++ achieves very high performance with a PCK@50 of 97.30%. Our codes are available inhttps://github.com/pridy999/metafi_pose_estimation. Yunjiao Zhou, Shenghai Yuan 0001, Han Zou, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Internet Things J. | 1 |