Qiming Cao

dblp:282/1756 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-3329-3239ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Privacy-Preserving and Heterogeneity-aware Split Federated Learning via Probabilistic Masking
abstract
Split Federated Learning (SFL) has emerged as an efficient alternative to traditional Federated Learning (FL) by reducing client-side computation through model partitioning. However, exchanging of intermediate activations and model updates introduces significant privacy risks, especially from data reconstruction attacks that recover original inputs from intermediate representations. Existing defenses using noise injection often degrade model performance. To overcome these challenges, we present PM-SFL, a scalable and privacy-preserving SFL framework that incorporates Probabilistic Mask training to add structured randomness without relying on explicit noise. This mitigates data reconstruction risks while maintaining model utility. To address data heterogeneity, PM-SFL employs personalized mask learning that tailors submodel structures to each client's local data. For system heterogeneity, we introduce a layer-wise knowledge compensation mechanism, enabling clients with varying resources to participate effectively under adaptive model splitting. Theoretical analysis confirms its privacy protection, and experiments on image and wireless sensing tasks demonstrate that PM-SFL consistently improves accuracy, communication efficiency, and robustness to privacy attacks, with particularly strong performance under data and system heterogeneity.
Feijie Wu, Chenglin Miao, Tianchun Li, Qiming Cao, Jing Gao 0004, Lu Su 0001
KDD (1)6
2025 RAM-Hand: Robust Acoustic Multi-Hand Pose Reconstruction Using a Microphone Array
abstract
Using 3D hand poses as the input of user interfaces can enable many novel human-computer interaction applications. However, conventional solutions for precisely reconstructing the hand poses are either vision-based, which are compute-intensive and may cause privacy issues, or wearable devices-based, which are intrusive to users. In this paper, we propose RAM-Hand, a Robust Acoustic 3D Multi-Hand pose reconstruction system built on a microphone array. Our RAM-Hand system can support multiple hands and is designed to be highly adaptable to new scenarios even when training data is limited. Specifically, it should robustly accommodate variations in environment, subject, and hand positions. To achieve this, on one hand, we propose a customized signal processing pipeline to segment multiple hands' reflections and extract the features corresponding to each hand, then feed those features into a transformer-based neural network for precise pose reconstruction. On the other hand, to tackle the challenge that the training data is limited, we propose a series of data augmentation methods to generate virtual training data, and utilize contrastive learning to ensure our model behaves well on new subjects. We conduct extensive experiments on a real-world microphone array testbed to evaluate the performance of the proposed system. The results show that our RAM-Hand system can localize each hand joint with an average error of 10.71 mm, handle multiple hands, and generalize well to the above mentioned new scenarios.
Henglin Pu, Qiming Cao, Tianci Liu 0003, Zhengxin Jiang, Hongfei Xue, Lu Su 0001
SenSys3
2024 Towards Robust mmWave-based Human Activity Recognition using Large Simulated Dataset for Model Pretraining
abstract
Human activity recognition (HAR) is crucial for real-world applications such as healthcare, surveillance, and smart homes. Among sensing technologies, millimeter wave (mmWave) sensors stand out due to their contactless nature, high sensitivity, and ability to operate in low-light environments while preserving privacy. However, the scarcity of mmWave sensing data limits the generalizability of mmWave-based HAR systems. To address this, we propose mmAP, a data augmentation and pretraining framework that synthesizes a large mmWave dataset using human mesh data, followed by pretraining a robust and general mmWave heatmap encoder using a multi-modal masked autoencoder framework using the synthesized data. We enhance the model’s robustness with heatmap-specific data perturbations and perform task-specific fine-tuning on a small real-world dataset. The experiment results over the baseline demonstrate the effectiveness of the proposed mmAP framework.
Vinay Joshi, Shengkai Xu, Qiming Cao, Yi Zhu 0012, Pu Wang 0001, Hongfei Xue
IEEE Big Data3
2024 mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment
abstract
Millimeter-wave (mmWave) based human activity recognition (HAR) systems have demonstrated promising performance in various applications, leveraging the power of deep neural networks. However, these systems are suffering from the scarcity of available mmWave data for model training. To address this challenge, we explore the possibility of transferring knowledge from large AI models built on massive text and visual data to enhance the generalizability of mmWave-based HAR models. Towards this end, we introduce mmCLIP, a novel system that aligns mmWave signal space and text space to facilitate zero-shot recognition for unseen activities. To enable this alignment, we employ cross-modality signal synthesis to augment mmWave signal data using large human mesh datasets and design an activity attribute decomposition and recomposition approach to characterize the semantic interconnections among activities. We conducted extensive experiments to demonstrate the effectiveness of our proposed framework.
Qiming Cao, Hongfei Xue, Tianci Liu 0003, Haoyu Wang 0004, Xincheng Zhang, Lu Su 0001
SenSys1
2024 Towards Efficient Heterogeneous Multi-Modal Federated Learning with Hierarchical Knowledge Disentanglement
abstract
Multi-modal sensing systems are becoming increasingly common in real-world applications like human activity recognition (HAR). To enable knowledge sharing among individuals, Federated Learning (FL) offers a solution as a distributed machine learning paradigm that retains user data locally, thereby safeguarding privacy. However, existing heterogeneous multi-modal Federated Learning (MMFL) solutions have yet to fully utilize all the potential knowledge-sharing opportunities, as they fail to capture fundamental common knowledge that is independent of both modality and client. In this paper, we propose Federated Hierarchical Knowledge Disentanglement (FedHKD), a new sensing system for heterogeneous multi-modal federated learning. FedHKD introduces a multi-stage training paradigm based on hierarchical knowledge disentanglement at both the modality and client levels. This design enhances collaboration among modality-heterogeneous clients while maintaining low storage overhead and high adaptation flexibility to new sensing modalities. Our evaluation of two public real-world multi-modal HAR datasets and a self-collected dataset demonstrates that FedHKD outperforms state-of-the-art baselines by up to 4.85% in accuracy while saving up to 2.29× in storage. Additionally, when adapting to new sensing modalities, it reduces communication overhead by up to 4.62×.
Haoyu Wang 0004, Feijie Wu, Tianci Liu 0003, Qiming Cao, Lu Su 0001
SenSys5
2024 Towards Smartphone-based 3D Hand Pose Reconstruction Using Acoustic Signals
abstract
Accurately reconstructing 3D hand poses is a pivotal element for numerous Human-Computer Interaction applications. In this work, we propose SonicHand, the first smartphone-based 3D hand pose reconstruction system using purely inaudible acoustic signals. SonicHand incorporates signal processing techniques and a deep learning framework to address a series of challenges. First, it encodes the topological information of the hand skeleton as prior knowledge and utilizes a deep learning model to realistically and smoothly reconstruct the hand poses. Second, the system employs adversarial training to enhance the generalization ability of our system to be deployed in a new environment or for a new user. Third, we adopt a hand tracking method based on channel impulse response estimation. It enables our system to handle the scenario where the hand performs gestures while moving arbitrarily as a whole. We conduct extensive experiments on a smartphone testbed to demonstrate the effectiveness and robustness of our system from various dimensions. The experiments involve 10 subjects performing up to 12 different hand gestures in three distinctive environments. When the phone is held in one of the user’s hands, the proposed system can track joints with an average error of 18.64 mm.
Chenglin Miao, Qiming Cao, Haoyu Wang 0004, Ke Sun 0012, Hongfei Xue, Lu Su 0001
ACM Trans. Sens. Networks5
2023 Towards Generalized mmWave-based Human Pose Estimation through Signal Augmentation
abstract
The unprecedented advance of wireless human sensing is enabled by the proliferation of the deep learning techniques, which, however, rely heavily on the completeness and representativeness of the data patterns contained in the training set. Thus, deep learning based wireless human perception models usually fail when the human subject is conducting activities that are unseen during the model training. To address this problem, we propose a novel wireless signal augmentation framework, named mmGPE, for Generalized mmWave-based Pose Estimation. In mmGPE, we adopt a physical simulator to generate mmWave FMCW signals. However, due to the imperfect simulation of the physical world, there is a big gap between the signals generated by the physical simulator and the real-world signals collected by the mmWave radar. To tackle this challenge, we propose to integrate the physical signal simulation with deep learning techniques. Specifically, we develop a deep learning-based signal refiner in mmGPE that is capable of bridging the gap and generating realistic signal data. Through extensive evaluations on a COTS mmWave testbed, our mmGPE system demonstrates high accuracy in generating human meshes for unseen activities.
Hongfei Xue, Qiming Cao, Chenglin Miao, Yan Ju, Haochen Hu, Aidong Zhang 0001, Lu Su 0001
MobiCom2
2022 M4esh: mmWave-Based 3D Human Mesh Construction for Multiple Subjects
abstract
The recent proliferation of various wireless sensing systems and applications demonstrates the advantages of radio frequency (RF) signals over traditional camera-based solutions that are faced with various challenges, such as occlusions and poor lighting conditions. Towards the ultimate goal of imaging human body using RF signals, researchers have been exploring the possibility of constructing the human mesh, a structure capturing not only the pose but also the shape of the human body, from RF signals. In this paper, we introduce M4esh, a novel system that utilizes commercial millimeter wave (mmWave) radar for multi-subject 3D human mesh construction. Our M4esh system can detect and track the subjects on a 2D energy map by predicting the subject bounding boxes on the map, and tackle the subjects' mutual occlusion through utilizing the location, velocity and size information of the subjects' bounding boxes from the previous frames as a clue to estimate the bounding box in the current frame. Through extensive experiments on a real-world COTS millimeter-wave testbed, we show that our proposed M4esh system can accurately localize the subjects and generate their human meshes, which demonstrate the superior effectiveness of the proposed M4esh system.
Hongfei Xue, Qiming Cao, Yan Ju, Haochen Hu, Haoyu Wang 0004, Aidong Zhang 0001, Lu Su 0001
SenSys2