Haocong Rao

dblp:269/4593 · DBLP profile ↗
← Back
19ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0002-9576-2379ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 8 first-author · 8 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 mmPred: Radar-based Human Motion Prediction in the Dark
abstract
Existing Human Motion Prediction (HMP) methods based on RGB(D) cameras are sensitive to lighting conditions and raise privacy concerns, limiting their real-world applications such as firefighting and elderly care. Motivated by the robustness and privacy-preserving nature of millimeter-wave (mmWave) radar, this work introduces radar as a novel sensing modality for HMP for the first time. Nevertheless, radar signals often suffer from specular reflections and multipath effects, resulting in noisy and temporally inconsistent measurements, such as body-part miss-detection. To address these radar-specific artifacts, we propose mmPred, the first diffusion-based framework tailored for radar-based HMP. mmPred introduces a dual-domain historical motion representation to guide the generation process, combining a Time-domain Pose Refinement (TPR) branch for fine-grained details and a Frequency-domain Dominant Motion (FDM) branch for capturing global motion trends and suppressing frame-level inconsistency. Furthermore, we design a Global Skeleton-relational Transformer (GST) as the diffusion backbone to model global inter-joint cooperation, enabling corrupted joints to dynamically aggregate information from others. Extensive experiments show that mmPred achieves state-of-the-art performance, outperforming existing methods by 8.6% on mmBody and 22% on mm-Fi.
Junqiao Fan, Haocong Rao, Jianfei Yang 0001, Lihua Xie 0001
AAAI2
2026 Text2CSG: Generating CAD models from natural language via constructive solid geometry
Luo Zhang 0002, Gaochao Song, Haocong Rao, Zhengyu Wen, Jiangbei Hu, Ying He 0001
Comput. Aided Des.3
2025 Motif Guided Graph Transformers with Combinatorial Skeleton Prototype Learning for Skeleton-Based Person Re-Identification
abstract
Person re-identification (re-ID) via 3D skeleton data is a challenging task with significant value in many scenarios. Existing skeleton-based methods typically assume virtual motion relations between all joints, and adopt average joint or sequence representations for learning. However, they rarely explore key body structure and motion such as gait to focus on more important body joints or limbs, while lacking the ability to fully mine valuable spatial-temporal sub-patterns of skeletons to enhance model learning. This paper presents a generic Motif guided graph transformer with Combinatorial skeleton prototype learning (MoCos) that exploits structure-specific and gait-related body relations as well as combinatorial features of skeleton graphs to learn effective skeleton representations for person re-ID. In particular, motivated by the locality within joints' structure and the body-component collaboration in gait, we first propose the motif guided graph transformer (MGT) that incorporates hierarchical structural motifs and gait collaborative motifs, which simultaneously focuses on multi-order local joint correlations and key cooperative body parts to enhance skeleton relation learning. Then, we devise the combinatorial skeleton prototype learning (CSP) that leverages random spatial-temporal combinations of joint nodes and skeleton graphs to generate diverse sub-skeleton and sub-tracklet representations, which are contrasted with the most representative features (prototypes) of each identity to learn class-related semantics and discriminative skeleton representations. Extensive experiments validate the superior performance of MoCos over existing state-of-the-art models. We further show its generality under RGB-estimated skeletons, different graph modeling, and unsupervised scenarios.
Haocong Rao, Chunyan Miao
AAAI1
2025 SMamDiff: Spatial Mamba for Stochastic Human Motion Prediction
abstract
With intelligent room-side sensing and service robots widely deployed, human motion prediction (HMP) is essential for safe, proactive assistance. However, many existing HMP methods either produce a single, deterministic forecast that ignores uncertainty or rely on probabilistic models that sacrifice kinematic plausibility. Diffusion models improve the accuracy-diversity trade-off but often depend on multi-stage pipelines that are costly for edge deployment. This work focuses on how to ensure spatial-temporal coherence within a single-stage diffusion model for HMP. We introduce SMamDiff, a Spatial Mam ba-based Diff usion model with two novel designs: (i) a residual-DCT motion encoding that subtracts the last observed pose before a temporal DCT, reducing the first DC component ($f=0$) dominance and highlighting informative higher-frequency cues so the model learns how joints move rather than where they are; and (ii) a stickman-drawing spatial-mamba module that processes joints in an ordered, joint-by-joint manner, making later joints condition on earlier ones to induce long-range, cross-joint dependencies. On Human3.6M and HumanEva, these coherence mechanisms deliver state-of-the-art results among single-stage probabilistic HMP methods while using less latency and memory than multi-stage diffusion baselines.
Junqiao Fan, Haocong Rao
CloudCom3
2025 A survey of artificial intelligence in gait-based neurodegenerative disease diagnosis
Haocong Rao, Minlin Zeng, Xuejiao Zhao, Chunyan Miao
Neurocomputing1
2025 Deformable Locality-Coordination Graph Motifs for 3D Skeleton Based Person Re-Identification
abstract
Existing 3D skeleton based person re-identification (re-ID) approaches typically model skeletons as graphs to capture body relations and motion. However, they often rely onfixedjoint’s connections such as adjacency for relation modeling, while lacking a flexible and specific focus on key body joints or parts ofdifferent levelsto capture various local relations (“locality”) and limb relations (“coordination”). In this letter, we propose Deformable Locality-Coordination graph Motifs (DL-CM) that can guide the body relation learning to particularly capture multi-orderlocalityandcoordinationof key gait-specific body parts to enhance person re-ID performance. Specifically, we first devise Deformable Locality Motifs (DLM), which are applicable to deformed skeleton graphs at different levels, to simultaneously focus on different-order neighbors’ relations for body structure and pattern learning. Then, we propose Deformable Coordination Motifs (DCM) to concurrently capture local and global coordination of different-level limbs in deformed graphs, so as to facilitate learning discriminative gait patterns for person re-ID. Extensive experiments on four public benchmarks demonstrate the effectiveness of DL-CM on state-of-the-art models and different-level graph representations to improve person re-ID performance.
Haocong Rao, Chunyan Miao
IEEE Signal Process. Lett.1
2024 Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-identification
Haocong Rao, Cyril Leung, Chunyan Miao
Int. J. Comput. Vis.1
2023 TranSG: Transformer-Based Skeleton Graph Prototype Contrastive Learning with Structure-Trajectory Prompted Reconstruction for Person Re-Identification
abstract
Person re-identification (re-ID) via 3D skeleton data is an emerging topic with prominent advantages. Existing methods usually design skeleton descriptors with raw body joints or perform skeleton sequence representation learning. However, they typically cannot concurrently model different body-component relations, and rarely explore useful semantics from fine-grained representations of body joints. In this paper, we propose a generic Transformer-based Skeleton Graph prototype contrastive learning (TranSG) approach with structure-trajectory prompted reconstruction to fully capture skeletal relations and valuable spatial-temporal semantics from skeleton graphs for person re-ID. Specifically, we first devise the Skeleton Graph Transformer (SGT) to simultaneously learn body and motion relations within skeleton graphs, so as to aggregate key correlative node features into graph representations. Then, we propose the Graph Prototype Contrastive learning (GPC) to mine the most typical graph features (graph prototypes) of each identity, and contrast the inherent similarity between graph representations and different prototypes from both skeleton and sequence levels to learn discriminative graph representations. Last, a graph Structure-Trajectory Prompted Reconstruction (STPR) mechanism is proposed to exploit the spatial and temporal contexts of graph nodes to prompt skeleton graph reconstruction, which facilitates capturing more valuable patterns and graph semantics for person re-ID. Empirical evaluations demonstrate that TranSG significantly outperforms existing state-of-the-art methods. We further show its generality under different graph modeling, RGB-estimated skeletons, and unsupervised scenarios. Our codes are available at https://github.com/Kali-Hac/TranSG.
Haocong Rao, Chunyan Miao
CVPR1
2023 Prototypical Contrast and Reverse Prediction: Unsupervised Skeleton Based Action Recognition
abstract
We focus on unsupervised representation learning for skeleton based action recognition. Existing unsupervised approaches usually learn action representations by motion prediction but they lack the ability to fully learn inherent semantic similarity. In this paper, we propose a novel framework named Prototypical Contrast and Reverse Prediction (PCRP) to address this challenge. Different from plain motion prediction, PCRP performs reverse motion prediction based on encoder-decoder structure to extract more discriminative temporal pattern, and derives action prototypes by clustering to explore the inherent action similarity within the action encoding. Specifically, we regard action prototypes as latent variables and formulate PCRP as an expectation-maximization (EM) task. PCRP iteratively runs (1) E-step as to determine the distribution of action prototypes by clustering action encoding from the encoder while estimating concentration around prototypes, and (2) M-step as optimizing the model by minimizing the proposed ProtoMAE loss, which helps simultaneously pull the action encoding closer to its assigned prototype by contrastive learning and perform reverse motion prediction task. Besides, the sorting can also serve as a temporal task similar as reverse prediction in the proposed framework. Extensive experiments on N-UCLA, NTU 60, and NTU 120 dataset present that PCRP outperforms main stream unsupervised methods and even achieves superior performance over many supervised methods. The codes are available at:https://github.com/LZUSIAT/PCRP.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
IEEE Trans. Multim.2
2022 SimMC: Simple Masked Contrastive Learning of Skeleton Representations for Unsupervised Person Re-Identification
abstract
Recent advances in skeleton-based person re-identification (re-ID) obtain impressive performance via either hand-crafted skeleton descriptors or skeleton representation learning with deep learning paradigms. However, they typically require skeletal pre-modeling and label information for training, which leads to limited applicability of these methods. In this paper, we focus on unsupervised skeleton-based person re-ID, and present a generic Simple Masked Contrastive learning (SimMC) framework to learn effective representations from unlabeled 3D skeletons for person re-ID. Specifically, to fully exploit skeleton features within each skeleton sequence, we first devise a masked prototype contrastive learning (MPC) scheme to cluster the most typical skeleton features (skeleton prototypes) from different subsequences randomly masked from raw sequences, and contrast the inherent similarity between skeleton features and different prototypes to learn discriminative skeleton representations without using any label. Then, considering that different subsequences within the same sequence usually enjoy strong correlations due to the nature of motion continuity, we propose the masked intra-sequence contrastive learning (MIC) to capture intra-sequence pattern consistency between subsequences, so as to encourage learning more effective skeleton representations for person re-ID. Extensive experiments validate that the proposed SimMC outperforms most state-of-the-art skeleton-based methods. We further show its scalability and efficiency in enhancing the performance of existing models. Our codes are available at https://github.com/Kali-Hac/SimMC.
Haocong Rao, Chunyan Miao
IJCAI1
2022 A Self-Supervised Gait Encoding Approach With Locality-Awareness for 3D Skeleton Based Person Re-Identification
abstract
Person re-identification (Re-ID) via gait features within 3D skeleton sequences is a newly-emerging topic with several advantages. Existing solutions either rely on hand-crafted descriptors or supervised gait representation learning. This paper proposes a self-supervised gait encoding approach that can leverage unlabeled skeleton data to learn gait representations for person Re-ID. Specifically, we first create self-supervision by learning to reconstruct unlabeled skeleton sequences reversely, which involves richer high-level semantics to obtain better gait representations. Other pretext tasks are also explored to further improve self-supervised learning. Second, inspired by the fact that motion's continuity endows adjacent skeletons in one skeleton sequence and temporally consecutive skeleton sequences with higher correlations (referred as locality in 3D skeleton data), we propose a locality-aware attention mechanism and a locality-aware contrastive learning scheme, which aim to preserve locality-awareness on intra-sequence level and inter-sequence level respectively during self-supervised learning. Last, with context vectors learned by our locality-aware attention mechanism and contrastive learning scheme, a novel feature named Constrastive Attention-based Gait Encodings (CAGEs) is designed to represent gait effectively. Empirical evaluations show that our approach significantly outperforms skeleton-based counterparts by 15-40 percent Rank-1 accuracy, and it even achieves superior performance to numerous multi-modal methods with extra RGB or depth information. Our codes are available at https://github.com/Kali-Hac/Locality-Awareness-SGE.
Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Yi Guo 0007, Jun Cheng 0002, Xinwang Liu 0002, Bin Hu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Revisiting -Reciprocal Distance Re-Ranking for Skeleton-Based Person Re-Identification
abstract
Person re-identification (re-ID) as a retrieval task often utilizes a re-ranking model to improve performance. Existing re-ranking methods are typically designed for conventional person re-ID with RGB images, while skeleton representation re-ranking for skeleton-based person re-ID still remains to be explored. To fill this gap, we revisit the$k$-reciprocal distance re-ranking model in this letter, and propose a generic re-ranking method that exploits the salient skeleton features to perform$k$-reciprocal distance encoding for skeleton-based person re-ID re-ranking. In particular, we devise the skeleton sequence pooling to aggregate the most salient features of skeletons within a sequence, and combine both original Euclidean distance and$k$-reciprocal distance to re-rank the skeleton sequence representations for person re-ID. Furthermore, we propose the context-based Rank-1 voting that jointly exploits the initial ranking list and re-ranking list to vote for the top candidate to enhance the Rank-1 matching. Extensive experiments on three public benchmarks demonstrate that our approach can effectively re-rank different state-of-the-art skeleton representations and significantly improve their person re-ID performance.
Haocong Rao, Chunyan Miao
IEEE Signal Process. Lett.1
2021 Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification
abstract
Skeleton-based person re-identification (Re-ID) is an emerging open topic providing great value for safety-critical applications. Existing methods typically extract hand-crafted features or model skeleton dynamics from the trajectory of body joints, while they rarely explore valuable relation information contained in body structure or motion. To fully explore body relations, we construct graphs to model human skeletons from different levels, and for the first time propose a Multi-level Graph encoding approach with Structural-Collaborative Relation learning (MG-SCR) to encode discriminative graph features for person Re-ID. Specifically, considering that structurally-connected body components are highly correlated in a skeleton, we first propose a multi-head structural relation layer to learn different relations of neighbor body-component nodes in graphs, which helps aggregate key correlative features for effective node representations. Second, inspired by the fact that body-component collaboration in walking usually carries recognizable patterns, we propose a cross-level collaborative relation layer to infer collaboration between different level components, so as to capture more discriminative skeleton graph features. Finally, to enhance graph dynamics encoding, we propose a novel self-supervised sparse sequential prediction task for model pre-training, which facilitates encoding high-level graph semantics for person Re-ID. MG-SCR outperforms state-of-the-art skeleton-based methods, and it achieves superior performance to many multi-modal methods that utilize extra RGB or depth features. Our codes are available at https://github.com/Kali-Hac/MG-SCR.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
IJCAI1
2021 SM-SGE: A Self-Supervised Multi-Scale Skeleton Graph Encoding Framework for Person Re-Identification
abstract
Person re-identification via 3D skeletons is an emerging topic with great potential in security-critical applications. Existing methods typically learn body and motion features from the body-joint trajectory, whereas they lack a systematic way to model body structure and underlying relations of body components beyond the scale of body joints. In this paper, we for the first time propose a Self-supervised Multi-scale Skeleton Graph Encoding (SM-SGE) framework that comprehensively models human body, component relations, and skeleton dynamics from unlabeled skeleton graphs of various scales to learn an effective skeleton representation for person Re-ID. Specifically, we first devise multi-scale skeleton graphs with coarse-to-fine human body partitions, which enables us to model body structure and skeleton dynamics at multiple levels. Second, to mine inherent correlations between body components in skeletal motion, we propose a multi-scale graph relation network to learn structural relations between adjacent body-component nodes and collaborative relations among nodes of different scales, so as to capture more discriminative skeleton graph features. Last, we propose a novel multi-scale skeleton reconstruction mechanism to enable our framework to encode skeleton dynamics and high-level semantics from unlabeled skeleton graphs, which encourages learning a discriminative skeleton representation for person Re-ID. Extensive experiments show that SM-SGE outperforms most state-of-the-art skeleton-based methods. We further demonstrate its effectiveness on 3D skeleton data estimated from large-scale RGB videos. Our codes are open at https://github.com/Kali-Hac/SM-SGE.
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
ACM Multimedia1
2021 Attention-Based Multilevel Co-Occurrence Graph Convolutional LSTM for 3-D Action Recognition
abstract
Action recognition is essential for many human-centered applications in the Internet of Things (IoT). Especially, in the Internet of Medical Things (IoMT), action recognition shows great importance in surgical assistance, patient monitoring, etc. Recently, 3-D skeleton sequence-based action recognition draws broad attention. It is a challenging task that needs effective modeling on intraframe skeleton representations and interframe temporal dynamics. Standard long short-term memory (LSTM)-based models are widely used for sequence modeling due to its long-term memory, yet they are unable to fully model the relationship between different body joints or persons to extract crucial co-occurrence features from different levels. To handle this shortcoming, we propose an attention-based multilevel co-occurrence graph convolutional LSTM (AMCGC-LSTM). By integrating graph convolutional networks (GCNs) into LSTM, the proposed model is capable of leveraging body structural information from skeletons and strengthening the multilevel co-occurrence (MC) feature learning. Specifically, we first design the spatial attention module for feature enhancement of key joints from skeleton inputs. Second, we design MC memory units coupled with GCN to automatically model the spatial relationship between joints, and simultaneously capture the co-occurrence features from different joints, persons, and frames. Finally, we construct aggregated features of MCs (AFMCs) from MC memory units to better represent the intraframe action context encoding, and leverage a concurrent LSTM (Co-LSTM) to further model their temporal dynamics for action recognition. Our model significantly outperforms mainstream methods on NTU RGB+D 60/120 data set, mutual action subset of NTU RGB+D 60/120 data set, and Northewestern-UCLA data set.
Haocong Rao, Hong Peng 0003, Xin Jiang 0004, Yi Guo 0007, Xiping Hu, Bin Hu 0001
IEEE Internet Things J.2
2021 Internet-of-Things-Enabled Data Fusion Method for Sleep Healthcare Applications
abstract
The Internet of Medical Things (IoMT) aims to exploit the Internet-of-Things (IoT) techniques to provide better medical treatment scheme for patients with smart, automatic, timely, and emotion-aware clinical services. One of the IoMT instances is applying IoT techniques to sleep-aware smartphones or wearable devices' applications to provide better sleep healthcare services. As we all know, sleep is vital to our daily health. What is more, studies have shown a strong relationship between sleep difficulties and various diseases such as COVID-19. Therefore, leveraging IoT techniques to develop a longer lifetime sleep healthcare IoMT system, with a tradeoff between data transferring/processing speed and battery energy efficiency, to provide longer time services for bad sleep condition persons, especially the COVID-19 patients or survivors, is a meaningful research topic. In this study, we propose an IoT-enabled sleep data fusion networks (SDFN) module with a star topology Bluetooth network to fuse data of sleep-aware applications. A machine learning model is built to detect sleep events through an audio signal. We design two data reprocessing mechanisms running on our IoT devices to alleviate the data jam problem and save the IoT devices' battery energy. The experiments manifest that the presented module and mechanisms can save the energy of the system and alleviate the data jam problem of the device.
Fan Yang 0082, Qilu Wu, Xiping Hu, Jiancong Ye, Haocong Rao, Bin Hu 0001
IEEE Internet Things J.6
2021 Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
Haocong Rao, Xiping Hu, Jun Cheng 0002, Bin Hu 0001
Inf. Sci.1
2020 Multi-Level Co-Occurrence Graph Convolutional LSTM for Skeleton-Based Action Recognition
abstract
Human action recognition plays an important role in e-health applications, such as surgical skill analysis, patient monitoring, and automatic nursing systems. Recently, skeleton-based action recognition gains massive attention. It is an essential yet challenging task that requires effectively modeling the intra-frame skeleton representation and inter-frame temporal dynamics. Traditional Long Short-Term Memory (LSTM) based methods mainly capture long-term action context information from global level, yet they cannot fully model the relationship between different joints or persons to mine crucial co-occurrence features from different levels. To overcome this drawback, we propose a general end-to-end Multi-level Co-occurrence Graph Convolutional LSTM (MCGC-LSTM). By incorporating graph convolutional networks (GCN) into LSTM, our model can not only better exploit body structural information from skeletons but also enhance the multi-level co-occurrence feature learning. Specifically, we first devise multi-level co-occurrence (MC) memory units coupled with GCN to automatically model the spatial relationship between joints, and simultaneously capture the co-occurrence features from different joints, persons, and frames. Then we construct aggregated features of multi-level co-occurrences (AFMC) from MC memory units to better represent the intra-frame action context encoding, and leverage a concurrent LSTM (Co-LSTM) to further model their temporal dynamics for action recognition. Experiments show that our proposed model significantly outperforms mainstream methods on NTU RGB+D 120 dataset and Northwestern-UCLA dataset.
Haocong Rao, Xiping Hu, Bin Hu 0001
HealthCom2
2020 Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification
abstract
Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supervised learning paradigms. Unlike previous methods, we for the first time propose a generic gait encoding approach that can utilize unlabeled skeleton data to learn gait representations in a self-supervised manner. Specifically, we first propose to introduce self-supervision by learning to reconstruct input skeleton sequences in reverse order, which facilitates learning richer high-level semantics and better gait representations. Second, inspired by the fact that motion's continuity endows temporally adjacent skeletons with higher correlations (“locality”), we propose a locality-aware attention mechanism that encourages learning larger attention weights for temporally adjacent skeletons when reconstructing current skeleton, so as to learn locality when encoding gait. Finally, we propose Attention-based Gait Encodings (AGEs), which are built using context vectors learned by locality-aware attention, as final gait representations. AGEs are directly utilized to realize effective person Re-ID. Our approach typically improves existing skeleton-based methods by 10-20% Rank-1 accuracy, and it achieves comparable or even superior performance to multi-modal methods with extra RGB or depth information.
Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Huang Da, Jun Cheng 0002, Bin Hu 0001
IJCAI1