VLDB 2026 Research / reviewers in the wild / expert
Xuecai Hu
dblp:222/7974
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-0483-0418ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visually-Guided Policy Optimization for Multimodal ReasoningabstractZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu, Yong Wang, Yanlin Wang, Man Zhang, Xiangxiang Chu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zengbin Wang, Liang Lin 0004, Xuecai Hu, Xiangxiang Chu |
ACL (1) | 4 |
| 2026 | Spatio-Temporal Multi-Granularity for Skeleton-Based Depression Risk RecognitionabstractAs the prevalence of depression continues to rise, the timely and accurate recognition of its early signs is crucial for effective prevention and intervention. However, current clinical diagnostic methods are limited by the absence of objective biomarkers and inefficiencies in early recognition. Recent research has revealed a significant correlation between gait patterns and depression risk, suggesting that gait analysis could serve as a promising tool for early diagnosis. Depression-associated gait characteristics are defined by two key aspects: (1) they are dynamic, reflecting temporal abnormalities in movement, and (2) they manifest across both localized body regions and broader global movement patterns of the body. Based on these insights, we propose a novel Spatio-temporal Multi-granularity Network (STM-Net) for depression risk recognition. In the temporal domain, we present a Multi-grain Temporal Focus (MTF) module, designed to capture the rich dynamic temporal information embedded in the gait cycle of individuals with depression. In the spatial domain, we introduce a Multi-grain Spatial Focus (MSF) module, which effectively captures spatial features and their interactions in depression-related body regions through joint-level and part-level attention mechanisms. Extensive experimental results demonstrate that STM-Net achieves state-of-the-art performance on a large open-source dataset. Xuecai Hu, Li Yao 0002, Yongzhen Huang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Seeing from Magic Mirror: Contrastive Learning from Reconstruction for Pose-based Gait RecognitionabstractWhile recent advancements in supervised gait recognition have yielded promising results, these approaches rely heavily on annotated walking data, limiting their generalizability to complex environments. This paper presents a self-supervised gait recognition framework using human poses as input to address this challenge, focusing on high-quality pretrained data and self-supervised learning strategies. We first introduce StreamGait, a large-scale, unlabelled dataset that captures in-the-wild distributions of walking sequences. This dataset is curated from Internet livestreams across diverse geographic and environmental scenarios, reflecting variations in real-world camera angles, weather, and pedestrian behavior. Our framework, MirrorGait, conducts self-supervised learning by integration with 2D-to-3D pose reconstruction to synthesize multi-view perspectives for effective 3D-aware contrastive learning. With specific designs of temporal position embedding and gait partition head on a Transformer backbone, the encoder can readily adapt to the periodic and fine-grained nature of gait. Extensive experiments on three widely used gait datasets, Gait3D, GREW, and OUMVLP-Pose, demonstrate that our method, with minimal fine-tuning on the pretrained model, achieves state-of-the-art performance among pose-based gait recognition approaches. The dataset, code, and models are available at https://github.com/BNU-IVC/StreamGait. Shibei Meng, Saihui Hou, Xuecai Hu, Junzhou Huang, Yongzhen Huang |
ACM Multimedia | 4 |
| 2025 | Multimodal depression recognition based on gait and rating scale
Xuecai Hu, Yongzhen Huang |
Expert Syst. Appl. | 3 |
| 2025 | From FastPoseGait to GPGait++: Bridging the Past and Future for Pose-Based Gait RecognitionabstractRecent studies on pose-based gait recognition have underscored the potential of utilizing such fundamental data to achieve superior outcomes. Nonetheless, the development of current pose-based methods faces significant obstacles due to several critical issues: (1) Misaligned Settings, which results in a lack of thorough and unbiased comparative analysis. (2) Inferior Performance, which causes diminished focus on pose-based gait representations. (3) Limited Generalization, which hinders the effective application in real-world scenarios. Focused on tackling the aforementioned challenges, our study introduces a comprehensive benchmark and a versatile approach to bridge the past and future for pose-based gait recognition. First, we revisit previous pose-based methods and make great efforts to establish a unified framework, FastPoseGait, aiming at a fair and comprehensive comparison investigation with consistent experimental settings and a more stable training process. Then, within this framework, we propose GPGait++, a generalized pose-based gait recognition method featuring a human-oriented input and part-aware modeling, intended to enhance the generalization ability and discriminative power across diverse environments and camera viewpoints. Experiments on six public gait recognition datasets reveal that our unified framework significantly enhances the performance of previous approaches, and GPGait++ exhibits state-of-the-art cross-domain capabilities compared to existing pose-based methods, marking a significant advancement in the field of pose-based gait recognition. Shibei Meng, Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal DenoiserabstractRecently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods decompose the task into bone length and bone direction prediction based on the human anatomical skeleton to explicitly incorporate more human body prior constraints, the performance of these methods is significantly lower than that of the SOTA diffusion-based methods. This can be attributed to the tree structure of the human skeleton. Direct application of the disentangled method could amplify the accumulation of hierarchical errors, propagating through each hierarchy. Meanwhile, the hierarchical information has not been fully explored by the previous methods. To address these problems, a Disentangled Diffusion-based 3D human Pose Estimation method with Hierarchical Spatial and Temporal Denoiser is proposed, termed DDHPose. In our approach: (1) We disentangle the 3d pose and diffuse the bone length and bone direction during the forward process of the diffusion model to effectively model the human pose prior. A disentanglement loss is proposed to supervise diffusion model learning. (2) For the reverse process, we propose Hierarchical Spatial and Temporal Denoiser (HSTDenoiser) to improve the hierarchical modelling of each joint. Our HSTDenoiser comprises two components: the Hierarchical-Related Spatial Transformer (HRST) and the Hierarchical-Related Temporal Transformer (HRTT). HRST exploits joint spatial information and the influence of the parent joint on each joint for spatial modeling, while HRTT utilizes information from both the joint and its hierarchical adjacent joints to explore the hierarchical temporal correlations among joints. Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets show that our method outperforms the SOTA disentangled-based, non-disentangled based, and probabilistic approaches by 10.0%, 2.0%, and 1.3%, respectively. Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 0002, Yongzhen Huang |
AAAI | 2 |
| 2024 | POPDG: Popular 3D Dance Generation with PopDanceSetabstractGenerating dances that are both lifelike and well-aligned with music continues to be a challenging task in the cross-modal domain. This paper introduces PopDanceSet, the first dataset tailored to the preferences of young audiences, enabling the generation of aesthetically oriented dances. And it surpasses the${\it AIST}++{\it dataset}$in music genre di-versity and the intricacy and depth of dance movements. Moreover, the proposed POPDG model within the iD-DPMframework enhances dance diversity and, through the Space Augmentation Algorithm, strengthens spatial physi-cal connections between human body joints, ensuring that increased diversity does not compromise generation qual-ity. A streamlined Alignment Module is also designed to improve the temporal alignment between dance and mu-sic. Extensive experiments show that POPDG achieves SOTA results on two datasets. Furthermore, the paper also expands on current evaluation metrics. The dataset and code are available at https://github.com/Luke-Luol/POPDG. Zhenye Luo, Xuecai Hu, Yongzhen Huang, Li Yao 0002 |
CVPR | 3 |
| 2024 | Cut Out the Middleman: Revisiting Pose-Based Gait Recognition
Saihui Hou, Shibei Meng, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
ECCV (31) | 4 |
| 2024 | Depression risk recognition based on gait: A benchmark
Saihui Hou, Xuecai Hu, Yongzhen Huang |
Neurocomputing | 5 |
| 2024 | Adaptive Knowledge Transfer for Weak-Shot Gait RecognitionabstractMost works for cloth-changing gait recognition assume that the sequences of different clothes are accessible for each subject in the training set, which, however, is almost impossible for real applications. In practice, the collection of gait sequences is usually aided by person re-identification which is more likely to cluster the cloth-consistent sequences for a subject, and it is laborious to merge the cloth-changing clusters with the same identity. As a result, the training set is usually comprised of two subsets, i.e., a fully-annotated base set where the cloth-changing sequences are available for each subject, and a weakly-annotated wild set where the sequences of different clothes for a subject are assigned diverse labels. In this work, we formulate a problem named Weak-Shot Gait Recognition which seeks to learn discriminative features from the mixture of base set and wild set. Furthermore, we propose an effective method called Adaptive Knowledge Transfer to deal with the weak-shot issue. In particular, we define the knowledge as the ability to judge whether two sequences come from the same subject or not, and we take an adaptive way to mine the useful information from wild set. For the experimental study, we build three weak-shot benchmarks based on CASIA-B, Outdoor-Gait, and CASIA-E respectively. Extensive experiments show that our method can bring consistent improvement. For example, under the cloth-changing condition of on the weak-shot CASIA-B, our method exceeds a naïve baseline by 7.10%. Saihui Hou, Xuecai Hu, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Integral Pose Learning via Appearance Transfer for Gait RecognitionabstractGait recognition plays an important role in video surveillance and security by identifying humans based on their unique walking patterns. The existing gait recognition methods have achieved competitive accuracy with shape and motion patterns under limited-covariate conditions. However, when extreme appearance changes distort discriminative features, gait recognition yields unsatisfactory results under cross-covariate conditions. In this work, we first indicate that the integral pose in each silhouette maintains an appearance-unrelated discriminative identity. However, the monotonous appearance variables in a gait database cause gait models to have difficulty extracting integral poses. Therefore, we propose an Appearance-transferable Disentangling and Generative Network (GaitApp) to generate gait silhouettes with rich appearances and invariant poses. Specifically, GaitApp leverages multi-branch cooperation to disentangle pose features and appearance features, and transfers the appearance information from one subject to another. By simulating a person constantly changing appearances under limited-covariate conditions, downstream models enable to extract integral discriminative pose features. Extensive experiments demonstrate that our method allows representative gait models to stand at a new altitude, further promoting the exploration to cross-covariate gait recognition. All the code is available at https://github.com/Hpjhpjhs/GaitApp.git. Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu 0008, Xuecai Hu, Yongzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Dynamic Aggregated Network for Gait RecognitionabstractGait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exterior factors in real scenes, such as carrying conditions, wearing overcoats, and diverse viewing angles. Recently, various deep learning-based gait recognition methods have achieved promising results, but they tend to extract one of the salient features using fixed-weighted convolutional networks, do not well consider the relationship within gait features in key regions, and ignore the aggregation of complete motion patterns. In this paper, we propose a new perspective that actual gait features include global motion patterns in multiple key regions, and each global motion pattern is composed of a series of local motion patterns. To this end, we propose a Dynamic Aggregation Network (DANet) to learn more discriminative gait features. Specifically, we create a dynamic attention mechanism between the features of neighboring pixels that not only adaptively focuses on key regions but also generates more expressive local motion patterns. In addition, we develop a selfattention mechanism to select representative local motion patterns and further learn robust global motion patterns. Extensive experiments on three popular public gait datasets, i.e., CASIA-B, OUMVLP, and Gait3D, demonstrate that the proposed method can provide substantial improvements over the current state-of-the-art methods.1 Ying Fu 0001, Dezhi Zheng, Chunshui Cao, Xuecai Hu, Yongzhen Huang |
CVPR | 5 |
| 2023 | GPGait: Generalized Pose-based Gait RecognitionabstractRecent works on pose-based gait recognition have demonstrated the potential of using such simple information to achieve results comparable to silhouette-based methods. However, the generalization ability of pose-based methods on different datasets is undesirably inferior to that of silhouette-based ones, which has received little attention but hinders the application of these methods in real-world scenarios. To improve the generalization ability of pose-based methods across datasets, we propose a Generalized Pose-based Gait recognition (GPGait) framework. First, a Human-Oriented Transformation (HOT) and a series of Human-Oriented Descriptors (HOD) are proposed to obtain a unified pose representation with discriminative multi-features. Then, given the slight variations in the unified representation after HOT and HOD, it becomes crucial for the network to extract local-global relationships between the keypoints. To this end, a Part-Aware Graph Convolutional Network (PAGCN) is proposed to enable efficient graph partition and local-global spatial feature extraction. Experiments on four public gait recognition datasets, CASIA-B, OUMVLP-Pose, Gait3D and GREW, show that our model demonstrates better and more stable cross-domain capabilities compared to existing skeleton-based methods, achieving comparable recognition results to silhouette-based ones. Code is available at https://github.com/BNU-IVC/FastPoseGait. Shibei Meng, Saihui Hou, Xuecai Hu, Yongzhen Huang |
ICCV | 4 |
| 2021 | Meta-USR: A Unified Super-Resolution Network for Multiple Degradation ParametersabstractRecent research on single image super-resolution (SISR) has achieved great success due to the development of deep convolutional neural networks. However, most existing SISR methods merely focus on super-resolution of a single fixed integer scale factor. This simplified assumption does not meet the complex conditions for real-world images which often suffer from various blur kernels or various levels of noise. More importantly, previous methods lack the ability to cope with arbitrary degradation parameters (scale factors, blur kernels, and noise levels) with a single model. A few methods can handle multiple degradation factors, e.g., noninteger scale factors, blurring, and noise, simultaneously within a single SISR model. In this work, we propose a simple yet powerful method termed meta-USR which is the first unified super-resolution network for arbitrary degradation parameters with meta-learning. In Meta-USR, a meta-restoration module (MRM) is proposed to enhance the traditional upscale module with the capability to adaptively predict the weights of the convolution filters for various combinations of degradation parameters. Thus, the MRM can not only upscale the feature maps with arbitrary scale factors but also restore the SR image with different blur kernels and noise levels. Moreover, the lightweight MRM can be placed at the end of the network, which makes it very efficient for iteratively/repeatedly searching the various degradation factors. We evaluate the proposed method through extensive experiments on several widely used benchmark data sets on SISR. The qualitative and quantitative experimental results show the superiority of our Meta-USR. Xuecai Hu, Zhang Zhang 0001, Caifeng Shan, Zilei Wang, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Meta-SR: A Magnification-Arbitrary Network for Super-ResolutionabstractRecent research on super-resolution has achieved great success due to the development of deep convolutional neural networks (DCNNs). However, super-resolution of arbitrary scale factor has been ignored for a long time. Most previous researchers regard super-resolution of differentscale factors as independent tasks. They train a specific model for each scale factor which is inefficient in computing, and prior work only take the super-resolution of several integer scale factors into consideration. In this work,we propose a novel method called Meta-SR to firstly solve super-resolution of arbitrary scale factor (including non-integer scale factors) with a single model. In our Meta-SR,the Meta-Upscale Module is proposed to replace the traditional upscale module. For arbitrary scale factor, the Meta-Upscale Module dynamically predicts the weights of the up-scale filters by taking the scale factor as input and use these weights to generate the HR image of arbitrary size. For any low-resolution image, our Meta-SR can continuously zoomin it with arbitrary scale factor by only using a single model.We evaluated the proposed method through extensive experiments on widely used benchmark datasets on single image super-resolution. The experimental results show the superiority of our Meta-Upscale. Xuecai Hu, Haoyuan Mu, Xiangyu Zhang 0005, Zilei Wang, Tieniu Tan, Jian Sun 0001 |
CVPR | 1 |
| 2019 | Focal Boundary Guided Salient Object DetectionabstractThe performance of salient object segmentation has been significantly advanced by using deep convolutional networks. However, these networks often produce blob-like saliency maps without accurate object boundaries. This is caused by the limited spatial resolution of their feature maps after multiple pooling operations, and might hinder downstream applications that require precise object shapes. To address this issue, we propose a novel deep model-Focal Boundary Guided (Focal- BG) network. Our model is designed to jointly learn to segment salient object masks and detect salient object boundaries. Our key idea is that additional knowledge about object boundaries can help to precisely identify the shape of the object. Moreover, our model incorporates a refinement pathway to refine the mask prediction, and makes use of the focal loss to facilitate the learning of the hard boundary pixels. To evaluate our model, we conduct extensive experiments. Our Focal-BG network consistently outperforms state-of-the-art methods on five major benchmarks. We provide a detailed analysis of these results and demonstrate that our joint modeling of salient object boundary and mask helps to better capture shape details, especially in the vicinity of object boundaries. Yupei Wang, Xin Zhao 0012, Xuecai Hu, Yin Li 0003, Kaiqi Huang |
IEEE Trans. Image Process. | 3 |
| 2018 | Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel FusionabstractShadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin. Yupei Wang, Xin Zhao 0012, Yin Li 0003, Xuecai Hu, Kaiqi Huang |
IJCAI | 4 |