VLDB 2026 Research / reviewers in the wild / expert
Xuehao Gao
dblp:319/1436
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-3168-5770ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge distillation-based distributed dynamic 3D Gaussian splatting for large scale scene reconstruction
Sicheng Fei, Xuehao Gao, Jinwen Hu, Xiaolei Hou, Dingwen Zhang |
Expert Syst. Appl. | 2 |
| 2026 | Hierarchical Diffusion for Sparse-to-Full Human Motion ReconstructionabstractHuman motion generation from sparse observations is an ill-posed problem in AR/VR, where head-mounted devices often capture only head and wrist trajectories. Prior methods usually reconstruct full-body motion in a single stage, forcing inference over a vast solution space and producing inaccurate lower-body motion, weak temporal coherence, and implausible sequences that degrade avatar embodiment. We presentMAGE, aMulti-stageAvatarGEnerator based on hierarchical diffusion. Instead of predicting 22-joint motion at once, MAGE progressively refines motion from a coarse 6-part representation to full joints. Each stage injects stage-specific motion priors and uses intermediate predictions to constrain subsequent refinement, reducing ambiguity and stabilizing dynamics. Experiments on large-scale motion datasets show that MAGE improves reconstruction accuracy, temporal smoothness, and perceptual realism over state-of-the-art baselines, enabling more reliable full-body animation from minimal AR/VR sensing while preserving real-time interaction. Fangyu Du, Yang Yang 0066, Xuehao Gao, Hongye Hou |
IEEE Signal Process. Lett. | 3 |
| 2025 | MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross AttentionabstractAchieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnic generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across facial similarity. Qi Xie 0009, Yongjia Ma, Donglin Di, Xuehao Gao, Xun Yang 0001 |
MMAsia | 4 |
| 2025 | STRIDER: Navigation via Instruction-Aligned Structural Decision Space OptimizationabstractThe Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents’ actions align with both spatial structure and task intent over long-horizon execution. Existing methods often fail to achieve robust navigation due to a lack of structured decision-making and insufficient integration of feedback from previous actions. To address these challenges, we propose STRIDER (Instruction-Aligned Structural Decision Space Optimization), a novel framework that systematically optimizes the agent’s decision space by integrating spatial layout priors and dynamic task feedback. Our approach introduces two key innovations: 1) a Structured Waypoint Generator that constrains the action space through spatial structure, and 2) a Task-Alignment Regulator that adjusts behavior based on task progress, ensuring semantic alignment throughout navigation. Extensive experiments on the R2R-CE and RxR-CE benchmarks demonstrate that STRIDER significantly outperforms strong SOTA across key metrics; in particular, it improves Success Rate (SR) from 29\% to 35\%, a relative gain of 20.7\%. Such results highlight the importance of spatially constrained decision-making and feedback-guided execution in improving navigation fidelity for zero-shot VLN-CE. Diqi He, Xuehao Gao, Hao Li 0075, Junwei Han 0001, Dingwen Zhang |
NeurIPS | 2 |
| 2025 | Jointly Understand Your Command and Intention: Reciprocal Co-Evolution Between Scene-Aware 3D Human Motion Synthesis and AnalysisabstractAs two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these two paired processes into a Co-Evolving Synthesis-Analysis (CESA) pipeline and mutually benefit their learning. Specifically, scene aware text-to-human synthesis generates diverse indoor motion samples from the same textual description to enrich human scene interaction intra-class diversity, thus significantly benefiting training a robust human motion analysis system. Reciprocally, human motion analysis would enforce semantic scrutiny on each synthesized motion sample to ensure its semantic consistency with the given textual description, thus improving realistic motion synthesis. Considering that real-world indoor human motions are goal-oriented and path-guided, we propose a cascaded generation strategy that factorizes text-driven scene-specific human motion generation into three stages: goal inferring, path planning, and pose synthesizing. Coupling CESA with this powerful cascaded motion synthesis model, we jointly improve realistic human motion synthesis and robust human motion analysis in 3D scenes. Xuehao Gao, Yang Yang 0066, Shaoyi Du, Guo-Jun Qi, Junwei Han 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion ModelabstractText-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these methods conduct diffusion processes either on the raw data distribution or the low-dimensional latent space, which typically suffer from the problem of modality inconsistency or detail-scarce. To tackle this problem, we propose a novel Basic-to-Advanced Hierarchical Diffusion Model, named B2A-HDM, to collaboratively exploit low-dimensional and high-dimensional diffusion models for high quality detailed motion synthesis. Specifically, the basic diffusion model in low-dimensional latent space provides the intermediate denoising result that to be consistent with the textual description, while the advanced diffusion model in high-dimensional latent space focuses on the following detail-enhancing denoising process. Besides, we introduce a multi-denoiser framework for the advanced diffusion model to ease the learning of high-dimensional model and fully explore the generative potential of the diffusion model. Quantitative and qualitative experiment results on two text-to-motion benchmarks (HumanML3D and KIT-ML) demonstrate that B2A-HDM can outperform existing state-of-the-art methods in terms of fidelity, modality consistency, and diversity. Zhenyu Xie, Yang Wu 0001, Xuehao Gao, Zhongqian Sun, Wei Yang 0019, Xiaodan Liang |
AAAI | 3 |
| 2024 | Full-Dimensional Optimizable Network: A Channel, Frame and Joint-Specific Network Modeling for Skeleton-Based Action RecognitionabstractRecent human action recognition systems widely adopt graph convolution networks to extract spatial-temporal movement patterns. In graph convolution layers, inter-joint and inter-frame dependencies dominate spatial and temporal feature aggregation and thus are pivotal to representation learning. To enrich learned motion patterns, a powerful feature extractor should introduce its information propagation flexibility into three dimensions: (1) inferring different inter-joint correlations at different frames; (2) inferring different inter-frame correlations at different joints; (3) inferring different inter-joint and interframe correlations at different channels. In this paper, we take a closer look at effective feature aggregation in a skeleton sequence and propose a novel full-dimensional optimizable network with Channel, Frame and Joint-specific Network (CFJ-s Net) modeling for improving action recognition. By promoting dynamic information flows within different channels, frames, and joints, CFJ-s Net significantly extracts richer body posture features and trajectory features from a skeleton sequence. As verified on three large-scale datasets, NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA, CFJ-s Net achieves substantial improvements over state-of-the-art methods. Yang Yang 0066, Xuehao Gao, Shaoyi Du |
IJCNN | 3 |
| 2024 | Lightweight Graph Convolutional Network For Efficient Skeleton Based Action RecognitionabstractGraph convolutional network (GCN) has been widely used by skeleton based action recognition algorithms and achieves remarkable performance. However, recent GCN based State-Of-The-Art (SOTA) models for skeleton based action recognition tend to become increasingly sophisticated and over-parameterized. The low efficiency in model training and inference poses a challenge for their practical implementation in real-world scenarios. To address this issue, we construct a GCN based lightweight model for skeleton based action recognition, termed LightGCN. In this work we introduce an efficient convolutional neural network (CNN) structure to our temporal convolutional (TC) layer to extract temporal dynamics, effectively reducing model complexity. Furthermore, we propose a novel attention module that first extends the multi-spectral channel attention mechanism to the field of skeleton based action recognition, which preserves not only the lowest frequency information, but also useful information encoded by other frequency components, reducing the information loss during the channel compression. In order to further reduce the model complexity, we design a new compound scaling strategy to expand the model’s width and depth to different extent. This strategy enables the model to achieve an excellent balance between complexity and accuracy. On the two large-scale datasets, i.e., NTU RGB+D 60 and 120, our proposed LightGCN achieves 92.5% accuracy on the cross-subject benchmark of NTU 60 dataset, outperforming previous SOTA lightweight models and most heavyweight models, while needing 24.54% fewer parameters and 24.76% fewer flops than EfficientGCN-B4, which is the SOTA lightweight model. Yang Yang 0066, Xuehao Gao |
IJCNN | 3 |
| 2024 | Using Rotation-Invariant Point and Line Features for Image MatchingabstractIn recent years, convolutional neural networks (CNNs) have outperformed traditional approaches in image matching tasks. However, suffering from poor robustness against object rotations, conventional CNNs tend to extract angle-specific feature representations from given images. To this end, group CNNs improve conventional CNNs with symmetric group theory and thus benefit their rotation equivariance for powerful feature learning. Nonetheless, how to enrich extracted features with better discriminability is an under-explored challenge for group CNNs. In this paper, we propose a powerful rotation-invariant image matching method that combines point and line features to jointly improve rotation equivariance and discriminability. Specifically, we first characterize richer features from images by detecting their keypoints and lines. Then, we employ a group convolutional backbone to extract rotation-invariant descriptors from detected keypoints and lines. Finally, we develop inter-image and intra-image attention strategies to integrate point-level and line-level features from two images, significantly facilitating the two-image matching task. Extensive experiments verify that our method achieves state-of-the-art matching accuracy among existing methods on varying rotation image datasets and also shows competitive results when transferred to real-world image matching. Wenpeng Zheng, Yang Yang 0066, Xuehao Gao |
IJCNN | 3 |
| 2024 | Dig into Detailed Structures: Key Context Encoding and Semantic-based Decoding for Point Cloud Completion
Hongye Hou, Xuehao Gao, Yang Yang 0066 |
ACM Multimedia | 2 |
| 2024 | Path-Guided Motion Prediction with Multi-view Scene Perception
Zongyun Li, Yang Yang 0069, Xuehao Gao |
PRCV (7) | 3 |
| 2024 | Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion PredictionabstractInferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to human motion prediction, most approaches focus on pose-driven prediction and inferring human motion in isolation from the contextual environment, thus leaving the body location movement in the scene behind. However, real-world human movements are goal-directed and highly influenced by the spatial layout of their surrounding scenes. In this paper, instead of planning future human motion in a "dark" room, we propose a Multi-Condition Latent Diffusion network (MCLD) that reformulates the human motion prediction task as a multi-condition joint inference problem based on the given historical 3D body motion and the current 3D scene contexts. Specifically, instead of directly modeling joint distribution over the raw motion sequences, MCLD performs a conditional diffusion process within the latent embedding space, characterizing the cross-modal mapping from the past body movement and current scene context condition embeddings to the future human motion embedding. Extensive experiments on large-scale human motion prediction datasets demonstrate that our MCLD achieves significant improvements over the state-of-the-art methods on both realistic and diverse predictions. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du, Guo-Jun Qi |
IEEE Trans. Image Process. | 1 |
| 2024 | Learning Heterogeneous Spatial-Temporal Context for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have been widely used and achieved fruitful progress in the skeleton-based action recognition task. In GCNs, node interaction modeling dominates the context aggregation and, therefore, is crucial for a graph-based convolution kernel to extract representative features. In this article, we introduce a closer look at a powerful graph convolution formulation to capture rich movement patterns from these skeleton-based graphs. Specifically, we propose a novel heterogeneous graph convolution (HetGCN) that can be considered as the middle ground between the extremes of (2 + 1)-D and 3-D graph convolution. The core observation of HetGCN is that multiple information flows are jointly intertwined in a 3-D convolution kernel, including spatial, temporal, and spatial-temporal cues. Since spatial and temporal information flows characterize different cues for action recognition, HetGCN first dynamically analyzes pairwise interactions between each node and its cross-space-time neighbors and then encourages heterogeneous context aggregation among them. Considering the HetGCN as a generic convolution formulation, we further develop it into two specific instantiations (i.e., intra-scale and inter-scale HetGCN) that significantly facilitate cross-space-time and cross-scale learning on skeleton graphs. By integrating these modules, we propose a strong human action recognition system that outperforms state-of-the-art methods with the accuracy of 93.1% on NTU-60 cross-subject (X-Sub) benchmark, 88.9% on NTU-120 X-Sub benchmark, and 38.4% on kinetics skeleton. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion GenerationabstractIn this article, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up generation objectives by grouping body joints of detailed skeletons in close semantic proximity together and then replacing each of such joint group with a single body-part node. Such an operation recursively abstracts a human pose to coarser and coarser skeletons at multiple granularity levels. With gradually increasing the abstraction level, human motion becomes more and more concise and stable, significantly benefiting the cross-modal motion synthesis task. The whole text-driven human motion synthesis problem is then divided into multiple abstraction levels and solved with a multi-stage generation framework with a cascaded latent diffusion model: an initial generator first generates the coarsest human motion guess from a given text description; then, a series of successive generators gradually enrich the motion details based on the textual description and the previous synthesized results. Notably, we further integrate GUESS with the proposed dynamic multi-condition fusion mechanism to dynamically balance the cooperative effects of the given textual condition and synthesized coarse motion prompt in different generation stages. Extensive experiments on large-scale datasets verify that GUESS outperforms existing state-of-the-art methods by large margins in terms of accuracy, realisticness, and diversity. Xuehao Gao, Yang Yang 0066, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, Yang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion PredictionabstractEncouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into the frequency space. In this paper, we introduce two closer looks at effective frequency representation learning for robust motion prediction and summarize them as: decompose more and aggregate better. Motivated by these two insights, we develop two powerful units that factorize the frequency representation learning task with a novel decomposition-aggregation two-stage strategy: (1) frequency decomposition unit unweaves multi-view frequency representations from an input body motion by embedding its frequency features into multiple spaces; (2) feature aggregation unit deploys a series of intra-space and inter-space feature aggregation layers to collect comprehensive frequency representations from these spaces for robust human motion prediction. As evaluated on large-scale datasets, we develop a strong baseline model for the human motion prediction task that outperforms state-of-the-art methods by large margins: 8%∼12% on Human3.6M, 3%∼7% on CMU MoCap, and 7%∼10% on 3DPW. Xuehao Gao, Shaoyi Du, Yang Wu 0001, Yang Yang 0066 |
CVPR | 1 |
| 2023 | Glimpse and focus: Global and local-scale graph convolution network for skeleton-based action recognition
Xuehao Gao, Shaoyi Du, Yang Yang 0066 |
Neural Networks | 1 |
| 2023 | Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action RecognitionabstractIn this paper, we propose a simple yet effective self-supervised method called spatio-temporal contrastive learning (ST-CL) for 3D skeleton-based action recognition. ST-CL acquires action-specific features by regarding the spatio-temporal continuity of motion tendency as the supervisory signal. To yield effective representations, ST-CL first designs some novel contrastive proxy tasks by providing different spatio-temporal observation scenes for the same 3D action and pulling them together in the embedding space. Second, three key components are devised in the action encoding to efficiently extract representations in contrastive tasks: (1) Information Representation introduces the awareness of joint type when analyzing motion dynamics. (2) Non-local GCN learns a data-driven graph topology structure and promotes a spatial message passing among long-range joints in each frame. (3) Multi-Scale TCN makes larger receptive fields for capturing richer longe-range temporal dynamics amomg adjacent frames. In ST-CL, these effective proxy tasks yield useful representations and efficient action encoding further enhances the representation capacity. As validated on four large-scale datasets, ST-CL is a strong baseline with high performance and efficiency for the contrastive learning study of the skeleton data. Compared to previous self-supervised methods, the proposed ST-CL achieves significant improvement consistently with a smaller model size and better training efficiency. Xuehao Gao, Yang Yang 0066, Maosen Li, Jin-Gang Yu, Shaoyi Du |
IEEE Trans. Multim. | 1 |
| 2022 | Motion Guided Attention Learning for Self-Supervised 3D Human Action Recognitionabstract3D human action recognition has received increasing attention due to its potential application in video surveillance equipment. To guarantee satisfactory performance, previous studies are mainly based on supervised methods, which have to add a large amount of manual annotation costs. In addition, general deep networks for video sequences suffer from heavy computational costs, thus cannot satisfy the basic requirement of embedded systems. In this paper, a novel Motion Guided Attention Learning (MG-AL) framework is proposed, which formulates the action representation learning as a self-supervised motion attention prediction problem. Specifically, MG-AL is a lightweight network. A set of simple motion priors (e.g., intra-joint variance, inter-frame deviation, intra-joint variance, and cross-joint covariance), which minimizes additional parameters and computational overhead, is regarded as a supervisory signal to guide the attention generation. The encoder is trained via predicting multiple self-attention tasks to capture action-specific feature representations. Extensive evaluations are performed on three challenging benchmark datasets (NTU-RGB+D 60, NTU-RGB+D 120 and NW-UCLA). The proposed method achieves superior performance compared to state-of-the-art methods, while having a very low computational cost. Yang Yang 0066, Xuehao Gao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |