Xinshun Wang

dblp:344/5333 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2024
0009-0001-8035-3687ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2024 GCNext: Towards the Unity of Graph Convolutions for Human Motion Prediction
abstract
The past few years has witnessed the dominance of Graph Convolutional Networks (GCNs) over human motion prediction. Various styles of graph convolutions have been proposed, with each one meticulously designed and incorporated into a carefully-crafted network architecture. This paper breaks the limits of existing knowledge by proposing Universal Graph Convolution (UniGC), a novel graph convolution concept that re-conceptualizes different graph convolutions as its special cases. Leveraging UniGC on network-level, we propose GCNext, a novel GCN-building paradigm that dynamically determines the best-fitting graph convolutions both sample-wise and layer-wise. GCNext offers multiple use cases, including training a new GCN from scratch or refining a preexisting GCN. Experiments on Human3.6M, AMASS, and 3DPW datasets show that, by incorporating unique module-to-network designs, GCNext yields up to 9x lower computational cost than existing GCN methods, on top of achieving state-of-the-art performance. Our code is available at https://github.com/BradleyWang0416/GCNext.
Xinshun Wang, Qiongjie Cui, Chen Chen 0001
AAAI1
2024 Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
abstract
In-context learning provides a new perspective for multi-task modeling for vision and NLP. Under this setting, the model can perceive tasks from prompts and accomplish them without any extra task-specific head predictions or model fine-tuning. However, skeleton sequence modeling via in-context learning remains unexplored. Directly applying existing in-context models from other areas onto skeleton sequences fails due to the similarity between inter-frame and cross-task poses, which makes it exceptionally hard to perceive the task correctly from a subtle context. To address this challenge, we propose Skeleton-in-Context (SiC), an effective framework for in-context skeleton sequence modeling. Our SiC is able to handle multiple skeleton-based tasks simultaneously after a single training process and accomplish each task from context according to the given prompt. It can further generalize to new, unseen tasks according to customized prompts. To facilitate context perception, we additionally propose a task-unified prompt, which adaptively learns tasks of different natures, such as partial joint-level generation, sequence-level prediction, or 2D-to-3D motion prediction. We conduct extensive experiments to evaluate the effectiveness of our SiC on multiple tasks, including motion prediction, pose estimation, joint completion, and future pose estimation. We also evaluate its generalization capability on unseen tasks such as motion-in-between. These experiments show that our model achieves state-of-the-art multi-task performance and even outperforms single-task methods on certain tasks.
Xinshun Wang, Zhongbin Fang, Xia Li 0005, Xiangtai Li, Chen Chen 0001, Mengyuan Liu 0001
CVPR1
2024 CHAMP: A Large-Scale Dataset for Skeleton-Based Composite HumAn Motion Prediction
abstract
Skeleton-based human motion prediction task aims to forecast future skeleton frames conditioned by observed skeleton sequence. Different from previous methods that focus on human motion prediction for atomic actions, we observe that people are witnessed to perform composite actions which consist of atomic actions that simultaneously happen. Considering the large number of action types, it is more laborious to collect composite actions than atomic actions. This paper presents a practical composite human motion prediction task, whose training data just contains atomic actions meanwhile the test data contains both atomic actions and composite actions. To evaluate this task, we collect a large-scale Composite HumAn Motion Prediction (CHAMP) dataset, whose training data has 16 types of atomic actions and test data has 50 types of composite actions. Despite the success of previous human motion prediction methods using Graph Convolutional Networks (GCN), these methods achieve inferior performances on our CHAMP dataset due to the huge domain gap between the training and test data. To solve this problem, we present a composite human motion prediction framework containing three modules. First, a Composite Motion Synthesis (CMS) module is designed to generate synthesized composite human actions from atomic actions. Second, a Composite GCN module is presented to predict human motion by modeling different human body parts. Third, a human body partition policy network is used to choose the best partition strategy for both the CMS and Composite GCN modules. Extensive experiments on the CHAMP dataset verify the effectiveness of our framework which obviously outperforms GCN-based methods.
Mengyuan Liu 0001, Xinshun Wang, Can Wang 0006
IEEE Trans. Circuits Syst. Video Technol.3
2024 Dynamic Dense Graph Convolutional Network for Skeleton-Based Human Motion Prediction
abstract
Graph Convolutional Networks (GCN) which typically follows a neural message passing framework to model dependencies among skeletal joints has achieved high success in skeleton-based human motion prediction task. Nevertheless, how to construct a graph from a skeleton sequence and how to perform message passing on the graph are still open problems, which severely affect the performance of GCN. To solve both problems, this paper presents a Dynamic Dense Graph Convolutional Network (DD-GCN), which constructs a dense graph and implements an integrated dynamic message passing. More specifically, we construct a dense graph with 4D adjacency modeling as a comprehensive representation of motion sequence at different levels of abstraction. Based on the dense graph, we propose a dynamic message passing framework that learns dynamically from data to generate distinctive messages reflecting sample-specific relevance among nodes in the graph. Extensive experiments on benchmark Human 3.6M and CMU Mocap datasets verify the effectiveness of our DD-GCN which obviously outperforms state-of-the-art GCN-based methods, especially when using long-term and our proposed extremely long-term protocol.
Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001
IEEE Trans. Image Process.1
2024 Temporal Decoupling Graph Convolutional Network for Skeleton-Based Gesture Recognition
abstract
Skeleton-based gesture recognition methods have achieved high success using Graph Convolutional Network (GCN), which commonly uses an adjacency matrix to model the spatial topology of skeletons. However, previous methods use the same adjacency matrix for skeletons from different frames, which limits the flexibility of GCN to model temporal information. To solve this problem, we propose a Temporal Decoupling Graph Convolutional Network (TD-GCN), which applies different adjacency matrices for skeletons from different frames. The main steps of each convolution layer in our proposed TD-GCN are as follows. To extract deep spatiotemporal information from skeleton joints, we first extract high-level spatiotemporal features from skeleton data. Then, channel-dependent and temporal-dependent adjacency matrices corresponding to different channels and frames are calculated to capture the spatiotemporal dependencies between skeleton joints. Finally, to fuse topology information from neighbor skeleton joints, spatiotemporal features of skeleton joints are fused based on channel-dependent and temporal-dependent adjacency matrices. To the best of our knowledge, we are the first to use temporal-dependent adjacency matrices for temporal-sensitive topology learning from skeleton joints. The proposed TD-GCN effectively improves the modeling ability of GCN and achieves state-of-the-art results on gesture datasets including SHREC'17 Track and DHG-14/28.
Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001
IEEE Trans. Multim.2
2023 Learning Snippet-to-Motion Progression for Skeleton-based Human Motion Prediction
abstract
Existing Graph Convolutional Networks to achieve human motion prediction largely adopt a one-step scheme, which output the prediction straight from history input, failing to exploit human motion patterns. We observe that human motions have transitional patterns and can be split into snippets representative of each transition. Each snippet can be reconstructed from its starting and ending poses referred to as the transitional poses. We propose a snippet-to-motion multi-stage framework that breaks motion prediction into sub-tasks easier to accomplish. Each sub-task integrates three modules: transitional pose prediction, snippet reconstruction, and snippet-to-motion prediction. Specifically, we propose to first predict only the transitional poses. Then we use them to reconstruct the corresponding snippets, obtaining a close approximation to the true motion sequence. Finally we refine them to produce the final prediction output. To implement the network, we propose a novel unified graph modeling, which allows for direct and effective feature propagation compared to existing approaches which rely on separate space-time modeling. Extensive experiments on Human 3.6M, CMU Mocap and 3DPW datasets verify the effectiveness of our method which achieves state-of-the-art performance.
Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001
MMAsia1
2023 Graph-Guided MLP-Mixer for Skeleton-Based Human Motion Prediction
abstract
In recent years, Graph Convolutional Networks (GCNs) have been widely used in human motion prediction, but their performance remains unsatisfactory. Recently, MLP-Mixer, initially developed for vision tasks, has been leveraged into human motion prediction as a promising alternative to GCNs, which achieves both better performance and better efficiency than GCNs.
Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001
MMAsia1