VLDB 2026 Research / reviewers in the wild / expert
Yu-Hui Wen
dblp:245/3506
· DBLP profile ↗
25ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-6195-9782ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music SynchronizationabstractRealistic choreography demands simultaneous attention to rhythm and motivation. Prevailing automated dance generation methods mainly depend on musical input, overlooking the motivations that drive meaningful dance creation. Inspired by the motivation choreography, we aim to articulate dance motivations through textual guidance. However, the absence of high-quality datasets concurrently containing music, textual descriptions, and motion data presents a challenge in achieving accurate fine-grained textual control. To address this limitation, we present MotivDance, a novel framework integrating fine-grained textual guidance with music to synthesize semantically coherent dance sequences. Our approach first synthesizes text-guided key poses as motivations. We then introduce an Adaptive Keyframe Locator that dynamically positions these motivations within the musical context through beat-aware synchronization and cross-modal latent space alignment. Finally, a Transformer-based U-Net diffusion model performs the motion in-betweening while preserving motivational integrity. Extensive qualitative and quantitative experiments demonstrate that MotivDance effectively integrates music with fine-grained text control to generate high-fidelity dance motions. Yu-Hui Wen, Liping Jing |
AAAI | 2 |
| 2026 | Co-speech holistic 3D motion generation with style from videoabstractSpeech-driven 3D motion generation has garnered increasing research attention. However, it faces significant challenges in achieving style controllability, primarily due to the scarcity of motion style annotations. To address this, we propose a novel diffusion-based framework for co-speech holistic motion generation that enables example-based style control from videos. Our approach integrates hierarchical speech encoding with rhythm-aware denoising to produce natural and synchronized gestures and expressions. For effective style guidance, we introduce a contrastive style encoder that captures discriminative style representations from reference clips without explicit labeling, enabling generalization to motion styles unseen during training. Furthermore, we design a neural mapper that aligns 2D and 3D gesture features in a shared embedding space, facilitating direct style extraction from in-the-wild videos and seamless transfer to 3D motion. Extensive experiments and user studies show that our proposed approach achieves state-of-the-art performance in both qualitative and quantitative evaluations, offering a flexible solution for controllable motion generation. Yayu Zhang, Yu-Hui Wen, Liping Jing, Jian Yu 0001 |
Graph. Model. | 2 |
| 2026 | A Text-to-3D Framework for Joint Generation of CG-Ready Humans and Compatible GarmentsabstractCreating detailed 3D human avatars with fitted garments traditionally requires specialized expertise and labor-intensive workflows. While recent advances in generative AI have enabled text-to-3D human and clothing synthesis, existing methods fall short in offering accessible, integrated pipelines for generating CG-ready 3D avatars with physically compatible outfits; here we use the term CG-ready for models following a technical aesthetic common in computer graphics (CG) and adopt standard CG polygonal meshes and strands representations (rather than radiance field representations like NeRF and 3DGS) that can be directly integrated into conventional CG pipelines and support downstream tasks such as physical simulation. To bridge this gap, we introduce Tailor, an integrated text-to-3D framework that generates high-fidelity, customizable 3D avatars dressed in simulation-ready garments. Tailor consists of three stages. (1) Semantic Parsing: we employ a large language model to interpret textual descriptions and translate them into parameterized human avatars and semantically matched garment templates. (2) Geometry-Aware Garment Generation: we propose topology-preserving deformation with novel geometric losses to generate body-aligned garments under text control. (3) Consistent Texture Synthesis: we propose a novel multi-view diffusion process optimized for garment texturing, which enforces view consistency, preserves photorealistic details, and optionally supports symmetric texture generation common in garments. Through comprehensive quantitative and qualitative evaluations, we demonstrate that Tailor outperforms state-of-the-art methods in fidelity, usability, and diversity. Zhiyao Sun, Yu-Hui Wen, Ho-Jui Fang, Matthieu Lin, Tian Lv, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Weighted Poisson-disk Resampling on Large-Scale Point CloudsabstractFor large-scale point cloud processing, resampling takes the important role of controlling the point number and density while keeping the geometric consistency. However, current methods cannot balance such different requirements. Particularly with large-scale point clouds, classical methods often struggle with decreased efficiency and accuracy. To address such issues, we propose a weighted Poisson-disk (WPD) resampling method to improve the usability and efficiency for the processing. We first design an initial Poisson resampling with a voxel-based estimation strategy. It is able to estimate a more accurate radius of the Poisson-disk while maintaining high efficiency. Then, we design a weighted tangent smoothing step to further optimize the Voronoi diagram for each point. At the same time, sharp features are detected and kept in the optimized results with isotropic property. Finally, we achieve a resampling copy from the original point cloud with the specified point number, uniform density, and high-quality geometric consistency. Experiments show that our method significantly improves the performance of large-scale point cloud resampling for different applications, and provides a highly practical solution. Xianhe Jiao, Chenlei Lv, Junli Zhao, Ran Yi 0002, Yu-Hui Wen, Zhenkuan Pan 0001, Zhongke Wu, Yong-Jin Liu 0001 |
AAAI | 5 |
| 2025 | VAction: A Lightweight and Integrated VR Training System for Authentic Film-Shooting Experience
Che Qu, Minjing Yu, Chao Zhou 0012, Yuntao Wang 0001, Yu-Hui Wen, Yuanchun Shi, Yong-Jin Liu 0001 |
CHI | 6 |
| 2025 | Indoor Scene Reconstruction With Fine-Grained Details Using Hybrid Representation and Normal Prior EnhancementabstractThe reconstruction of indoor scenes from multi-view RGB images is challenging due to the coexistence of flat and texture-less regions alongside delicate and fine-grained regions. Recent methods leverage neural radiance fields aided by predicted surface normal priors to recover the scene geometry. These methods excel in producing complete and smooth results for floor and wall areas. However, they struggle to capture complex surfaces with high-frequency structures due to the inadequate neural representation and the inaccurately predicted normal priors. This work aims to reconstruct high-fidelity surfaces with fine-grained details by addressing the above limitations. To improve the capacity of the implicit representation, we propose a hybrid architecture to represent low-frequency and high-frequency regions separately. To enhance the normal priors, we introduce a simple yet effective image sharpening and denoising technique, coupled with a network that estimates the pixel-wise uncertainty of the predicted surface normal vectors. Identifying such uncertainty can prevent our model from being misled by unreliable surface normal supervisions that hinder the accurate reconstruction of intricate geometries. Experiments on the benchmark datasets show that our method outperforms existing methods in terms of reconstruction quality. Furthermore, the proposed method also generalizes well to real-world indoor scenarios captured by our hand-held mobile phones. Yubin Hu 0001, Matthieu Lin, Yu-Hui Wen, Wang Zhao 0001, Yong-Jin Liu 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion ModelabstractOcclusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon. Yubin Hu 0001, Wang Zhao 0001, Matthieu Lin, Yu-Hui Wen, Ying He 0001, Yong-Jin Liu 0001 |
AAAI | 6 |
| 2024 | ECAvatar: 3D Avatar Facial Animation with Controllable Identity and Emotion
Minjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun, Tian Lv, Jenny Sheng, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001 |
ACM Multimedia | 8 |
| 2024 | AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular VideosabstractWe introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accurate, consistent and flexible modeling of 3D planes. We derive differentiable rasterization on top of AlphaTablets to efficiently render 3D planes into images, and propose a novel bottom-up pipeline for 3D planar reconstruction from monocular videos. Starting with 2D superpixels and geometric cues from pre-trained models, we initialize 3D planes as AlphaTablets and optimize them via differentiable rendering. An effective merging scheme is introduced to facilitate the growth and refinement of AlphaTablets. Through iterative optimization and merging, we reconstruct complete and accurate 3D planes with solid surfaces and clear boundaries. Extensive experiments on the ScanNet dataset demonstrate state-of-the-art performance in 3D planar reconstruction, underscoring the great potential of AlphaTablets as a generic 3D plane representation for various applications. Wang Zhao 0001, Shaohui Liu, Yubin Hu 0001, Yushi Bai, Yu-Hui Wen, Yong-Jin Liu 0001 |
NeurIPS | 6 |
| 2024 | Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 7 |
| 2024 | Automatic tooth arrangement with joint features of point and mesh representations via diffusion probabilistic models
Changsong Lei, Mengfei Xia, Shaofeng Wang, Yaqian Liang, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 6 |
| 2024 | Gaussian in the Dark: Real-Time View Synthesis From Inconsistent Dark Images Using Gaussian SplattingabstractAbstract 3D Gaussian Splatting has recently emerged as a powerful representation that can synthesize remarkable novel views using consistent multi‐view images as input. However, we notice that images captured in dark environments where the scenes are not fully illuminated can exhibit considerable brightness variations and multi‐view inconsistency, which poses great challenges to 3D Gaussian Splatting and severely degrades its performance. To tackle this problem, we propose Gaussian‐DK. Observing that inconsistencies are mainly caused by camera imaging, we represent a consistent radiance field of the physical world using a set of anisotropic 3D Gaussians, and design a camera response module to compensate for multi‐view inconsistencies. We also introduce a step‐based gradient scaling strategy to constrain Gaussians near the camera, which turn out to be floaters, from splitting and cloning. Experiments on our proposed benchmark dataset demonstrate that Gaussian‐DK produces high‐quality renderings without ghosting and floater artifacts and significantly outperforms existing methods. Furthermore, we can also synthesize light‐up images by controlling exposure levels that clearly show details in shadow areas. Zhen-Hui Dong, Yubin Hu 0001, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Graph. Forum | 4 |
| 2024 | Continuously Controllable Facial Expression Editing in Talking Face VideosabstractRecently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable expressions, which is a strong demand in the industry. The challenge is that speech-related expressions and emotion-related expressions are often highly coupled. Meanwhile, traditional image-to-image translation methods cannot work well in our application due to the coupling of expressions with other attributes such as poses, i.e., translating the expression of the character in each frame may simultaneously change the head pose due to the bias of the training data distribution. In this paper, we propose a high-quality facial expression editing method for talking face videos, allowing the user to control the target emotion in the edited video continuously. We present a new perspective for this task as a special case of motion information editing, where we use a 3DMM to capture major facial movements and an associated texture map modeled by a StyleGAN to capture appearance details. Both representations (3DMM and texture map) contain emotional information and can be continuously modified by neural networks and easily smoothed by averaging in coefficient/latent spaces, making our method simple yet effective. We also introduce a mouth shape preservation loss to control the trade-off between lip synchronization and the degree of exaggeration of the edited expression. Extensive experiments and a user study show that our method achieves state-of-the-art performance across various evaluation criteria. Zhiyao Sun, Yu-Hui Wen, Tian Lv, Yanan Sun 0006, Yaoyuan Wang, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion ModelsabstractThe generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods either employ a deterministic model for speech-to-motion mapping or encode the style using a one-hot encoding scheme. Notably, the one-hot encoding approach fails to capture the complexity of the style and thus limits generalization ability. In this paper, we propose DiffPoseTalk, a generative framework based on the diffusion model combined with a style encoder that extracts style embeddings from short reference videos. During inference, we employ classifier-free guidance to guide the generation process based on the speech and style. In particular, our style includes the generation of head poses, thereby enhancing user perception. Additionally, we address the shortage of scanned 3D talking face data by training our model on reconstructed 3DMM parameters from a high-quality, in-the-wild audio-visual dataset. Extensive experiments and user study demonstrate that our approach outperforms state-of-the-art methods. The code and dataset are at https://diffposetalk.github.io. Zhiyao Sun, Tian Lv, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, Yong-Jin Liu 0001 |
ACM Trans. Graph. | 6 |
| 2024 | PVP-Recon: Progressive View Planning via Warping Consistency for Sparse-View Surface ReconstructionabstractNeural implicit representations have revolutionized dense multi-view surface reconstruction, yet their performance significantly diminishes with sparse input views. A few pioneering works have sought to tackle this challenge by leveraging additional geometric priors or multi-scene generalizability. However, they are still hindered by the imperfect choice of input views, using images under empirically determined viewpoints. We propose PVP-Recon , a novel and effective sparse-view surface reconstruction method that progressively plans the next best views to form an optimal set of sparse viewpoints for image capturing. PVP-Recon starts initial surface reconstruction with as few as 3 views and progressively adds new views which are determined based on a novel warping score that reflects the information gain of each newly added view. This progressive view planning progress is interleaved with a neural SDF-based reconstruction module that utilizes multi-resolution hash features, enhanced by a progressive training scheme and a directional Hessian loss. Quantitative and qualitative experiments on three benchmark datasets show that our system achieves high-quality reconstruction with a constrained input budget and outperforms existing baselines. Matthieu Lin, Jenny Sheng, Ruoyu Fan, Yiheng Han, Yubin Hu 0001, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001, Wenping Wang 0001 |
ACM Trans. Graph. | 9 |
| 2024 | Keyframe Control of Music-Driven 3D Dance GenerationabstractFor 3D animators, choreography with artificial intelligence has attracted more attention recently. However, most existing deep learning methods mainly rely on music for dance generation and lack sufficient control over generated dance motions. To address this issue, we introduce the idea of keyframe interpolation for music-driven dance generation and present a novel transition generation technique for choreography. Specifically, this technique synthesizes visually diverse and plausible dance motions by using normalizing flows to learn the probability distribution of dance motions conditioned on a piece of music and a sparse set of key poses. Thus, the generated dance motions respect both the input musical beats and the key poses. To achieve a robust transition of varying lengths between the key poses, we introduce a time embedding at each timestep as an additional condition. Extensive experiments show that our model generates more realistic, diverse, and beat-matching dance motions than the compared state-of-the-art methods, both qualitatively and quantitatively. Our experimental results demonstrate the superiority of the keyframe-based control for improving the diversity of the generated dance motions. Yu-Hui Wen, Xiao Liu 0040, Yong-Jin Liu 0001, Lin Gao 0004, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Invertible Residual Neural Networks with Conditional Injector and Interpolator for Point Cloud UpsamplingabstractPoint clouds obtained by LiDAR and other sensors are usually sparse and irregular. Low-quality point clouds have serious influence on the final performance of downstream tasks. Recently, a point cloud upsampling network with normalizing flows has been proposed to address this problem. However, the network heavily relies on designing specialized architectures to achieve invertibility. In this paper, we propose a novel invertible residual neural network for point cloud upsampling, called PU-INN, which allows unconstrained architectures to learn more expressive feature transformations. Then, we propose a conditional injector to improve nonlinear transformation ability of the neural network while guaranteeing invertibility. Furthermore, a lightweight interpolator is proposed based on semantic similarity distance in the latent space, which can intuitively reflect the interpolation changes in Euclidean space. Qualitative and quantitative results show that our method outperforms the state-of-the-art works in terms of distribution uniformity, proximity-to-surface accuracy, 3D reconstruction quality, and computation efficiency. Aihua Mao, Yaqi Duan, Yu-Hui Wen, Zihui Du, Hongmin Cai, Yong-Jin Liu 0001 |
IJCAI | 3 |
| 2023 | Generation of virtual digital human for customer service industry
Yanan Sun 0006, Zhiyao Sun, Yu-Hui Wen, Tian Lv, Minjing Yu, Ran Yi 0002, Lin Gao 0004, Yong-Jin Liu 0001 |
Comput. Graph. | 3 |
| 2023 | Motif-GCNs With Local and Non-Local Temporal Blocks for Skeleton-Based Action RecognitionabstractRecent works have achieved remarkable performance for action recognition with human skeletal data by utilizing graph convolutional models. Existing models mainly focus on developing graph convolutional operations to encode structural properties of a skeletal graph, whose topology is manually predefined and fixed over all action samples. Some recent works further take sample-dependent relationships among joints into consideration. However, the complex relationships between arbitrary pairwise joints are difficult to learn and the temporal features between frames are not fully exploited by simply using traditional convolutions with small local kernels. In this paper, we propose a motif-based graph convolution method, which makes use of sample-dependent latent relations among non-physically connected joints to impose a high-order locality and assigns different semantic roles to physical neighbors of a joint to encode hierarchical structures. Furthermore, we propose a sparsity-promoting loss function to learn a sparse motif adjacency matrix for latent dependencies in non-physical connections. For extracting effective temporal information, we propose an efficient local temporal block. It adopts partial dense connections to reuse temporal features in local time windows, and enrich a variety of information flow by gradient combination. In addition, we introduce a non-local temporal block to capture global dependencies among frames. Our model can capture local and non-local relationships both spatially and temporally, by integrating the local and non-local temporal blocks into the sparse motif-based graph convolutional networks (SMotif-GCNs). Comprehensive experiments on four large-scale datasets show that our model outperforms the state-of-the-art methods. Our code is publicly available at https://github.com/wenyh1616/SAMotif-GCN. Yu-Hui Wen, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | PD-Flow: A Point Cloud Denoising Framework with Normalizing Flows
Aihua Mao, Zihui Du, Yu-Hui Wen, Jun Xuan, Yong-Jin Liu 0001 |
ECCV (3) | 3 |
| 2022 | Audio-Driven Stylized Gesture Generation with Flow-Based Model
Yu-Hui Wen, Yanan Sun 0006, Ying He 0001, Yaoyuan Wang, Weihua He, Yong-Jin Liu 0001 |
ECCV (5) | 2 |
| 2022 | Generating Smooth and Facial-Details-Enhanced Talking Head Video: A Perspective of Pre and Post ProcessesabstractTalking head video generation has received increasing attention recently. So far the quality (especially the facial details) of the videos output from state-of-the-art deep learning methods is limited by either the quality of training data or the performance of generators, and needs to be further improved. In this paper, we propose a data pre- and post- processing strategy based on a key observation: generating talking head video from multi-modal input is a challenging problem and generating smooth video with fine facial details makes the problem even harder. Then we propose to decompose the problem solution into a main deep model, a pre- and a post- processing. The main deep model generates a reasonably good talking face video, with the aid of a pre-process, which also contributes to a post-process for restoring smooth and fine facial details in the final video. In particular, our main deep model reconstructs a 3D face from an input reference frame, and then uses an AudioNet to generate a sequence of facial expression coefficients with an input audio clip. To ensure final facial details in the generated video, we sample the original texture from the reference frame in the pre-process with the aid of reconstructed 3D face and a predefined UV map. Accordingly, in the post-process, we smooth the expression coefficients of adjacent frames to alleviate jitters and apply a pretrained face restoration module to recover the fine facial details. Experimental results and ablation study show the advantage of our proposed method. Tian Lv, Yu-Hui Wen, Zhiyao Sun, Zipeng Ye, Yong-Jin Liu 0001 |
ACM Multimedia | 2 |
| 2021 | Autoregressive Stylized Motion Synthesis With Generative FlowabstractMotion style transfer is an important problem in many computer graphics and computer vision applications, including human animation, games, and robotics. Most existing deep learning methods for this problem are supervised and trained by registered motion pairs. In addition, these methods are often limited to yielding a deterministic output, given a pair of style and content motions. In this paper, we propose an unsupervised approach for motion style transfer by synthesizing stylized motions autoregressively using a generative flow model $\mathcal{M}$. $\mathcal{M}$ is trained to maximize the exact likelihood of a collection of unlabeled motions, based on an autoregressive context of poses in previous frames and a control signal representing the movement of a root joint. Thanks to invertible flow transformations, latent codes that encode deep properties of motion styles are efficiently inferred by $\mathcal{M}$. By combining the latent codes (from an input style motion S) with the autoregressive context and control signal (from an input content motion C), $\mathcal{M}$ outputs a stylized motion which transfers style from S to C. Moreover, our model is probabilistic and is able to generate various plausible motions with a specific style. We evaluate the proposed model on motion capture datasets containing different human motion styles. Experiment results show that our model outperforms the state-of-the-art methods, despite not requiring manually labeled training data. Yu-Hui Wen, Hongbo Fu 0001, Lin Gao 0004, Yanan Sun 0006, Yong-Jin Liu 0001 |
CVPR | 1 |
| 2020 | View planning in robot active vision: A survey of systems, algorithms, and applicationsabstractRapid development of artificial intelligence motivates researchers to expand the capabilities of intelligent and autonomous robots. In many robotic applications, robots are required to make planning decisions based on perceptual information to achieve diverse goals in an efficient and effective way. The planning problem has been investigated in active robot vision, in which a robot analyzes its environment and its own state in order to move sensors to obtain more useful information under certain constraints. View planning, which aims to find the best view sequence for a sensor, is one of the most challenging issues in active robot vision. The quality and efficiency of view planning are critical for many robot systems and are influenced by the nature of their tasks, hardware conditions, scanning states, and planning strategies. In this paper, we first summarize some basic concepts of active robot vision, and then review representative work on systems, algorithms and applications from four perspectives: object reconstruction, scene reconstruction, object recognition, and pose estimation. Finally, some potential directions are outlined for future work. Yu-Hui Wen, Wang Zhao 0001, Yong-Jin Liu 0001 |
Comput. Vis. Media | 2 |
| 2019 | Graph CNNs with Motif and Variable Temporal Block for Skeleton-Based Action RecognitionabstractHierarchical structure and different semantic roles of joints in human skeleton convey important information for action recognition. Conventional graph convolution methods for modeling skeleton structure consider only physically connected neighbors of each joint, and the joints of the same type, thus failing to capture highorder information. In this work, we propose a novel model with motif-based graph convolution to encode hierarchical spatial structure, and a variable temporal dense block to exploit local temporal information over different ranges of human skeleton sequences. Moreover, we employ a non-local block to capture global dependencies of temporal domain in an attention mechanism. Our model achieves improvements over the stateof-the-art methods on two large-scale datasets. Yu-Hui Wen, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia |
AAAI | 1 |