VLDB 2026 Research / reviewers in the wild / expert
Shihong Xia
dblp:38/897
· DBLP profile ↗
68ranked-venue papers
5as first author
22since 2021 · last 2025
0000-0002-7228-9646ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 11 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base ModelabstractTraining a generic 3D face reconstruction model in a self-supervised manner using large-scale, in-the-wild 2D face image datasets enhances robustness to varying lighting conditions and occlusions while allowing the model to capture animatable wrinkle details across diverse facial expressions. However, a generic model often fails to adequately represent the unique characteristics of specific individuals. In this paper, we propose a method to train a generic base model and then transfer it to yield person-specific models by integrating lightweight adapters within the large-parameter ViT-MAE base model. These person-specific models excel at capturing individual facial shapes and detailed features while preserving the robustness and prior knowledge of detail variations from the base model. During training, we introduce a silhouette vertex re-projection loss to address boundary "landmark marching" issues on the 3D face caused by pose variations. Additionally, we employ an innovative teacher-student loss to leverage the inherent strengths of UNet in feature boundary localization for training our detail MAE. Quantitative and qualitative experiments demonstrate that our approach achieves state-of-the-art performance in face alignment, detail accuracy, and richness. The source code is available at https://github.com/danielmao2000/person-specific-animatable-face. Yuxiang Mao, Zhenfeng Fan, ZhiJie Zhang, Shihong Xia |
CVPR | 5 |
| 2025 | Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Face Animation
Yuxiang Mao, Shihong Xia |
ICXR | 6 |
| 2025 | HMVLM: Human Motion-Vision-Language Model via MoE LoRAabstractThe expansion of instruction-tuning data has enabled foundation language models to exhibit improved instruction adherence and superior performance across diverse downstream tasks. Semantically-rich 3D human motion is being progressively integrated with these foundation models to enhance multimodal understanding and cross-modal generation capabilities. However, the modality gap between human motion and text raises unresolved concerns about catastrophic forgetting during this integration. In addition, developing autoregressive-compatible pose representations that preserve generalizability across heterogeneous downstream tasks remains a critical technical barrier. To address these issues, we propose the Human Motion-Vision-Language Model (HMVLM), a unified framework based on the Mixture of Expert Low-Rank Adaption(MoE LoRA) strategy. The framework leverages the gating network to dynamically allocate LoRA expert weights based on the input prompt, enabling synchronized fine-tuning of multiple tasks. To mitigate catastrophic forgetting during instruction-tuning, we introduce a novel zero expert that preserves the pre-trained parameters for general linguistic tasks. For pose representation, we implement body-part-specific tokenization by partitioning the human body into different joint groups, enhancing the spatial resolution of the representation. Experiments show that our method effectively alleviates knowledge forgetting during instruction-tuning and achieves remarkable performance across diverse human motion downstream tasks. Lei Hu 0008, Yongjing Ye, Shihong Xia |
NeurIPS | 3 |
| 2024 | Diffusion-based Human Motion Style Transfer with Semantic GuidanceabstractAbstract 3D Human motion style transfer is a fundamental problem in computer graphic and animation processing. Existing AdaIN‐based methods necessitate datasets with balanced style distribution and content/style labels to train the clustered latent space. However, we may encounter a single unseen style example in practical scenarios, but not in sufficient quantity to constitute a style cluster for AdaIN‐based methods. Therefore, in this paper, we propose a novel two‐stage framework for few‐shot style transfer learning based on the diffusion model. Specifically, in the first stage, we pre‐train a diffusion‐based text‐to‐motion model as a generative prior so that it can cope with various content motion inputs. In the second stage, based on the single style example, we fine‐tune the pre‐trained diffusion model in a few‐shot manner to make it capable of style transfer. The key idea is regarding the reverse process of diffusion as a motion‐style translation process since the motion styles can be viewed as special motion variations. During the fine‐tuning for style transfer, a simple yet effective semantic‐guided style transfer loss coordinated with style example reconstruction loss is introduced to supervise the style transfer in CLIP semantic space. The qualitative and quantitative evaluations demonstrate that our method can achieve state‐of‐the‐art performance and has practical applications. The source code is available at https://github.com/hlcdyy/diffusion-based-motion-style-transfer . Lei Hu 0008, Yongjing Ye, Shihong Xia |
Comput. Graph. Forum | 5 |
| 2024 | Mining collaborative spatio-temporal clues for face forgery detection
Bo Ding 0006, Zhenfeng Fan, Zejun Zhao, Shihong Xia |
Multim. Tools Appl. | 4 |
| 2024 | Pose-Aware Attention Network for Flexible Motion Retargeting by Body PartabstractMotion retargeting is a fundamental problem in computer graphics and computer vision. Existing approaches usually have many strict requirements, such as the source-target skeletons needing to have the same number of joints or share the same topology. To tackle this problem, we note that skeletons with different structure may have some common body parts despite the differences in joint numbers. Following this observation, we propose a novel, flexible motion retargeting framework. The key idea of our method is to regard the body part as the basic retargeting unit rather than directly retargeting the whole body motion. To enhance the spatial modeling capability of the motion encoder, we introduce a pose-aware attention network (PAN) in the motion encoding phase. The PAN is pose-aware since it can dynamically predict the joint weights within each body part based on the input pose, and then construct a shared latent space for each body part by feature pooling. Extensive experiments show that our approach can generate better motion retargeting results both qualitatively and quantitatively than state-of-the-art methods. Moreover, we also show that our framework can generate reasonable results even for a more challenging retargeting scenario, like retargeting between bipedal and quadrupedal skeletons because of the body part retargeting strategy and PAN. Lei Hu 0008, Chongyang Zhong, Boyuan Jiang, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Unpaired Multi-domain Attribute Translation of 3D Facial Shapes with a Square and Symmetric Geometric MapabstractWhile impressive progress has recently been made in image-oriented facial attribute translation, shape-oriented 3D facial attribute translation remains an unsolved issue. This is primarily limited by the lack of 3D generative models and ineffective usage of 3D facial data. We propose a learning framework for 3D facial attribute translation to relieve these limitations. Firstly, we customize a novel geometric map for 3D shape representation and embed it in an end-to-end generative adversarial network. The geometric map represents 3D shapes symmetrically on a square image grid, while preserving the neighboring relationship of 3D vertices in a local least-square sense. This enables effective learning for the latent representation of data with different attributes. Secondly, we employ a unified and unpaired learning framework for multi-domain attribute translation. It not only makes effective usage of data correlation from multiple domains, but also mitigates the constraint for hardly accessible paired data. Finally, we propose a hierarchical architecture for the discriminator to guarantee robust results against both global and local artifacts. We conduct extensive experiments to demonstrate the advantage of the proposed framework over the state-of-the-art in generating high-fidelity facial shapes. Given an input 3D facial shape, the proposed framework is able to synthesize novel shapes of different attributes, which covers some downstream applications, such as expression transfer, gender translation, and aging. Code at https://github.com/NaughtyZZ/3D_facial_shape_attribute_translation_ssgmap. Zhenfeng Fan, Chongyang Zhong, Min Cao 0005, Shihong Xia |
ICCV | 6 |
| 2023 | Probabilistic Triangulation for Uncalibrated Multi-View 3D Human Pose Estimationabstract3D human pose estimation has been a long-standing challenge in computer vision and graphics, where multi-view methods have significantly progressed but are limited by the tedious calibration processes. Existing multi-view methods are restricted to fixed camera pose and therefore lack generalization ability. This paper presents a novel Probabilistic Triangulation module that can be embedded in a calibrated 3D human pose estimation method, generalizing it to uncalibration scenes. The key idea is to use a probability distribution to model the camera pose and iteratively update the distribution from 2D features instead of using camera pose. Specifically, We maintain a camera pose distribution and then iteratively update this distribution by computing the posterior probability of the camera pose through Monte Carlo sampling. This way, the gradients can be directly back-propagated from the 3D pose estimation to the 2D heatmap, enabling end-to-end training. Extensive experiments on Human3.6M and CMU Panoptic demonstrate that our method outperforms other uncalibration methods and achieves comparable results with state-of-the-art calibration methods. Thus, our method achieves a trade-off between estimation accuracy and generalizability. Our code is in https://github.com/bymaths/probabilistictriangulation Boyuan Jiang, Lei Hu 0008, Shihong Xia |
ICCV | 3 |
| 2023 | AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention MechanismabstractGenerating 3D human motion based on textual descriptions has been a research focus in recent years. It requires the generated motion to be diverse, natural, and conform to the textual description. Due to the complex spatio-temporal nature of human motion and the difficulty in learning the cross-modal relationship between text and motion, text-driven motion generation is still a challenging problem. To address these issues, we propose AttT2M, a two-stage method with multi-perspective attention mechanism: body-part attention and global-local motion-text attention. The former focuses on the motion embedding perspective, which means introducing a body-part spatio-temporal encoder into VQ-VAE to learn a more expressive discrete latent space. The latter is from the cross-modal perspective, which is used to learn the sentence-level and word-level motion-text cross-modal relationship. The text-driven motion is finally generated with a generative transformer. Extensive experiments conducted on HumanML3D and KIT-ML demonstrate that our method outperforms the current state-of-the-art works in terms of qualitative and quantitative evaluation, and achieve fine-grained synthesis and action2motion. Our code is in https://github.com/ZcyMonkey/AttT2M. Chongyang Zhong, Lei Hu 0008, Shihong Xia |
ICCV | 4 |
| 2023 | Towards Fine-Grained Optimal 3D Face Dense Registration: An Iterative Dividing and Diffusing Method
Zhenfeng Fan, Silong Peng, Shihong Xia |
Int. J. Comput. Vis. | 3 |
| 2023 | Motif-GCNs With Local and Non-Local Temporal Blocks for Skeleton-Based Action RecognitionabstractRecent works have achieved remarkable performance for action recognition with human skeletal data by utilizing graph convolutional models. Existing models mainly focus on developing graph convolutional operations to encode structural properties of a skeletal graph, whose topology is manually predefined and fixed over all action samples. Some recent works further take sample-dependent relationships among joints into consideration. However, the complex relationships between arbitrary pairwise joints are difficult to learn and the temporal features between frames are not fully exploited by simply using traditional convolutions with small local kernels. In this paper, we propose a motif-based graph convolution method, which makes use of sample-dependent latent relations among non-physically connected joints to impose a high-order locality and assigns different semantic roles to physical neighbors of a joint to encode hierarchical structures. Furthermore, we propose a sparsity-promoting loss function to learn a sparse motif adjacency matrix for latent dependencies in non-physical connections. For extracting effective temporal information, we propose an efficient local temporal block. It adopts partial dense connections to reuse temporal features in local time windows, and enrich a variety of information flow by gradient combination. In addition, we introduce a non-local temporal block to capture global dependencies among frames. Our model can capture local and non-local relationships both spatially and temporally, by integrating the local and non-local temporal blocks into the sparse motif-based graph convolutional networks (SMotif-GCNs). Comprehensive experiments on four large-scale datasets show that our model outperforms the state-of-the-art methods. Our code is publicly available at https://github.com/wenyh1616/SAMotif-GCN. Yu-Hui Wen, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | 3D Face Reconstruction and Gaze Tracking in the HMD for Virtual InteractionabstractWith the rapid development of virtual reality (VR) technology, VR headsets, a.k.a. Head-Mounted Displays (HMDs), are widely available, allowing immersive 3D content to be viewed. A natural need for truly immersive VR is to allow bidirectional communication: the user should be able to interact with the virtual world using facial expressions and eye gaze, in addition to traditional means of interaction. The typical application scenario includes VR virtual conferencing and virtual roaming, where ideally users are able to see other users’ expressions and have eye contact with them in the virtual world. In addition, eye gaze also provides a natural means of interaction with virtual objects. Despite significant achievements in recent years for reconstruction of 3D faces from RGB or RGB-D images, it remains a challenge to reliably capture and reconstruct 3D facial expressions including eye gaze when the user is wearing an HMD, because the majority of the face is occluded, especially those areas around the eyes which are essential for recognizing facial expressions and eye gaze. In this paper, we introduce a novel real-time system that is able to capture and reconstruct 3D faces wearing HMDs, and robustly recover eye gaze. We further propose a novel method to map eye gaze directions to the 3D virtual world, which provides a novel and useful interactive mode in VR. We compare our method with state-of-the-art techniques both qualitatively and quantitatively, and demonstrate the effectiveness of our system using live capture. Yukun Lai, Shihong Xia, Paul L. Rosin, Lin Gao 0004 |
IEEE Trans. Multim. | 3 |
| 2023 | Multiscale Mesh Deformation Component Analysis With Attention-Based AutoencodersabstractDeformation component analysis is a fundamental problem in geometry processing and shape understanding. Existing approaches mainly extract deformation components in local regions at a similar scale while deformations of real-world objects are usually distributed in a multi-scale manner. In this article, we propose a novel method to exact multiscale deformation components automatically with a stacked attention-based autoencoder. The attention mechanism is designed to learn to softly weight multi-scale deformation components in active deformation regions, and the stacked attention-based autoencoder is learned to represent the deformation components at different scales. Quantitative and qualitative evaluations show that our method outperforms state-of-the-art methods. Furthermore, with the multiscale deformation components extracted by our method, the user can edit shapes in a coarse-to-fine fashion which facilitates effective modeling of new shapes. Jie Yang 0038, Lin Gao 0004, Qingyang Tan, Yihua Huang 0002, Shihong Xia, Yukun Lai |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Spatio-Temporal Gating-Adjacency GCN for Human Motion PredictionabstractPredicting future motion based on historical motion sequence is a fundamental problem in computer vision, and it has wide applications in autonomous driving and robotics. Some recent works have shown that Graph Convolutional Networks(GCN) are instrumental in modeling the relationship between different joints. However, considering the variants and diverse action types in human motion data, the cross-dependency of the spatio-temporal relationships will be difficult to depict due to the decoupled modeling strategy, which may also exacerbate the problem of insufficient generalization. Therefore, we propose the Spatio-Temporal Gating-Adjacency GCN(GAGCN) to learn the complex spatio-temporal dependencies over diverse action types. Specifically, we adopt gating networks to enhance the generalization of GCN via the trainable adaptive adjacency matrix obtained by blending the candidate spatio-temporal adjacency matrices. Moreover, GAGCN addresses the cross-dependency of space and time by balancing the weights of spatio-temporal modeling and fusing the decoupled spatio-temporal features. Extensive experiments on Human 3.6M, AMASS, and 3DPW demonstrate that GAGCN achieves state-of-the-art performance in both short-term and long-term predictions. Chongyang Zhong, Lei Hu 0008, Yongjing Ye, Shihong Xia |
CVPR | 5 |
| 2022 | Learning Uncoupled-Modulation CVAE for 3D Action-Conditioned Human Motion Synthesis
Chongyang Zhong, Lei Hu 0008, Shihong Xia |
ECCV (21) | 4 |
| 2022 | Neural3Points: Learning to Generate Physically Realistic Full-body Motion for Virtual Reality UsersabstractAbstract Animating an avatar that reflects a user's action in the VR world enables natural interactions with the virtual environment. It has the potential to allow remote users to communicate and collaborate in a way as if they met in person. However, a typical VR system provides only a very sparse set of up to three positional sensors, including a head‐mounted display (HMD) and Optionally two hand‐held controllers, making the estimation of the user's full‐body movement a difficult problem. In this work, we present a data‐driven physics‐based method for predicting the realistic full‐body movement of the user according to the transformations of these VR trackers and simulating an avatar character to mimic such user actions in the virtual world in realtime. We train our system using reinforcement learning with carefully designed pretraining processes to ensure the success of the training and the quality of the simulation. We demonstrate the effectiveness of the method with an extensive set of examples. Yongjing Ye, Libin Liu 0002, Lei Hu 0008, Shihong Xia |
Comput. Graph. Forum | 4 |
| 2022 | Spatial-temporal modeling for prediction of stylized human motion
Chongyang Zhong, Lei Hu 0008, Shihong Xia |
Neurocomputing | 3 |
| 2022 | Active Colorization for Cartoon Line DrawingsabstractIn the animation industry, the colorization of raw sketch images is a vitally important but very time-consuming task. This article focuses on providing a novel solution that semiautomatically colorizes a set of images using a single colorized reference image. Our method is able to provide coherent colors for regions that have similar semantics to those in the reference image. An active-learning-based framework is used to match local regions, followed by mixed-integer quadratic programming (MIQP) which considers the spatial contexts to further refine the matching results. We efficiently utilize user interactions to achieve high accuracy in the final colorized images. Experiments show that our method outperforms the current state-of-the-art deep learning based colorization method in terms of color coherency with the reference image. The region matching framework could potentially be applied to other applications, such as color transfer. Jia-Qi Zhang, Lin Gao 0004, Shihong Xia, Min Shi 0005 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Sequential 3D Human Pose Estimation Using Adaptive Point Cloud Sampling Strategyabstract3D human pose estimation is a fundamental problem in artificial intelligence, and it has wide applications in AR/VR, HCI and robotics. However, human pose estimation from point clouds still suffers from noisy points and estimated jittery artifacts because of handcrafted-based point cloud sampling and single-frame-based estimation strategies. In this paper, we present a new perspective on the 3D human pose estimation method from point cloud sequences. To sample effective point clouds from input, we design a differentiable point cloud sampling method built on density-guided attention mechanism. To avoid the jitter caused by previous 3D human pose estimation problems, we adopt temporal information to obtain more stable results. Experiments on the ITOP dataset and the NTU-RGBD dataset demonstrate that all of our contributed components are effective, and our method can achieve state-of-the-art performance. Lei Hu 0008, Xiaoming Deng 0001, Shihong Xia |
IJCAI | 4 |
| 2021 | Sparse Data Driven Mesh DeformationabstractExample-based mesh deformation methods are powerful tools for realistic shape editing. However, existing techniques typically combine all the example deformation modes, which can lead to overfitting, i.e., using an overly complicated model to explain the user-specified deformation. This leads to implausible or unstable deformation results, including unexpected global changes outside the region of interest. To address this fundamental limitation, we propose a sparse blending method that automatically selects a smaller number of deformation modes to compactly describe the desired deformation. This along with a suitably chosen deformation basis including spatially localized deformation modes leads to significant advantages, including more meaningful, reliable, and efficient deformations because fewer and localized deformation modes are applied. To cope with large rotations, we develop a simple but effective representation based on polar decomposition of deformation gradients, which resolves the ambiguity of large global rotations using an as-consistent-as-possible global optimization. This simple representation has a closed form solution for derivatives, making it efficient for our sparse localized representation and thus ensuring interactive performance. Experimental results show that our method outperforms state-of-the-art data-driven mesh deformation methods, for both quality of results and efficiency. Lin Gao 0004, Yukun Lai, Jie Yang 0038, Ling-Xiao Zhang, Shihong Xia, Leif Kobbelt |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Combining Recurrent Neural Networks and Adversarial Training for Human Motion Synthesis and ControlabstractThis paper introduces a new generative deep learning network for human motion synthesis and control. Our key idea is to combine recurrent neural networks (RNNs) and adversarial training for human motion modeling. We first describe an efficient method for training an RNN model from prerecorded motion data. We implement RNNs with long short-term memory (LSTM) cells because they are capable of addressing the nonlinear dynamics and long term temporal dependencies present in human motions. Next, we train a refiner network using an adversarial loss, similar to generative adversarial networks (GANs), such that refined motion sequences are indistinguishable from real mocap data using a discriminative network. The resulting model is appealing for motion synthesis and control because it is compact, contact-aware, and can generate an infinite number of naturally looking motions with infinite lengths. Our experiments show that motions generated by our deep learning model are always highly realistic and comparable to high-quality motion capture data. We demonstrate the power and effectiveness of our models by exploring a variety of applications, ranging from random motion synthesis, online/offline motion control, and motion filtering. We show the superiority of our generative model by comparison against baseline models. Jinxiang Chai, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Realtime and Accurate 3D Eye Gaze Capture with DCNN-Based Iris and Pupil SegmentationabstractThis paper presents a realtime and accurate method for 3D eye gaze tracking with a monocular RGB camera. Our key idea is to train a deep convolutional neural network(DCNN) that automatically extracts the iris and pupil pixels of each eye from input images. To achieve this goal, we combine the power of Unet\cite{ronneberger2015u-net:} and Squeezenet\cite{iandola2017squeezenet:} to train an efficient convolutional neural network for pixel classification. In addition, we track the 3D eye gaze state in the Maximum A Posteriori (MAP) framework, which sequentially searches for the most likely state of the 3D eye gaze at each frame. When eye blinking occurs, the eye gaze tracker can obtain an inaccurate result. We further extend the convolutional neural network for eye close detection in order to improve the robustness and accuracy of the eye gaze tracker. Our system runs in realtime on desktop PCs and smart phones. We have evaluated our system on live videos and Internet videos, and our results demonstrate that the system is robust and accurate for various genders, races, lighting conditions, poses, shapes and facial expressions. A comparison against Wang et al.[3] shows that our method advances the state of the art in 3D eye tracking using a single RGB camera. Jinxiang Chai, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | DeepFaceDrawing: deep generation of face images from sketchesabstractRecent deep image-to-image translation techniques allow fast generation of face images from freehand sketches. However, existing solutions tend to overfit to sketches, thus requiring professional sketches or even edge maps as input. To address this issue, our key idea is to implicitly model the shape space of plausible face images and synthesize a face image in this space to approximate an input sketch. We take a local-to-global approach. We first learn feature embeddings of key face components, and push corresponding parts of input sketches towards underlying component manifolds defined by the feature vectors of face component samples. We also propose another deep neural network to learn the mapping from the embedded component features to realistic images with multi-channel feature maps as intermediate results to improve the information flow. Our method essentially uses input sketches as soft constraints and is thus able to produce high-quality face images even from rough and/or incomplete sketches. Our tool is easy to use even for non-artists, while still supporting fine-grained control of shape details. Both qualitative and quantitative evaluations show the superior generation ability of our system to existing and alternative solutions. The usability and expressiveness of our system are confirmed by a user study. Wanchao Su, Lin Gao 0004, Shihong Xia, Hongbo Fu 0001 |
ACM Trans. Graph. | 4 |
| 2020 | Weakly Supervised Adversarial Learning for 3D Human Pose Estimation from Point CloudsabstractPoint clouds-based 3D human pose estimation that aims to recover the 3D locations of human skeleton joints plays an important role in many AR/VR applications. The success of existing methods is generally built upon large scale data annotated with 3D human joints. However, it is a labor-intensive and error-prone process to annotate 3D human joints from input depth images or point clouds, due to the self-occlusion between body parts as well as the tedious annotation process on 3D point clouds. Meanwhile, it is easier to construct human pose datasets with 2D human joint annotations on depth images. To address this problem, we present a weakly supervised adversarial learning framework for 3D human pose estimation from point clouds. Compared to existing 3D human pose estimation methods from depth images or point clouds, we exploit both the weakly supervised data with only annotations of 2D human joints and fully supervised data with annotations of 3D human joints. In order to relieve the human pose ambiguity due to weak supervision, we adopt adversarial learning to ensure the recovered human pose is valid. Instead of using either 2D or 3D representations of depth images in previous methods, we exploit both point clouds and the input depth image. We adopt 2D CNN to extract 2D human joints from the input depth image, 2D human joints aid us in obtaining the initial 3D human joints and selecting effective sampling points that could reduce the computation cost of 3D human pose regression using point clouds network. The used point clouds network can narrow down the domain gap between the network input i.e. point clouds and 3D joints. Thanks to weakly supervised adversarial learning framework, our method can achieve accurate 3D human pose from point clouds. Experiments on the ITOP dataset and EVAL dataset demonstrate that our method can achieve state-of-the-art performance efficiently. Lei Hu 0008, Xiaoming Deng 0001, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Graph CNNs with Motif and Variable Temporal Block for Skeleton-Based Action RecognitionabstractHierarchical structure and different semantic roles of joints in human skeleton convey important information for action recognition. Conventional graph convolution methods for modeling skeleton structure consider only physically connected neighbors of each joint, and the joints of the same type, thus failing to capture highorder information. In this work, we propose a novel model with motif-based graph convolution to encode hierarchical spatial structure, and a variable temporal dense block to exploit local temporal information over different ranges of human skeleton sequences. Moreover, we employ a non-local block to capture global dependencies of temporal domain in an attention mechanism. Our model achieves improvements over the stateof-the-art methods on two large-scale datasets. Yu-Hui Wen, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia |
AAAI | 5 |
| 2019 | Data-driven weight optimization for real-time mesh deformation
Yu-Jie Yuan, Yukun Lai, Tong Wu 0009, Shihong Xia, Lin Gao 0004 |
Graph. Model. | 4 |
| 2019 | Temporal Upsampling of Depth Maps Using a Hybrid CameraabstractIn recent years, consumer-level depth cameras have been adopted for various applications. However, they often produce depth maps at only a moderately high frame rate (approximately 30 frames per second), preventing them from being used for applications such as digitizing human performance involving fast motion. On the other hand, low-cost, high-frame-rate video cameras are available. This motivates us to develop a hybrid camera that consists of a high-frame-rate video camera and a low-frame-rate depth camera and to allow temporal interpolation of depth maps with the help of auxiliary color images. To achieve this, we develop a novel algorithm that reconstructs intermediate depth maps and estimates scene flow simultaneously. We test our algorithm on various examples involving fast, non-rigid motions of single or multiple objects. Our experiments show that our scene flow estimation method is more precise than a tracking-based method and the state-of-the-art techniques. Mingzhe Yuan, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | Mesh-Based Autoencoders for Localized Deformation Component AnalysisabstractSpatially localized deformation components are very useful for shape analysis and synthesis in 3D geometry processing. Several methods have recently been developed, with an aim to extract intuitive and interpretable deformation components. However, these techniques suffer from fundamental limitations especially for meshes with noise or large-scale deformations, and may not always be able to identify important deformation components.In this paper we propose a novel mesh-based autoencoder architecture that is able to cope with meshes with irregular topology. We introduce sparse regularization in this framework, which along with convolutional operations, helps localize deformations.Our framework is capable of extracting localized deformation components from mesh data sets with large-scale deformations and is robust to noise. It also provides a nonlinear approach to reconstruction of meshes using the extracted basis, which is more effective than the current linear combination approach. Extensive experiments show that our method outperforms state-of-the-art methods in both qualitative and quantitative evaluations. Qingyang Tan, Lin Gao 0004, Yukun Lai, Jie Yang 0038, Shihong Xia |
AAAI | 5 |
| 2018 | SF-Net: Learning Scene Flow from RGB-D Images with CNNs
Yi-Ling Qiao, Lin Gao 0004, Yukun Lai, Mingzhe Yuan, Shihong Xia |
BMVC | 6 |
| 2018 | Variational Autoencoders for Deforming 3D Mesh Modelsabstract3D geometric contents are becoming increasingly popular. In this paper, we study the problem of analyzing deforming 3D meshes using deep neural networks. Deforming 3D meshes are flexible to represent 3D animation sequences as well as collections of objects of the same category, allowing diverse shapes with large-scale non-linear deformations. We propose a novel framework which we call mesh variational autoencoders (mesh VAE), to explore the probabilistic latent space of 3D surfaces. The framework is easy to train, and requires very few training examples. We also propose an extended model which allows flexibly adjusting the significance of different latent variables by altering the prior distribution. Extensive experiments demonstrate that our general framework is able to learn a reasonable representation for a collection of deformable shapes, and produce competitive results for a variety of applications, including shape generation, shape interpolation, shape space embedding and shape exploration, outperforming state-of-the-art methods. Qingyang Tan, Lin Gao 0004, Yukun Lai, Shihong Xia |
CVPR | 4 |
| 2018 | Real-Time 3D Face Reconstruction and Gaze Tracking for Virtual RealityabstractWith the rapid development of virtual reality (VR) technology, VR glasses, a.k.a. Head-Mounted Displays (HMDs) are widely available, allowing immersive 3D content to be viewed. A natural need for truly immersive VR is to allow bidirectional communication: the user should be able to interact with the virtual world using facial expressions and eye gaze, in addition to traditional means of interaction. Typical application scenarios include VR virtual conferencing and virtual roaming, where ideally users are able to see other users' expressions and have eye contact with them in the virtual world. Despite significant achievements in recent years for reconstruction of 3D faces from RGB or RGB- D images, it remains a challenge to reliably capture and reconstruct 3D facial expressions including eye gaze when the user is wearing VR glasses, because the majority of the face is occluded, especially those areas around the eyes which are essential for recognizing facial expressions and eye gaze. In this paper, we introduce a novel real-time system that is able to capture and reconstruct 3D faces wearing HMDs and robustly recover eye gaze. We demonstrate the effectiveness of our system using live capture and more results are shown in the accompanying video. Lin Gao 0004, Yukun Lai, Paul L. Rosin, Shihong Xia |
VR | 5 |
| 2018 | Cascaded 3D Full-Body Pose Regression from Single Depth Image at 100 FPSabstractThere are increasing real-time live applications in virtual reality, where it plays an important role in capturing and retargetting 3D human pose. But it is still challenging to estimate accurate 3D pose from consumer imaging devices such as depth camera. This paper presents a novel cascaded 3D full-body pose regression method to estimate accurate pose from a single depth image at 100 fps. The key idea is to train cascaded regressors based on Gradient Boosting algorithm from pre-recorded human motion capture database. By incorporating hierarchical kinematics model of human pose into the learning procedure, we can directly estimate accurate 3D joint angles instead of joint positions. The biggest advantage of this model is that the bone length can be preserved during the whole 3D pose estimation procedure, which leads to more effective features and higher pose estimation accuracy. Our method can be used as an initialization procedure when combining with tracking methods. We demonstrate the power of our method on a wide range of synthesized human motion data from CMU mocap database, Human3.6M dataset and real human movements data captured in real time. In our comparison against previous 3D pose estimation methods and commercial system such as Kinect 2017, we achieve the state-of-the-art accuracy. Shihong Xia, Le Su |
VR | 1 |
| 2018 | Biharmonic deformation transfer with automatic key point selection
Jie Yang 0038, Lin Gao 0004, Yukun Lai, Paul L. Rosin, Shihong Xia |
Graph. Model. | 5 |
| 2018 | Automatic unpaired shape deformation transferabstractTransferring deformation from a source shape to a target shape is a very useful technique in computer graphics. State-of-the-art deformation transfer methods require either point-wise correspondences between source and target shapes, or pairs of deformed source and target shapes with corresponding deformations. However, in most cases, such correspondences are not available and cannot be reliably established using an automatic algorithm. Therefore, substantial user effort is needed to label the correspondences or to obtain and specify such shape sets. In this work, we propose a novel approach to automatic deformation transfer between two unpaired shape sets without correspondences. 3D deformation is represented in a high-dimensional space. To obtain a more compact and effective representation, two convolutional variational autoencoders are learned to encode source and target shapes to their latent spaces. We exploit a Generative Adversarial Network (GAN) to map deformed source shapes to deformed target shapes, both in the latent spaces, which ensures the obtained shapes from the mapping are indistinguishable from the target shapes. This is still an under-constrained problem, so we further utilize a reverse mapping from target shapes to source shapes and incorporate cycle consistency loss, i.e. applying both mappings should reverse to the input shape. This VAE-Cycle GAN (VC-GAN) architecture is used to build a reliable mapping between shape spaces. Finally, a similarity constraint is employed to ensure the mapping is consistent with visual similarity, achieved by learning a similarity neural network that takes the embedding vectors from the source and target latent spaces and predicts the light field distance between the corresponding shapes. Experimental results show that our fully automatic method is able to obtain high-quality deformation transfer results with unpaired data sets, comparable or better than existing methods where strict correspondences are required. Lin Gao 0004, Jie Yang 0038, Yi-Ling Qiao, Yukun Lai, Paul L. Rosin, Weiwei Xu 0003, Shihong Xia |
ACM Trans. Graph. | 7 |
| 2017 | Data-Driven Shape Interpolation and Morphing EditingabstractAbstract Shape interpolation has many applications in computer graphics such as morphing for computer animation. In this paper, we propose a novel data‐driven mesh interpolation method. We adapt patch‐based linear rotational invariant coordinates to effectively represent deformations of models in a shape collection, and utilize this information to guide the synthesis of interpolated shapes. Unlike previous data‐driven approaches, we use a rotation/translation invariant representation which defines the plausible deformations in a global continuous space. By effectively exploiting the knowledge in the shape space, our method produces realistic interpolation results at interactive rates, outperforming state‐of‐the‐art methods for challenging cases. We further propose a novel approach to interactive editing of shape morphing according to the shape distribution. The user can explore the morphing path and select example models intuitively and adjust the path with simple interactions to edit the morphing sequences. This provides a useful tool to allow users to generate desired morphing with little effort. We demonstrate the effectiveness of our approach using various examples. Lin Gao 0004, Yukun Lai, Shihong Xia |
Comput. Graph. Forum | 4 |
| 2017 | Rigidity controllable as-rigid-as-possible shape deformation
Lin Gao 0004, Yukun Lai, Shihong Xia |
Graph. Model. | 4 |
| 2017 | A Survey on Human Performance Capture and Animation
Shihong Xia, Lin Gao 0004, Yukun Lai, Mingzhe Yuan, Jinxiang Chai |
J. Comput. Sci. Technol. | 1 |
| 2017 | Templateless Non-Rigid Reconstruction and Motion Tracking With a Single RGB-D CameraabstractWe present a novel templateless approach for nonrigid reconstruction and motion tracking using a single RGB-D camera. Without any template prior, our system achieves accurate reconstruction and tracking for considerably deformable objects. To robustly register the input sequence of partial depth scans with dynamic motion, we propose an efficient local-to-global hierarchical optimization framework inspired by the idea of traditional structure-from-motion. Our proposed framework mainly consists of two stages, local nonrigid bundle adjustment and global optimization. To eliminate error accumulation during the nonrigid registration of loop motion sequences, we split the full sequence into several segments and apply local nonrigid bundle adjustment to align each segment locally. Global optimization is then adopted to combine all segments and handle the drift problem through loop-closure constraint. By fitting to the input partial data, a deforming 3D model sequence of dynamic objects is finally generated. Experiments on both synthetic and real test data sets and comparisons with state of the art demonstrate that our approach can handle considerable motions robustly and efficiently, and reconstruct high-quality 3D model sequences without drift. Kangkan Wang, Guofeng Zhang 0001, Shihong Xia |
IEEE Trans. Image Process. | 3 |
| 2017 | Toward accurate real-time marker labeling for live optical motion captureabstractMarker labeling plays an important role in optical motion capture pipeline especially in real-time applications; however, the accuracy of online marker labeling is still unclear. This paper presents a novel accurate real-time online marker labeling algorithm for simultaneously dealing with missing and ghost markers. We first introduce a soft graph matching model that automatically labels the markers by using Hungarian algorithm for finding the global optimal matching. The key idea is to formulate the problem in a combinatorial optimization framework. The objective function minimizes the matching cost, which simultaneously measures the difference of markers in the model and data graphs as well as their local geometrical structures consisting of edge constraints. To achieve high subsequent marker labeling accuracy, which may be influenced by limb occlusions or self-occlusions, we also propose an online high-quality full-body pose reconstruction process to estimate the positions of missing markers. We demonstrate the power of our approach by capturing a wide range of human movements and achieve the state-of-the-art accuracy by comparing against alternative methods and commercial system like VICON. Shihong Xia, Le Su, Xinyu Fei |
Vis. Comput. | 1 |
| 2016 | Efficient and Flexible Deformation Representation for Data-Driven Surface ModelingabstractEffectively characterizing the behavior of deformable objects has wide applicability but remains challenging. We present a new rotation-invariant deformation representation and a novel reconstruction algorithm to accurately reconstruct the positions and local rotations simultaneously. Meshes can be very efficiently reconstructed from our representation by matrix pre-decomposition, while, at the same time, hard or soft constraints can be flexibly specified with only positions of handles needed. Our approach is thus particularly suitable for constrained deformations guided by examples, providing significant benefits over state-of-the-art methods. Based on this, we further propose novel data-driven approaches to mesh deformation and non-rigid registration of deformable objects. Both problems are formulated consistently as finding an optimized model in the shape space that satisfies boundary constraints, either specified by the user, or according to the scan. By effectively exploiting the knowledge in the shape space, our method produces realistic deformation results in real-time and produces high quality registrations from a template model to a single noisy scan captured using a low-quality depth camera, outperforming state-of-the-art methods. Lin Gao 0004, Yukun Lai, Dun Liang, Shihong Xia |
ACM Trans. Graph. | 5 |
| 2016 | Data-driven inverse dynamics for human motionabstractInverse dynamics is an important and challenging problem in human motion modeling, synthesis and simulation, as well as in robotics and biomechanics. Previous solutions to inverse dynamics are often noisy and ambiguous particularly when double stances occur. In this paper, we present a novel inverse dynamics method that accurately reconstructs biomechanically valid contact information, including center of pressure, contact forces, torsional torques and internal joint torques from input kinematic human motion data. Our key idea is to apply statistical modeling techniques to a set of preprocessed human kinematic and dynamic motion data captured by a combination of an optical motion capture system, pressure insoles and force plates. We formulate the data-driven inverse dynamics problem in a maximum a posteriori (MAP) framework by estimating the most likely contact information and internal joint torques that are consistent with input kinematic motion data. We construct a low-dimensional data-driven prior model for contact information and internal joint torques to reduce ambiguity of inverse dynamics for human motion. We demonstrate the accuracy of our method on a wide variety of human movements including walking, jumping, running, turning and hopping and achieve state-of-the-art accuracy in our comparison against alternative methods. In addition, we discuss how to extend the data-driven inverse dynamics framework to motion editing, filtering and motion control. Xiaolei Lv, Jinxiang Chai, Shihong Xia |
ACM Trans. Graph. | 3 |
| 2016 | Realtime 3D eye gaze animation using a single RGB cameraabstractThis paper presents the first realtime 3D eye gaze capture method that simultaneously captures the coordinated movement of 3D eye gaze, head poses and facial expression deformation using a single RGB camera. Our key idea is to complement a realtime 3D facial performance capture system with an efficient 3D eye gaze tracker. We start the process by automatically detecting important 2D facial features for each frame. The detected facial features are then used to reconstruct 3D head poses and large-scale facial deformation using multi-linear expression deformation models. Next, we introduce a novel user-independent classification method for extracting iris and pupil pixels in each frame. We formulate the 3D eye gaze tracker in the Maximum A Posterior (MAP) framework, which sequentially infers the most probable state of 3D eye gaze at each frame. The eye gaze tracker could fail when eye blinking occurs. We further introduce an efficient eye close detector to improve the robustness and accuracy of the eye gaze tracker. We have tested our system on both live video streams and the Internet videos, demonstrating its accuracy and robustness under a variety of uncontrolled lighting conditions and overcoming significant differences of races, genders, shapes, poses and expressions across individuals. Congyi Wang, Fuhao Shi, Shihong Xia, Jinxiang Chai |
ACM Trans. Graph. | 3 |
| 2015 | Realtime style transfer for unlabeled heterogeneous human motionabstractThis paper presents a novel solution for realtime generation of stylistic human motion that automatically transforms unlabeled, heterogeneous motion data into new styles. The key idea of our approach is an online learning algorithm that automatically constructs a series of local mixtures of autoregressive models (MAR) to capture the complex relationships between styles of motion. We construct local MAR models on the fly by searching for the closest examples of each input pose in the database. Once the model parameters are estimated from the training data, the model adapts the current pose with simple linear transformations. In addition, we introduce an efficient local regression model to predict the timings of synthesized poses in the output style. We demonstrate the power of our approach by transferring stylistic human motion for a wide variety of actions, including walking, running, punching, kicking, jumping and transitions between those behaviors. Our method achieves superior performance in a comparison against alternative methods. We have also performed experiments to evaluate the generalization ability of our data-driven model as well as the key components of our system. Shihong Xia, Congyi Wang, Jinxiang Chai, Jessica K. Hodgins |
ACM Trans. Graph. | 1 |
| 2014 | Automatic Gait Motion Capture with Missing-Marker FillingsabstractAlthough marker-based optical motion capture has been a useful method for computer animation during the past decades, automatic and robust motion tracking from multiple video sequences is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to track and identify the reconstructed 3D points after image matching process? How to handle the heavy occlusion problem? This paper gives a careful investigation of the above issues. In particular, we propose a novel way to track and identify proper markers, and a new method of filling missing markers by taking account of the human model constraints. Experiments are presented to show its accuracy and robustness. Xiaoming Deng 0001, Shihong Xia, Wenzhong Wang, Liang Chang 0001, Hongan Wang |
ICPR | 2 |
| 2012 | Planning interactive task for intelligent charactersabstractABSTRACT Motion planning is an important problem in character animation and interactive simulation. However, few planning methods have considered domain‐specific knowledge that governs the agent's behaviors, and none of them is capable of planning the interactive task in which the agent interacts with the objects in the virtual environment. This paper presents a novel method to plan the interactive task based on Q‐learning for intelligent characters. The approach can be described as a three‐phase framework: data preprocessing phase, controller learning phase, and motion‐synthesis phase. In the data preprocessing phase, we abstract the motion clips as high‐level behaviors and construct the interactive behavior graph (IBG) to define the interactive capabilities of the agent in terms of interactive features. For the controller training phase, with IBG, Q‐learning algorithm is employed to train the control policy in the discrete domain with interactive features. In the motion‐synthesis phase, the optimal motion sequences can be generated by following the policy to accomplish the interactive task finally. The experimental results demonstrate that the uniform framework can generate reasonable and realistic motion sequences to plan interactive task in complex environment. Copyright © 2012 John Wiley & Sons, Ltd. Dan Zong, Chunpeng Li, Shihong Xia |
Comput. Animat. Virtual Worlds | 3 |
| 2011 | Exploring Non-Linear Relationship of Blendshape Facial AnimationabstractAbstract Human face is a complex biomechanical system and non‐linearity is a remarkable feature of facial expressions. However, in blendshape animation, facial expression space is linearized by regarding linear relationship between blending weights and deformed face geometry. This results in the loss of reality in facial animation. To synthesize more realistic facial animation, aforementioned relationship should be non‐linear to allow the greatest generality and fidelity of facial expressions. Unfortunately, few existing works pay attention to the topic about how to measure the non‐linear relationship. In this paper, we propose an optimization scheme that automatically explores the non‐linear relationship of blendshape facial animation from captured facial expressions. Experiments show that the explored non‐linear relationship is consistent with the non‐linearity of facial expressions soundly and is able to synthesize more realistic facial animation than the linear one. Xuecheng Liu, Shihong Xia, Yiwen Fan |
Comput. Graph. Forum | 2 |
| 2010 | Motion track: Visualizing variations of human motion dataabstractThis paper proposes a novel visualization approach, which can depict the variations between different human motion data. This is achieved by representing the time dimension of each animation sequence with a sequential curve in a locality-preserving reference 2D space, called the motion track representation. The principal advantage of this representation over standard representations of motion capture data - generally either a keyframed timeline or a 2D motion map in its entirety - is that it maps the motion differences along the time dimension into parallel perceptible spatial dimensions but at the same time captures the primary content of the source data. Latent semantic differences that are difficult to be visually distinguished can be clearly displayed, favoring effective summary, clustering, comparison and analysis of motion database. Yueqi Hu, Shuangyuan Wu, Shihong Xia, Jinghua Fu, Wei Chen 0001 |
PacificVis | 3 |
| 2010 | Parallelizing continuum crowdsabstractIn this paper, we present a novel parallelizing method for crowd simulators constructed with a continuum model rather than an agent-based model. The basic idea is to partition a crowded virtual environment into some districts, each of which keeps its own dynamic continuum fields and has several transitional blocks to make individuals keep continuum motion from one district to another. Our method makes continuum models to be parallelizable while preserving their existing superiority of generating smooth motion. Moreover, for most of large-scale applications, our partitioning method effectively simplifies the complexity of simulation. Experiments show that our method has achieved super-linear speedup and could employ more than one hundred worker processors to simulate 1 million people in an area of 672,400m2. Tianlu Mao, Hao Jiang 0013, Jian Li 0057, Shihong Xia |
VRST | 5 |
| 2010 | Continuum crowd simulation in complex environments
Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
Comput. Graph. | 5 |
| 2009 | Crowds flow in complex environmentabstractThis paper presents a hybrid approach based on the continuum model proposed by Treuille et al.. Compared to the original method, our solution is well suited for complex environment. We first present an environment structure and a corresponding discretization scheme that help us to organize and simulate crowds in large-scale scenarios. Second, additional discomforts around obstacles are auto-generated for keeping a certain distance between pedestrians and obstacles which is psychologically plausible, and it could obtain smoother trajectory when people move around many obstacles. Thirdly, we propose a technique for density conversion; the density field is dynamically affected by each individual so that it could be adapted to different grid resolution. The experiment results demonstrate that our hybrid solution can perform plausible crowds flow in complex dynamic environments. Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
CAD/Graphics | 5 |
| 2009 | Evaluating simplified air force models for cloth simulationabstractAny cloth simulation system needs the aerodynamics model to describe the dynamic behavior of the cloth interacting with the air. Different air force model has different simplification treatment with different approximation degree. In this paper, we present a quantitative evaluation method for simplified air force models, and experimentally investigate 5 different air force models which are all commonly used in cloth simulations. The results show that the lift component, which was usually neglected in air force models, actually plays an important role in the simulation. Moreover, when the air field is simplified as the linear global wind field, the air force model containing the linear drag component and simple upward lift component matches the real cloth motion better than the complicated nonlinear ones. Tianlu Mao, Shihong Xia |
CAD/Graphics | 2 |
| 2009 | Learning local models for 2D human motion trackingabstractWe present a novel approach to tracking 2D human motion in uncalibrated monocular videos. Human motion usually exhibits time-varying patterns, and we propose to use locally learnt prior models to capture this characteristics. For each input image, our method automatically learns a local probability density model and a local dynamical model from a set of training examples that are close matches to the input. We evaluate the image likelihood by matching a deformable 2D human body model to the input images. The local models and the image likelihood are integrated to optimize the pose for the current input. Experiments on both synthetic and real videos demonstrate the effectiveness of our method. Wenzhong Wang, Xiaoming Deng 0001, Xianjie Qiu, Shihong Xia |
ICIP | 4 |
| 2009 | A semantic environment model for crowd simulation in multilayered complex environmentabstractSimulating crowds in complex environment is fascinating and challenging, however, modeling of the environment is always neglected in the past, which is one of the essential problems in crowd simulation especially for multilayered complex environment. This paper presents a semantic model for representing the complex environment, where the semantic information is described with a three-tier framework: a geometric level, a semantic level and an application level. Each level contains different maps for different purposes and our approach greatly facilitates the interactions between individuals and virtual environment. And then a modified continuum crowd method is designed to fit the proposed virtual environment model so that realistic behaviors of large dense crowds could be simulated in multilayered complex environments such as buildings and subway stations. Finally, we implement this method and test it in two complex synthetic urban spaces. The experiment results demonstrate that the semantic environment model can provide sufficient and accurate information for crowd simulation in multilayered complex environment. Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
VRST | 5 |
| 2009 | Indexing and retrieval of human motion data by a hierarchical treeabstractFor the convenient reuse of large-scale 3D motion capture data, browsing and searching methods for the data should be explored. In this paper, an efficient indexing and retrieval approach for human motion data is presented based on a novel similarity metric. We divide the human character model into three partitions to reduce the spatial complexity and measure the temporal similarity of each partition by self-organizing map and Smith--Waterman algorithm. The overall similarity between two motion clips can be achieved by integrating the similarities of the separate body partitions. Then the hierarchical clustering method is implemented, which can not only cluster the motion data accurately, but also discover the relationships between different motion types by a binary tree structure. With our typical cluster locating algorithm and motion motif mining method, fast and accurate retrieval can be performed. The experiment results show the effectiveness of our approach. Shuangyuan Wu, Shihong Xia |
VRST | 3 |
| 2009 | Efficient motion data indexing and retrieval with local similarity measure of motion strings
Shuangyuan Wu, Shihong Xia, Chunpeng Li |
Vis. Comput. | 2 |
| 2008 | Facial animation by optimized blendshapes from motion capture dataabstractAbstract This paper presents a labor‐saving method to construct optimal facial animation blendshapes from given blendshape sketches and facial motion capture data. At first, a mapping function is established between target “Marker Face” and performer's face by RBF interpolating selected feature points. Sketched blendshapes are transferred to performer's “Marker Face” by using motion vector adjustment technique. Then, the blendshapes of performer's “Marker Face” are optimized according to the facial motion capture data. At last, the optimized blendshapes are inversely transferred to target facial model. Apart from that, this paper also proposes a method of computing blendshape weights from facial motion capture data more accurately. Experiments show that expressive facial animation can be acquired. Copyright © 2008 John Wiley & Sons, Ltd. Xuecheng Liu, Tianlu Mao, Shihong Xia |
Comput. Animat. Virtual Worlds | 3 |
| 2007 | Articulated Skeleton FittingabstractAs it widely exists in applications as diverse as biomechanics, motion capture and robot control, the fundamental problem of articulated skeleton fitting, which is to fit an articulated skeleton to k frames of corresponding raw joint points under the fixed-bone-length constraint, is proposed in this paper. With different parametrization of the articulated skeleton, three types of potential solutions to the above problem are presented and compared. Special cases for this problem, including that generated under joint trajectory smoothness constraints and from weighted joints, are also discussed and analyzed. Based on the joint location parametrization and the sparse matrix techniques, we propose an efficient and accurate articulated skeleton fitting method (JLP2). Elaborate experiments and interesting applications in motion capture are presented to show its effectiveness. Gaojin Wen, Shihong Xia, Chunpeng Li |
CAD/Graphics | 3 |
| 2007 | Pose Synthesis Using the Inverse of Jacobian Matrix Learned from ExamplesabstractThis paper presents a method of pose synthesis based on a low-dimensional space and a set of characteristics of motion learned from examples. This method consists of two phases: learning and synthesis. In the learning phase, a low-dimensional and discrete representation of the space of natural poses is constructed by using a self organizing map (SOM). Meanwhile, a set of matrices is extracted from the motion data. These matrices describe how the poses change with the end-effectors' positions, and play a key role in synthesizing natural looking results. In the synthesis phase, a lightweight algorithm based on the learned parameters is used. The synthesis process is very efficient because there is no time-consuming calculation, like numeric optimization or matrix inverting. Compared with other methods, our method not only can produce natural looking poses in real-time, but also works well with constraints positioned in a larger range. We apply our method in applications of interactive pose editing, real-time motion modification, and pose reconstruction from image. The results have proven the robustness and effectiveness of our method Chunpeng Li, Shihong Xia |
VR | 2 |
| 2007 | CrowdViewer: from simple script to large-scale virtual crowdsabstractVisualization of large-scale virtual crowds is ubiquitous in many applications of computer graphics. For reasons of efficiency in modeling, animating and rendering, it is difficult to populate scenes with a large number of individually animated virtual characters in real-time applications. In this paper, we present an effective and readily usable solution to this problem. It accepts simple script which includes motion state and position information of each individual at each time step. Supported by material database and motion database, various human models are generated from model templates, and then driven by an agile on-line animation approach. A developed point-based rendering approach is presented to accelerate rendering. We test our system with script including 30,000 people evacuating from a sports arena. The results demonstrate that our approach provides a very effective way to visualize large-scale crowds with high visual realism in real-time. Tianlu Mao, Bo Shu, Shihong Xia |
VRST | 4 |
| 2006 | A Video-Driven Approach to Continuous Human Motion Synthesis
Xianjie Qiu, Shihong Xia |
Computer Graphics International | 4 |
| 2006 | Motion Editing with the State Feedback Dynamic Model
Dengming Zhu, Shihong Xia |
Computer Graphics International | 3 |
| 2006 | Motion synthesis in motion reconstruction based on videoabstractWe propose a framework to reconstruct human motion based on monocular camera video and motion database. In this framework, we use silhouettes for motion estimation based on a set of discriminative features and search motion database to find out the exact motion clips that meet with the video content. To eliminate the discontinuities between motion clips, we adopt a seamless motion stitch method using multiresolution analysis technique. We verify the effectiveness of our method by reconstructing trampoline sports video as an example. The reconstruction results are visually comparable to those motions obtained by a commercial motion capture system in the premise that similar motions are included in the motion database. Xianjie Qiu, Shihong Xia |
MMM | 4 |
| 2006 | A robust method for analyzing the physical correctness of motion capture dataabstractThe physical correctness of motion capture data is important for human motion analysis and athlete training. However, until now there is little work that wholly explores this problem of analyzing the physical correctness of motion capture data. In this paper, we carefully discuss this problem and solve two major issues in it. Firstly, a new form of Newton-Euler equations encoded by quaternions and Euler angles which are very fit for analyzing the motion capture data are proposed. Secondly, a robust optimization method is proposed to correct the motion capture data to satisfy the physical constraints. We demonstrate the advantage of our method with several experiments. Shihong Xia, Dengming Zhu |
VRST | 2 |
| 2006 | From motion capture data to character animationabstractIn this paper, we propose a practical and systematical solution to the mapping problem that is from 3D marker position data recorded by optical motion capture systems to joint trajectories together with a matching skeleton based on least-squares fitting techniques. First, we preprocess the raw data and estimate the joint centers based on related efficient techniques. Second, a skeleton of fixed length which precisely matching the joint centers are generated by an articulated skeleton fitting method. Finally, we calculate and rectify joint angles with a minimum angle modification technique. We present the results for our approach as applied to several motion-capture behaviors, which demonstrates the positional accuracy and usefulness of our method. Gaojin Wen, Shihong Xia, Dengming Zhu |
VRST | 3 |
| 2006 | Least-squares fitting of multiple M-dimensional point sets
Gaojin Wen, Shihong Xia, Dengming Zhu |
Vis. Comput. | 3 |
| 2005 | Total least squares fitting of point sets in m-DabstractThe absolute orientation technique, minimizing the mean squared error between two matched point sets under similarity transformations, has numerously applied in the areas of photogrammetry, robotics, object motion analysis as well as object pose estimation following recognition. Based on it, in this paper, a total least squares fitting algorithm, which generates a fixed point set from k corresponding original point sets and minimizes the mean squared error between the fixed point sets and these k point sets, is proposed and proved. Experiments and interesting applications are also presented to show its efficiency, accuracy and robustness. Gaojin Wen, Dengming Zhu, Shihong Xia |
Computer Graphics International | 3 |
| 2005 | Inferring 3D Body Pose from Uncalibrated Video
Xianjie Qiu, Shihong Xia, Yong-Chao Sun |
RoboCup | 3 |
| 2005 | A novel framework for athlete training based on interactive motion editing and silhouette analysisabstractThere are mainly two Hi-Tech methods for athlete training. One method is based on virtual reality, where the athlete can learn and improve performance mainly through using virtual equipments to interact with the virtual environment. Another method is based on video analysis, where improvements can be made by comparing the videos of the trainees with those of excellent trainers. In this paper, we present a novel framework for athlete training, which can circumvent difficulties the current methods faced in practical applications. For retargeting the example motion to personalized virtual athlete, the coach interactively sets motion constraints with his experience based on motion warping and motion verification techniques. The display of the simulated motion is adjusted semi-automatically to create the reference virtual video with the same viewpoint as the real one. The moment invariants of both virtual and real athlete's silhouette are computed, and motion analysis result is presented subsequently. This method is more suitable for gymnastic athlete training because of without virtual equipment and more instructive having the same viewpoint in video analysis. Finally, an application of the proposed techniques to trampoline training is implemented. Shihong Xia, Xianjie Qiu |
VRST | 1 |