EDBT 2026 Demo / reviewers in the wild / expert
Xiaosong Yang
dblp:14/5743
· DBLP profile ↗
89ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0003-3815-0584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 4 first-author · 24 since 2021Artificial intelligence and machine learning · 13 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speaker-Invariant Emotion Representations with Gradient Reversal
Kavisha Jayathunge, Xiaosong Yang |
ICPR (7) | 2 |
| 2026 | Capability of large language models in assisting GPs with diagnoses
Ruibin Wang, Abdul Rehman 0007, Rupert Page, Hailing Li, Xiaokun Wang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Appl. Intell. | 7 |
| 2025 | Motion Style Transfer: Methods, Challenges, and Future Directions
Siyao Du, Boyuan Cheng, Xiaosong Yang |
CASA | 5 |
| 2025 | Unsupervised Salient Object Detection with Pseudo-Labels Refinement
Yanfeng Zheng, Pengjie Wang 0001, Xiaosong Yang |
CASA | 4 |
| 2025 | PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
Yingjie Xi, Jian J. Zhang 0001, Xiaosong Yang |
ACM Multimedia | 3 |
| 2025 | Automating visual narratives: Learning cinematic camera perspectives from 3D human interactionabstractCinematic camera control is essential for guiding audience attention and conveying narrative intent, yet current data-driven methods largely rely on predefined visual datasets and handcrafted rules, limiting generalization and creativity. This paper introduces a novel diffusion-based framework that generates camera trajectories directly from two-character 3D motion sequences, eliminating the need for paired video–camera annotations. The approach leverages Toric features to encode spatial relations between characters and conditions the diffusion process through a dual-stream motion encoder and interaction module, enabling the camera to adapt dynamically to evolving character interactions. A new dataset linking character motion with camera parameters is constructed to train and evaluate the model. Experiments demonstrate that our method outperforms strong baselines in both quantitative metrics and perceptual quality, producing camera motions that are smooth, temporally coherent, and compositionally consistent with cinematic conventions. This work opens new opportunities for automating virtual cinematography in animation, gaming, and interactive media. Boyuan Cheng, Shang Ni, Xiaosong Yang |
Comput. Graph. | 4 |
| 2025 | Enhanced collapsible linear blocks for arbitrary sized image super-resolutionabstractAbstract Image up-scaling and super-resolution (SR) techniques have been a hot research topic for many years due to its large impact in the field of medical imaging, surveillance etc. Especially single image super-resolution (SISR) become very popular because of the fast development of deep convolution neural network (DCNN) and the low requirement on the input. They are achieving outstanding performance. However, there are still problems in the state-of-the-art works, especially from two perspectives: 1. failed at exploiting the hierarchical characteristics from the input, resulting in loss of information and artifacts in the final high resolution (HR) image; 2. failed to handle arbitrary-sized images; the existing research works are focused on fixed size input images. To address these challenges, this paper proposed a residual dense network (RDN) and multi-scale sub-pixel convolution network (MSSPCN) which are integrated into a Collapsible Linear Block Super Efficient Super-Resolution (SESR) network. The RDNs aims to tackle the first challenge, carrying the hierarchical features from end-to-end. An adaptive cropping strategy (ACS) technique is introduced before feature extraction targeting at the image size challenge. The novelty of this work is extracting the hierarchical features and integrating RDNs with MSSPCNs. The proposed network can upscale any arbitrary-sized image (1080p) to ×2 (4K) and ×4 (8K). To secure ground truth for evaluation, this paper follows the opposite flow, generating the input LR images by down-sampling the given HR images (ground truth). To evaluate the performance, the proposed algorithm is compared with eight state-of-the-art algorithms, both quantitatively and qualitatively. The results are verified on six benchmark datasets. The extensive experiments justify that the proposed architecture performs better than other methods and upscales the images satisfactorily. Prathap Soma, Xiaosong Yang, Jian Chang 0001, Jian J. Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2025 | FaTNET: Feature-alignment transformer network for human pose transfer
Chengzhi Yuan, Lin Gao 0004, Weiwei Xu 0003, Xiaosong Yang, Pengjie Wang 0001 |
Pattern Recognit. | 5 |
| 2025 | Unsupervised Salient Object Detection on Light Field With High-Quality Synthetic LabelsabstractMost current Light Field Salient Object Detection (LFSOD) methods require full supervision with labor-intensive pixel-level annotations. Unsupervised Light Field Salient Object Detection (ULFSOD) has gained attention due to this limitation. However, existing methods use traditional handcrafted techniques to generate noisy pseudo-labels, which degrades the performance of models trained on them. To mitigate this issue, we present a novel learning-based approach to synthesize labels for ULFSOD. We introduce a prominent focal stack identification module that utilizes light field information (focal stack, depth map, and RGB color image) to generate high-quality pixel-level pseudo-labels, aiding network training. Additionally, we propose a novel model architecture for LFSOD, combining a multi-scale spatial attention module for focal stack information with a cross fusion module for RGB and focal stack integration. Through extensive experiments, we demonstrate that our pseudo-label generation method significantly outperforms existing methods in label quality. Our proposed model, trained with our labels, shows significant improvement on ULFSOD, achieving new state-of-the-art scores across public benchmarks. Yanfeng Zheng, Zhong Luo, Ying Cao 0001, Xiaosong Yang, Weiwei Xu 0003, Zheng Lin 0005, Pengjie Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior ModelingabstractLong-term user behavior sequences are a goldmine for businesses to explore users’ interests to improve Click-Through Rate (CTR). However, it is very challenging to accurately capture users’ long-term interests from their long-term behavior sequences and give quick responses from the online serving systems. To meet such requirements, existing methods “inadvertently” destroy two basic requirements in long-term sequence modeling:R1) make full use of the entire sequence to keep the information as much as possible;R2) extract information from the most relevant behaviors to keep high relevance between learned interests and current target items. The performance of online serving systems is significantly affected by incomplete and inaccurate user interest information obtained by existing methods. To this end, we propose an efficient two-stage long-term sequence modeling approach, named asEfficieNtClustering based twO-stage interest moDEling (ENCODE), consisting of offline extraction stage and online inference stage. It not only meets the aforementioned two basic requirements but also achieves a desirable balance between online service efficiency and precision. Specifically, in the offline extraction stage, ENCODE clusters the entire behavior sequence and extracts accurate interests. To reduce the overhead of the clustering process, we design a metric learning-based dimension reduction algorithm that preserves the relative pairwise distances of behaviors in the new feature space. While in the online inference stage, ENCODE takes the off-the-shelf user interests to predict the associations with target items. Besides, to further ensure the relevance between user interests and target items, we adopt the same relevance metric throughout the whole pipeline of ENCODE. The extensive experiment and comparison with SOTA on both industrial and public datasets have demonstrated the effectiveness and efficiency of our proposed ENCODE. Yuhang Zheng 0003, Yinfu Feng, Yunan Ye, Rong Xiao 0005, Long Chen 0016, Xiaosong Yang, Jun Xiao 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | Taming High-Resolution Auxiliary G-Buffers for Deep Supersampling of Rendered ContentabstractHigh-resolution images come with rich color information and texture details. Due to the rapid upgrading of display devices and rendering technologies, high-resolution real-time rendering faces the computational overhead challenge. To address this, the current mainstream solution is to render at a lower resolution and then upsample to the target resolution by supersampling techniques. However, while many prior supersampling approaches have attempted to exploit rich rendered data such as color, depth, motion vectors at low resolution, there is little discussion on how to harness high-frequency information that is readily available in the high-resolution (HR) G-buffers of modern renders. In this article, we seek to investigate how to fully leverage information from HR G-buffers to maximize the visual quality of supersampling results. We propose a neural network for real-time supersampling of rendered content, which is based on several core designs, including gated G-buffers encoder, G-buffers attended encoder and reflection-aware loss. These designs are especially made for the sake of effectively using HR G-buffers, enabling faithful recovery of a variety of high-frequency scene details from low-resolution, highly aliased inputs. Furthermore, a simple occlusion-aware blender is proposed to efficiently rectify dis-occluded features in the warped previous frame, allowing us to better exploit history information to improve temporal stability. The experiments show that our method, equipped with strong ability to harness HR G-buffer information, significantly improves the visual fidelity of high-resolution reconstructions upon previous state-of-the-art methods, even for challenging $4 \times 4$4×4 upsampling, while still being compute-efficient. Pengjie Wang 0001, Chengzhi Yuan, Jie Guo 0001, Xiaosong Yang, Houjie Li, Ian Stephenson, Jian Chang 0001, Ying Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Existence Is Chaos: Enhancing 3D Human Motion Prediction with Uncertainty ConsiderationabstractHuman motion prediction is consisting in forecasting future body poses from historically observed sequences. It is a longstanding challenge due to motion's complex dynamics and uncertainty. Existing methods focus on building up complicated neural networks to model the motion dynamics. The predicted results are required to be strictly similar to the training samples with L2 loss in current training pipeline. However, little attention has been paid to the uncertainty property which is crucial to the prediction task. We argue that the recorded motion in training data could be an observation of possible future, rather than a predetermined result. In addition, existing works calculate the predicted error on each future frame equally during training, while recent work indicated that different frames could play different roles. In this work, a novel computationally efficient encoder-decoder model with uncertainty consideration is proposed, which could learn proper characteristics for future frames by a dynamic function. Experimental results on benchmark datasets demonstrate that our uncertainty consideration approach has obvious advantages both in quantity and quality. Moreover, the proposed method could produce motion sequences with much better quality that avoids the intractable shaking artefacts. We believe our work could provide a novel perspective to consider the uncertainty quality for the general motion prediction task and encourage the studies in this field. The code will be available in https://github.com/Motionpre/Adaptive-Salient-Loss-SAGGB. Ningyu Zhang 0001, Xiaosong Yang, Jun Xiao 0001 |
AAAI | 4 |
| 2024 | Survey on Multi-person 3D Reconstruction from Monocular View
Jingyao Cai, Boyuan Cheng, Yingjie Xi, Xiaosong Yang |
CGI (2) | 4 |
| 2024 | GAMAFlow: Estimating 3D Scene Flow via Grouped Attention and Global Motion AggregationabstractThe estimation of 3D motion fields, known as scene flow estimation, is an essential task in autonomous driving and robotic navigation. Existing learning-based methods either predict scene flow through flow-embedding layers or rely on local search methods to establish soft correspondences. However, these methods often neglect distant points which, in fact, represent the true matching elements. To address this challenge, we introduce GAMAFlow, a point-voxel architecture that models local motion and global motion to predict scene flow iteratively. In particular, GAMAFlow integrates the advantages of (i) the point Transformer with Grouped Attention and (ii) global Motion Aggregation to boost the efficacy of point-voxel correlation. Such an approach facilitates learning long-distance dependencies between current frame and next frame. Experiments illustrate the performance gains achieved by GAMAFlow compared to existing works on both FlyingThings3D and KITTI benchmarks. Zhiqi Li 0002, Xiaosong Yang, Jian J. Zhang 0001 |
ICASSP | 2 |
| 2024 | Forecasting Distillation: Enhancing 3D Human Motion Prediction with Guidance RegularizationabstractHuman motion prediction aims to forecast future body poses from historically observed sequences, which is challenging due to motion’s complex dynamics. Existing methods mainly focus on dedicated network structures to model the spatial and temporal dependencies. The predicted results are required to be strictly similar to the training samples with ℓ2loss in the current training pipeline. It needs to be pointed out that most approaches predict the next frame conditioned on the previously predicted sequence, where a small error in the initial frame could be accumulated significantly. In addition, recent work indicated that different stages could play different roles. Hence, this paper considers a new direction by introducing a model learning framework with motion guidance regularization to reduce uncertainty. The guidance information is extracted from a designed Fusion Feature Extraction network (FE-Net) while knowledge distilling is conducted through intermediate supervision to improve the multi-stage prediction network during training. Incorporated with baseline models, our guidance design exhibits clear performance gains in terms of 3D mean per joint position error (MPJPE) on benchmark datasets Human3.6M, CMU Mocap, and 3DPW datasets, respectively. Related code will be available on https://github.com/tempAnonymous2024/MotionPredict-GuidanceReg. Yawen Du, Yinmin Li, Xiaosong Yang |
IJCNN | 4 |
| 2024 | Harmony Everything! Masked Autoencoders for Video Harmonization
Yuhang Li 0011, Jincen Jiang, Xiaosong Yang, Youdong Ding, Jian J. Zhang 0001 |
ACM Multimedia | 3 |
| 2024 | Multi-colour sketch-based image retrieval with an explicable feature embedding
Shuangbu Wang, Yu Xia 0012, Kun Qian 0009, Xiaosong Yang, Lihua You, Jian J. Zhang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | FrseGAN: Free-style editable facial makeup transfer based on GAN combined with transformerabstractAbstract Makeup in real life varies widely and is personalized, presenting a key challenge in makeup transfer. Most previous makeup transfer techniques divide the face into distinct regions for color transfer, frequently neglecting details like eyeshadow and facial contours. Given the successful advancements of Transformers in various visual tasks, we believe that this technology holds large potential in addressing pose, expression, and occlusion differences. To explore this, we propose novel pipeline which combines well‐designed Convolutional Neural Network with Transformer to leverage the advantages of both networks for high‐quality facial makeup transfer. This enables hierarchical extraction of both local and global facial features, facilitating the encoding of facial attributes into pyramid feature maps. Furthermore, a Low‐Frequency Information Fusion Module is proposed to address the problem of large pose and expression variations which exist between the source and reference faces by extracting makeup features from the reference and adapting them to the source. Experiments demonstrate that our method produces makeup faces that are visually more detailed and realistic, yielding superior results. Pengjie Wang 0001, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 3 |
| 2024 | TMSDNet: Transformer with multi-scale dense network for single and multi-view 3D reconstructionabstractAbstract 3D reconstruction is a long‐standing problem. Recently, a number of studies have emerged that utilize transformers for 3D reconstruction, and these approaches have demonstrated strong performance. However, transformer‐based 3D reconstruction methods tend to establish the transformation relationship between the 2D image and the 3D voxel space directly using transformers or rely solely on the powerful feature extraction capabilities of transformers. They ignore the crucial role played by deep multi‐scale representation of the object in the voxel feature domain, which can provide extensive global shape and local detail information about the object in a multi‐scale manner. In this article, we propose a novel framework TMSDNet (transformer with multi‐scale dense network) for single‐view and multi‐view 3D reconstruction with transformer to solve this problem. Based on our well‐designed combined‐transformer Block, which is canonical encoder–decoder architecture, voxel features with spatial order can be extracted from the input image, which are used to further extract multi‐scale global features in parallel using a multi‐scale residual attention module. Furthermore, a residual dense attention block is introduced for deep local features extraction and adaptive fusion. Finally, the reconstructed objects are produced with the voxel reconstruction block. Experiment results on the benchmarks such as ShapeNet and Pix3D datasets demonstrate that TMSDNet outperforms the existing state‐of‐the‐art reconstruction methods substantially. Xiaoqiang Zhu, Xinsheng Yao, Junjie Zhang 0002, Lihua You, Xiaosong Yang, Jian J. Zhang 0001, Dan Zeng 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2024 | DFIE3D: 3D-Aware Disentangled Face Inversion and Editing via Facial-Contrastive LearningabstractRecent advances in NeRF-based 3D-aware GANs have achieved outstanding performance, especially in the realm of human facial representations, making projection of facial images back into their latent space superior and preferable compared to 2D GAN inversion. However, the direct application of 2DGAN inversion techniques to 3DGAN raises challenges due to potential appearance distortions and geometric inconsistences. To tackle these issues, this work presents a novel integrated framework that combines a composite inversion pipeline in both the SS and W+ spaces and integrates a contrastive-based training strategy, ensuring proficient disentanglement within the module. Moreover, we design a facial semantic manipulation technique based on dimensional analysis of the latent code, which is fully compatible with the proposed 3DGAN inversion pipeline. Comprehensive experimental validations substantiate the effectiveness of the proposed approach in executing 3d-aware face inversion and semantic editing tasks, presenting a robust technological solution for a diverse array of digital human modeling applications in the downstream. Xiaoqiang Zhu, Lihua You, Xiaosong Yang, Jian Chang 0001, Jian J. Zhang 0001, Dan Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Few-shot anime pose transferabstractAbstract In this paper, we propose a few-shot method for pose transfer of anime characters—given a source image of an anime character and a target pose, we transfer the pose of the target to the source character. Despite recent advances in pose transfer on real people images, these methods typically require large numbers of training images of different person under different poses to achieve reasonable results. However, anime character images are expensive to obtain they are created with a lot of artistic authoring. To address this, we propose a meta-learning framework for few-shot pose transfer, which can well generalize to an unseen character given just a few examples of the character. Further, we propose fusion residual blocks to align the features of the source and target so that the appearance of the source character can be well transferred to the target pose. Experiments show that our method outperforms leading pose transfer methods, especially when the source characters are not in the training set. Pengjie Wang 0001, Chengzhi Yuan, Houjie Li, Wen Tang 0004, Xiaosong Yang |
Vis. Comput. | 6 |
| 2023 | Deep Learning for Scene Flow Estimation on Point Clouds: A Survey and Prospective TrendsabstractAbstract Aiming at obtaining structural information and 3D motion of dynamic scenes, scene flow estimation has been an interest of research in computer vision and computer graphics for a long time. It is also a fundamental task for various applications such as autonomous driving. Compared to previous methods that utilize image representations, many recent researches build upon the power of deep analysis and focus on point clouds representation to conduct 3D flow estimation. This paper comprehensively reviews the pioneering literature in scene flow estimation based on point clouds. Meanwhile, it delves into detail in learning paradigms and presents insightful comparisons between the state‐of‐the‐art methods using deep learning for scene flow estimation. Furthermore, this paper investigates various higher‐level scene understanding tasks, including object tracking, motion segmentation, etc. and concludes with an overview of foreseeable research trends for scene flow estimation. Zhiqi Li 0002, Honghua Chen, Jian J. Zhang 0001, Xiaosong Yang |
Comput. Graph. Forum | 5 |
| 2023 | PerimetryNet: A multiscale fine grained deep network for three-dimensional eye gaze estimation using visual field analysisabstractAbstract Three‐dimensional gaze estimation aims to reveal where a person is looking, which plays an important role in identifying users' point‐of‐interest in terms of the direction, attention and interactions. Appearance‐based gaze estimation methods could provide relatively unconstrained gaze tracking from commodity hardware. Inspired by medical perimetry test, we have proposed a multiscale framework with visual field analysis branch to improve estimation accuracy. The model is based on the feature pyramids and predicts vision field to help gaze estimation. In particular, we analysis the effect of the multiscale component and the visual field branch on challenging benchmark datasets: MPIIGaze and EYEDIAP. Based on these studies, our proposed PerimetryNet significantly outperforms state‐of‐the‐art methods. In addition, the multiscale mechanism and visual field branch can be easily applied to existing network architecture for gaze estimation. Related code would be available at public repository https://github.com/gazeEs/PerimetryNet . Shuqing Yu, Shuowen Zhou, Xiaosong Yang, Chao Wu 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2023 | Spatiotemporal Learning Transformer for Video-Based Human Pose EstimationabstractMulti-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance. Di Gai, Runyang Feng, Weidong Min, Xiaosong Yang, Pengxiang Su, Qi Wang 0061 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A mixed reality framework for microsurgery simulation with visual-tactile perception
Hai-Ning Liang, Lingyun Yu 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Vis. Comput. | 4 |
| 2022 | Shifting Perspective to See Difference: A Novel Multi-view Method for Skeleton based Action RecognitionabstractSkeleton-based human action recognition is a longstanding challenge due to its complex dynamics. Some fine-grain details of the dynamics play a vital role in classification. The existing work largely focuses on designing incremental neural networks with more complicated adjacent matrices to capture the details of joints relationships. However, they still have difficulties distinguishing actions that have broadly similar motion patterns but belong to different categories. Interestingly, we found that the subtle differences in motion patterns can be significantly amplified and become easy for audience to distinct through specified view directions, where this property haven't been fully explored before. Drastically different from previous work, we boost the performance by proposing a conceptually simple yet effective Multi-view strategy that recognizes actions from a collection of dynamic view features. Specifically, we design a novel Skeleton-Anchor Proposal (SAP) module which contains a Multi-head structure to learn a set of views. For feature learning of different views, we introduce a novel Angle Representation to transform the actions under different views and feed the transformations into the baseline model. Our module can work seamlessly with the existing action classification model. Incorporated with baseline models, our SAP module exhibits clear performance gains on many challenging benchmarks. Moreover, comprehensive experiments show that our model consistently beats down the state-of-the-art and remains effective and robust especially when dealing with corrupted data. Related code will be available on https://github.com/ideal-idea/SAP Ruijie Hou, Yanran Li, Ningyu Zhang 0001, Xiaosong Yang |
ACM Multimedia | 5 |
| 2021 | TsFPS: An Accurate and Flexible 6DoF Tracking System with Fiducial Platonic SolidsabstractWe present a vision-based system for real-time pose tracking of the rigid object, it can not only estimate a single pose in six degrees of freedom (6DoF), but also suitable for recovering compound movements. The system is comprised of a monocular camera, and a series of 3D printed platonic solids with squared fiducial markers attached on each single face, which is easy to setup and extend, extra cameras are allowed to incorporate into the pipeline for meeting different requirements. The system realizes object tracking by estimating the pose of the fiducial platonic solid (FPS) which can be fixed onto the surface of the target object. Different sizes and shapes of the platonic solids are allowed to combine with each other to adapt to different application scenarios, this strategy provides enormous flexibility and applicability to our system. In order to track the motion of the fiducial platonic solid accurately, a robust algorithm that combines the fiducial constraint and the statistical constraint is introduced, which is able to handle illumination changes, motion blur and partial occlusion. We evaluate the performance of the proposed approach with qualitative and quantitative experiments, in addition, a couple of mixed reality (MR) applications are developed for demonstrating the effectiveness of the system. Xiaosong Yang, Jian J. Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Bas-relief modelling from enriched detail and geometry with deep normal transfer
Meili Wang 0001, Li Wang 0105, Tao Jiang 0020, Juncong Lin, Mingqiang Wei, Xiaosong Yang, Taku Komura, Jian J. Zhang 0001 |
Neurocomputing | 7 |
| 2021 | KeyFrame extraction for human motion capture data via multiple binomial fittingabstractAbstract In this paper, we make two contributions. The first is to propose a new keyframe extraction algorithm, which reduces the keyframe redundancy and reduces the motion sequence reconstruction error. Secondly, a new motion sequence reconstruction method is proposed, which further reduces the error of motion sequence reconstruction. Specifically, we treated the input motion sequence as curves, then the binomial fitting was extended to obtain the points where the slope changes dramatically in the vicinity. Then we took these points as inputs to obtain keyframes by density clustering. Finally, the motion curves were segmented by keyframes and the segmented curves were fitted by binomial formula again to obtain the binomial parameters for motion reconstruction. Experiments show that our methods outperform existing techniques, in terms of reconstruction error. Chenxu Xu, Yanran Li, Xuequan Lu, Meili Wang 0001, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 6 |
| 2021 | Driver Yawning Detection Based on Subtle Facial Action RecognitionabstractVarious investigations have shown that driver fatigue is the main cause of traffic accidents. Research on the use of computer vision techniques to detect signs of fatigue from facial actions, such as yawning, has demonstrated good potential. However, accurate and robust detection of yawning is difficult because of the complicated facial actions and expressions of drivers in the real driving environment. Several facial actions and expressions have the same mouth deformation as yawning. Thus, a novel approach to detecting yawning based on subtle facial action recognition is proposed in this study to alleviate the abovementioned problems. A 3D deep learning network with a low time sampling characteristic is proposed for subtle facial action recognition. This network uses 3D convolutional and bidirectional long short-term memory networks for spatiotemporal feature extraction and adopts SoftMax for classification. A keyframe selection algorithm is designed to select the most representative frame sequence from subtle facial actions. This algorithm rapidly eliminates redundant frames using image histograms with low computation cost and detects outliers by median absolute deviation. A series of experiments are also conducted on YawDD benchmark and self-collected datasets. Compared with several state-of-the-art methods, the proposed method has high yawning detection rates and can effectively distinguish yawning from similar facial actions. Hao Yang 0027, Li Liu 0010, Weidong Min, Xiaosong Yang, Xin Xiong 0016 |
IEEE Trans. Multim. | 4 |
| 2020 | Symmetric Dilated Convolution for Surgical Gesture Recognition
Jinglu Zhang, Yinyu Nie, Yao Lyu, Hailin Li, Jian Chang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
MICCAI (3) | 6 |
| 2020 | Hybrid features for skeleton-based action recognition based on network fusionabstractAbstract In recent years, the topic of skeleton‐based human action recognition has attracted significant attention from researchers and practitioners in graphics, vision, animation, and virtual environments. The most fundamental issue is how to learn an effective and accurate representation from spatiotemporal action sequences towards improved performance, and this article aims to address the aforementioned challenge. In particular, we design a novel method of hybrid features' extraction based on the construction of multistream networks and their organic fusion. First, we train a convolution neural networks (CNN) model to learn CNN‐based features with the raw skeleton coordinates and their temporal differences serving as input signals. The attention mechanism is injected into the CNN model to weigh more effective and important information. Then, we employ long short‐term memory (LSTM) to obtain long‐term temporal features from action sequences. Finally, we generate the hybrid features by fusing the CNN and LSTM networks, and we classify action types with the hybrid features. The extensive experiments are performed on several large‐scale publically available databases, and promising results demonstrate the efficacy and effectiveness of our proposed framework. Zhangmeng Chen, JunJun Pan, Xiaosong Yang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | Densely connected GCN model for motion predictionabstractAbstract Human motion prediction is a fundamental problem in understanding human natural movements. This task is very challenging due to the complex human body constraints and diversity of action types. Due to the human body being a natural graph, graph convolutional network (GCN)‐based models perform better than the traditional recurrent neural network (RNN)‐based models on modeling the natural spatial and temporal dependencies lying in the motion data. In this paper, we develop the GCN‐based models further by adding densely connected links to increase their feature utilizations and address oversmoothing problem. More specifically, the GCN block is used to learn the spatial relationships between the nodes and each feature map of the GCN block propagates directly to every following block as input rather than residual linked. In this way, the spatial dependency of human motion data is exploited more sufficiently and the features of different level of scale are fused more efficiently. Extensive experiments demonstrate our model achieving the state‐of‐the‐art results on CMU dataset. Yanran Li, Lingteng Qiu, Li Wang 0105, Fangde Liu, Sebastian Iulian Poiana, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 7 |
| 2020 | Sketch-based modeling with a differentiable rendererabstractAbstract Sketch‐based modeling aims to recover three‐dimensional (3D) shape from two‐dimensional line drawings. However, due to the sparsity and ambiguity of the sketch, it is extremely challenging for computers to interpret line drawings of physical objects. Most conventional systems are restricted to specific scenarios such as recovering for specific shapes, which are not conducive to generalize. Recent progress of deep learning methods have sparked new ideas for solving computer vision and pattern recognition issues. In this work, we present an end‐to‐end learning framework to predict 3D shape from line drawings. Our approach is based on a two‐steps strategy, it converts the sketch image to its normal image, then recover the 3D shape subsequently. A differentiable renderer is proposed and incorporated into this framework, it allows the integration of the rendering pipeline with neural networks. Experimental results show our method outperforms the state‐of‐art, which demonstrates that our framework is able to cope with the challenges in single sketch‐based 3D shape modeling. Ruibin Wang, Tao Jiang 0020, Li Wang 0105, Yanran Li, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2020 | EditorialabstractThis special issue contains 28 full papers selected from the Computer Animation and Social Agents 2020 Conference (CASA2020). This conference was founded by the Computer Graphics Society in 1988 in Geneva and is the oldest conference on Computer Animation in the world. It has been held in many countries around the world and in recent years in Beijing, China (2018), Paris, France (2019) and this year in Bournemouth, United Kingdom. Because of the Covid-19 pandemic, this year, the conference will be held online through the Youtube Channel. The best paper award will be announced on the conference website after the conference. We would like to thank the authors for sharing their research findings by submitting papers to CASA2020. We are very grateful to the Program Committee members for reviewing the papers and to all the people who have contributed to the success of CASA2020 in Bournemouth. The conference is organized by Bournemouth University under the guidance of the Computer Graphics Society (CGS). Conference co-chairs Jian Jun Zhang (Bournemouth University, UK) Nadia Magnenat Thalmann (University of Geneva, Switzerland and Nanyang Technological University, Singapore) Program co-chairs Daniel Thalmann (EPFL, Switzerland) Xiaosong Yang (Bournemouth University, UK) Weiwei Xu (Zhejiang University, China) Publicity chair Jian Chang (Bournemouth University, UK) Local chair Feng Tian (Bournemouth University, UK) International program committee Nadine Aburumman, Brunel University, UK Norman Badler, University of Pennsylvania, USA Selim Balcisoy, Sabanci University, Turkey Loic Barthe, IRIT—Université de Toulouse, France Jan Bender, RWTH Aachen University, Germany Raphaëlle Chaine, LIRIS Université Lyon 1, France Jian Chang, Bournemouth University, UK Fred Charles, Bournemouth University, UK Parag Chaudhuri, Indian Institute of Technology, Bombay, India Marc Christie, INRIA, France Justin Dauwels, Nanyang Technological University, Singapore Shujie Deng, King's College London, UK Zhigang Deng, University of Houston, USA Etienne de Sevin, SANPSY University of Bordeaux, France Petros Faloutsos, York University, Canada Christos Gatzidis, Bournemouth University, UK Ugur Gudukbay, Bilkent University, Turkey Shihui Guo, Xiamen University, China Xiaohu Guo, The University of Texas at Dallas, USA James Hahn, George Washington University, USA Carlo Harvey, Birmingham City University, UK Gaoqi He, East China Normal University, China Ying He, Nanyang Technological University, Singapore Kemao Qian, Nanyang Technological University, Singapore Ruizhen Hu, Shenzhen University, China Jinyuan Jia, Tongji University, China Tao Jiang, University of Surrey, UK Xiaogang Jin, Zhejiang University, China Marcelo Kallmann, University of California, Merced, USA Prem Kalra, IIT Delhi, India Dongwann Kang, Seoul National University of Science and Technology, Korea Mubbasir Kapadia, Rutgers University, USA Min H. Kim, Korea Advanced Institute of Science and Technology, Korea Scott King, Texas A&M University—Corpus Christi, USA Taesoo Kwon, Hanyang University, China Sung-Hee Lee, Korea Advanced Institute of Science and Technology, Korea Wonsook Lee, University of Ottawa, Canada Tsai-Yen Li, National Chengchi University, Taiwan Guoliang Luo, East China Jiaotong University, China Chongyang Ma, Snap Inc., USA Anderson Maciel, Universidade Federal do Rio Grande do Sul, Brazil Nadia Magnenat Thalmann, University Of Geneva, Switzerland Shigeo Morishima, Waseda University, Japan Soraia Musse, Pontificia Universidade Catolica do Roi Grande do Sul, PUCRS, Brazil Rahul Narain, Indian Institute of Technology, Delhi, India Junjun Pan, Beihang University, China Nuria Pelechano, Universitat Politècnica de Catalunya, Spain Julien Pettre, INRIA, France Nicolas Pronost, Université Claude Bernard Lyon 1, France Kun Qian, King's College London, UK Craig Schroeder, University of California, Riverside, USA Ari Shapiro, Embody Digital, USA Hubert P. H. Shum, Northumbria University, UK Shinjiro Sueda, Texas A&M University, USA Daniel Thalmann, Ecole Polytechnique Fédérale de Lausanne, Switzerland Feng Tian, Bournemouth University, UK Yiying Tong, Michigan State University, USA Meili Wang, Northwest A&F University, China Zhao Wang, Zhejiang University, China Enhua Wu, University of Macau & ISCAS, China Zhongke Wu, Beijing Normal University, China Weiwei Xu, Zhejiang University, China Yachun Fan, Beijing Normal University, China Bailin Yang, Zhejiang Gongshang University, China Yin Yang, University of New Mexico, USA Xiaosong Yang, Bournemouth University, UK Yuting Ye, Oculus Research, USA Lihua You, Bournemouth University, UK Hongchuan Yu, Bournemouth University, UK Zerrin Yumak, Utrecht University, Netherlands Wenshu Zhang, Cardiff Metropolitan University Jian Zhang, Bournemouth University, UK Jianmin Zheng, Nanyang Technological University, Singapore Jian J. Zhang 0001, Nadia Magnenat-Thalmann, Daniel Thalmann, Xiaosong Yang, Weiwei Xu 0003, Jian Chang 0001, Feng Tian 0009 |
Comput. Animat. Virtual Worlds | 4 |
| 2020 | Fast character modeling with sketch-based PDE surfacesabstractAbstract Virtual characters are 3D geometric models of characters. They have a lot of applications in multimedia. In this paper, we propose a new physics-based deformation method and efficient character modelling framework for creation of detailed 3D virtual character models. Our proposed physics-based deformation method uses PDE surfaces. Here PDE is the abbreviation of Partial Differential Equation, and PDE surfaces are defined as sculpting force-driven shape representations of interpolation surfaces. Interpolation surfaces are obtained by interpolating key cross-section profile curves and the sculpting force-driven shape representation uses an analytical solution to a vector-valued partial differential equation involving sculpting forces to quickly obtain deformed shapes. Our proposed character modelling framework consists of global modeling and local modeling. The global modeling is also called model building, which is a process of creating a whole character model quickly with sketch-guided and template-based modeling techniques. The local modeling produces local details efficiently to improve the realism of the created character model with four shape manipulation techniques. The sketch-guided global modeling generates a character model from three different levels of sketched profile curves called primary, secondary and key cross-section curves in three orthographic views. The template-based global modeling obtains a new character model by deforming a template model to match the three different levels of profile curves. Four shape manipulation techniques for local modeling are investigated and integrated into the new modelling framework. They include: partial differential equation-based shape manipulation, generalized elliptic curve-driven shape manipulation, sketch assisted shape manipulation, and template-based shape manipulation. These new local modeling techniques have both global and local shape control functions and are efficient in local shape manipulation. The final character models are represented with a collection of surfaces, which are modeled with two types of geometric entities: generalized elliptic curves (GECs) and partial differential equation-based surfaces. Our experiments indicate that the proposed modeling approach can build detailed and realistic character models easily and quickly. Lihua You, Xiaosong Yang, JunJun Pan, Tong-Yee Lee, Shaojun Bian, Kun Qian 0009, Zulfiqar Habib, Allah Bux Sargano, Ismail Khalid Kazmi, Jian J. Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Photographic style transferabstractImage style transfer has attracted much attention in recent years. However, results produced by existing works still have lots of distortions. This paper investigates the CNN-based artistic style transfer work specifically and finds out the key reasons for distortion coming from twofold: the loss of spatial structures of content image during content-preserving process and unexpected geometric matching introduced by style transformation process. To tackle this problem, this paper proposes a novel approach consisting of a dual-stream deep convolution network as the loss network and edge-preserving filters as the style fusion model. Our key contribution is the introduction of an additional similarity loss function that constrains both the detail reconstruction and style transfer procedures. The qualitative evaluation shows that our approach successfully suppresses the distortions as well as obtains faithful stylized results compared to state-of-the-art methods. Li Wang 0105, Xiaosong Yang, Shi-Min Hu 0001, Jian J. Zhang 0001 |
Vis. Comput. | 3 |
| 2019 | Single-image Mesh Reconstruction and Pose Estimation via Generative Normal MapabstractWe present a unified learning framework for recovering both 3D mesh and camera pose of the object from a single image. Our approach learns to recover outer shape and surface geometric details of the mesh without relying on 3D supervision. We adopt multi-view normal maps as the 2D supervision so that the silhouette and geometric details information can be transferred to neural network. A normal mismatch based objective function is introduced to train the network, and the camera pose is parameterized into the objective, it integrates pose estimation with the mesh reconstruction in a same optimization procedure. We demonstrate the abilities of the proposed approach in generating 3D mesh and estimating camera pose with qualitative and quantitative experiments. Li Wang 0105, Tao Jiang 0020, Yanran Li, Xiaosong Yang, Jian J. Zhang 0001 |
CASA | 5 |
| 2019 | Fine-Grained Color Sketch-Based Image Retrieval
Yu Xia 0012, Shuangbu Wang, Yanran Li, Lihua You, Xiaosong Yang, Jian J. Zhang 0001 |
CGI | 5 |
| 2019 | Efficient and realistic character animation through analytical physics-based skin deformation
Shaojun Bian, Zhigang Deng 0001, Ehtzaz Chaudhry, Lihua You, Xiaosong Yang, Hassan Ugail, Xiaogang Jin 0001, Zhidong Xiao, Jian J. Zhang 0001 |
Graph. Model. | 5 |
| 2019 | Huber- L1-based non-isometric surface registrationabstractNon-isometric surface registration is an important task in computer graphics and computer vision. It, however, remains challenging to deal with noise from scanned data and distortion from transformation. In this paper, we propose a Huber- $$L_1$$ -based non-isometric surface registration and solve it by the alternating direction method of multipliers. With a Huber- $$L_1$$ -regularized model constrained on the transformation variation and position difference, our method is robust to noise and produces piecewise smooth results while still preserving fine details on the target. The introduced as-similar-as-possible energy is able to handle different size of shapes with little stretching distortion. Extensive experimental results have demonstrated that our method is more accurate and robust to noise in comparison with the state-of-the-arts. Tao Jiang 0020, Xiaosong Yang, Jian J. Zhang 0001, Feng Tian 0009, Shuang Liu 0006, Kun Qian 0009 |
Vis. Comput. | 2 |
| 2019 | Efficient convolutional hierarchical autoencoder for human motion predictionabstractHuman motion prediction is a challenging problem due to the complicated human body constraints and high-dimensional dynamics. Recent deep learning approaches adopt RNN, CNN or fully connected networks to learn the motion features which do not fully exploit the hierarchical structure of human anatomy. To address this problem, we propose a convolutional hierarchical autoencoder model for motion prediction with a novel encoder which incorporates 1D convolutional layers and hierarchical topology. The new network is more efficient compared to the existing deep learning models with respect to size and speed. We train the generic model on Human3.6M and CMU benchmark and conduct extensive experiments. The qualitative and quantitative results show that our model outperforms the state-of-the-art methods in both short-term prediction and long-term prediction. Yanran Li, Xiaosong Yang, Meili Wang 0001, Sebastian Iulian Poiana, Ehtzaz Chaudhry, Jian J. Zhang 0001 |
Vis. Comput. | 3 |
| 2018 | Fast photographic style transfer based on convolutional neural networksabstractThe techniques for photographic style transfer have been researched for a long time, which explores effective ways to transfer the style features of a reference photo onto another content photograph. Recent works based on convolutional neural networks present an effective solution for style transfer, especially for paintings. The artistic style transformation results are visually appealing, however, the photorealism is lost because of content-mismatching and distortions even when both input images are photographic. To tackle this challenge, this paper introduces a similarity loss function and a refinement method into the style transfer network. The similarity loss function can solve the content-mismatching problem, however, the distortion and noise artefacts may still exist in the stylized results due to the content-style trade-off. Hence, we add a post-processing refinement step to reduce the artefacts. The robustness and effectiveness of our approach has been evaluated through extensive experiments which show that our method can obtain finer content details and less artefacts than state-of-the-art methods, and transfer style faithfully. In addition, our approach is capable of processing photographic style transfer in almost real-time, which makes it a potential solution for video style transfer. Li Wang 0105, Xiaosong Yang, Jian J. Zhang 0001 |
CGI | 3 |
| 2018 | Motion Capture Data Completion via Truncated Nuclear Norm RegularizationabstractThe objective of motion capture (mocap) data completion is to recover missing measurement of the body markers from mocap. It becomes increasingly challenging as the missing ratio and duration of mocap data grow. Traditional approaches usually recast this problem as a low-rank matrix approximation problem based on the nuclear norm. However, the nuclear norm defined as the sum of all the singular values of a matrix is not a good approximation to the rank of mocap data. This paper proposes a novel approach to solve mocap data completion problem by adopting a new matrix norm, called truncated nuclear norm. An efficient iterative algorithm is designed to solve this problem based on the augmented Lagrange multiplier. The convergence of the proposed method is proved mathematically under mild conditions. To demonstrate the effectiveness of the proposed method, various comparative experiments are performed on synthetic data and mocap data. Compared to other methods, the proposed method is more efficient and accurate. Shuang Liu 0006, Xiaosong Yang, Gaohang Yu, Jian J. Zhang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2017 | Rib-reinforced Shell StructureabstractAbstract Shell structures are extensively used in engineering due to their efficient load‐carrying capacity relative to material volume. However, large‐span shells require additional supporting structures to strengthen fragile regions. The problem of designing optimal stiffeners is therefore becoming a major challenge for shell applications. To address it, we propose a computational framework to design and optimize rib layout on arbitrary shell to improve the overall structural stiffness and mechanical performance. The essential of our method is to place ribs along the principal stress lines which reflect paths of material continuity and indicates trajectories of internal forces. Given a surface and user‐specified external loads, we perform a Finite Element Analysis. Using the resulting principal stress field, we generate a quad‐mesh whose edges align with this cross field. Then we extract an initial rib network from the quad‐mesh. After simplifying rib network by removing ribs with little contribution, we perform a rib flow optimization which allows ribs to swing on surface to further adjust rib distribution. Finally, we optimize rib cross‐section to maximally reduce material usage while achieving certain structural stiffness requirements. We demonstrate that our rib‐reinforced shell structures achieve good static performances. And experimental results by 3D printed objects show the effectiveness of our method. Anzong Zheng, Lihua You, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Graph. Forum | 4 |
| 2017 | Robust facial landmark detection and tracking across poses and expressions for in-the-wild monocular videoabstractWe present a novel approach for automatically detecting and tracking facial landmarks across poses and expressions from in-the-wild monocular video data, e.g., YouTube videos and smartphone recordings. Our method does not require any calibration or manual adjustment for new individual input videos or actors. Firstly, we propose a method of robust 2D facial landmark detection across poses, by combining shape-face canonical-correlation analysis with a global supervised descent method. Since 2D regression-based methods are sensitive to unstable initialization, and the temporal and spatial coherence of videos is ignored, we utilize a coarse-todense 3D facial expression reconstruction method to refine the 2D landmarks. On one side, we employ an in-the-wild method to extract the coarse reconstruction result and its corresponding texture using the detected sparse facial landmarks, followed by robust pose, expression, and identity estimation. On the other side, to obtain dense reconstruction results, we give a face tracking flow method that corrects coarse reconstruction results and tracks weakly textured areas; this is used to iteratively update the coarse face model. Finally, a dense reconstruction result is estimated after it converges. Extensive experiments on a variety of video sequences recorded by ourselves or downloaded from YouTube show the results of facial landmark detection and tracking under various lighting conditions, for various head poses and facial expressions. The overall performance and a comparison with state-of-art methods demonstrate the robustness and effectiveness of our method. Shuang Liu 0006, Yongqiang Zhang 0003, Xiaosong Yang, Daming Shi 0001, Jian J. Zhang 0001 |
Comput. Vis. Media | 3 |
| 2017 | Essential techniques for laparoscopic surgery simulationabstractAbstract Laparoscopic surgery is a complex minimum invasive operation that requires long learning curve for the new trainees to have adequate experience to become a qualified surgeon. With the development of virtual reality technology, virtual reality‐based surgery simulation is playing an increasingly important role in the surgery training. The simulation of laparoscopic surgery is challenging because it involves large non‐linear soft tissue deformation, frequent surgical tool interaction and complex anatomical environment. Current researches mostly focus on very specific topics (such as deformation and collision detection) rather than a consistent and efficient framework. The direct use of the existing methods cannot achieve high visual/haptic quality and a satisfactory refreshing rate at the same time, especially for complex surgery simulation. In this paper, we proposed a set of tailored key technologies for laparoscopic surgery simulation, ranging from the simulation of soft tissues with different properties, to the interactions between surgical tools and soft tissues to the rendering of complex anatomical environment. Compared with the current methods, our tailored algorithms aimed at improving the performance from accuracy, stability and efficiency perspectives. We also abstract and design a set of intuitive parameters that can provide developers with high flexibility to develop their own simulators. Copyright © 2016 John Wiley & Sons, Ltd. Kun Qian 0009, Junxuan Bai, Xiaosong Yang, JunJun Pan, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | Supervised coordinate descent method with a 3D bilinear model for face alignment and trackingabstractAbstract Face alignment and tracking play important roles in facial performance capture. Existing data‐driven methods for monocular videos suffer from large variations of pose and expression. In this paper, we propose an efficient and robust method for this task by introducing a novel supervised coordinate descent method with 3D bilinear representation. Instead of learning the mapping between the whole parameters and image features directly with a cascaded regression framework in current methods, we learn individual sets of parameters mappings separately step by step by a coordinate descent mean. Because different parameters make different contributions to the displacement of facial landmarks, our method is more discriminative to current whole‐parameter cascaded regression methods. Benefiting from a 3D bilinear model learned from public databases, the proposed method can handle the head pose changes and extreme expressions out of plane better than other 2D‐based methods. We present the reliable result of face tracking under various head poses and facial expressions on challenging video sequences collected online. The experimental results show that our method outperforms state‐of‐art data‐driven methods. Yongqiang Zhang 0003, Shuang Liu 0006, Xiaosong Yang, Jian J. Zhang 0001, Daming Shi 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | A human motion feature based on semi-supervised learning of GMM
Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Hanzhi Zhang, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Multim. Syst. | 6 |
| 2017 | Consistent as-similar-as-possible non-isometric surface registrationabstractNon-isometric surface registration, aiming to align two surfaces with different sizes and details, has been widely used in computer animation industry. Various existing surface registration approaches have been proposed for accurate template fitting; nevertheless, two challenges remain. One is how to avoid the mesh distortion and fold over of surfaces during transformation. The other is how to reduce the amount of landmarks that have to be specified manually. To tackle these challenges simultaneously, we propose a consistent as-similar-as-possible (CASAP) surface registration approach. With a novel defined energy, it not only achieves the consistent discretization for the surfaces to produce accurate result, but also requires a small number of landmarks with little user effort only. Besides, CASAP is constrained as-similar-as-possible so that angles of triangle meshes are preserved and local scales are allowed to change. Extensive experimental results have demonstrated the effectiveness of CASAP in comparison with the state-of-the-art approaches. Tao Jiang 0020, Kun Qian 0009, Shuang Liu 0006, Xiaosong Yang, Jian J. Zhang 0001 |
Vis. Comput. | 5 |
| 2016 | Sign-Correlation Partition Based on Global Supervised Descent Method for Face Alignment
Yongqiang Zhang 0003, Shuang Liu 0006, Xiaosong Yang, Daming Shi 0001, Jian J. Zhang 0001 |
ACCV (3) | 3 |
| 2016 | A 3D human motion refinement method based on sparse motion bases selectionabstractMotion capture (MOCAP) is an important technique that is widely used in many areas such as computer animation, film industry, physical training and so on. Even with professional MOCAP system, the missing marker problems always occur. Motion refinement is an essential preprocessing step for MOCAP data based applications. Although many existing approaches for motion refinement have been developed, it is still a challenging task due to the complexity and diversity of human motion. A data driven based motion refinement method is proposed in this paper, which modifies the traditional sparse coding process for special task of motion recovery from missing parts. Meanwhile, the objective function is derived by taking both statistical and kinematical property of motion data into account. Poselet model and moving window grouping are applied in the proposed method to achieve a fine-grained feature representation, which preserves the embedded spatial-temporal kinematic information. 5 motion dictionaries are learnt for each kind of poselet from training data in parallel. The motion refine problem is finally solved as an ℓ1-minimization problem. Compared with several state-of-art motion refine methods, the experimental result shows that our approach outperforms the competitors. Yinfu Feng, Shuang Liu 0006, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001 |
CASA | 5 |
| 2016 | 3D Body Shapes Estimation from Dressed-Human SilhouettesabstractAbstract Estimation of 3D body shapes from dressed‐human photos is an important but challenging problem in virtual fitting. We propose a novel automatic framework to efficiently estimate 3D body shapes under clothes. We construct a database of 3D naked and dressed body pairs, based on which we learn how to predict 3D positions of body landmarks (which further constrain a parametric human body model) automatically according to dressed‐human silhouettes. Critical vertices are selected on 3D registered human bodies as landmarks to represent body shapes, so as to avoid the time‐consuming vertices correspondences finding process for parametric body reconstruction. Our method can estimate 3D body shapes from dressed‐human silhouettes within 4 seconds, while the fastest method reported previously need 1 minute. In addition, our estimation error is within the size tolerance for clothing industry. We dress 6042 naked bodies with 3 sets of common clothes by physically based cloth simulation technique. To the best of our knowledge, We are the first to construct such a database containing 3D naked and dressed body pairs and our database may contribute to the areas of human body shapes estimation and cloth simulation. Dan Song 0006, Ruofeng Tong 0001, Jian Chang 0001, Xiaosong Yang, Min Tang 0001, Jian J. Zhang 0001 |
Comput. Graph. Forum | 4 |
| 2016 | Video-Based Classification of Driving Behavior Using a Hierarchical Classification System with Multiple FeaturesabstractDriver fatigue and inattention have long been recognized as one of the main contributing factors in traffic accidents. Therefore, the development of intelligent driver assistance systems, which provides automatic monitoring of driver's vigilance, is an urgent and challenging task. This paper presents a novel system for video-based driving behavior recognition. The fundamental idea is to monitor driver's hand movements and to use these as predictors for safe/unsafe driving behavior. In comparison to previous work, the proposed method utilizes hierarchical classification and treats driving behavior in terms of a spatio-temporal reference framework as opposed to a static image. The approach was verified using the Southeast University Driving-Posture Dataset, a dataset comprised of video clips covering aspects of driving such as: normal driving, responding to a cell phone call, eating and smoking. After pre-processing for illumination variations and motion sequence segmentation, eight classes of behavior were identified. The overall prediction accuracy obtained using the proposed approach was [Formula: see text] when using a hierarchical classification approach. The proposed approach was able to clearly identify two dangerous driving behaviors, Responding to a cellphone call and Eating, with recognition rates of 92.39% and 92.29% respectively. Chao Yan 0003, Frans Coenen, Yong Yue 0001, Xiaosong Yang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2016 | Real-time facial expression transfer with single video cameraabstractAbstract Facial expression transfer has been actively researched in the past few years. Existing methods either suffer from depth ambiguity or require special hardware. We present a novel marker‐less, real‐time facial transfer method that requires only a single video camera. We develop a robust model, which is adaptive to user‐specific facial data. It computes expression variances in real time and rapidly transfers them onto a target character either from images or videos. Our method can be applied to videos without prior camera calibration and focal adjustment. It enables realistic online facial expression editing and performance transferring in many scenarios such as video conference, news broadcasting, lip‐syncing for song performances and so on. With low computational cost and hardware requirement, our method tracks a single user at an average of 38fps and runs smoothly even in web browsers. Copyright © 2016 John Wiley & Sons, Ltd. Shuang Liu 0006, Xiaosong Yang, Zhidong Xiao, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2016 | Energized soft tissue dissection in surgery simulationabstractAbstract With the development of virtual reality technology, surgery simulation has become an effective way to train the operation skills for surgeons. Soft tissue dissection, as one of the most frequently performed operations in surgery, is indispensable to an immersive and high‐fidelity surgery simulator. Energized dissection tools are much more commonly used than the traditional sharp scalpels for patient safety. Unfortunately, the interaction of such tools with the soft tissues has been largely ignored in the research of surgical simulators. In this paper, we have proposed an energized soft tissue dissection model. We categorize the soft tissues into three types (fascia, membrane, and fat) and simulate their physical property accordingly. The dissection algorithm we propose employs an edge‐based structure, which offers an effective mechanism for the generation of incisions dissected with energized tools. The mesh topology will not be changed when it is dissected by an energized tool, rather it is controlled by the heat transfer model. Our dissection method is highly compatible and efficient to the physically based simulation resolved by a pre‐factorized linear system. Copyright © 2016 John Wiley & Sons, Ltd. Kun Qian 0009, Tao Jiang 0020, Meili Wang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2016 | Adaptive multi-view feature selection for human motion retrievalabstractHuman motion retrieval plays an important role in many motion data based applications. In the past, many researchers tended to use a single type of visual feature as data representation. Because different visual feature describes different aspects about motion data, and they have dissimilar discriminative power with respect to one particular class of human motion, it led to poor retrieval performance. Thus, it would be beneficial to combine multiple visual features together for motion data representation. In this article, we present an Adaptive Multi-view Feature Selection (AMFS) method for human motion retrieval. Specifically, we first use a local linear regression model to automatically learn multiple view-based Laplacian graphs for preserving the local geometric structure of motion data. Then, these graphs are combined together with a non-negative view-weight vector to exploit the complementary information between different features. Finally, in order to discard the redundant and irrelevant feature components from the original high-dimensional feature representation, we formulate the objective function of AMFS as a general trace ratio optimization problem, and design an effective algorithm to solve the corresponding optimization problem. Extensive experiments on two public human motion database, i.e., HDM05 and MSR Action3D, demonstrate the effectiveness of the proposed AMFS over the state-of-art methods for motion data retrieval. The scalability with large motion dataset, and insensitivity with the algorithm parameters, make our method can be widely used in real-world applications. Yinfu Feng, Tian Qi, Xiaosong Yang, Jian J. Zhang 0001 |
Signal Process. | 4 |
| 2015 | An Adaptive Spherical Collision Detection and Resolution Method for Deformable Object SimulationabstractCollision detection and resolution are of great importance to physically based animation. Real time responses are essential for many applications, which largely rely on the efficiency of localising the potentially colliding geometry and calculating the polygon intersections. It is an extremely heavy computation task using the existing polygon based methods, especially for deformable objects. To improve this issue, we present an implicit circumsphere based collision detection and resolution method for deformable objects which takes into consideration both local geometry features and the material properties. Our method approximates the mesh in question with an implicit circumsphere surface, which is used to perform finest level collision detection and resolution instead of the original polygonal mesh. The dynamic deformation as a result of collision is determined by both the geometry and the material properties of the surface. Due to the simplicity of sphere overlap test, our method is not only computationally efficient, but also stable and comparatively accurate, outperforming the existing methods in overall performance. Our implicit circumsphere method can also provide better prevention to collision tunnelling than existing methods. Besides, this method is compatible with all existing broad phase and narrow phase collision query techniques. Kun Qian 0009, Xiaosong Yang, Jian J. Zhang 0001, Meili Wang 0001 |
CAD/Graphics | 2 |
| 2015 | Virtual reality based laparoscopic surgery simulationabstractWith the development of computer graphic and haptic devices, training surgeons with virtual reality technology has proven to be very effective in surgery simulation. Many successful simulators have been deployed for training medical students. However, due to the various unsolved technical issues, the laparoscopic surgery simulation has not been widely used. Such issues include modeling of complex anatomy structure, large soft tissue deformation, frequent surgical tools interactions, and the rendering of complex material under the illumination of headlight. A successful laparoscopic surgery simulator should integrate all these required components in a balanced and efficient manner to achieve both visual/haptic quality and a satisfactory refreshing rate. In this paper, we propose an efficient framework integrating a set of specially tailored and designed techniques, ranging from deformation simulation, collision detection, soft tissue dissection and rendering. We optimize all the components based on the actual requirement of laparoscopic surgery in order to achieve an improved overall performance of fidelity and responding speed. Kun Qian 0009, Junxuan Bai, Xiaosong Yang, JunJun Pan, Jian J. Zhang 0001 |
VRST | 3 |
| 2015 | Dehydration of core/shell fruitsabstractDehydrated core/shell fruits, such as jujubes, raisins and plums, show very complex buckles and wrinkles on their exocarp. It is a challenging task to model such complicated patterns and their evolution in a virtual environment even for professional animators. This paper presents a unified physically-based approach to simulate the morphological transformation for the core/shell fruits in the dehydration process. A finite element method (FEM), which is based on the multiplicative decomposition of the deformation gradient into an elastic part and a dehydrated part, is adopted to model the morphological evolution. In the method, the dehydration pattern can be conveniently controlled through physically prescribed parameters according to the geometry and material of the real fruits. The effects of the parameters on the final dehydrated surface patterns are investigated and summarized in detail. Experiments on jujubes, wolfberries, raisins and plums are given, which demonstrate the efficacy of the method. Yin Liu 0004, Xiaosong Yang, Biaosong Chen, Jian J. Zhang 0001, Hongwu Zhang |
Comput. Graph. | 2 |
| 2015 | Efficient semi-supervised multiple feature fusion with out-of-sample extension for 3D model retrieval
Mingming Ji, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Neurocomputing | 5 |
| 2015 | Efficient sketch-based creation of detailed character models through data-driven mesh deformationsabstractAbstract Creation of detailed character models is a very challenging task in animation production. Sketch‐based character model creation from a 3D template provides a promising solution. However, how to quickly find correct correspondences between user's drawn sketches and the 3D template model, how to efficiently deform the 3D template model to exactly match user's drawn sketches, and realize real‐time interactive modeling is still an open topic. In this paper, we propose a new approach and develop a user interface to effectively tackle this problem. Our proposed approach includes using user's drawn sketches to retrieve a most similar 3D template model from our dataset and marrying human's perception and interactions with computer's highly efficient computing to extract occluding and silhouette contours of the 3D template model and find correct correspondences quickly. We then combine skeleton‐based deformation and mesh editing to deform the 3D template model to fit user's drawn sketches and create new and detailed 3D character models. The results presented in this paper demonstrate the effectiveness and advantages of our proposed approach and usefulness of our developed user interface. Copyright © 2015 John Wiley & Sons, Ltd. Ismail Khalid Kazmi, Lihua You, Xiaosong Yang, Xiaogang Jin 0001, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2015 | Advanced ordinary differential equation based head modelling for Chinese marionette art preservationabstractAbstract Puppetry has been a popular art form for many centuries in different cultures, which becomes a valuable and fascinating heritage assert. Traditional Chinese marionette art with over 2000 years history is one of the most representative forms offering a mixture of stage performance of singing, dancing, music, poetry, opera, story narrative and action. Apart from a set of string rules, which controls the dynamics, head carving skill is another important pillar in this art form. This paper addresses the heritage preservation of the marionette head carving by digitalizing the head models with a novel modelling technique using ordinary differential equations (ODEs). The technique has been specially tailored to suit the modelling complexity and the need of accurate description of shapes. It offers smoothly sewing ODE swept patches to represent the distinct features of a marionette head with sharp variance of local geometry. Such features otherwise are difficult to model and capture accurately, which may require a great effort and tedious handcrafting of an experienced modeller, when using other representation forms like polygons. Copyright © 2015 John Wiley & Sons, Ltd. Hui Liang 0004, Jian Chang 0001, Xiaosong Yang, Lihua You, Shaojun Bian, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2015 | Sparse motion bases selection for human motion denoisingabstractHuman motion denoising is an indispensable step of data preprocessing for many motion data based applications. In this paper, we propose a data-driven based human motion denoising method that sparsely selects the most correlated subset of motion bases for clean motion reconstruction. Meanwhile, it takes the statistic property of two common noises, i.e., Gaussian noise and outliers, into account in deriving the objective functions. In particular, our method firstly divides each human pose into five partitions termed as poselets to gain a much fine-grained pose representation. Then, these poselets are reorganized into multiple overlapped poselet groups using a lagged window moving across the entire motion sequence to preserve the embedded spatial–temporal motion patterns. Afterward, five compacted and representative motion dictionaries are constructed in parallel by means of fast K-SVD in the training phase; they are used to remove the noise and outliers from noisy motion sequences in the testing phase by solving ℓ 1 -minimization problems. Extensive experiments show that our method outperforms its competitors. More importantly, compared with other data-driven based method, our method does not need to specifically choose the training data , it can be more easily applied to real-world applications. Jun Xiao 0001, Yinfu Feng, Mingming Ji, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang |
Signal Process. | 4 |
| 2015 | Mining Spatial-Temporal Patterns and Structural Sparsity for Human Motion Data DenoisingabstractMotion capture is an important technique with a wide range of applications in areas such as computer vision, computer animation, film production, and medical rehabilitation. Even with the professional motion capture systems, the acquired raw data mostly contain inevitable noises and outliers. To denoise the data, numerous methods have been developed, while this problem still remains a challenge due to the high complexity of human motion and the diversity of real-life situations. In this paper, we propose a data-driven-based robust human motion denoising approach by mining the spatial-temporal patterns and the structural sparsity embedded in motion data. We first replace the regularly used entire pose model with a much fine-grained partlet model as feature representation to exploit the abundant local body part posture and movement similarities. Then, a robust dictionary learning algorithm is proposed to learn multiple compact and representative motion dictionaries from the training data in parallel. Finally, we reformulate the human motion denoising problem as a robust structured sparse coding problem in which both the noise distribution information and the temporal smoothness property of human motion have been jointly taken into account. Compared with several state-of-the-art motion denoising methods on both the synthetic and real noisy motion data, our method consistently yields better performance than its counterparts. The outputs of our approach are much more stable than that of the others. In addition, it is much easier to setup the training dataset of our method than that of the other data-driven-based methods. Yinfu Feng, Mingming Ji, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2014 | Locomotion Skills for Insects with Sample-based ControllerabstractAbstract Natural‐looking insect animation is very difficult to simulate. The fast movement and small scale of insects often challenge the standard motion capture techniques. As for the manual key‐framing or physics‐driven methods, significant amounts of time and efforts are necessary due to the delicate structure of the insect, which prevents practical applications. In this paper, we address this challenge by presenting a two‐level control framework to efficiently automate the modeling and authoring of insects’ locomotion. On the top level, we design a Triangle Placement Engine to automatically determine the location and orientation of insects’ foot contacts, given the user‐defined trajectory and settings, including speed, load, path and terrain etc. On the low‐level, we relate the Central Pattern Generator to the triangle profiles with the assistance of a Controller Look‐Up Table to fast simulate the physically‐based movement of insects. With our approach, animators can directly author insects’ behavior among a wide range of locomotion repertoire, including walking along a specified path or on an uneven terrain, dynamically adjusting to external perturbations and collectively transporting prey back to the nest. Shihui Guo, Jian Chang 0001, Xiaosong Yang, Wencheng Wang 0001, Jian J. Zhang 0001 |
Comput. Graph. Forum | 3 |
| 2014 | Exploiting temporal stability and low-rank structure for motion capture data refinementabstractInspired by the development of the matrix completion theories and algorithms, a low-rank based motion capture (mocap) data refinement method has been developed, which has achieved encouraging results. However, it does not guarantee a stable outcome if we only consider the low-rank property of the motion data. To solve this problem, we propose to exploit the temporal stability of human motion and convert the mocap data refinement problem into a robust matrix completion problem, where both the low-rank structure and temporal stability properties of the mocap data as well as the noise effect are considered. An efficient optimization method derived from the augmented Lagrange multiplier algorithm is presented to solve the proposed model. Besides, a trust data detection method is also introduced to improve the degree of automation for processing the entire set of the data and boost the performance. Extensive experiments and comparisons with other methods demonstrate the effectiveness of our approaches on both predicting missing data and de-noising. Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001, Rong Song |
Inf. Sci. | 4 |
| 2014 | Human motion retrieval based on freehand sketchabstractABSTRACT In this paper, we present an integrated framework of human motion retrieval based on freehand sketch. With some simple rules, the user can acquire a desired motion by sketching several key postures. To retrieve efficiently and accurately by sketch, the 3D postures are projected onto several 2D planes. The limb direction feature is proposed to represent the input sketch and the projected‐postures. Furthermore, a novel index structure based on k‐d tree is constructed to index the motions in the database, which speeds up the retrieval process. With our posture‐by‐posture retrieval algorithm, a continuous motion can be got directly or generated by using a pre‐computed graph structure. What's more, our system provides an intuitive user interface. The experimental results demonstrate the effectiveness of our method. © 2014 The Authors.Computer Animation and Virtual Worldspublished by John Wiley & Sons, Ltd. Zhangpeng Tang, Jun Xiao 0001, Yinfu Feng, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 4 |
| 2014 | Real-time motion data annotation via action stringabstractABSTRACT Even though there is an explosive growth of motion capture data, there is still a lack of efficient and reliable methods to automatically annotate all the motions in a database. Moreover, because of the popularity of mocap devices in home entertainment systems, real‐time human motion annotation or recognition becomes more and more imperative. This paper presents a new motion annotation method that achieves both the aforementioned two targets at the same time. It uses a probabilistic pose feature based on the Gaussian Mixture Model to represent each pose. After training a clustered pose feature model, a motion clip could be represented as an action string. Then, a dynamic programming‐based string matching method is introduced to compare the differences between action strings. Finally, in order to achieve the real‐time target, we construct a hierarchical action string structure to quickly label each given action string. The experimental results demonstrate the efficacy and efficiency of our method. Copyright © 2014 John Wiley & Sons, Ltd. Qi Tian 0001, Jun Xiao 0001, Yueting Zhuang, Hanzhi Zhang, Xiaosong Yang, Jian J. Zhang 0001, Yinfu Feng |
Comput. Animat. Virtual Worlds | 5 |
| 2013 | Shape modeling for animated characters using ordinary differential equations
Ehtzaz Chaudhry, Lihua You, Xiaogang Jin 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Graph. | 4 |
| 2013 | A semantic feature for human motion retrievalabstractABSTRACT With the explosive growth of motion capture data, it becomes very imperative in animation production to have an efficient search engine to retrieve motions from large motion repository. However, because of the high dimension of data space and complexity of matching methods, most of the existing approaches cannot return the result in real time. This paper proposes a high level semantic feature in a low dimensional space to represent the essential characteristic of different motion classes. On the basis of the statistic training of Gauss Mixture Model, this feature can effectively achieve motion matching on both global clip level and local frame level. Experiment results show that our approach can retrieve similar motions with rankings from large motion database in real‐time and also can make motion annotation automatically on the fly. Copyright © 2013 John Wiley & Sons, Ltd. Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2013 | Motion Adaptation With Motor Invariant TheoryabstractBipedal walking is not fully understood. Motion generated from methods employed in robotics literature is stiff and is not nearly as energy efficient as what we observe in nature. In this paper, we propose validity conditions for motion adaptation from biological principles in terms of the topology of the dynamic system. This allows us to provide a closed-form solution to the problem of motion adaptation to environmental perturbations. We define both global and local controllers that improve structural and state stability, respectively. Global control is achieved by coupling the dynamic system with a neural oscillator, which preserves the periodic structure of the motion primitive and ensures stability by entrainment. A group action derived from Lie group symmetry is introduced as a local control that transforms the underlying state space while preserving certain motor invariants. We verify our method by evaluating the stability and energy consumption of a synthetic passive dynamic walker and compare this with motion data of a real walker. We also demonstrate that our method can be applied to a variety of systems. Fangde Liu, Richard Southern, Shihui Guo, Xiaosong Yang, Jian J. Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2013 | Automatic cage construction for retargeted muscle fitting
Xiaosong Yang, Jian Chang 0001, Richard Southern, Jian J. Zhang 0001 |
Vis. Comput. | 1 |
| 2012 | An RBF-Based Reparameterization Method for Constrained Texture MappingabstractTexture mapping has long been used in computer graphics to enhance the realism of virtual scenes. However, to match the 3D model feature points with the corresponding pixels in a texture image, surface parameterization must satisfy specific positional constraints. However, despite numerous research efforts, the construction of a mathematically robust, foldover-free parameterization that is subject to positional constraints continues to be a challenge. In the present paper, this foldover problem is addressed by developing radial basis function (RBF)-based reparameterization. Given initial 2D embedding of a 3D surface, the proposed method can reparameterize 2D embedding into a foldover-free 2D mesh, satisfying a set of user-specified constraint points. In addition, this approach is mesh free. Therefore, generating smooth texture mapping results is possible without extra smoothing optimization. Hongchuan Yu, Tong-Yee Lee, I-Cheng Yeh 0001, Xiaosong Yang, Wenxi Li, Jian J. Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2011 | Solid modelling based on sixth order partial differential equations
Lihua You, Jian Chang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Aided Des. | 3 |
| 2011 | Tensor-Based Feature Representation with Application to Multimodal Face RecognitionabstractIn this paper, a novel feature representation to multimodal face recognition is proposed, which possesses three properties: completeness, robustness and compactness. This feature descriptor allows all information of an object to be reproduced and its representation is invariant to rigid motion. In order to effectively take advantage of the proposed feature descriptor, we amend our previous ND-PCA scheme with multidirectional decomposition technique, and provide the estimation of the upper bound error of the amended classifier. It is proved to be linear optimal compared to other linear classifiers. To investigate the numerical performance of the presented feature descriptor, we apply it to both multiple modal and single modal samples, and the revised ND-PCA classifier is performed on the resulting feature representations. The experiments of verification and identification are carried out on two different gallery-probe face databases in order for the results to be evaluated by ROC and CMC curves independently. Hongchuan Yu, Jian J. Zhang 0001, Xiaosong Yang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2011 | A fast hybrid computation model for rectum deformation
Jian Chang 0001, Xiaosong Yang, Jun J. Pan, Wenxi Li, Jian J. Zhang 0001 |
Vis. Comput. | 2 |
| 2010 | Shape manipulation using physically based wire deformationsabstractAbstract This paper develops an efficient, physically based shape manipulation technique. It defines a 3D model with profile curves, and uses spine curves generated from the profile curves to control the motion and global shape of 3D models. Profile and spine curves are changed into profile and spine wires by specifying proper material and geometric properties together with external forces. The underlying physics is introduced to deform profile and spine wires through the closed form solution to ordinary differential equations for axial and bending deformations. With the proposed approach, global shape changes are achieved through manipulating spine wires, and local surface details are created by deforming profile wires. A number of examples are presented to demonstrate the applications of our proposed approach in shape manipulation. Copyright © 2010 John Wiley & Sons, Ltd. Lihua You, Xiaosong Yang, X. Y. You, Xiaogang Jin 0001, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Automatic rigging for animation characters with 3D silhouetteabstractAbstract Animating an articulated 3D character requires the specification of its interior skeleton structure which defines how the skin surface is deformed during animation. Currently this task is to a large extent accomplished manually, which consumes a large amount of animators' time. This paper presents an automatic rigging method making use of a new geometry entity called the 3D silhouette. The first step is to extract a coarse 3D curve skeleton and some skeletal joints of a character. This curve skeleton is then refined with a perpendicular silhouette. According to the connectivity of the skeletal joints, the hierarchical animation skeleton is finally constructed. By avoiding complicated computation such as voxelization and pruning, this method is simple and efficient, much faster than existing methods. It proves very useful for quick animation production, with applications including games design and prototype graphical systems. Copyright © 2009 John Wiley & Sons, Ltd. JunJun Pan, Xiaosong Yang, Philip J. Willis, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Fast simulation of skin slidingabstractAbstract Skin sliding is the phenomenon of the skin moving over underlying layers of fat, muscle and bone. Due to the complex interconnections between these separate layers and their differing elasticity properties, it is difficult to model and expensive to compute. We present a novel method to simulate this phenomenon at real‐time by remeshing the surface based on a parameter space resampling. In order to evaluate the surface parametrization, we borrow a technique from structural engineering known as the force density method (FDM)which solves for an energy minimizing form with a sparse linear system. Our method creates a realistic approximation of skin sliding in real‐time, reducing texture distortions in the region of the deformation. In addition it is flexible, simple to use, and can be incorporated into any animation pipeline. Copyright © 2009 John Wiley & Sons, Ltd. Xiaosong Yang, Richard Southern, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2008 | Dynamic skin deformation with characteristic curvesabstractAbstract By simulating motion and deformation of a set of 3D characteristic curves defining surfaces, we introduce a new skin deformation technique to animate skin deformation of character models. This technique consists of two parts, representation of characteristic curves and dynamic skin deformation. The first part is to extract characteristic curves of an existing model, that is a polygon model or a NURBS model, represent the characteristic curves with time‐dependent 3D trigonometric curves, and relate the surface points of the model to these trigonometric curves. The advantage of such a treatment is that it can transform an original 3D problem into a 2D problem, making it much quicker and easier to process and compute. In the second part, a vector‐valued dynamic fourth‐order differential equation is employed to govern the behaviour of the time‐dependent 3D trigonometric curves. These dynamic differential equations incorporate both the time component and the physical properties of the material, which are given by the user. Thus, the character model can be made to animate (deform) following the user‐specified behaviour parameters. Our experiments demonstrate that this technique is able to produce realistic skin deformations efficiently and avoid the undesirable skinning problems suffered by some existing skinning techniques. Copyright © 2008 John Wiley & Sons, Ltd. Lihua You, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2007 | Boundary Constrained Swept Surfaces for Modelling and AnimationabstractAbstract Due to their simplicity and intuitiveness, swept surfaces are widely used in many surface modelling applications. In this paper, we present a versatile swept surface technique called the boundary constrained swept surfaces. The most distinct feature is its ability to satisfy boundary constraints, including the shape and tangent conditions at the boundaries of a swept surface. This permits significantly varying surfaces to be both modelled and smoothly assembled, leading to the construction of complex objects. The representation, similar to an ordinary swept surface, is analytical in nature and thus it is light in storage cost and numerically very stable to compute. We also introduce a number of useful shape manipulation tools, such as sculpting forces, to deform a surface both locally and globally. In addition to being a complementary method to the mainstream surface modelling and deformation techniques, we have found it very effective in automatically rebuilding existing complex models. Model reconstruction is arguably one of the most laborious and expensive tasks in modelling complex animated characters. We demonstrate how our technique can be used to automate this process. Lihua You, Xiaosong Yang, M. Pachulski, Jian J. Zhang 0001 |
Comput. Graph. Forum | 2 |
| 2007 | Bar-net driven skinning for character animationabstractAbstract In this paper we present a physically motivated technique for the deformation of animated characters, called thebar‐net driven skinning. We use a bar‐network (bar‐net) as a deforming mechanism. This technique can be used similarly to a conventional skinning tool, but can also make a skin surface behave in a physically plausible manner due to the inherent physical properties of the network. A bar‐net is a structure commonly used in structural engineering. Its shape depends on the structural and material properties and the forces acting upon it. Computing the rest shape of an arbitrary bar‐net is a time‐consuming non‐linear problem. In order to speed up the computation and also for such a bar‐net to be used intuitively to help computer animation, we have defined a set of properties that a desirable bar‐net should satisfy. This allows a bar‐net shape finding problem to be solved using linear equations. We adopt a two‐layer structure for the representation of a skin surface, including a coarse mesh and a fine mesh. To deform a skin surface, we couple a bar‐net to its coarse mesh, which in turn deforms the fine mesh when the coupled bar‐net is deformed. The fine surface mesh can be of different forms, including Nurbs, subdivision surfaces and polygons. Copyright © 2007 John Wiley & Sons, Ltd. Jian J. Zhang 0001, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 2 |
| 2006 | Curve skeleton skinning for human and creature charactersabstractAbstract The skeleton driven skinning technique is still the most popular method for animating deformable human and creature characters. Albeit an industryde factodue to its computational performance and intuitiveness, it suffers from problems like collapsing elbow and candy wrapper joint. To remedy these problems, one needs to formulate the non‐linear relationship between the skeleton and the skin shape of a character properly, which however proves mathematically very challenging. Placing additional joints where the skin bends increases the sampling rate and is an ad hoc way of approximating this non‐linear relationship. In this paper, we propose a method that is able to accommodate the inherent non‐linear relationships between the movement of the skeleton and the skin shape. We use the so‐called curve skeletons along with the joint‐based skeletons to animate the skin shape. Since the deformation follows the tangent of the curve skeleton and also due to higher sampling rates received from the curve points, collapsing skin and other undesirable skin deformation problems are avoided. The curve skeleton retains the advantages of the current skeleton driven skinning. It is easy to use and allows full control over the animation process. As a further enhancement, it is also fairly simple to build realistic muscle and fat bulge effect. A practical implementation in the form of a Maya plug‐in is created to demonstrate the viability of the technique. Copyright © 2006 John Wiley & Sons, Ltd. Xiaosong Yang, Arun Somasekharan, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2006 | Automatic muscle generation for character skin deformationabstractAbstract As skin shape depends on the underlying anatomical structure, the anatomy‐based techniques usually afford greater realism than the traditional skeleton‐driven approach. On the downside, however, it is against the current animation workflow, as the animator has to model many individual muscles before the final skin layer arrives, resulting in an unintuitive modelling process. In this paper, we present a new anatomy‐based technique that allows the animator to start from an already modelled character. Muscles having visible influence on the skin shape at the rest pose are extracted automatically by studying the surface geometry of the skin. The extracted muscles are then used to deform the skin in areas where there exist complex deformations. The remaining skin areas, unaffected or hardly affected by the muscles, are handled by the skeleton‐driven technique, allowing both techniques to play their strengths. In order for the extracted muscles to produce realistic local skin deformation during animation, muscle bulging and special movements are both represented. Whereas the former ensues volume preservation, the latter allows a muscle not only to deform along a straight path, but also to slide and bend around joints and bones, resulting in the production of sophisticated muscle movements and deformations. Copyright © 2006 John Wiley & Sons, Ltd. Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2005 | Fast mesh-free deformationsabstractMesh-free deformation is an effective method for the simulation of deformable objects and characters in computer animation. It bears the merits of flexible control, good accuracy and easy implementation. Similar to other physically-based deformation techniques, however, computational cost remains a pressing issue, especially for large scale problems. Based on our previous work, in this paper we investigate the underlying structure of the numerical approach in order to speed up the computation. Three algorithms are presented and their efficiency and convergence are analysed. Our results show that we are able to significantly reduce the computation time and memory usage with little loss in accuracy. The new technique is capable of achieving interactive frame rates with large models. Jian Chang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
CAD/Graphics | 2 |
| 2005 | Realistic Skeleton Driven Skin Deformation
Xiaosong Yang, Jian J. Zhang 0001 |
ICCSA (3) | 1 |
| 2005 | Motion Data Correction and Extrapolation Using Physical ConstraintsabstractOptimization techniques have proven to be a powerful approach for generating new motions. In this paper, we present a physically based optimization method to synthesize motions by using motion capture data as input. We assume that the captured motion data is physically plausible. We start by defining and estimating the physical properties of human characters. The procedure of motion synthesis is from coarse to fine according to the objective function and physical constraints. Our motion synthesis is like a motion editing method, which is appropriate for motion correction and extrapolation. By this means, we can correct and eliminate unrealistic motion data. Zhidong Xiao, Xiaosong Yang, Jian J. Zhang 0001 |
IV | 2 |
| 2002 | A CSCW Method for Designing CRT Correcting LensabstractThe design of CRT correcting lens is a key and bottleneck technology in CRT production. Here we use a CSCW method with an improved algorithm to design the CRT correcting lens. The cooperation relation and cooperation workflow of color CRT design is given. The design procedure becomes simpler and more efficient by using CSCW in our project. This work is done on the Windows 2000 platform with the support of Visual C++ and Visual SourceSafe. Yuxiu Cao, Xipeng Tong, Xiaosong Yang |
CSCWD | 3 |