VLDB 2026 Research / reviewers in the wild / expert
Jinxiang Chai
dblp:62/1586
· DBLP profile ↗
47ranked-venue papers
7as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A causal convolutional neural network for multi-subject motion modeling and generationabstractInspired by the success of WaveNet in multi-subject speech synthesis, we propose a novel neural network based on causal convolutions for multi-subject motion modeling and generation. The network can capture the intrinsic characteristics of the motion of different subjects, such as the influence of skeleton scale variation on motion style. Moreover, after fine-tuning the network using a small motion dataset for a novel skeleton that is not included in the training dataset, it is able to synthesize high-quality motions with a personalized style for the novel skeleton. The experimental results demonstrate that our network can model the intrinsic characteristics of motions well and can be applied to various motion modeling and synthesis tasks. Shuaiying Hou, Congyi Wang, Wenlin Zhuang, Yangang Wang 0001, Hujun Bao, Jinxiang Chai, Weiwei Xu 0003 |
Comput. Vis. Media | 7 |
| 2022 | Music2Dance: DanceNet for Music-Driven Dance GenerationabstractSynthesize human motions from music (i.e., music to dance) is appealing and has attracted lots of research interests in recent years. It is challenging because of the requirement for realistic and complex human motions for dance, but more importantly, the synthesized motions should be consistent with the style, rhythm, and melody of the music. In this article, we propose a novel autoregressive generative model, DanceNet, to take the style, rhythm, and melody of music as the control signals to generate 3D dance motions with high realism and diversity. Due to the high long-term spatio-temporal complexity of dance, we propose the dilated convolution to improve the receptive field, and adopt the gated activation unit as well as separable convolution to enhance the fusion of motion features and control signals. To boost the performance of our proposed model, we capture several synchronized music-dance pairs by professional dancers and build a high-quality music-dance pair dataset. Experiments have demonstrated that the proposed method can achieve state-of-the-art results. Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang 0001, Ming Shao, Si-Yu Xia |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Live speech portraits: real-time photorealistic talking-head animationabstractTo the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural network that extracts deep audio features along with a manifold projection to project the features to the target person's speech space. In the second stage, we learn facial dynamics and motions from the projected audio features. The predicted motions include head poses and upper body motions, where the former is generated by an autoregressive probabilistic model which models the head pose distribution of the target person. Upper body motions are deduced from head poses. In the final stage, we generate conditional feature maps from previous predictions and send them with a candidate image set to an image-to-image translation network to synthesize photorealistic renderings. Our method generalizes well to wild audio and successfully synthesizes high-fidelity personalized facial details, e.g., wrinkles, teeth. Our method also allows explicit control of head poses. Extensive qualitative and quantitative evaluations, along with user studies, demonstrate the superiority of our method over state-of-the-art techniques. Yuanxun Lu, Jinxiang Chai, Xun Cao |
ACM Trans. Graph. | 2 |
| 2021 | Combining Recurrent Neural Networks and Adversarial Training for Human Motion Synthesis and ControlabstractThis paper introduces a new generative deep learning network for human motion synthesis and control. Our key idea is to combine recurrent neural networks (RNNs) and adversarial training for human motion modeling. We first describe an efficient method for training an RNN model from prerecorded motion data. We implement RNNs with long short-term memory (LSTM) cells because they are capable of addressing the nonlinear dynamics and long term temporal dependencies present in human motions. Next, we train a refiner network using an adversarial loss, similar to generative adversarial networks (GANs), such that refined motion sequences are indistinguishable from real mocap data using a discriminative network. The resulting model is appealing for motion synthesis and control because it is compact, contact-aware, and can generate an infinite number of naturally looking motions with infinite lengths. Our experiments show that motions generated by our deep learning model are always highly realistic and comparable to high-quality motion capture data. We demonstrate the power and effectiveness of our models by exploring a variety of applications, ranging from random motion synthesis, online/offline motion control, and motion filtering. We show the superiority of our generative model by comparison against baseline models. Jinxiang Chai, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Realtime and Accurate 3D Eye Gaze Capture with DCNN-Based Iris and Pupil SegmentationabstractThis paper presents a realtime and accurate method for 3D eye gaze tracking with a monocular RGB camera. Our key idea is to train a deep convolutional neural network(DCNN) that automatically extracts the iris and pupil pixels of each eye from input images. To achieve this goal, we combine the power of Unet\cite{ronneberger2015u-net:} and Squeezenet\cite{iandola2017squeezenet:} to train an efficient convolutional neural network for pixel classification. In addition, we track the 3D eye gaze state in the Maximum A Posteriori (MAP) framework, which sequentially searches for the most likely state of the 3D eye gaze at each frame. When eye blinking occurs, the eye gaze tracker can obtain an inaccurate result. We further extend the convolutional neural network for eye close detection in order to improve the robustness and accuracy of the eye gaze tracker. Our system runs in realtime on desktop PCs and smart phones. We have evaluated our system on live videos and Internet videos, and our results demonstrate that the system is robust and accurate for various genders, races, lighting conditions, poses, shapes and facial expressions. A comparison against Wang et al.[3] shows that our method advances the state of the art in 3D eye tracking using a single RGB camera. Jinxiang Chai, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | A Survey on Human Performance Capture and Animation
Shihong Xia, Lin Gao 0004, Yukun Lai, Mingzhe Yuan, Jinxiang Chai |
J. Comput. Sci. Technol. | 5 |
| 2017 | Motion Capture With Ellipsoidal Skeleton Using Multiple Depth CamerasabstractThis paper introduces a novel motion capturing framework which works by minimizing the fitting error between an ellipsoid based skeleton and the input point cloud data captured by multiple depth cameras. The novelty of this method comes from that it uses the ellipsoids equipped with the spherical harmonics encoded displacement and normal functions to capture the geometry details of the tracked object. This method is also integrated with a mechanism to avoid collisions of bones during the motion capturing process. The method is implemented parallelly with CUDA on GPU and has a fast running speed without dedicated code optimization. The errors of the proposed method on the data from Berkeley Multimodal Human Action Database (MHAD) are within a reasonable range compared with the ground truth results. Our experiment shows that this method succeeds on many challenging motions which are failed to be reported by Microsoft Kinect SDK and not tested by existing works. In the comparison with the state-of-art marker-less depth camera based motion tracking work our method shows advantages in both robustness and input data modality. Liang Shuai, Chao Li 0021, Xiaohu Guo, B. Prabhakaran 0001, Jinxiang Chai |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Data-driven inverse dynamics for human motionabstractInverse dynamics is an important and challenging problem in human motion modeling, synthesis and simulation, as well as in robotics and biomechanics. Previous solutions to inverse dynamics are often noisy and ambiguous particularly when double stances occur. In this paper, we present a novel inverse dynamics method that accurately reconstructs biomechanically valid contact information, including center of pressure, contact forces, torsional torques and internal joint torques from input kinematic human motion data. Our key idea is to apply statistical modeling techniques to a set of preprocessed human kinematic and dynamic motion data captured by a combination of an optical motion capture system, pressure insoles and force plates. We formulate the data-driven inverse dynamics problem in a maximum a posteriori (MAP) framework by estimating the most likely contact information and internal joint torques that are consistent with input kinematic motion data. We construct a low-dimensional data-driven prior model for contact information and internal joint torques to reduce ambiguity of inverse dynamics for human motion. We demonstrate the accuracy of our method on a wide variety of human movements including walking, jumping, running, turning and hopping and achieve state-of-the-art accuracy in our comparison against alternative methods. In addition, we discuss how to extend the data-driven inverse dynamics framework to motion editing, filtering and motion control. Xiaolei Lv, Jinxiang Chai, Shihong Xia |
ACM Trans. Graph. | 2 |
| 2016 | Realtime 3D eye gaze animation using a single RGB cameraabstractThis paper presents the first realtime 3D eye gaze capture method that simultaneously captures the coordinated movement of 3D eye gaze, head poses and facial expression deformation using a single RGB camera. Our key idea is to complement a realtime 3D facial performance capture system with an efficient 3D eye gaze tracker. We start the process by automatically detecting important 2D facial features for each frame. The detected facial features are then used to reconstruct 3D head poses and large-scale facial deformation using multi-linear expression deformation models. Next, we introduce a novel user-independent classification method for extracting iris and pupil pixels in each frame. We formulate the 3D eye gaze tracker in the Maximum A Posterior (MAP) framework, which sequentially infers the most probable state of 3D eye gaze at each frame. The eye gaze tracker could fail when eye blinking occurs. We further introduce an efficient eye close detector to improve the robustness and accuracy of the eye gaze tracker. We have tested our system on both live video streams and the Internet videos, demonstrating its accuracy and robustness under a variety of uncontrolled lighting conditions and overcoming significant differences of races, genders, shapes, poses and expressions across individuals. Congyi Wang, Fuhao Shi, Shihong Xia, Jinxiang Chai |
ACM Trans. Graph. | 4 |
| 2015 | A Suggestive Interface for Sketch-based Character PosingabstractWe present a user-friendly suggestive interface for sketch-based character posing. Our interface provides suggestive information on the sketching canvas in succession by combining image retrieval technique with 3D character posing, while the user is drawing. The system highlights the canvas region where the user should draw on and constrains the user's sketches in a reasonable solution space. This is based on an efficient image descriptor, which is used to measure the distance between the user's sketch and 2D views of 3D poses. In order to achieve faster query response, local sensitive hashing is involved in our system. In addition, sampling-based optimization algorithm is adopted to synthesize and optimize the retrieved 3D pose to match the user's sketches the best. Experiments show that our interface can provide smooth suggestive information to improve the reality of sketching poses and shorten the time required for 3D posing. Pei Lv, Pengjie Wang 0001, Weiwei Xu 0003, Jinxiang Chai |
Comput. Graph. Forum | 4 |
| 2015 | Video-audio driven real-time facial animationabstractWe present a real-time facial tracking and animation system based on a Kinect sensor with video and audio input. Our method requires no user-specific training and is robust to occlusions, large head rotations, and background noise. Given the color, depth and speech audio frames captured from an actor, our system first reconstructs 3D facial expressions and 3D mouth shapes from color and depth input with a multi-linear model. Concurrently a speaker-independent DNN acoustic model is applied to extract phoneme state posterior probabilities (PSPP) from the audio frames. After that, a lip motion regressor refines the 3D mouth shape based on both PSPP and expression weights of the 3D mouth shapes, as well as their confidences. Finally, the refined 3D mouth shape is combined with other parts of the 3D face to generate the final result. The whole process is fully automatic and executed in real time. The key component of our system is a data-driven regresor for modeling the correlation between speech data and mouth shapes. Based on a precaptured database of accurate 3D mouth shapes and associated speech audio from one speaker, the regressor jointly uses the input speech and visual features to refine the mouth shape of a new actor. We also present an improved DNN acoustic model. It not only preserves accuracy but also achieves real-time performance. Our method efficiently fuses visual and acoustic information for 3D facial performance capture. It generates more accurate 3D mouth motions than other approaches that are based on audio or video input only. It also supports video or audio only input for real-time facial animation. We evaluate the performance of our system with speech and facial expressions captured from different actors. Results demonstrate the efficiency and robustness of our method. Feng Xu 0005, Jinxiang Chai, Xin Tong 0001, Qiang Huo |
ACM Trans. Graph. | 3 |
| 2015 | Realtime style transfer for unlabeled heterogeneous human motionabstractThis paper presents a novel solution for realtime generation of stylistic human motion that automatically transforms unlabeled, heterogeneous motion data into new styles. The key idea of our approach is an online learning algorithm that automatically constructs a series of local mixtures of autoregressive models (MAR) to capture the complex relationships between styles of motion. We construct local MAR models on the fly by searching for the closest examples of each input pose in the database. Once the model parameters are estimated from the training data, the model adapts the current pose with simple linear transformations. In addition, we introduce an efficient local regression model to predict the timings of synthesized poses in the output style. We demonstrate the power of our approach by transferring stylistic human motion for a wide variety of actions, including walking, running, punching, kicking, jumping and transitions between those behaviors. Our method achieves superior performance in a comparison against alternative methods. We have also performed experiments to evaluate the generalization ability of our data-driven model as well as the key components of our system. Shihong Xia, Congyi Wang, Jinxiang Chai, Jessica K. Hodgins |
ACM Trans. Graph. | 3 |
| 2014 | Automatic acquisition of high-fidelity facial performances using monocular videosabstractThis paper presents a facial performance capture system that automatically captures high-fidelity facial performances using uncontrolled monocular videos ( e.g ., Internet videos). We start the process by detecting and tracking important facial features such as the nose tip and mouth corners across the entire sequence and then use the detected facial features along with multilinear facial models to reconstruct 3D head poses and large-scale facial deformation of the subject at each frame. We utilize per-pixel shading cues to add fine-scale surface details such as emerging or disappearing wrinkles and folds into large-scale facial deformation. At a final step, we iterate our reconstruction procedure on large-scale facial geometry and fine-scale facial details to further improve the accuracy of facial reconstruction. We have tested our system on monocular videos downloaded from the Internet, demonstrating its accuracy and robustness under a variety of uncontrolled lighting conditions and overcoming significant shape differences across individuals. We show our system advances the state of the art in facial performance capture by comparing against alternative methods. Fuhao Shi, Hsiang-Tao Wu, Xin Tong 0001, Jinxiang Chai |
ACM Trans. Graph. | 4 |
| 2014 | Controllable high-fidelity facial performance transferabstractRecent technological advances in facial capture have made it possible to acquire high-fidelity 3D facial performance data with stunningly high spatial-temporal resolution. Current methods for facial expression transfer, however, are often limited to large-scale facial deformation. This paper introduces a novel facial expression transfer and editing technique for high-fidelity facial performance data. The key idea of our approach is to decompose high-fidelity facial performances into high-level facial feature lines, large-scale facial deformation and fine-scale motion details and transfer them appropriately to reconstruct the retargeted facial animation in an efficient optimization framework. The system also allows the user to quickly modify and control the retargeted facial sequences in the spatial-temporal domain. We demonstrate the power of our approach by transferring and editing high-fidelity facial animation data from high-resolution source models to a wide range of target models, including both human faces and non-human faces such as "monster" and "dog". Feng Xu 0005, Jinxiang Chai, Xin Tong 0001 |
ACM Trans. Graph. | 2 |
| 2014 | Leveraging depth cameras and wearable pressure sensors for full-body kinematics and dynamics captureabstractWe present a new method for full-body motion capture that uses input data captured by three depth cameras and a pair of pressure-sensing shoes. Our system is appealing because it is low-cost, non-intrusive and fully automatic, and can accurately reconstruct both full-body kinematics and dynamics data. We first introduce a novel tracking process that automatically reconstructs 3D skeletal poses using input data captured by three Kinect cameras and wearable pressure sensors. We formulate the problem in an optimization framework and incrementally update 3D skeletal poses with observed depth data and pressure data via iterative linear solvers. The system is highly accurate because we integrate depth data from multiple depth cameras, foot pressure data, detailed full-body geometry, and environmental contact constraints into a unified framework. In addition, we develop an efficient physics-based motion reconstruction algorithm for solving internal joint torques and contact forces in the quadratic programming framework. During reconstruction, we leverage Newtonian physics, friction cone constraints, contact pressure information, and 3D kinematic poses obtained from the kinematic tracking process to reconstruct full-body dynamics data. We demonstrate the power of our approach by capturing a wide range of human movements and achieve state-of-the-art accuracy in our comparison against alternative systems. Peizhao Zhang, Kristin Siu, Jianjie Zhang, C. Karen Liu, Jinxiang Chai |
ACM Trans. Graph. | 5 |
| 2013 | Accurate and Robust 3D Facial Capture Using a Single RGBD CameraabstractThis paper presents an automatic and robust approach that accurately captures high-quality 3D facial performances using a single RGBD camera. The key of our approach is to combine the power of automatic facial feature detection and image-based 3D nonrigid registration techniques for 3D facial reconstruction. In particular, we develop a robust and accurate image-based nonrigid registration algorithm that incrementally deforms a 3D template mesh model to best match observed depth image data and important facial features detected from single RGBD images. The whole process is fully automatic and robust because it is based on single frame facial registration framework. The system is flexible because it does not require any strong 3D facial priors such as blend shape models. We demonstrate the power of our approach by capturing a wide range of 3D facial expressions using a single RGBD camera and achieve state-of-the-art accuracy by comparing against alternative methods. Yen-Lin Chen, Hsiang-Tao Wu, Fuhao Shi, Xin Tong 0001, Jinxiang Chai |
ICCV | 5 |
| 2013 | Video-based hand manipulation capture through composite motion controlabstractThis paper describes a new method for acquiring physically realistic hand manipulation data from multiple video streams. The key idea of our approach is to introduce a composite motion control to simultaneously model hand articulation, object movement, and subtle interaction between the hand and object. We formulate video-based hand manipulation capture in an optimization framework by maximizing the consistency between the simulated motion and the observed image data. We search an optimal motion control that drives the simulation to best match the observed image data. We demonstrate the effectiveness of our approach by capturing a wide range of high-fidelity dexterous manipulation data. We show the power of our recovered motion controllers by adapting the captured motion data to new objects with different properties. The system achieves superior performance against alternative methods such as marker-based motion capture and kinematic hand motion tracking. Yangang Wang 0001, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu 0005, Qionghai Dai, Jinxiang Chai |
ACM Trans. Graph. | 7 |
| 2013 | Robust realtime physics-based motion control for human graspingabstractThis paper presents a robust physics-based motion control system for realtime synthesis of human grasping. Given an object to be grasped, our system automatically computes physics-based motion control that advances the simulation to achieve realistic manipulation with the object. Our solution leverages prerecorded motion data and physics-based simulation for human grasping. We first introduce a data-driven synthesis algorithm that utilizes large sets of prerecorded motion data to generate realistic motions for human grasping. Next, we present an online physics-based motion control algorithm to transform the synthesized kinematic motion into a physically realistic one. In addition, we develop a performance interface for human grasping that allows the user to act out the desired grasping motion in front of a single Kinect camera. We demonstrate the power of our approach by generating physics-based motion control for grasping objects with different properties such as shapes, weights, spatial orientations, and frictions. We show our physics-based motion control for human grasping is robust to external perturbations and changes in physical quantities. Wenping Zhao, Jianjie Zhang, Jianyuan Min, Jinxiang Chai |
ACM Trans. Graph. | 4 |
| 2012 | Motion graphs++: a compact generative model for semantic motion analysis and synthesisabstractThis paper introduces a new generative statistical model that allows for human motion analysis and synthesis at both semantic and kinematic levels. Our key idea is to decouple complex variations of human movements into finite structural variations and continuous style variations and encode them with a concatenation of morphable functional models. This allows us to model not only a rich repertoire of behaviors but also an infinite number of style variations within the same action. Our models are appealing for motion analysis and synthesis because they are highly structured, contact aware , and semantic embedding . We have constructed a compact generative motion model from a huge and heterogeneous motion database (about two hours mocap data and more than 15 different actions). We have demonstrated the power and effectiveness of our models by exploring a wide variety of applications, ranging from automatic motion segmentation, recognition, and annotation, and online/offline motion synthesis at both kinematics and behavior levels to semantic motion editing. We show the superiority of our model by comparing it with alternative methods. Jianyuan Min, Jinxiang Chai |
ACM Trans. Graph. | 2 |
| 2012 | Accurate realtime full-body motion capture using a single depth cameraabstractWe present a fast, automatic method for accurately capturing full-body motion data using a single depth camera. At the core of our system lies a realtime registration process that accurately reconstructs 3D human poses from single monocular depth images, even in the case of significant occlusions. The idea is to formulate the registration problem in a Maximum A Posteriori (MAP) framework and iteratively register a 3D articulated human body model with monocular depth cues via linear system solvers. We integrate depth data, silhouette information, full-body geometry, temporal pose priors, and occlusion reasoning into a unified MAP estimation framework. Our 3D tracking process, however, requires manual initialization and recovery from failures. We address this challenge by combining 3D tracking with 3D pose detection. This combination not only automates the whole process but also significantly improves the robustness and accuracy of the system. Our whole algorithm is highly parallel and is therefore easily implemented on a GPU. We demonstrate the power of our approach by capturing a wide range of human movements in real time and achieve state-of-the-art accuracy in our comparison against alternative systems such as Kinect [2012]. Xiaolin K. Wei, Peizhao Zhang, Jinxiang Chai |
ACM Trans. Graph. | 3 |
| 2011 | Realtime human motion control with a small number of inertial sensorsabstractThis paper introduces an approach to performance animation that employs a small number of motion sensors to create an easy-to-use system for an interactive control of a full-body human character. Our key idea is to construct a series of online local dynamic models from a prerecorded motion database and utilize them to construct full-body human motion in a maximum a posteriori framework (MAP). We have demonstrated the effectiveness of our system by controlling a variety of human actions, such as boxing, golf swinging, and table tennis, in real time. Given an appropriate motion capture database, the results are comparable in quality to those obtained from a commercial motion capture system with a full set of motion sensors (e.g., XSens [2009]); however, our performance animation system is far less intrusive and expensive because it requires a small of motion sensors for full body control. We have also evaluated the performance of our system by leave-one-out-experiments and by comparing with two baseline algorithms. Huajun Liu, Xiaolin K. Wei, Jinxiang Chai, Inwoo Ha, Taehyun Rhee |
SI3D | 3 |
| 2011 | EditorialabstractThis special issue contains 28 papers selected from the Computer Animation and Social Agents 2011 (CASA'2011) Conference. This conference was founded by the Computer Graphics Society in 1988 in Geneva and has, since then, been held in various countries. Last year in France, this year in China and it will be held next year in Singapore. This year we received 163 papers and selected only 28 of those for this special issue; meaning an acceptance rate of only 17.2%. It is needless to say that these are of high quality; all having been reviewed by at least three reviewers. Animation techniques: Motion Control, Motion Capture and Retargeting, Path Planning, Physics-Based Animation, Image-Based Animation, Behavioral Animation, Artificial Life, Deformation, Facial Animation, Multi-Resolution and Multi-Scale Models, Knowledge-Based Animation and Motion Synthesis. Social agents: Social Agents and Avatars, Emotion and Personality, Virtual Humans, Autonomous Actors, AI-Based Animation, Social and Conversational Agents, Inter-Agent Communication, Social Behavior, Gesture Generation and Crowd Simulation. Other related: Animation Compression and Transmission, Semantics and Ontologies for Virtual Humans and Virtual Environments, Animation Analysis and Structuring, Anthropometric Virtual Human Models, Acquisition and Reconstruction of Animation Data, Level of Details, Semantic Representation of Motion and Animation, Medical Simulation, Cultural Heritage, Interaction for Virtual Humans, Augmented Reality and Virtual Reality, Computer Games and Online Virtual Worlds. We would like to thank Prof. Yueting Zhuang from Zhejiang University in China, Prof. Daniel Thalmann from EPFL in Switzerland and Prof. Enhua Wu from University of Macao in China, the co-conference chairs of CASA'2011, for their strong support in the conference. We also like to thank the organizing co-chairs, Prof. Jieqing Feng from Zhejiang University in China and Prof. Leiting Chen, from University of Electronic Science and Technology of China, for their intense collaboration as well as the international program committee for their strong commitment and the external referees. We would like also to thank the authors for having submitted a paper to CASA'2011. Conference Co-Chairs Yueting Zhuang Daniel Thalmann Enhua Wu Program Co-Chairs Zhigeng Pan Nadia Magnenat-Thalmann Jinxiang Chai International Program Committee Abdennour El-Rhalibi Ahmad Nasri Ana Paiva Anton Nijholt Arie Kaufman Arjan Egges Bing-Yu Chen Carlos Martinho Catherine Pelachaud Chris Joslin Courty Nicolas Daniel Thalmann Dinesh Manocha Dinesh Pai Donald House Dumont Georges Elisabeth Andre Enhua Wu Fabian Di Fiore Feng Dong Florence Bertails Franck Multon Grisoni Laurent Hans-Peter Seidel Herwin Welbergen Hujun Bao Hwan-Gue Cho Hyewon Seo Igor Pandzic J.P. Lewis James Hahn Jean-Paul Laumond Jian Zhang Jieqing Feng Jinhui Yu Jinxiang Chai John Patterson Jos Stam Julien Pettre Kangkang Yin Kuffner James Kulpa Richard Lau Rynson Lee Tong-Yee Louis-Philippe Morency Luiz Velho Magnenat-Thalmann Marc Cavazza Marcelo Kallmann Marie-Paule Cani Mark Overmars Martin Jean-Claude Massimo Bergamasco Matthias Teschner Michael Gleicher Min-Hyung Choi Nancy Amato Neeharika Adabala Ning Wang Norman Badler Paolo Petta Petros Faloutsos Philip Willis Porcher-Nedel Luciana Prem Kalra Qunsheng Peng Raupp-Musse Soraia Rick Parent Stephane Donikian Stephane Redon Sung-Yong Shin Taku Komura Tolga Capin Van-Reeth Frank Weidong Geng Weiwei Xu William Baxter Wonsook Lee Xiaogang Jin Yangsheng Wang Ying-Qing Xu Yiying Tong Yizhou Yu Yu Qizhi Yueting Zhuang Zhigang Deng Zsofi Ruttkay Nadia Magnenat-Thalmann, Jinxiang Chai |
Comput. Animat. Virtual Worlds | 3 |
| 2011 | Leveraging motion capture and 3D scanning for high-fidelity facial performance acquisitionabstractThis paper introduces a new approach for acquiring high-fidelity 3D facial performances with realistic dynamic wrinkles and fine-scale facial details. Our approach leverages state-of-the-art motion capture technology and advanced 3D scanning technology for facial performance acquisition. We start the process by recording 3D facial performances of an actor using a marker-based motion capture system and perform facial analysis on the captured data, thereby determining a minimal set of face scans required for accurate facial reconstruction. We introduce a two-step registration process to efficiently build dense consistent surface correspondences across all the face scans. We reconstruct high-fidelity 3D facial performances by combining motion capture data with the minimal set of face scans in the blendshape interpolation framework. We have evaluated the performance of our system on both real and synthetic data. Our results show that the system can capture facial performances that match both the spatial resolution of static face scans and the acquisition speed of motion capture systems. Hao-Da Huang, Jinxiang Chai, Xin Tong 0001, Hsiang-Tao Wu |
ACM Trans. Graph. | 2 |
| 2011 | Physically valid statistical models for human motion generationabstractThis article shows how statistical motion priors can be combined seamlessly with physical constraints for human motion modeling and generation. The key idea of the approach is to learn a nonlinear probabilistic force field function from prerecorded motion data with Gaussian processes and combine it with physical constraints in a probabilistic framework. In addition, we show how to effectively utilize the new model to generate a wide range of natural-looking motions that achieve the goals specified by users. Unlike previous statistical motion models, our model can generate physically realistic animations that react to external forces or changes in physical quantities of human bodies and interaction environments. We have evaluated the performance of our system by comparing against ground-truth motion data and alternative methods. Xiaolin K. Wei, Jianyuan Min, Jinxiang Chai |
ACM Trans. Graph. | 3 |
| 2010 | Synthesis and editing of personalized stylistic human motionabstractThis paper presents a generative human motion model for synthesis, retargeting, and editing of personalized human motion styles. We first record a human motion database from multiple actors performing a wide variety of motion styles for particular actions. We then apply multilinear analysis techniques to construct a generative motion model of the form x = g(a, e) for particular human actions, where the parameters a and e control "identity" and "style" variations of the motion x respectively. The new modular representation naturally supports motion generalization to new actors and/or styles. We demonstrate the power and flexibility of the multilinear motion models by synthesizing personalized stylistic human motion and transferring the stylistic motions from one actor to another. We also show the effectiveness of our model by editing stylistic motion in style and/or identity space. Jianyuan Min, Huajun Liu, Jinxiang Chai |
SI3D | 3 |
| 2010 | VideoMocap: modeling physically realistic human motion from monocular video sequencesabstractThis paper presents a video-based motion modeling technique for capturing physically realistic human motion from monocular video sequences. We formulate the video-based motion modeling process in an image-based keyframe animation framework. The system first computes camera parameters, human skeletal size, and a small number of 3D key poses from video and then uses 2D image measurements at intermediate frames to automatically calculate the "in between" poses. During reconstruction, we leverage Newtonian physics, contact constraints, and 2D image measurements to simultaneously reconstruct full-body poses, joint torques, and contact forces. We have demonstrated the power and effectiveness of our system by generating a wide variety of physically realistic human actions from uncalibrated monocular video sequences such as sports video footage. Xiaolin K. Wei, Jinxiang Chai |
ACM Trans. Graph. | 2 |
| 2010 | Example-Based Human Motion DenoisingabstractWith the proliferation of motion capture data, interest in removing noise and outliers from motion capture data has increased. In this paper, we introduce an efficient human motion denoising technique for the simultaneous removal of noise and outliers from input human motion data. The key idea of our approach is to learn a series of filter bases from precaptured motion data and use them along with robust statistics techniques to filter noisy motion data. Mathematically, we formulate the motion denoising process in a nonlinear optimization framework. The objective function measures the distance between the noisy input and the filtered motion in addition to how well the filtered motion preserves spatial-temporal patterns embedded in captured human motion data. Optimizing the objective function produces an optimal filtered motion that keeps spatial-temporal patterns in captured motion data. We also extend the algorithm to fill in the missing values in input motion data. We demonstrate the effectiveness of our system by experimenting with both real and simulated motion data. We also show the superior performance of our algorithm by comparing it with three baseline algorithms and to those in state-of-art motion capture data processing software such as Vicon Blade. Hui Lou, Jinxiang Chai |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2009 | 3D Reconstruction of Human Motion and Skeleton from Uncalibrated Monocular Video
Yen-Lin Chen, Jinxiang Chai |
ACCV (1) | 2 |
| 2009 | Modeling 3D human poses from uncalibrated monocular imagesabstractThis paper introduces an efficient algorithm that reconstructs 3D human poses as well as camera parameters from a small number of 2D point correspondences obtained from uncalibrated monocular images. This problem is challenging because 2D image constraints (e.g. 2D point correspondences) are often not sufficient to determine 3D poses of an articulated object. The key idea of this paper is to identify a set of new constraints and use them to eliminate the ambiguity of 3D pose reconstruction. We also develop an optimization process to simultaneously reconstruct both human poses and camera parameters from various forms of reconstruction constraints. We demonstrate the power and effectiveness of our system by evaluating the performance of the algorithm on both real and synthetic data. We show the algorithm can accurately reconstruct 3D poses and camera parameters from a wide variety of real images, including internet photos and key frames extracted from monocular video sequences. Xiaolin K. Wei, Jinxiang Chai |
ICCV | 2 |
| 2009 | Flexible registration of human motion data with parameterized motion modelsabstractThis paper presents an efficient model-based approach for automatic human motion registration, which builds temporal correspondences between structurally similar but distinctive motion examples. The key idea of the model-based registration process is to construct a parameterized motion model from a set of preregistered motion examples. With such a model, we can register an input motion with the parameterized motion model by continuously deforming the model to best match the input motion. We formulate the registration process in a gradient-based nonlinear optimization framework by minimizing an objective function that measures differences between the input motion and deforming motion. We also develop a multi-resolution optimization process to efficiently estimate the model parameters as well as the temporal correspondences between the input motion and deforming motion. We demonstrate the performance of our approach by testing the algorithm on difficult motion sequences and comparing with alternative approaches. Yen-Lin Chen, Jianyuan Min, Jinxiang Chai |
SI3D | 3 |
| 2009 | Face poser: Interactive modeling of 3D facial expressions using facial priorsabstractThis article presents an intuitive and easy-to-use system for interactively posing 3D facial expressions. The user can model and edit facial expressions by drawing freeform strokes, by specifying distances between facial points, by incrementally editing curves on the face, or by directly dragging facial points in 2D screen space. Designing such an interface for 3D facial modeling and editing is challenging because many unnatural facial expressions might be consistent with the user's input. We formulate the problem in a maximum a posteriori framework by combining the user's input with priors embedded in a large set of facial expression data. Maximizing the posteriori allows us to generate an optimal and natural facial expression that achieves the goal specified by the user. We evaluate the performance of our system by conducting a thorough comparison of our method with alternative facial modeling techniques. To demonstrate the usability of our system, we also perform a user study of our system and compare with state-of-the-art facial expression modeling software (Poser 7). Manfred Lau, Jinxiang Chai, Ying-Qing Xu, Harry Shum |
ACM Trans. Graph. | 2 |
| 2009 | Interactive generation of human animation with deformable motion modelsabstractThis article presents a new motion model deformable motion models for human motion modeling and synthesis. Our key idea is to apply statistical analysis techniques to a set of precaptured human motion data and construct a low-dimensional deformable motion model of the form x = M (α, γ), where the deformable parameters α and γ control the motion's geometric and timing variations, respectively. To generate a desired animation, we continuously adjust the deformable parameters' values to match various forms of user-specified constraints. Mathematically, we formulate the constraint-based motion synthesis problem in a Maximum A Posteriori (MAP) framework by estimating the most likely deformable parameters from the user's input. We demonstrate the power and flexibility of our approach by exploring two interactive and easy-to-use interfaces for human motion generation: direct manipulation interfaces and sketching interfaces. Jianyuan Min, Yen-Lin Chen, Jinxiang Chai |
ACM Trans. Graph. | 3 |
| 2008 | A hybrid camera for motion deblurring and depth map super-resolutionabstractWe present a hybrid camera that combines the advantages of a high resolution camera and a high speed camera. Our hybrid camera consists of a pair of low-resolution high-speed (LRHS) cameras and a single high-resolution low-speed (HRLS) camera. The LRHS cameras are able to capture fast-motion with little motion blur. They also form a stereo pair and provide a low-resolution depth map. The HRLS camera provides a high spatial resolution but also introduces severe motion blur when capturing fast moving objects. We develop efficient algorithms to simultaneously motion-deblur the HRLS image and reconstruct a high resolution depth map. Our method estimates the motion flow in the LRHS pair and then warps the flow field to the HRLS camera to estimate the point spread function (PSF).We then deblur the HRLS image and use the resulting image to enhance the low-resolution depth map using joint bilateral filters. We demonstrate the hybrid camera in depth map super-resolution and motion deblurring with spatially varying kernels. Experiments show that our framework is robust and highly effective. Feng Li 0005, Jingyi Yu 0001, Jinxiang Chai |
CVPR | 3 |
| 2008 | Interactive Tracking of 2D Generic Objects with Spacetime Optimization
Xiaolin K. Wei, Jinxiang Chai |
ECCV (1) | 2 |
| 2007 | Constraint-based motion optimization using a statistical dynamic modelabstractIn this paper, we present a technique for generating animation from a variety of user-defined constraints. We pose constraint-based motion synthesis as a maximum a posterior (MAP) problem and develop an optimization framework that generates natural motion satisfying user constraints. The system automatically learns a statistical dynamic model from motion capture data and then enforces it as a motion prior. This motion prior, together with user-defined constraints, comprises a trajectory optimization problem. Solving this problem in the low-dimensional space yields optimal natural motion that achieves the goals specified by the user. We demonstrate the effectiveness of this approach by generating whole-body and facial motion from a variety of spatial-temporal constraints. Jinxiang Chai, Jessica K. Hodgins |
ACM Trans. Graph. | 1 |
| 2006 | A Closed-Form Solution to Non-Rigid Shape and Motion Recovery
Jing Xiao 0006, Jinxiang Chai, Takeo Kanade |
Int. J. Comput. Vis. | 2 |
| 2005 | Performance animation from low-dimensional control signalsabstractThis paper introduces an approach to performance animation that employs video cameras and a small set of retro-reflective markers to create a low-cost, easy-to-use system that might someday be practical for home use. The low-dimensional control signals from the user's performance are supplemented by a database of pre-recorded human motion. At run time, the system automatically learns a series of local models from a set of motion capture examples that are a close match to the marker locations captured by the cameras. These local models are then used to reconstruct the motion of the user as a full-body animation. We demonstrate the power of this approach with real-time control of six different behaviors using two video cameras and a small set of retro-reflective markers. We compare the resulting animation to animation from commercial motion capture equipment with a full set of markers. Jinxiang Chai, Jessica K. Hodgins |
ACM Trans. Graph. | 1 |
| 2004 | A Closed-Form Solution to Non-rigid Shape and Motion Recovery
Jing Xiao 0006, Jinxiang Chai, Takeo Kanade |
ECCV (4) | 2 |
| 2002 | Rendering by Manifold Hopping
Harry Shum, Lifeng Wang 0001, Jinxiang Chai, Xin Tong 0001 |
Int. J. Comput. Vis. | 3 |
| 2002 | Layered lumigraph with LOD controlabstractAbstract The rendering performance of an image‐based rendering (IBR) system is determined by the number of images and the amount of geometrical information used. In this paper, we propose a layered lumigraph representation that is configured for optimized rendering performance based on the rendering platform (e.g., processor speed, memory) and output image resolution. The layered lumigraph is produced by classifying all pixels into a number of depth layers. Based on prior work on plenoptic sampling analysis, the layered lumigraph is constructed to achieve the same rendering quality along the minimum sampling curve by balancing the number of images and depth layers. For a given rendering platform, the best rendering performance can be obtained by choosing the optimal number of images and depth layers. Moreover, the layered lumigraph is capable of level‐of‐detail (LOD) control using the same image geometry trade‐off. Therefore, the layered lumigraph fully exploits the inherent constraints between the number of images, depth complexity, and output resolution. Finally, a backward warping technique is designed to efficiently render the layered lumigraph by taking advantage of texture mapping hardware. Copyright © 2002 John Wiley & Sons, Ltd. Xin Tong 0001, Jinxiang Chai, Harry Shum |
Comput. Animat. Virtual Worlds | 2 |
| 2002 | Interactive control of avatars animated with human motion dataabstractReal-time control of three-dimensional avatars is an important problem in the context of computer games and virtual environments. Avatar animation and control is difficult, however, because a large repertoire of avatar behaviors must be made available, and the user must be able to select from this set of behaviors, possibly with a low-dimensional input device. One appealing approach to obtaining a rich set of avatar behaviors is to collect an extended, unlabeled sequence of motion data appropriate to the application. In this paper, we show that such a motion database can be preprocessed for flexibility in behavior and efficient search and exploited for real-time avatar control. Flexibility is created by identifying plausible transitions between motion segments, and efficient search through the resulting graph structure is obtained through clustering. Three interface techniques are demonstrated for controlling avatar motion using this data structure: the user selects from a set of available choices, sketches a path through an environment, or acts out a desired motion in front of a video camera. We demonstrate the flexibility of the approach through four different applications and compare the avatar motion to directly recorded human motion. Jehee Lee, Jinxiang Chai, Paul S. A. Reitsma, Jessica K. Hodgins, Nancy S. Pollard |
ACM Trans. Graph. | 2 |
| 2001 | Handling Occlusions in Dense Multi-view StereoabstractWhile stereo matching was originally formulated as the recovery of 3D shape from a pair of images, it is now generally recognized that using more than two images can dramatically improve the quality of the reconstruction. Unfortunately, as more images are added, the prevalence of semi-occluded regions (pixels visible in some but not all images) also increases. We propose some novel techniques to deal with this problem. Our first idea is to use a combination of shiftable windows and a dynamically selected subset of the neighboring images to do the matches. Our second idea is to explicitly label occluded pixels within a global energy minimization framework, and to reason about visibility within this framework so that only truly visible pixels are matched. Experimental results show a dramatic improvement using the first idea over conventional multibaseline stereo, especially when used in conjunction with a global energy minimization technique. These results also show that explicit occlusion labeling and visibility reasoning do help, but not significantly, if the spatial and temporal selection is applied first. Sing Bing Kang, Richard Szeliski, Jinxiang Chai |
CVPR (1) | 3 |
| 2000 | Parallel Projections for Stereo ReconstructionabstractThis paper proposes a novel technique to computing geometric information from images captured under parallel projections. Parallel images are desirable for stereo reconstruction because parallel projection significantly reduces foreshortening. As a result, correlation based matching becomes more effective. Since parallel projection cameras are not commonly available, we construct parallel images by rebinning a large sequence of perspective images. Epipolar geometry, depth recovery and projective invariant for both 1D and 2D parallel stereos are studied. From the uncertainty analysis of depth reconstruction, it is shown that parallel stereo is superior to both conventional perspective stereo and the recently developed multiperspective stereo for vision reconstruction, in that uniform reconstruction error is obtained in parallel stereo. Traditional stereo reconstruction techniques, e.g. multi-baseline stereo, can still be applicable to parallel stereo without any modifications because epipolar lines in a parallel stereo are perfectly straight. Experimental results further confirm the performance of our approach. Jinxiang Chai, Harry Shum |
CVPR | 1 |
| 2000 | Plenoptic samplingabstractThis paper studies the problem of plenoptic sampling in image-based rendering (IBR). From a spectral analysis of light field signals and using the sampling theorem, we mathematically derive the analytical functions to determine the minimum sampling rate for light field rendering. The spectral support of a light field signal is bounded by the minimum and maximum depths only, no matter how complicated the spectral support might be because of depth variations in the scene. The minimum sampling rate for light field rendering is obtained by compacting the replicas of the spectral support of the sampled light field within the smallest interval. Given the minimum and maximum depths, a reconstruction filter with an optimal and constant depth can be designed to achieve anti-aliased light field rendering. Plenoptic sampling goes beyond the minimum number of images needed for anti-aliased light field rendering. More significantly, it utilizes the scene depth information to determine the minimum sampling curve in the joint image and geometry space. The minimum sampling curve quantitatively describes the relationship among three key elements in IBR systems: scene complexity (geometrical and textural information), the number of image samples, and the output resolution. Therefore, plenoptic sampling bridges the gap between image-based rendering and traditional geometry-based rendering. Experimental results demonstrate the effectiveness of our approach. Jinxiang Chai, S. C. Chan 0001, Harry Shum, Xin Tong 0001 |
SIGGRAPH | 1 |
| 1998 | Robust Epipolar Geometry Estimation Using Genetic Algorithm
Jinxiang Chai, Songde Ma |
ACCV (1) | 1 |
| 1998 | An evolutionary framework for stereo correspondenceabstractIn this paper we propose an evolutionary framework to establish feature correspondence from two uncalibrated images. By minimizing a proposed cost function, we match the feature points, discard the outliers and recover the epipolar geometry in one step. The unifying framework is furnished by genetic algorithm. We also create a new genetic operator to exchange the information between match process and robust epipolar line estimation so that epipolar geometry constraint can be elegantly incorporated into match process. Experiments on synthetic and real images show that this approach is very effective and fast. Jinxiang Chai, Songde Ma |
ICPR | 1 |
| 1998 | Robust epipolar geometry estimation using genetic algorithm
Jinxiang Chai, Songde Ma |
Pattern Recognit. Lett. | 1 |