Zhigang Deng 0001

dblp:23/6641 · DBLP profile ↗
← Back
140ranked-venue papers
7as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 103 · 6 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 32 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 7 since 2021Systems, architecture and hardware · 4Computer networks · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Data-Driven Control of Insect Flapping Flight via Deep Reinforcement Learning
abstract
Modeling and simulating realistic insect flight pose unique challenges due to the complex interaction between multi-degree-of-freedom wing kinematics and highly precise aerodynamic forces. To solve this challenge, this article presents a bidirectional kinematics-aerodynamics coupled simulation framework for miniature insect flight. Our approach first models the kinematics of flying insects by parameterizing natural wingbeat cycles based on available real-world datasets. Subsequently, we compute aerodynamic forces utilizing an improved semi-empirical model, which extends from quasi-steady formulation by incorporating critical unsteady force components. To achieve closed-loop control for both kinematics and aerodynamics, we employ deep reinforcement learning to train a virtual insect to adaptively adjust flapping strategies in response to dynamic flight states. Finally, an integrated controller enables the simulated insect to autonomously regulate the wing motion and perform complex tasks such as visual obstacle avoidance. Extensive experiments and comparisons demonstrate that our framework can effectively generate physically plausible and autonomous insect flight across a variety of scenarios.
Tingsong Lu, Yuming Fang 0001, Camille Le Roy, Xiaogang Jin 0001, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.6
2025 CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
abstract
Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging.
Trong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti, Tien-Phat Nguyen, Viet-Khoa Vo-Ho, Cuong Tran 0010, Yuki Ikebe, Anh Totti Nguyen, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le
ICCV13
2025 Enhancing Gaze Prediction in Multi-Party Conversations via Speaker-Aware Multimodal Adaptation
Meng-Chen Lee, Zhigang Deng 0001
ICMI2
2025 Learning Multimodal Motion Cues for Online End-of-Turn Prediction in Multi-Party Dialogue
Meng-Chen Lee, Zhigang Deng 0001
ICMI2
2025 MAARTA:Multi-agentic Adaptive Radiology Teaching Assistant
Akash Awasthi, Brandon V. Chung, Anh M. Vu, T. Hoang Ngan Le, Rishi Agrawal, Zhigang Deng 0001, Carol C. Wu, Hien Van Nguyen
MICCAI (5)6
2025 Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis
Trong-Thang Pham, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le
ACM Multimedia3
2025 GazeSearch: Radiology Findings Search Benchmark
abstract
Medical eye-tracking data is an important information source for understanding how radiologists visually inter-pret medical images. This information not only improves the accuracy of deep learning models for X-ray analysis but also their interpretability, enhancing transparency in decision-making. However, the current eye-tracking data is dispersed, unprocessed, and ambiguous, making it difficult to derive meaningful insights. Therefore, there is a need to create a new dataset with more focus and purposeful eye-tracking data, improving its utility for diagnostic applications. In this work, we propose a refinement method inspired by the target-present visual search challenge: there is a specific finding and fixations are guided to locate it. After re-fining the existing eye-tracking datasets, we transform them into a curated visual search dataset, called Gazesearch. specifically for radiology findings, where each fixation sequence is purposefully aligned to the task of locating a particular finding. Subsequently, we introduce a scan path prediction baseline, called ChestSearch, specifically tailored to Gazesearch. Finally, we employ the newly introduced Gazesearch as a benchmark to evaluate the performance of current state-of-the-art methods, offering a comprehensive assessment for visual search in the medical imaging domain. Code is available at https://github.com/UARK-AICV/GazeSearch.
Trong-Thang Pham, Tien-Phat Nguyen, Yuki Ikebe, Akash Awasthi, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le
WACV5
2025 ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists' intentions
Trong-Thang Pham, Jacob Brecheisen, Carol C. Wu, Hien Van Nguyen, Zhigang Deng 0001, Donald A. Adjeroh, Gianfranco Doretto, Arabinda Choudhary, T. Hoang Ngan Le
Artif. Intell. Medicine5
2025 Structural chain of thoughts for radiology education
Akash Awasthi, Brandon Chung, Anh M. Vu, Saba Khan, T. Hoang Ngan Le, Zhigang Deng 0001, Rishi Agrawal, Carol C. Wu, Hien Van Nguyen
Knowl. Based Syst.6
2025 A Bio-Inspired Model for Bee Simulations
abstract
As eusocial creatures, bees display unique macro collective behavior and local body dynamics that hold potential applications in various fields, such as computer animation, robotics, and social behavior. Unlike birds and fish, bees fly in a low-aligned zigzag pattern. Additionally, bees rely on visual signals for foraging and predator avoidance, exhibiting distinctive local body oscillations, such as body lifting, thrusting, and swaying. These inherent features pose significant challenges to realistic bee simulations in practical animation applications. In this article, we present a bio-inspired model for bee simulations capable of replicating both macro collective behavior and local body dynamics of bees. Our approach utilizes a visually-driven system to simulate a bee's local body dynamics, incorporating obstacle perception and body rolling control for effective collision avoidance. Moreover, we develop an oscillation rule that captures the dynamics of the bee's local bodies, drawing on insights from biological research. Our model extends beyond simulating individual bees' dynamics; it can also represent bee swarms by integrating a fluid-based field with the bees' innate noise and zigzag motions. To fine-tune our model, we utilize pre-collected honeybee flight data. Through extensive simulations and comparative experiments, we demonstrate that our model can efficiently generate realistic low-aligned and inherently noisy bee swarms.
Wenxiu Guo, Yuming Fang 0001, Yang Tong, Tingsong Lu, Xiaogang Jin 0001, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.7
2024 Online Multimodal End-of-Turn Prediction for Three-party Conversations
abstract
Predicting end-of-turn in multiparty conversations is crucial to increase the usability and natural flow of spoken dialogue systems, offering substantial enhancements to conversational agents. We present a novel window-based method to predict end-of-turn moments in real-time in multiparty conversations, by leveraging the capabilities of cutting-edge pre-trained language models (PLMs) and recurrent neural networks (RNN). Our method fuses the distilBERT language model with a Gated Recurrent Unit (GRU) to accurately predict end-of-turn points in an online fashion. Our approach can significantly outperform conventional Inter-Pausal Unit (IPU)-based prediction methods that often overlook the nuances of overlap and interruption during dynamic conversations. Potential applications of this study are significant, particularly in the domains of virtual agents and human-robot interactions. Our accurate online end-of-turn prediction model can be facilitated to enhance the user experience in these applications, making them more natural and seamlessly integrated into real-world conversations.
Meng-Chen Lee, Zhigang Deng 0001
ICMI2
2024 Toward User-Aware Interactive Virtual Agents: Generative Multi-Modal Agent Behaviors in VR
abstract
Virtual agents serve as a vital interface within XR platforms. However, generating virtual agent behaviors typically rely on pre-coded actions or physics-based reactions. In this paper we present a learning-based multimodal agent behavior generation framework that adapts to users’ in-situ behaviors, similar to how humans interact with each other in the real world. By leveraging an in-house collected, dyadic conversational behavior dataset, we trained a conditional variational autoencoder (CVAE) model to achieve user-conditioned generation of virtual agents’ behaviors. Together with large language models (LLM), our approach can generate both the verbal and non-verbal reactive behaviors of virtual agents. Our comparative user study confirmed our method’s superiority over conventional animation graph-based baseline techniques, particularly regarding user-centric criteria. Thorough analyses of our results underscored the authentic nature of our virtual agents’ interactions and the heightened user engagement during VR interaction.
Bhasura S. Gunawardhana, Qi Sun 0003, Zhigang Deng 0001
ISMAR4
2024 A Computational Study on Sentence-based Next Speaker Prediction in Multiparty Conversations
abstract
In this paper we present a computational study to quantitatively examine the task of predicting the next speaker in multi-party conversations using machine learning models. To accomplish this, we create features that accurately represent information relevant to speaker changes in such conversations. We utilize sentence-based models, rather than the widely-used InterPausal Unit (IPU)-based models, and extend the definition of verbal backchanneling to include additional reactions that signify listeners’ attention or interest. Through extensive experiments with various machine learning models and inputs, we show that our sentence-based models outperform existing IPU-based models, with the best model achieving 61.39% accuracy. Our study provides design implications and recommendations for the development of virtual agents or humanoid robots with interactive social interaction capabilities.
Meng-Chen Lee, Angela W. Li, Zhigang Deng 0001
IVA3
2024 Learning facial expression-aware global-to-local representation for robust action unit detection
Rudong An, Aobo Jin, Wei Chen 0157, Wei Zhang 0219, Hao Zeng 0001, Zhigang Deng 0001, Yu Ding 0001
Appl. Intell.6
2024 Noise4Denoise: Leveraging noise for unsupervised point cloud denoising
abstract
Existing deep learning-based point cloud denoising methods are generally trained in a supervised manner that requires clean data as ground-truth labels. However, in practice, it is not always feasible to obtain clean point clouds. In this paper, we introduce a novel unsupervised point cloud denoising method that eliminates the need to use clean point clouds as groundtruth labels during training. We demonstrate that it is feasible for neural networks to only take noisy point clouds as input, and learn to approximate and restore their clean versions. In particular, we generate two noise levels for the original point clouds, requiring the second noise level to be twice the amount of the first noise level. With this, we can deduce the relationship between the displacement information that recovers the clean surfaces across the two levels of noise, and thus learn the displacement of each noisy point in order to recover the corresponding clean point. Comprehensive experiments demonstrate that our method achieves outstanding denoising results across various datasets with synthetic and real-world noise, obtaining better performance than previous unsupervised methods and competitive performance to current supervised methods.
Xiao Liu 0004, Hailing Zhou, Lei Wei 0002, Zhigang Deng 0001, M. Manzur Murshed, Xuequan Lu
Comput. Vis. Media5
2024 Color Theme Evaluation through User Preference Modeling
abstract
Color composition (or color theme) is a key factor to determine how well a piece of art work or graphical design is perceived by humans. Despite a few color harmony models have been proposed, their results are often less satisfactory since they mostly neglect the variations of aesthetic cognition among individuals and treat the influence of all ratings equally as if they were all rated by the same anonymous user. To overcome this issue, in this article we propose a new color theme evaluation model by combining a back propagation neural network and a kernel probabilistic model to infer both the color theme rating and the user aesthetic preference. Our experiment results show that our model can predict more accurate and personalized color theme ratings than state of the art methods. Our work is also the first-of-its-kind effort to quantitatively evaluate the correlation between user aesthetic preferences and color harmonies of five-color themes, and study such a relation for users with different aesthetic cognition.
Bailin Yang, Tianxiang Wei, Frederick W. B. Li, Xiaohui Liang 0001, Zhigang Deng 0001, Yili Fang
ACM Trans. Appl. Percept.5
2024 Detecting Facial Action Units From Global-Local Fine-Grained Expressions
abstract
Since Facial Action Unit (AU) annotations require domain expertise, common AU datasets only contain a limited number of subjects. As a result, a crucial challenge for AU detection is addressing identity overfitting. We find that AUs and facial expressions are highly associated, and existing facial expression datasets often contain a large number of identities. In this paper, we aim to utilize the expression datasets without AU labels to facilitate AU detection. Specifically, we develop a novel AU detection framework aided by the Global-Local facial Expressions Embedding, dubbed GLEE-Net. Our GLEE-Net consists of three branches to extract identity-independent expression features for AU detection. We introduce a global branch for modeling the overall facial expression while eliminating the impacts of identities. We also design a local branch focusing on specific local face regions. The combined output of global and local branches is firstly pre-trained on an expression dataset as an identity-independent expression embedding, and then finetuned on AU datasets. Therefore, we significantly alleviate the issue of limited identities. Furthermore, we introduce a 3D global branch that extracts expression coefficients through 3D face reconstruction to consolidate 2D AU descriptions. Finally, a Transformer-based multi-label classifier is employed to fuse all the representations for AU detection. Extensive experiments demonstrate that our method significantly outperforms the state-of-the-art on the widely-used DISFA, BP4D and BP4D+ datasets.
Wei Zhang 0219, Lincheng Li, Yu Ding 0001, Wei Chen 0157, Zhigang Deng 0001, Xin Yu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Large Scale Farm Scene Modeling from Remote Sensing Imagery
abstract
In this paper we propose a scalable framework for large-scale farm scene modeling that utilizes remote sensing data, specifically satellite images. Our approach begins by accurately extracting and categorizing the distributions of various scene elements from satellite images into four distinct layers: fields, trees, roads, and grasslands. For each layer, we introduce a set of controllable Parametric Layout Models (PLMs). These models are capable of learning layout parameters from satellite images, enabling them to generate complex, large-scale farm scenes that closely reproduce reality across multiple scales. Additionally, our framework provides intuitive control for users to adjust layout parameters to simulate different stages of crop growth and planting patterns. This adaptability makes our model an excellent tool for graphics and virtual reality applications. Experimental results demonstrate that our approach can rapidly generate a variety of realistic and highly detailed farm scenes with minimal inputs.
Zhiqi Xiao, Hao Jiang 0013, Zhigang Deng 0001, Wenwei Han
ACM Trans. Graph.3
2024 Agent-based crowd simulation: an in-depth survey of determining factors for heterogeneous behavior
Saba Khan, Zhigang Deng 0001
Vis. Comput.2
2023 Tele-Mentoring Using Augmented Reality: A Feasibility Study to Assess Teaching of Laparoscopic Suturing Skills
abstract
The work assesses the efficacy of computer based remote tele-mentoring system (i.e. when the mentor and mentee are physically separated) for teaching minimally invasive surgical skills. The visual cues used for tele-mentoring comprises real-time virtual surgical instruments' motion augmented onto the operative field and remotely controlled by the mentor. In the feasibility study, the surgical task of laparoscopic intracorporeal suturing was simulated among 18 mentor-mentee pairs. Three modes of mentoring were used. Mode-I included traditional learning using pre-recorded videos (in absence of a mentor). Mode-II used traditional in-person hands-on mentoring. In Mode-III, a tele-mentoring prototype was used that connected a mentee with a remote mentor. Error count and duration were recorded for a learning stage followed by a testing stage for the three modes. The results show the error count for Mode-III reduces significantly as compared to Mode-I in the learning stage. Similarly, the error count for Mode-III also reduces significantly as compared to Mode-I in the testing stage. The errors count for Mode-III were equivalent to that of Mode-II for both learning and teaching stages. Furthermore, in Mode-III the duration reduces from learning to testing stage exhibiting the learning effect. Thus, computer based remote tele-mentoring is effective and more convenient to demonstrate surgical sub-steps consisting of tool-tissue interaction facilitating surgical skill transfer.
Dehlela Shabir, Shidin Balakrishnan, Jhasketan Padhan, Julien Abinahed, Elias Yaacoub, Amr Mohamed 0001, Zhigang Deng 0001, Abdulla Al-Ansari, Panagiotis Tsiamyrtzis, Nikhil V. Navkar
CBMS7
2023 Double Doodles: Sketching Animation in Immersive Environment With 3+6 DOFs Motion Gestures
abstract
We present "Double Doodles'' to make full use of two sequential inputs of a VR controller with 9 DOFs in total, 3 DOFs of the first input sequence for the generation of motion paths and 6 DOFs of the second input sequence for motion gestures. While engineering our system, we take ergonomics into consideration and design a set of user-defined motion gestures to describe character motions. We employ a real-time deep learning-based approach for highly accurate motion gesture classification. We then integrate our approach into a prototype system, and it allows users to directly create character animations in VR environments using motion gestures with a VR controller, followed by animation preview and animation interactive editing. Finally, we evaluate the feasibility and effectiveness of our system through a user study, demonstrating the usefulness of our system for visual storytelling dedicated to amateurs, as well as for providing fast drafting tools for artists.
Ruizhao Chen, Zhigang Deng 0001, Lili Wang 0006, Lizhuang Ma
ACM Multimedia3
2023 Foreword to the special section on motion, interaction, and games, 2022
Aline Normoyle, Zhigang Deng 0001
Comput. Graph.2
2023 Fuzzy-based indoor scene modeling with differentiated examples
abstract
Well-designed indoor scenes incorporate interior design knowledge, which has been an essential prior for most indoor scene modeling methods. However, the layout qualities of indoor scene datasets are often uneven, and most existing data-driven methods do not differentiate indoor scene examples in terms of quality. In this work, we aim to explore an approach that leverages datasets with differentiated indoor scene examples for indoor scene modeling. Our solution conducts subjective evaluations on lightweight datasets having various room configurations and furniture layouts, via pairwise comparisons based on fuzzy set theory. We also develop a system to use such examples to guide indoor scene modeling using user-specified objects. Specifically, we focus on object groups associated with certain human activities, and define room features to encode the relations between the position and direction of an object group and the room configuration. To perform indoor scene modeling, given an empty room, our system first assesses it in terms of the user-specified object groups, and then places associated objects in the room guided by the assessment results. A series of experimental results and comparisons to state-of-the-art indoor scene synthesis methods are presented to validate the usefulness and effectiveness of our approach.
Qiang Fu 0004, Shuhan He, Hongbo Fu 0001, Xueming Li 0002, Zhigang Deng 0001
Comput. Vis. Media5
2023 Effects of spatial constraints and ages on children's upper limb performance in mid-air gesture interaction
Fei Lyu 0001, Huijing Li, Qiang Fu 0004, Jin Huang 0009, Zhigang Deng 0001
Int. J. Hum. Comput. Stud.6
2023 A Music-Driven Deep Generative Adversarial Model for Guzheng Playing Animation
abstract
To date relatively few efforts have been made on the automatic generation of musical instrument playing animations. This problem is challenging due to the intrinsically complex, temporal relationship between music and human motion as well as the lacking of high quality music-playing motion datasets. In this article, we propose a fully automatic, deep learning based framework to synthesize realistic upper body animations based on novel guzheng music input. Specifically, based on a recorded audiovisual motion capture dataset, we delicately design a generative adversarial network (GAN) based approach to capture the temporal relationship between the music and the human motion data. In this process, data augmentation is employed to improve the generalization of our approach to handle a variety of guzheng music inputs. Through extensive objective and subjective experiments, we show that our method can generate visually plausible guzheng-playing animations that are well synchronized with the input guzheng music, and it can significantly outperform the state-of-the-art methods. In addition, through an ablation study, we validate the contributions of the carefully-designed modules in our framework.
Changjie Fan, Gongzheng Li, Zeng Zhao, Zhigang Deng 0001, Yu Ding 0001
IEEE Trans. Vis. Comput. Graph.6
2022 A Comparative Study on Single-handed Keyboards on Large-screen Mobile Devices
abstract
Many questions regarding single-hand text entry on modern smartphones (in particular, large-screen smartphones) remain under-explored, such as, (i) will the existing prevailing single-handed keyboards fit for large-screen smartphone users? and (ii) will individual-customization improve single-handed keyboard performance? In this paper we study single-handed typing behaviors on several representative keyboards on large-screen mobile devices. We found that, (i) the user-adaptable-shape curved keyboard performs best among all the studied keyboards; (ii) users’ familiarity with the Qwerty layout plays a significant role at the beginning, but after several sessions of training, the user-adaptable curved keyboard can have the best learning curve and performs best; (iii) generally the statistical decoding algorithms via spatial and language models can well handle the input noise from single-handed typing.
Zhigang Deng 0001
AVI2
2022 Dynamic Guidance Virtual Fixtures for Guiding Robotic Interventions: Intraoperative MRI-guided Transapical Cardiac Intervention Paradigm
abstract
The advent of intraoperative real-time image guidance has led to the emergence of new surgical interventional paradigms including image-guided robot assistance. Most often the use of an intraoperative imaging modality is limited to visual perception of the area of procedure. In this work, we propose a system for performing interventions with real-time Magnetic Resonance Imaging (rtMRI). The described computational core, processes on-the-fly rtMRI and generates dynamic guidance virtual fixture that in turn is used to update visualization and a force-feedback interface. The system was experimentally tested by applying it to a simulated Transapical Aortic Valve Implantation with a virtual robotic manipulator. The study results demonstrate significant improvement in the surgical task by decreasing the duration of the procedure and increasing safety in the presence of cardiac and breathing motion.
Jhasketan Padhan, Nikolaos V. Tsekos, Abdulla Al-Ansari, Julien Abinahed, Zhigang Deng 0001, Nikhil V. Navkar
BIBE5
2022 Benchmarking Network Performance of Augmented Reality Based Surgical Telementoring Systems
abstract
Telementoring in surgery facilitates the transfer of surgical knowledge from the mentor to the mentee. Augmented Reality (AR) further assists this transfer by overlaying visual cues (e.g., in the form of virtual surgical instrument motion) generated by the mentor onto the operative field of the mentee. In this work, we present a benchmark for comparing such AR based surgical telementoring systems. The results compare the network performances of these systems across different types of surgery (open or minimally invasive), based on the locations of the mentor and the mentee (inter- or intra- country), and finally the underlying networking protocols (RTMP versus WebRTC).
Dehlela Shabir, Malek Anabtawi, Nihal Abdurahiman, May Trinh, Jhasketan Padhan, Abdulla Al-Ansari, Julien Abinahed, Zhigang Deng 0001, Elias Yaacoub, Amr Mohamed 0001, Nikhil V. Navkar
BIBE8
2022 A Practical AR-based Surgical Navigation System Using Optical See-through Head Mounted Display
abstract
The work presents a practical Augmented Reality (AR) based surgical navigation system using optical see-through head-mounted display as a standalone solution, without the need of additional tracking hardware. Specifically, we propose a fiducial marker-based instrument tracking, which entirely relies on the built-in hardware of the Microsoft HoloLens 2. The tracking algorithm computes the pose of the tracked object from the real-time image obtained from the on-board front-facing RGB camera. The estimated transformation is then transmitted back to the HoloLens for visualization. Our experimental evaluation shows that the system can achieve 0.81 mm / 1.52 degree in tracking accuracy and sub-millimeter alignment accuracy.
Mai Trinh, Nikhil V. Navkar, Zhigang Deng 0001
BIBE3
2022 A Practical Method for Butterfly Motion Capture
abstract
Simulating realistic butterfly motion has been a widely-known challenging problem in computer animation. Arguably, one of its main reasons is the difficulty of acquiring accurate flight motion of butterflies. In this paper we propose a practical yet effective, optical marker-based approach to capture and process the detailed motion of a flying butterfly. Specifically, we first capture the trajectories of the wings and thorax of a flying butterfly using optical marker-based motion tracking. After that, our method automatically fills the positions of missing markers by exploiting the continuity and relevance of neighboring frames, and improves the quality of the captured motion via noise filtering with optimized parameter settings. Through comparisons with existing motion processing methods, we demonstrate the effectiveness of our approach to obtain accurate flight motions of butterflies. Furthermore, we created and will release a first-of-its-kind butterfly motion capture dataset to research community.
Tingsong Lu, Yang Tong, Yuming Fang 0001, Zhigang Deng 0001
MIG5
2022 S2M-Net: Speech Driven Three-party Conversational Motion Synthesis Networks
abstract
In this paper we propose a novel conditional generative adversarial network (cGAN) architecture, called S2M-Net, to holistically synthesize realistic three-party conversational animations based on acoustic speech input together with speaker marking (i.e., the speaking time of each interlocutor). Specifically, based on a pre-collected three-party conversational motion dataset, we design and train the S2M-Net for three-party conversational animation synthesis. In the architecture, a generator contains a LSTM encoder to encode a sequence of acoustic speech features to a latent vector that is further fed into a transform unit to transform the latent vector into a gesture kinematics space. Then, the output of this transform unit is fed into a LSTM decoder to generate corresponding three-party conversational gesture kinematics. Meanwhile, a discriminator is implemented to check whether an input sequence of three-party conversational gesture kinematics is real or fake. To evaluate our method, besides quantitative and qualitative evaluations, we also conducted paired comparison user studies to compare it with the state of the art.
Aobo Jin, Qixin Deng, Zhigang Deng 0001
MIG3
2022 End-to-End 3D Face Reconstruction with Expressions and Specular Albedos from Single In-the-wild Images
abstract
Recovering 3D face models from in-the-wild face images has numerous potential applications. However, properly modeling complex lighting effects in reality, including specular lighting, shadows, and occlusions, from a single in-the-wild face image is still considered as a widely open research challenge. In this paper, we propose a convolutional neural network based framework to regress the face model from a single image in the wild. The outputted face model includes dense 3D shape, head pose, expression, diffuse albedo, specular albedo, and the corresponding lighting conditions. Our approach uses novel hybrid loss functions to disentangle face shape identities, expressions, poses, albedos, and lighting. Besides a carefully-designed ablation study, we also conduct direct comparison experiments to show that our method can outperform state-of-art methods both quantitatively and qualitatively.
Qixin Deng, Binh Huy Le, Aobo Jin, Zhigang Deng 0001
ACM Multimedia4
2022 Indoor layout programming via virtual navigation detectors
Qiang Fu 0004, Hongbo Fu 0001, Zhigang Deng 0001, Xueming Li 0002
Sci. China Inf. Sci.3
2022 R-CTM: A data-driven macroscopic simulation model for heterogeneous traffic
abstract
Abstract There is a well‐known trade‐off between computational efficiency and computational accuracy in the field of traffic simulation. In this article, we propose a novel recurrent neural network based model with an integrated attention mechanism, called R‐CTM, to simulate heterogeneous traffic flow with multiple types of vehicles. It can effectively extract the traffic flow patterns of spatial and temporal changes from training traffic data, which can be real‐world traffic data or synthetic traffic data via microscopic simulation models. Through experiments and comparisons, we show that it can significantly outperform the state of the art methods in terms of simulation accuracy. Besides accuracy, we also demonstrate its scalability: its runtime consumption does not linearly increase with respect to the spatial extent.
Zhigang Deng 0001, Tianlu Mao
Comput. Animat. Virtual Worlds2
2022 A Practical Model for Realistic Butterfly Flight Simulation
abstract
Butterflies are not only ubiquitous around the world but are also widely known for inspiring thrill resonance, with their elegant and peculiar flights. However, realistically modeling and simulating butterfly flights—in particular, for real-time graphics and animation applications—remains an under-explored problem. In this article, we propose an efficient and practical model to simulate butterfly flights. We first model a butterfly with parametric maneuvering functions, including wing-abdomen interaction. Then, we simulate dynamic maneuvering control of the butterfly through our force-based model, which includes both the aerodynamics force and the vortex force. Through many simulation experiments and comparisons, we demonstrate that our method can efficiently simulate realistic butterfly flight motions in various real-world settings.
Tingsong Lu, Yang Tong, Guoliang Luo, Xiaogang Jin 0001, Zhigang Deng 0001
ACM Trans. Graph.6
2022 Plausible 3D Face Wrinkle Generation Using Variational Autoencoders
abstract
Realistic 3D facial modeling and animation have been increasingly used in many graphics, animation, and virtual reality applications. However, generating realistic fine-scale wrinkles on 3D faces, in particular, on animated 3D faces, is still a challenging problem that is far away from being resolved. In this article we propose an end-to-end system to automatically augment coarse-scale 3D faces with synthesized fine-scale geometric wrinkles. By formulating the wrinkle generation problem as a supervised generation task, we implicitly model the continuous space of face wrinkles via a compact generative model, such that plausible face wrinkles can be generated through effective sampling and interpolation in the space. We also introduce a complete pipeline to transfer the synthesized wrinkles between faces with different shapes and topologies. Through many experiments, we demonstrate our method can robustly synthesize plausible fine-scale wrinkles on a variety of coarse-scale 3D faces with different shapes and expressions.
Qixin Deng, Luming Ma, Aobo Jin, Huikun Bi, Binh Huy Le, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.6
2021 A linear wave propagation-based simulation model for dense and polarized crowds
abstract
Abstract Fluid‐like motion and linear wave propagation behavior will emerge when we impose boundary constraints and polarized conditions on crowds. To this end, we present a Lagrangian hydrodynamics method to simulate the fluid‐like motion of crowd and a triggering approach to generate the linear stop‐and‐go wave behavior. Specifically, we impose a self‐propulsion force on the leading agents of the crowd to push the crowd to move forward and introduce a Smoothed Particle Hydrodynamics‐based model to simulate the dynamics of dense crowds. Besides, we present a motion signal propagation approach to trigger the rest of the crowd so that they respond to the immediate leaders linearly, which can lead to the linear stop‐and‐go wave effect of the fluid‐like motion for the crowd. Our experiments demonstrate that our model can simulate large‐scale dense crowds with linear wave propagation.
Guoliang Luo, Yang Tong, Xiaogang Jin 0001, Zhigang Deng 0001
Comput. Animat. Virtual Worlds5
2021 Emotion-Based Crowd Simulation Model Based on Physical Strength Consumption for Emergency Scenarios
abstract
Increasing attention is being given to the modeling and simulation of traffic flow and crowd movement, two phenomena that both deal with interactions between pedestrians and cars in many situations. In particular, crowd simulation is important for understanding mobility and transportation patterns. In this paper, we propose an emotion-based crowd simulation model integrating physical strength consumption. Inspired by the theory of “the devoted actor,” the movements of each individual in our model are determined by modeling the influence of physical strength consumption and the emotion of panic. In particular, human physical strength consumption is computed using a physics-based numerical method. Inspired by the James-Lange theory, panic levels are estimated by means of an enhanced emotional contagion model that leverages the inherent relationship between physical strength consumption and panic. To the best of our knowledge, our model is the first method integrating physical strength consumption into an emotion-based crowd simulation model by exploiting the relationship between physical strength consumption and emotion. We highlight the performance on different scenarios and compare the resulting behaviors with real-world video sequences. Our approach can reliably predict changes in physical strength consumption and panic levels of individuals in an emergency situation.
Mingliang Xu 0001, Chaochao Li, Pei Lv, Wei Chen 0001, Zhigang Deng 0001, Bing Zhou 0003, Dinesh Manocha
IEEE Trans. Intell. Transp. Syst.5
2021 Crowd Behavior Simulation With Emotional Contagion in Unexpected Multihazard Situations
abstract
Numerous research efforts have been conducted to simulate the crowd movements, while relatively few of them are specifically focused on multihazard situations. In this paper, we propose a novel crowd simulation method by modeling the generation and contagion of panic emotion under multihazard circumstances. In order to depict the effect from hazards and other agents to crowd movement, we first classify hazards into different types (transient and persistent, concurrent and nonconcurrent, and static and dynamic) based on their inherent characteristics. Second, we introduce the concept of perilous field for each hazard and further transform the critical level of the field to its invoked-panic emotion. After that, we propose an emotional contagion model to simulate the evolving process of panic emotion caused by multiple hazards. Finally, we introduce an emotional reciprocal velocity obstacles (RVOs) model to simulate the crowd behaviors by augmenting the traditional RVO model with emotional contagion, which for the first time combines the emotional impact and local avoidance together. Our experimental results demonstrate that the overall approach is robust, can better generate realistic crowds and the panic emotion dynamics in a crowd. Furthermore, it is recommended that our method can be applied to various complex multihazard environments.
Mingliang Xu 0001, Xiaozheng Xie, Pei Lv, Jianwei Niu 0002, Chaochao Li, Ruijie Zhu 0001, Zhigang Deng 0001, Bing Zhou 0003
IEEE Trans. Syst. Man Cybern. Syst.8
2021 Motion Planning for Convertible Indoor Scene Layout Design
abstract
We present a system for designing indoor scenes with convertible furniture layouts. Such layouts are useful for scenarios where an indoor scene has multiple purposes and requires layout conversion, such as merging multiple small furniture objects into a larger one or changing the locus of the furniture. We aim at planning the motion for the convertible layouts of a scene with the most efficient conversion process. To achieve this, our system first establishes object-level correspondences between the layout of a given source and that of a reference to compute a target layout, where the objects are re-arranged in the source layout with respect to the reference layout. After that, our system initializes the movement paths of objects between the source and target layouts based on various mechanical constraints. A joint space-time optimization is then performed to program a control stream of object translations, rotations, and stops, under which the movements of all objects are efficient and the potential object collisions are avoided. We demonstrate the effectiveness of our system through various design examples of multi-purpose, indoor scenes with convertible layouts.
Guoming Xiong, Qiang Fu 0004, Hongbo Fu 0001, Guoliang Luo, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.6
2020 How Can I See My Future? FvTraj: Using First-Person View for Pedestrian Trajectory Prediction
Huikun Bi, Ruisi Zhang, Tianlu Mao, Zhigang Deng 0001
ECCV (7)4
2020 Contour-based 3D Modeling through Joint Embedding of Shapes and Contours
abstract
In this paper, we propose a novel space that jointly embeds both 2D occluding contours and 3D shapes via a variational autoencoder (VAE) and a volumetric autoencoder. Given a dataset of 3D shapes, we extract their occluding contours via projections from random views and use the occluding contours to train the VAE. Then, the obtained continuous embedding space, where each point is a latent vector that represents an occluding contour, can be used to measure the similarity between occluding contours. After that, the volumetric autoencoder is trained to first map 3D shapes onto the embedding space through a supervised learning process and then decode the merged latent vectors of three occluding contours (from three different views) of a 3D shape to its 3D voxel representation. We conduct various experiments and comparisons to demonstrate the usefulness and effectiveness of our method for sketch-based 3D modeling and shape manipulation applications.
Aobo Jin, Qiang Fu 0004, Zhigang Deng 0001
I3D3
2020 Real-time Face Video Swapping From A Single Portrait
abstract
We present a novel high-fidelity real-time method to replace the face in a target video clip by the face from a single source portrait image. Specifically, we first reconstruct the illumination, albedo, camera parameters, and wrinkle-level geometric details from both the source image and the target video. Then, the albedo of the source face is modified by a novel harmonization method to match the target face. Finally, the source face is re-rendered and blended into the target video using the lighting and camera parameters from the target video. Our method runs fully automatically and at real-time rate on any target face captured by cameras or from legacy video. More importantly, unlike existing deep learning based methods, our method does not need to pre-train any models, i.e., pre-collecting a large image/video dataset of the source or target face for model training is not needed. We demonstrate that a high level of video-realism can be achieved by our method on a variety of human faces with different identities, ethnicities, skin colors, and expressions.
Luming Ma, Zhigang Deng 0001
I3D2
2020 A Survey on Visual Traffic Simulation: Models, Evaluations, and Applications in Autonomous Driving
abstract
Abstract Virtualized traffic via various simulation models and real‐world traffic data are promising approaches to reconstruct detailed traffic flows. A variety of applications can benefit from the virtual traffic, including, but not limited to, video games, virtual reality, traffic engineering and autonomous driving. In this survey, we provide a comprehensive review on the state‐of‐the‐art techniques for traffic simulation and animation. We start with a discussion on three classes of traffic simulation models applied at different levels of detail. Then, we introduce various data‐driven animation techniques, including existing data collection methods, and the validation and evaluation of simulated traffic flows. Next, we discuss how traffic simulations can benefit the training and testing of autonomous vehicles. Finally, we discuss the current states of traffic simulation and animation and suggest future research directions.
Qianwen Chao, Huikun Bi, Weizi Li, Tianlu Mao, Ming C. Lin, Zhigang Deng 0001
Comput. Graph. Forum7
2020 Curve Skeleton Extraction From 3D Point Clouds Through Hybrid Feature Point Shifting and Clustering
abstract
Abstract Curve skeleton is an important shape descriptor with many potential applications in computer graphics, visualization and machine intelligence. We present a curve skeleton expression based on the set of the cross‐section centroids from a point cloud model and propose a corresponding extraction approach. We first provide the substitution of a distance field for a 3D point cloud model, and then combine it with curvatures to capture hybrid feature points. By introducing relevant facets and points, we shift these hybrid feature points along the skeleton‐guided normal directions to approach local centroids, simplify them through a tensor‐based spectral clustering and finally connect them to form a primary connected curve skeleton. Furthermore, we refine the primary skeleton through pruning, trimming and smoothing. We compared our results with several state‐of‐the‐art algorithms including the rotational symmetry axis (ROSA) and L1‐medial methods for incomplete point cloud data to evaluate the effectiveness and accuracy of our method.
Xiaogang Jin 0001, Zhigang Deng 0001, Minhong Chen
Comput. Graph. Forum4
2020 Hyperspectral Inverse Skinning
abstract
Abstract In example‐based inverse linear blend skinning (LBS), a collection of poses (e.g. animation frames) are given, and the goal is finding skinning weights and transformation matrices that closely reproduce the input. These poses may come from physical simulation, direct mesh editing, motion capture or another deformation rig. We provide a re‐formulation of inverse skinning as a problem in high‐dimensional Euclidean space. The transformation matrices applied to a vertex across all poses can be thought of as a point in high dimensions. We cast the inverse LBS problem as one of finding a tight‐fitting simplex around these points (a well‐studied problem in hyperspectral imaging). Although we do not observe transformation matrices directly, the 3D position of a vertex across all of its poses defines an affine subspace, or flat. We solve a ‘closest flat’ optimization problem to find points on these flats, and then compute a minimum‐volume enclosing simplex whose vertices are the transformation matrices and whose barycentric coordinates are the skinning weights. We are able to create LBS rigs with state‐of‐the‐art reconstruction error and state‐of‐the‐art compression ratios for mesh animation sequences. Our solution does not consider weight sparsity or the rigidity of recovered transformations. We include observations and insights into the closest flat problem. Its ideal solution and optimal LBS reconstruction error remain an open problem.
Songrun Liu, Jianchao Tan, Zhigang Deng 0001, Yotam I. Gingold
Comput. Graph. Forum3
2020 Cover Image
abstract
The cover image is based on the Original Article Sketch-based Shape-constrained Fireworks Simulation in Head Mounted Virtual Reality by Xiaogang Jin et al., https://doi.org/10.1002/cav.1920.
Xiaoyu Cui, Ruifan Cai, Xiangjun Tang, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds4
2020 Sketch-based shape-constrained fireworks simulation in head-mounted virtual reality
abstract
Abstract In this paper we present a novel shape‐constrained fireworks simulation method with rich textures in an HMD (Helmet Mounted Display) virtual environment using sketched feature lines as input. Our approach first retrieves an object from a three‐dimensional (3D) model database using a sketch‐based 3D shape retrieval algorithm. Then, in order to approximate models with complex structures, we introduce a novel point sampling algorithm based on Gaussian curvatures, which stores not only the positions of the selected vertices but also the texture (UV) coordinates information for texture display. In addition, we introduce a multilevel explosion process so that the fireworks can dynamically form specific, visually pleasing shapes. Through our experiments, we demonstrate that our approach can produce better results than state‐of‐the‐art approaches.
Xiaoyu Cui, Ruifan Cai, Xiangjun Tang, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds4
2020 Low-Level Characterization of Expressive Head Motion Through Frequency Domain Analysis
abstract
For the purpose of understanding how head motions contribute to the perception of emotion in an utterance, we aim to examine the perception of emotion based on Fourier transform-based static and dynamic features of head motion. Our work is to conduct intra-related objective analysis and perceptual experiments on the link between the perception of emotion and the static/dynamic features. The objective analysis outcome shows that the static and dynamic features are effective in characterizing and recognizing emotions. The perceptual experiments enable us to collect human perception of emotion through head motion. The collected perceptual data shows that humans are unable to reliably perceive emotion from head motion alone but reveals that humans are sensitive to the static feature (in reference to the averaged up-down rotation angle) and the dynamic features (which reflect the fluidity and speed of movement). It also indicates that humans perceive emotion carried in head motion and the naturalness of head motion in two different channels. Our work contributes to the understanding and the characterization of head motion in expressive speech through low-level descriptions of motion features, instead of commonly used high-level motion style (e.g., head nods, shakes, tilts, and raises).
Yu Ding 0001, Lei Shi 0027, Zhigang Deng 0001
IEEE Trans. Affect. Comput.3
2020 Spatio-temporal Segmentation Based Adaptive Compression of Dynamic Mesh Sequences
abstract
With the recent advances in data acquisition techniques, the compression of various dynamic mesh sequence data has become an important topic in the computer graphics community. In this article, we present a new spatio-temporal segmentation-based approach for the adaptive compression of the dynamic mesh sequences. Given an input dynamic mesh sequence, we first compute an initial temporal cut to obtain a small subsequence by detecting the temporal boundary of dynamic behavior. Then, we apply a two-stage vertex clustering on the resulting subsequence to classify the vertices into groups with optimal intra-affinities. After that, we design a temporal segmentation step based on the variations of the principal components within each vertex group prior to performing a PCA-based compression. Furthermore, we apply an extra step on the lossless compression of the PCA bases and coefficients to gain more storage saving. Our approach can adaptively determine the temporal and spatial segmentation boundaries to exploit both temporal and spatial redundancies. We have conducted extensive experiments on different types of 3D mesh animations with various segmentation configurations. Our comparative studies show the advantages of our approach for the compression of 3D mesh animations.
Guoliang Luo, Zhigang Deng 0001, Xiaogang Jin 0001, Wenqiang Xie, Hyewon Seo
ACM Trans. Multim. Comput. Commun. Appl.2
2020 A Deep Learning-Based Framework for Intersectional Traffic Simulation and Editing
abstract
Most of existing traffic simulation methods have been focused on simulating vehicles on freeways or city-scale urban networks. However, relatively little research has been done to simulate intersectional traffic to date despite its broad potential applications. In this paper, we propose a novel deep learning-based framework to simulate and edit intersectional traffic. Specifically, based on an in-house collected intersectional traffic dataset, we employ the combination of convolution network (CNN) and recurrent network (RNN) to learn the patterns of vehicle trajectories in intersectional traffic. Besides simulating novel intersectional traffic, our method can be used to edit existing intersectional traffic. Through many experiments as well as comparative user studies, we demonstrate that the results by our method are visually indistinguishable from ground truth, and our method can outperform existing methods.
Huikun Bi, Tianlu Mao, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.4
2020 Dictionary-based Fidelity Measure for Virtual Traffic
abstract
Aiming at objectively measuring the realism of virtual traffic flows and evaluating the effectiveness of different traffic simulation techniques, this paper introduces a general, dictionary-based learning method to evaluate the fidelity of any traffic trajectory data. First, a traffic pattern dictionary that characterizes common patterns of real-world traffic behavior is built offline from pre-collected ground truth traffic data. The corresponding learning error is set as the benchmark of the dictionary-based traffic representation. With the aid of the constructed dictionary, the realism of input simulated traffic flow data can be evaluated by comparing its dictionary-based reconstruction error with the dictionary error benchmark. This evaluation metric can be robustly applied to any simulated traffic flow data; in other words, it is independent of how the traffic data are generated. We demonstrated the effectiveness and robustness of this metric through many experiments on real-world traffic data and various simulated traffic data, comparisons with the state-of-the-art entropy-based similarity metric for aggregate crowd motions, and perceptual evaluation studies.
Qianwen Chao, Zhigang Deng 0001, Yangxi Xiao, Dunbang He, Qiguang Miao, Xiaogang Jin 0001
IEEE Trans. Vis. Comput. Graph.2
2019 ODE-Driven Sketch-Based Organic Modelling
Ouwen Li, Zhigang Deng 0001, Shaojun Bian, Algirdas Noreika, Xiaogang Jin 0001, Ismail Khalid Kazmi, Lihua You, Jian J. Zhang 0001
CGI2
2019 Joint Prediction for Kinematic Trajectories in Vehicle-Pedestrian-Mixed Scenes
abstract
Trajectory prediction for objects is challenging and critical for various applications (e.g., autonomous driving, and anomaly detection). Most of the existing methods focus on homogeneous pedestrian trajectories prediction, where pedestrians are treated as particles without size. However, they fall short of handling crowded vehicle-pedestrian-mixed scenes directly since vehicles, limited with kinematics in reality, should be treated as rigid, non-particle objects ideally. In this paper, we tackle this problem using separate LSTMs for heterogeneous vehicles and pedestrians. Specifically, we use an oriented bounding box to represent each vehicle, calculated based on its position and orientation, to denote its kinematic trajectories. We then propose a framework called VP-LSTM to predict the kinematic trajectories of both vehicles and pedestrians simultaneously. In order to evaluate our model, a large dataset containing the trajectories of both vehicles and pedestrians in vehicle-pedestrian-mixed scenes is specially built. Through comparisons between our method with state-of-the-art approaches, we show the effectiveness and advantages of our method on kinematic trajectories prediction in vehicle-pedestrian-mixed scenes.
Huikun Bi, Zhong Fang, Tianlu Mao, Zhigang Deng 0001
ICCV5
2019 3D mesh animation compression based on adaptive spatio-temporal segmentation
abstract
With the recent advances of data acquisition techniques, the compression of various 3D mesh animation data has become an important topic in computer graphics community. In this paper, we present a new spatio-temporal segmentation-based approach for the compression of 3D mesh animations. Given an input mesh sequence, we first compute an initial temporal cut to obtain a small subsequence by detecting the temporal boundary of dynamic behavior. Then, we apply a two-stage vertex clustering on the resulting subsequence to classify the vertices into groups with optimal intra-affinities. After that, we design a temporal segmentation step based on the variations of the principle components within each vertex group prior to performing a PCA-based compression. Our approach can adaptively determine the temporal and spatial segmentation boundaries in order to exploit both temporal and spatial redundancies. We have conducted many experiments on different types of 3D mesh animations with various segmentation configurations. Our comparative studies show the competitive performance of our approach for the compression of 3D mesh animations.
Guoliang Luo, Zhigang Deng 0001, Xiaogang Jin 0001, Wenqiang Xie, Hyewon Seo
I3D2
2019 Real-time hierarchical facial performance capture
abstract
This paper presents a novel method to reconstruct high resolution facial geometry and appearance in real-time by capturing an individual-specific face model with fine-scale details, based on monocular RGB video input. Specifically, after reconstructing the coarse facial model from the input video, we subsequently refine it using shape-from-shading techniques, where illumination, albedo texture, and displacements are recovered by minimizing the difference between the synthesized face and the input RGB video. In order to recover wrinkle level details, we build a hierarchical face pyramid through adaptive subdivisions and progressive refinements of the mesh from a coarse level to a fine level. We both quantitatively and qualitatively evaluate our method through many experiments on various inputs. We demonstrate that our approach can produce results close to off-line methods and better than previous real-time methods.
Luming Ma, Zhigang Deng 0001
I3D2
2019 Single RGB-D Fitting: Total Human Modeling with an RGB-D Shot
abstract
Existing single shot based human modeling methods generally cannot model the complete pose details (e.g., head and hand positions) without non-trivial interactions. We explore the merits of both RGB and depth images and propose a new method called Single RGB-D Fitting (SRDF) to generate a realistic 3D human model with a single RGB-D shot from a consumer-grade depth camera. Specifically, the state-of-the-art deep learning techniques for RGB images are incorporated into SRDF, so that: 1) A compound skeleton detection method is introduced to obtain accurate 3D skeletons with refined hands based on the combination of depth and RGB images; and 2) an RGB image segmentation assisted point cloud pre-processing method is presented to obtain smooth foreground point clouds. In addition, several novel constraints are also introduced into the energy minimization model, including the shape continuity constraint, the keypoint-guided head pose prior constraint, and the penalty-enforced point cloud prior constraint. The energy model is optimized in a two-pass way so that a realistic shape can be estimated from coarse to fine. Through extensive experiments and comparisons with the state of the art methods, we demonstrate the effectiveness and efficiency of the proposed method.
Xianyong Fang, Jikui Yang, Jie Rao, Linbo Wang 0001, Zhigang Deng 0001
VRST5
2019 Real-Time Facial Expression Transformation for Monocular RGB Video
abstract
Abstract This paper describes a novel real‐time end‐to‐end system for facial expression transformation, without the need of any driving source. Its core idea is to directly generate desired and photo‐realistic facial expressions on top of input monocular RGB video. Specifically, an unpaired learning framework is developed to learn the mapping between any two facial expressions in the facial blendshape space. Then, it automatically transforms the source expression in an input video clip to a specified target expression through the combination of automated 3D face construction, the learned bi‐directional expression mapping and automated lip correction. It can be applied to new users without additional training. Its effectiveness is demonstrated through many experiments on faces from live and online video, with different identities, ages, speeches and expressions.
Luming Ma, Zhigang Deng 0001
Comput. Graph. Forum2
2019 A Color-Pair Based Approach for Accurate Color Harmony Estimation
abstract
Abstract Harmonious color combinations can stimulate positive user emotional responses. However, a widely open research question is: how can we establish a robust and accurate color harmony measure for the public and professional designers to identify the harmony level of a color theme or color set. Building upon the key discovery that color pairs play an important role in harmony estimation, in this paper we present a novel color‐pair based estimation model to accurately measure the color harmony. It first takes a two‐layer maximum likelihood estimation (MLE) based method to compute an initial prediction of color harmony by statistically modeling the pair‐wise color preferences from existing datasets. Then, the initial scores are refined through a back‐propagation neural network (BPNN) with a variety of color features extracted in different color spaces, so that an accurate harmony estimation can be obtained at the end. Our extensive experiments, including performance comparisons of harmony estimation applications, show the advantages of our method in comparison with the state of the art methods.
Bailin Yang, Tianxiang Wei, Xianyong Fang, Zhigang Deng 0001, Frederick W. B. Li, Xun Wang 0007
Comput. Graph. Forum4
2019 Efficient and realistic character animation through analytical physics-based skin deformation
Shaojun Bian, Zhigang Deng 0001, Ehtzaz Chaudhry, Lihua You, Xiaosong Yang, Hassan Ugail, Xiaogang Jin 0001, Zhidong Xiao, Jian J. Zhang 0001
Graph. Model.2
2019 3D articulated skeleton extraction using a single consumer-grade depth camera
Xuequan Lu, Zhigang Deng 0001, Jun Luo 0001, Wenzhi Chen, Sai-Kit Yeung, Ying He 0001
Comput. Vis. Image Underst.2
2019 Shape-constrained flying insects animation
abstract
Abstract During the past decades, high‐fidelity realistic simulations of various flying insects exhibiting collective behavior have been broadly used in entertainment industries and virtual reality applications. However, due to the intrinsic complexity and high computational cost, shape constrained simulation of collective behaviors remains a challenging topic. In this paper, we present a robust multi‐agent model for large‐scale controllable shape constrained simulation of flying insects. Specifically, we design an internal force model to biologically mimic an individual insect. We also propose an external force model based on a trade‐off mechanic to guide the insects smoothly deforming into a target shape. Our experimental results and comparative studies show our method is able to simulate realistic and dynamic flying insects with various user‐specified shape constraints.
Guoliang Luo, Yang Tong, Xiaogang Jin 0001, Zhigang Deng 0001
Comput. Animat. Virtual Worlds5
2019 Biologically inspired ant colony simulation
abstract
Abstract We present a unified biologically inspired approach to simulate ant colonies inspired by the key observation of collective behaviors of ants in nature. To generate the trajectories of virtual ants, we construct a motion controller to determine the motion states and the paths of virtual ants, considering dynamic internal and external interactions. The motion controller computes a target position for each ant at every time step according to its motion states. The motion states include four states: basic movement, the stop state, and two dynamic interactions (i.e., internal and external , respectively referring to interaction with neighbors for necessary information transfer about the destination, and interaction with surroundings such as food sources, nests, and obstacles) to represent basic exploration, casual or intentional stop, and purposeful movement, respectively. Based on the motion states, the motion controller plans an optimal path for each virtual ant. Through many simulation experiments, we demonstrate that our method is controllable, scalable, and flexible to simulate hybrid colonies with a large number of ants.
Jiaping Ren, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds4
2019 Screwing assembly oriented interactive model segmentation in HMD VR environment
abstract
Abstract Although different approaches of segmenting and assembling geometric models for 3D printing have been proposed, it is difficult to find any research studies, which investigate model segmentation and assembly in head‐mounted display (HMD) virtual reality (VR) environments for 3D printing. In this work, we propose a novel and interactive segmentation method for screwing assembly in the environments to tackle this problem. Our approach divides a large model into semantic parts with a screwing interface for repeated tight assembly. Specifically, after a user places the cutting interface, our algorithm computes the bounding box of the current part automatically for subsequent multicomponent semantic Boolean segmentations. Afterwards, the bolt is positioned with an improved K3M image thinning algorithm and is used for merging paired components with union and subtraction Boolean operations respectively. Moreover, we introduce a swept Boolean‐based rotation collision detection and location method to guarantee a collision‐free screwing assembly. Experiments show that our approach provides a new interactive multicomponent semantic segmentation tool that supports not only repeated installation and disassembly but also tight and aligned assembly.
Xiaoqiang Zhu, Shenshuai Chen, Xiangyang Wang 0003, Lihua You, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds9
2019 Motion-Aware Compression and Transmission of Mesh Animation Sequences
abstract
With the increasing demand in using 3D mesh data over networks, supporting effective compression and efficient transmission of meshes has caught lots of attention in recent years. This article introduces a novel compression method for 3D mesh animation sequences, supporting user-defined and progressive transmissions over networks. Our motion-aware approach starts with clustering animation frames based on their motion similarities, dividing a mesh animation sequence into fragments of varying lengths. This is done by a novel temporal clustering algorithm, which measures motion similarity based on the curvature and torsion of a space curve formed by corresponding vertices along a series of animation frames. We further segment each cluster based on mesh vertex coherence, representing topological proximity within an object under certain motion. To produce a compact representation, we perform intra-cluster compression based on Graph Fourier Transform (GFT) and Set Partitioning In Hierarchical Trees (SPIHT) coding. Optimized compression results can be achieved by applying GFT due to the proximity in vertex position and motion. We adapt SPIHT to support progressive transmission and design a mechanism to transmit mesh animation sequences with user-defined quality. Experimental results show that our method can obtain a high compression ratio while maintaining a low reconstruction error.
Bailin Yang, Luhong Zhang, Frederick W. B. Li, Xiaoheng Jiang, Zhigang Deng 0001, Meng Wang 0001, Mingliang Xu 0001
ACM Trans. Intell. Syst. Technol.5
2018 Unsupervised Articulated Skeleton Extraction From Point Set Sequences Captured by a Single Depth Camera
abstract
How to robustly and accurately extract articulated skeletons from point set sequences captured by a single consumer-grade depth camera still remains to be an unresolved challenge to date. To address this issue, we propose a novel, unsupervised approach consisting of three contributions (steps): (i) a non-rigid point set registration algorithm to first build one-to-one point correspondences among the frames of a sequence; (ii) a skeletal structure extraction algorithm to generate a skeleton with reasonable numbers of joints and bones; (iii) a skeleton joints estimation algorithm to achieve accurate joints. At the end, our method can produce a quality articulated skeleton from a single 3D point sequence corrupted with noise and outliers. The experimental results show that our approach soundly outperforms state of the art techniques, in terms of both visual quality and accuracy.
Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Zhigang Deng 0001, Wenzhi Chen
AAAI4
2018 An emotion evolution based model for collective behavior simulation
abstract
Current crowd simulation progresses still fall short of simulating many real-world collective behaviors. Arguably, one of the main reasons is that some essential qualities of human beings such as emotion have not been effectively modeled and incorporated into crowd simulation algorithms. In this paper, we propose a novel computational model for emotion evolution and demonstrate its applications for crowd simulation. Specifically, our approach is designed to tackle three major issues in the emotion evolution process: (i) how to perceive and evaluate emotion when individuals face emergency or external events, (ii) how to evolve the emotion during induction, and (iii) how specific actions of individuals in a crowd are impacted by emotion. Through many experiments, we demonstrate that our method can effectively simulate emergent dynamic collective patterns observed in real-world crowd footages.
Hao Jiang 0013, Zhigang Deng 0001, Xiangjun He, Tianlu Mao
I3D2
2018 Shadow traffic: A unified model for abnormal traffic behavior simulation
Mingliang Xu 0001, Fubao Zhu, Zhigang Deng 0001, Bing Zhou 0003
Comput. Graph.4
2018 Online Global Non-rigid Registration for 3D Object Reconstruction Using Consumer-level Depth Cameras
abstract
Abstract We investigate how to obtain high‐quality 360‐degree 3D reconstructions of small objects using consumer‐level depth cameras. For many homeware objects such as shoes and toys with dimensions around 0.06 – 0.4 meters, their whole projections, in the hand‐held scanning process, occupy fewer than 20% pixels of the camera's image. We observe that existing 3D reconstruction algorithms like KinectFusion and other similar methods often fail in such cases even under the close‐range depth setting. To achieve high‐quality 3D object reconstruction results at this scale, our algorithm relies on an online global non‐rigid registration, where embedded deformation graph is employed to handle the drifting of camera tracking and the possible nonlinear distortion in the captured depth data. We perform an automatic target object extraction from RGBD frames to remove the unrelated depth data so that the registration algorithm can focus on minimizing the geometric and photogrammetric distances of the RGBD data of target objects. Our algorithm is implemented using CUDA for a fast non‐rigid registration. The experimental results show that the proposed method can reconstruct high‐quality 3D shapes of various small objects with textures.
Jiamin Xu, Weiwei Xu 0003, Yin Yang 0002, Zhigang Deng 0001, Hujun Bao
Comput. Graph. Forum4
2018 Sketch-based shape-preserving tree animations
abstract
Abstract We present a novel and intuitive sketch‐based tree animation technique, targeting on generating a new type of special effect of smoothly transforming leafy trees into morphologically different new shapes. Both topological consistencies of branches and meaningful in‐between crown shapes are preserved during the transformation. Specifically, it takes a leafy tree and a user's sketch describing the silhouette of the desired crown shape under a certain viewpoint as the input. Based on a self‐adaptive multiscale cage tree representation, branches are locally transformed through a series of topology‐aware deformations, and the resulting tree conforms to the user‐designed shape, demonstrating better aesthetics compared to global single‐cage‐based methods. By interpolating the transformations, we are able to create visually pleasing shape‐preserving animations of trees transforming between two crown shapes. Our proposed framework also provides an efficient way to interactively edit leafy trees toward desired shapes, demonstrating its potential to leverage existing tree modeling frameworks by providing flexible and intuitive tree editing operations.
Luyuan Wang, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds3
2018 A fast garment fitting algorithm using skeleton-based error metric
abstract
Abstract We present a fast and automatic method to fit a given 3D garment onto a human model with various shapes and poses, without using a reference human model. Our approach uses a novel skeleton‐based error metric to find the pose that best fits the input garment. Specifically, we first generate the skeleton of the given human model and its corresponding skinning weights. Then, we iteratively rotate each bone to find its best position to fit the garment. After that, we rig the surface of the human model according to the transformations of the skeleton. Potential penetrations are resolved using collision handling and physically based simulation. Finally, we restore the human model back to the original pose in order to obtain the desired fitting result. Our experiment results show that besides its efficiency and automation, our method is about two orders of magnitudes faster than existing approaches, and it can handle various garments, including jacket, trousers, skirt, a suit of clothing, and even multilayered clothing.
Zhigang Deng 0001, Chen Liu 0012, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds2
2018 Realistic Data-Driven Traffic Flow Animation Using Texture Synthesis
abstract
We present a novel data-driven approach to populate virtual road networks with realistic traffic flows. Specifically, given a limited set of vehicle trajectories as the input samples, our approach first synthesizes a large set of vehicle trajectories. By taking the spatio-temporal information of traffic flows as a 2D texture, the generation of new traffic flows can be formulated as a texture synthesis process, which is solved by minimizing a newly developed traffic texture energy. The synthesized output captures the spatio-temporal dynamics of the input traffic flows, and the vehicle interactions in it strictly follow traffic rules. After that, we position the synthesized vehicle trajectory data to virtual road networks using a cage-based registration scheme, where a few traffic-specific constraints are enforced to maintain each vehicle's original spatial location and synchronize its motion in concert with its neighboring vehicles. Our approach is intuitive to control and scalable to the complexity of virtual road networks. We validated our approach through many experiments and paired comparison user studies.
Qianwen Chao, Zhigang Deng 0001, Jiaping Ren, Qianqian Ye, Xiaogang Jin 0001
IEEE Trans. Vis. Comput. Graph.2
2017 Perceptual enhancement of emotional mocap head motion: An experimental study
abstract
Motion capture (mocap) systems have been widely used to collect various human behavior data. Despite existing numerous research efforts on mocap motion processing and understanding, to the best of our knowledge, to date few works have been dedicated to the investigation into whether the mo-cap human behavior data can be further enhanced to improve its perception. In this work, we investigate whether and how it is feasible to consistently manipulate mocap emotional head motion to enhance its perceived expressiveness. Our study relies on a mocap audiovisual dataset acquired in a laboratory setting. Participants are invited to view the animation clips of a virtual talking character displaying the original mocap head motion or manipulated head motion, and then to rate their perceived expressiveness. Statistical analysis of the rated perceptions shows that humans are sensitive to the mean of head pitch rotation (called up-down rotation) in an utterance and that the expressiveness of emotion could be improved by adjusting the mean of head pitch rotation in an utterance.
Yu Ding 0001, Lei Shi 0027, Zhigang Deng 0001
ACII3
2017 A Multifaceted Study on Eye Contact based Speaker Identification in Three-party Conversations
abstract
To precisely understand human gaze behaviors in three-party conversations, this work is dedicated to look into whether the speaker can be reliably identified from the interlocutors in a three-party conversation on the basis of the interactive behaviors of eye contact, where speech signals are not provided. Derived from a pre-recorded, multimodal, and three-party conversational behavior dataset, a statistical framework is pro- posed to determine who is the speaker from the interactive behaviors of eye contact. Additionally, with the aid of virtual human technologies, a user study is conducted to study whether subjects are capable of distinguishing the speaker from the listeners according to the gaze behaviors of the interlocutors alone. Our results show that eye contact provides a reliable cue for the identification of the speaker in three-party conversations.
Yu Ding 0001, Meihua Xiao, Zhigang Deng 0001
CHI4
2017 Interactive cage generation for mesh deformation
abstract
Many previous efforts have been focused on generating optimal coordinates for cage deformation; cage generation for 3D models has been relatively understudied. We introduce an efficient complete pipeline to generate high quality cages for 3D models with arbitrary topological complexities, including high genus models and those without perceptible skeletal structures. Specifically, starting from user-specified cut slides, our method automatically optimizes the consistent, orthogonal orientations of cage cross sections. Then, through automated cage meshing and refining, it can further improve the cage quality by tackling the cage coverage issue and bounding the input model with a controllable tightness. Our experiments demonstrate this approach is efficient and robust to handle a variety of 3D models including human-like, animal, and high genus models.
Binh Huy Le, Zhigang Deng 0001
I3D2
2017 Evaluating Hex-mesh Quality Metrics via Correlation Analysis
abstract
Abstract Hexahedral (hex‐) meshes are important for solving partial differential equations (PDEs) in applications of scientific computing and mechanical engineering. Many methods have been proposed aiming to generate hex‐meshes with high scaled Jacobians. While it is well established that a hex‐mesh should be inversion‐free (i.e. have a positive Jacobian measured at every corner of its hexahedron), it is not well‐studied that whether the scaled Jacobian is the most effective indicator of the quality of simulations performed on inversion‐free hex‐meshes given the existing dozens of quality metrics for hex‐meshes. Due to the challenge of precisely defining the relations among metrics, studying the correlations among different quality metrics and their correlations with the stability and accuracy of the simulations is a first and effective approach to address the above question. In this work, we propose a correlation analysis framework to systematically study these correlations. Specifically, given a large hex‐mesh dataset, we classify the existing quality metrics into groups based on their correlations, which characterizes their similarity in measuring the quality of hex‐elements. In addition, we rank the individual metrics based on their correlations with the accuracy and stability metrics for simulations that solve a number of elliptic PDE problems. Our preliminary experiments suggest that metrics that assess the conditioning of the elements are more correlated to the quality of solving elliptic PDEs than the others. Furthermore, an inversion‐free hex‐mesh with higher average quality (measured by any quality metrics) usually leads to a more accurate and stable computation of elliptic PDEs. To support our correlation study and address the lack of a publicly available large hex‐mesh dataset with sufficiently varying quality metric values, we also propose a two‐level perturbation strategy to generate the desired dataset from a small number of meshes to exclude the influences of element numbers, vertex connectivity, and volume sizes to our study.
Xifeng Gao, Jin Huang 0001, Kaoji Xu, Zherong Pan, Zhigang Deng 0001, Guoning Chen
Comput. Graph. Forum5
2017 Hexahedral Meshing With Varying Element Sizes
abstract
Abstract Hexahedral (or Hex‐) meshes are preferred in a number of scientific and engineering simulations and analyses due to their desired numerical properties. Recent state‐of‐the‐art techniques can generate high‐quality hex‐meshes. However, they typically produce hex‐meshes with uniform element sizes and thus may fail to preserve small‐scale features on the boundary surface. In this work, we present a new framework that enables users to generate hex‐meshes with varying element sizes so that small features will be filled with smaller and denser elements, while the transition from smaller elements to larger ones is smooth, compared to the octree‐based approach. This is achieved by first detecting regions of interest (ROIs) of small‐scale features. These ROIs are then magnified using the as‐rigid‐as‐possible deformation with either an automatically determined or a user‐specified scale factor. A hex‐mesh is then generated from the deformed mesh using existing approaches that produce hex‐meshes with uniform‐sized elements. This initial hex‐mesh is then mapped back to the original volume before magnification to adjust the element sizes in those ROIs. We have applied this framework to a variety of man‐made and natural models to demonstrate its effectiveness.
Kaoji Xu, Xifeng Gao, Zhigang Deng 0001, Guoning Chen
Comput. Graph. Forum3
2017 Topologically consistent leafy tree morphing
abstract
Abstract We present a novel morphing technique to generate pleasing visual effects between 2 topologically varying trees while preserving the topological consistency and botanical meanings of any in‐between shapes as natural trees. Specifically, we first efficiently convert leafy trees into botanically inspired chain‐lobe representations in an automatic way. With the aid of branching‐pattern aware, one‐to‐many correspondences between branches and leaves, we hierarchically interpolate branches of in‐between trees while maintaining their topological consistencies. Finally, we simultaneously interpolate foliage, specifically every single leaf, during the morphing process, avoiding the generation of unpleasant “floating” leaves. We demonstrate the effectiveness of our approach by creating visually compelling tree morphing animations, even between cross‐species.
Luyuan Wang, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds3
2017 Robust structure simplification for hex re-meshing
abstract
We introduce a robust and automatic algorithm to simplify the structure and reduce the singularities of a hexahedral mesh. Our algorithm interleaves simplification operations to collapse sheets and chords of the base complex of the input mesh with a geometric optimization, which improves the elements quality. All our operations are guaranteed not to introduce elements with negative Jacobians, ensuring that our algorithm always produces valid hex-meshes, and not to increase the Hausdorff distance from the original shape more than a user-defined threshold, ensuring a faithful approximation of the input geometry. Our algorithm can improve meshes produced with any existing hexahedral meshing algorithm --- we demonstrate its effectiveness by processing a dataset of 194 hex-meshes created with octree-based, polycube-based, and field-aligned methods.
Xifeng Gao, Daniele Panozzo, Zhigang Deng 0001, Guoning Chen
ACM Trans. Graph.4
2017 Creative Virtual Tree Modeling Through Hierarchical Topology-Preserving Blending
abstract
We present a new method to efficiently generate a set of morphologically diverse and inspiring virtual trees through hierarchical topology-preserving blending, aiming to facilitate designers' creativity production. By maintaining the topological consistency of the tree branches, sequences of similar yet different trees and novel intermediate trees with encouragingly interesting structures are generated by performing inner-species and cross-species blending, respectively. Hierarchical fuzzy correspondences are automatically established between two or multiple trees based on the multi-scale topology tree representations. Fundamental blending tasks including morph, grow and wilt are introduced and organized into a tree-structured blending scheduler, which not only introduces the randomness into the blending procedure but also wisely schedules the tasks to generate topology-aware blending sequences, contributing to a variety of resulting trees that exhibit diversities in both geometry and topology. Most significantly, multiple batches of blending can be executed in parallel, resulting in a rapid creation of a large repository of diverse trees.
Xiaowei Xue, Xiaogang Jin 0001, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.4
2016 A Multidisciplinary, Multifaceted Approach to Improve the Computer Science based Game Design Education: Methodology and Assessment
abstract
In this paper, we introduce a multidisciplinary and multifaceted pedagogical approach to enhance game design education in computer science curriculum and assess its effectiveness using outcomes from Microsoft US and World Imagine Cup competitions in the game design category. We offer team project-based courses that cover multiple disciplines such as computer science, art and animation, game design, production, and business and entrepreneurship. Our students gain fundamental knowledge and skills from the multidisciplinary approach and utilize them to undergo a systematic game development process over two semesters. We also implement a unique grading system that includes ranking duels to promote the competitiveness among students which ultimately improves the quality of every game designed in our courses. We successfully demonstrate the effectiveness of our approach with results from the Microsoft Imagine Cup competitions - dozens of our student teams have been nationally and internationally recognized in the past eight consecutive years.
Chang Yun, Hesam Panahi, Zhigang Deng 0001
SIGCSE3
2016 Interactive mechanism modeling from multi-view images
abstract
In this paper, we present an interactive system for mechanism modeling from multi-view images. Its key feature is that the generated 3D mechanism models contain not only geometric shapes but also internal motion structures: they can be directly animated through kinematic simulation. Our system consists of two steps: interactive 3D modeling and stochastic motion parameter estimation. At the 3D modeling step, our system is designed to integrate the sparse 3D points reconstructed from multi-view images and a sketching interface to achieve accurate 3D modeling of a mechanism. To recover the motion parameters, we record a video clip of the mechanism motion and adopt stochastic optimization to recover its motion parameters by edge matching. Experimental results show that our system can achieve the 3D modeling of a range of mechanisms from simple mechanical toys to complex mechanism objects.
Weiwei Xu 0003, Zhigang Deng 0001, Yin Yang 0002, Kun Zhou 0001
ACM Trans. Graph.4
2016 Structured Volume Decomposition via Generalized Sweeping
abstract
In this paper, we introduce a volumetric partitioning strategy based on a generalized sweeping framework to seamlessly partition the volume of an input triangle mesh into a collection of deformed cuboids. This is achieved by a user-designed volumetric harmonic function that guides the decomposition of the input volume into a sequence of two-manifold level sets. A skeletal structure whose corners correspond to corner vertices of a 2D parameterization is extracted for each level set. Corners are placed so that the skeletal structure aligns with features of the input object. Then, a skeletal surface is constructed by matching the skeletal structures of adjacent level sets. The surface sheets of this skeletal surface partition the input volume into the deformed cuboids. The collection of cuboids does not exhibit T-junctions, significantly simplifying the hexahedral mesh generation process, and in particular, it simplifies fitting trivariate B-splines to the deformed cuboids. Intersections of the surface sheets of the skeletal surface correspond to the singular edges of the generated hex-meshes. We apply our technique to a variety of 3D objects and demonstrate the benefit of the structure decomposition in data fitting.
Xifeng Gao, Sai Deng, Elaine Cohen, Zhigang Deng 0001, Guoning Chen
IEEE Trans. Vis. Comput. Graph.5
2016 A Robust Scheme for Feature-Preserving Mesh Denoising
abstract
In recent years researchers have made noticeable progresses in mesh denoising, that is, recovering high-quality 3D models from meshes corrupted with noise (raw or synthetic). Nevertheless, these state of the art approaches still fall short for robustly handling various noisy 3D models. The main technical challenge of robust mesh denoising is to remove noise while maximally preserving geometric features. In particular, this issue becomes more difficult for models with considerable amount of noise. In this paper we present a novel scheme for robust feature-preserving mesh denoising. Given a noisy mesh input, our method first estimates an initial mesh, then performs feature detection, identification and connection, and finally, iteratively updates vertex positions based on the constructed feature edges. Through many experiments, we show that our approach can robustly and effectively denoise various input mesh models with synthetic noise or raw scanned noise. The qualitative and quantitative comparisons between our method and the selected state of the art methods also show that our approach can noticeably outperform them in terms of both quality and robustness.
Xuequan Lu, Zhigang Deng 0001, Wenzhi Chen
IEEE Trans. Vis. Comput. Graph.2
2015 Collective Crowd Formation Transform with Mutual Information-Based Runtime Feedback
abstract
Abstract This paper introduces a new crowd formation transform approach to achieve visually pleasing group formation transition and control. Its core idea is to transform crowd formation shapes with a least effort pair assignment using the Kuhn–Munkres algorithm, discover clusters of agent subgroups using affinity propagation and Delaunay triangulation algorithms and apply subgroup‐based social force model (SFM) to the agent subgroups to achieve alignment, cohesion and collision avoidance. Meanwhile, mutual information of the dynamic crowd is used to guide agents' movement at runtime. This approach combines both macroscopic (involving least effort position assignment and clustering) and microscopic (involving SFM) controls of the crowd transformation to maximally maintain subgroups' local stability and dynamic collective behaviour, while minimizing the overall effort (i.e. travelling distance) of the agents during the transformation. Through simulation experiments and comparisons, we demonstrate that this approach is efficient and effective to generate visually pleasing and smooth transformations and outperform several existing crowd simulation approaches including reciprocal velocity avoidances, optimal reciprocal collision avoidance and OpenSteer.
Yunpeng Wu, Yangdong Ye, Illés J. Farkas, Hao Jiang 0013, Zhigang Deng 0001
Comput. Graph. Forum6
2015 Spectral Animation Compression
Chao Wang 0088, Yang Liu 0013, Xiaohu Guo, Zichun Zhong, Binh Le, Zhigang Deng 0001
J. Comput. Sci. Technol.6
2015 Vehicle-pedestrian interaction for mixed traffic simulation
abstract
Abstract Simulation of real‐world traffic scenarios is widely needed in virtual environments. Different from many previous works on simulating vehicles or pedestrians separately, our approach aims to capture the realistic process of vehicle–pedestrian interaction for mixed traffic simulation. We model a decision‐making process for their interaction based on a gap acceptance judging criterion and then design a novel environmental feedback mechanism for both vehicles' and pedestrians' behavior‐control models to drive their motions. We demonstrate that our proposed method can soundly model vehicle–pedestrian interaction behaviors in a realistic and efficient manner and is convenient to be plugged into various traffic simulation systems. Copyright © 2015 John Wiley & Sons, Ltd.
Qianwen Chao, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds2
2015 An efficient lane model for complex traffic simulation
abstract
Abstract Traffic simulation heavily relies on lane model. This paper presents a novel method to model lanes based on the road axis under the Frenet frame. The road axis is generated from the geographic information system data after curve approximation, discretization, and compression. This lane model couples mileage information with three‐dimensional geometric information, so it offers an easy and fast position transformation from mileage to the Cartesian coordinate. It also keeps strictly consistent for mileage among neighboring lanes so that it facilitates lane‐change processing. Compared with existing methods that depict lanes as simple polylines or curves, the proposed lane model is more functional and more efficient, especially for complex traffic simulation with a large number of lane‐changes. Copyright © 2015 John Wiley & Sons, Ltd.
Tianlu Mao, Zhigang Deng 0001
Comput. Animat. Virtual Worlds3
2015 Expressive talking avatar synthesis and animation
Lei Xie 0001, Jia Jia 0001, Helen M. Meng, Zhigang Deng 0001
Multim. Tools Appl.4
2015 Hexahedral mesh re-parameterization from aligned base-complex
abstract
Recently, generating a high quality all-hex mesh of a given volume has gained much attention. However, little, if any, effort has been put into the optimization of the hex-mesh structure, which is equally important to the local element quality of a hex-mesh that may influence the performance and accuracy of subsequent computations. In this paper, we present a first and complete pipeline to optimize the global structure of a hex-mesh. Specifically, we first extract the base-complex of a hex-mesh and study the misalignments among its singularities by adapting the previously introduced hexahedral sheets to the base-complex. Second, we identify the valid removal base-complex sheets from the base-complex that contain misaligned singularities. We then propose an effective algorithm to remove these valid removal sheets in order. Finally, we present a structure-aware optimization strategy to improve the geometric quality of the resulting hex-mesh after fixing the misalignments. Our experimental results demonstrate that our pipeline can significantly reduce the number of components of a variety of hex-meshes generated by state-of-the-art methods, while maintaining high geometric quality.
Xifeng Gao, Zhigang Deng 0001, Guoning Chen
ACM Trans. Graph.2
2015 GPU-based polygonization and optimization for implicit surfaces
Xiaogang Jin 0001, Zhigang Deng 0001
Vis. Comput.3
2014 Inherent Noise-Aware Insect Swarm Simulation
abstract
Abstract Collective behaviour of winged insects is a wondrous and familiar phenomenon in the real world. In this paper, we introduce a highly efficient field‐based approach to simulate various insect swarms. Its core idea is to construct a smooth yet noise‐aware governing velocity field that can be further decomposed into two sub‐fields: (i) a divergence‐free curl‐noise field to model noise‐induced movements of individual insects in a swarm, and (ii) an enhanced global velocity field to control navigational paths in a complex environment along which all the insects in a swarm fly. Through simulation experiments and comparisons with existing crowd simulation approaches, we demonstrate that our approach is effective to simulate various insect swarm behaviours including aggregation, positive phototaxis, sedation, mass‐migrating, and so on. Besides its high efficiency, our approach is very friendly to parallel implementation on GPUs (e.g. the speedup achieved through GPU acceleration is higher than 50 if the number of simulated insects is more than 10 000 on an off‐the‐shelf computer). Our approach is the first multi‐agent modelling system that introduces curl‐noise into agents' velocity field and uses its non‐scattering nature to maintain non‐colliding movements in 3D crowd simulation.
Xinjie Wang 0003, Xiaogang Jin 0001, Zhigang Deng 0001, Linling Zhou
Comput. Graph. Forum3
2014 Crowd Simulation and Its Applications: Recent Advances
Mingliang Xu 0001, Hao Jiang 0013, Xiaogang Jin 0001, Zhigang Deng 0001
J. Comput. Sci. Technol.4
2014 AA-FVDM: An accident-avoidance full velocity difference model for animating realistic street-level traffic in rural scenes
abstract
ABSTRACT Most of existing traffic simulation efforts focus on urban regions with a coarse two‐dimensional representation; relatively few studies have been conducted to simulate realistic three‐dimensional traffic flows on a large, complex road web in rural scenes. In this paper, we present a novel agent‐based approach called accident‐avoidance full velocity difference model (abbreviated as AA‐FVDM) to simulate realistic street‐level rural traffics, on top of the existing FVDM. The main distinction between FVDM and AA‐FVDM is that FVDM cannot handle a critical real‐world traffic problem while AA‐FVDM settles this problem and retains the essence of FVDM. We also design a novel scheme to animate the lane‐changing maneuvering process (in particular, the execution course). Through numerous simulations, we demonstrate that besides addressing a previously unaddressed real‐world traffic problem, our AA‐FVDM method efficiently (in real time) simulates large‐scale traffic flows (tens of thousands of vehicles) with realistic, smooth effects. Furthermore, we validate our method using real‐world traffic data, and the validation results show that our method measurably outperforms state‐of‐the‐art traffic simulation methods.Copyright © 2013 John Wiley & Sons, Ltd.
Xuequan Lu, Wenzhi Chen, Zonghui Wang, Zhigang Deng 0001, Yangdong Ye
Comput. Animat. Virtual Worlds5
2014 A personality model for animating heterogeneous traffic behaviors
abstract
ABSTRACT How to automatically generate realistic and heterogeneous traffic behaviors has been a much needed yet challenging problem for numerous traffic simulation and urban planning applications. In this paper, we propose a novel approach to model heterogeneous traffic behaviors by adapting a well‐established personality trait model (i.e., Eysenck's PEN (psychoticism, extraversion and neuroticism) model) into widely used traffic simulation approaches. First, we collected a large amount of user feedback while users watch a variety of computer‐generated traffic simulation video clips. Then, we trained regression models to bridge low‐level traffic simulation parameters and high‐level perceived traffic behaviors (i.e., adjectives according to the PEN model and the three PEN traits). We also conducted an additional user study to validate the effectiveness and usefulness of our approach, in particular, high correlation coefficients and the Pearson values between users’ feedback and our model predictions prove the effectiveness of our approach. Furthermore, our approach can also produce interesting emergent traffic patterns including faster‐is‐slower effect and sticking‐in‐a‐pin‐wherever‐there‐is‐room effect. Copyright © 2014 John Wiley & Sons, Ltd.
Xuequan Lu, Zonghui Wang, Wenzhi Chen, Zhigang Deng 0001
Comput. Animat. Virtual Worlds5
2014 Flock morphing animation
abstract
ABSTRACT We propose a new animation technique, called flock morphing, to create special morphing effects between two arbitrary 3D objects by combining the features of 3D morphing and flock animation. Its core idea is first to tetrahedralize the source 3D mesh and regard each tetrahedron as an agent in a flock and then continually generate the flock morphing animation until the target mesh emerges, formed by the same set of tetrahedra. By applying plausible trajectory planning scheme and smooth deformation algorithm, we demonstrate that our proposed method can simultaneously achieve visually desired morphing effects. Copyright © 2014 John Wiley & Sons, Ltd.
Xinjie Wang 0003, Linling Zhou, Zhigang Deng 0001, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds3
2014 Multimodal joint information processing in human machine interaction: recent advances
Lei Xie 0001, Zhigang Deng 0001, Stephen J. Cox
Multim. Tools Appl.2
2014 Robust and accurate skeletal rigging from mesh sequences
abstract
We introduce an example-based rigging approach to automatically generate linear blend skinning models with skeletal structure. Based on a set of example poses, our approach can output its skeleton, joint positions, linear blend skinning weights, and corresponding bone transformations. The output can be directly used to set up skeleton-based animation in various 3D modeling and animation software as well as game engines. Specifically, we formulate the solving of a linear blend skinning model with a skeleton as an optimization with joint constraints and weight smoothness regularization, and solve it using an iterative rigging algorithm that (i) alternatively updates skinning weights, joint locations, and bone transformations, and (ii) automatically prunes redundant bones that can be generated by an over-estimated bone initialization. Due to the automatic redundant bone pruning, our approach is more robust than existing example-based rigging approaches. Furthermore, in terms of rigging accuracy, even with a single set of parameters, our approach can soundly outperform state of the art methods on various types of experimental datasets including humans, quadrupled animals, and highly deformable models.
Binh Huy Le, Zhigang Deng 0001
ACM Trans. Graph.2
2013 Implementation of a force-feedback interface for robotic assisted interventions with real-time MRI guidance
abstract
Efficient and intuitive interfacing of the interventionalist to the information and tools available from image-guided robotic assisted surgeries is required to achieve the full benefit of these technologies. Ongoing research has been performed into the use of forbidden region guided fixtures (FRVF) for human-in-the-loop control of image-guided procedures via haptic force-feedback devices (FFD). Although commercially available FFD provide sufficient degrees-of-freedom (DoF), collaborating clinicians, as well as the results of our previous work indicate that these systems are not completely intuitive for controlling fixed-point access interventional tool which have a remote center of motion. Within this context, we introduce a new FFD which is designed with the same DoF constraints as a fixed-point access interventional tool. The device is tested in a clinical simulation of a robot assisted trans-apical valve implantation under guidance from real-time magnetic resonance imaging. Pre-acquired real-time images are used in the clinical simulation to dynamically update the FRVF and therefore provide guiding forces to allow the operator to see the safe boundaries of operation via a visualization interface and physically feel them through the FFD. Inertial and gravity compensation and per DoF dynamic response of the physical prototype are validated and the frequency response of the system demonstrates it is adequate for tactile sensing. During clinical simulation the operator was successfully able to maneuver the tool within the safe path to the region of interest with the guidance of visual and force-feedback.
Nicholas C. von Sternberg, Atilla Kilicarslan, Nikhil V. Navkar, Zhigang Deng 0001, Karolos M. Grigoriadis, Nikolaos V. Tsekos
ICRA4
2013 A Text-Driven Conversational Avatar Interface for Instant Messaging on Mobile Devices
abstract
In this letter, we investigate the use of conversational avatars as a means to improve the user experience on instant messaging (IM) for mobile devices. We describe the design and implementation of an interface for IM featuring a 3-D facial avatar that is driven by the text messages being exchanged between chatting participants. Our design is affordable under the limited computational capacity of current generation mobiles. We evaluate user acceptance and reaction via user studies, by comparing it to a more conventional IM interface, and provide recommendations for the effective design of conversational avatar interfaces for mobile applications.
Mario Rincón Nigro, Zhigang Deng 0001
IEEE Trans. Hum. Mach. Syst.2
2013 Two-layer sparse compression of dense-weight blend skinning
abstract
Weighted linear interpolation has been widely used in many skinning techniques including linear blend skinning, dual quaternion blend skinning, and cage based deformation. To speed up performance, these skinning models typically employ a sparseness constraint, in which each 3D model vertex has a small fixed number of non-zero weights. However, the sparseness constraint also imposes certain limitations to skinning models and their various applications. This paper introduces an efficient two-layer sparse compression technique to substantially reduce the computational cost of a dense-weight skinning model, with insignificant loss of its visual quality. It can directly work on dense skinning weights or use example-based skinning decomposition to further improve its accuracy. Experiments and comparisons demonstrate that the introduced sparse compression model can significantly outperform state of the art weight reduction algorithms, as well as skinning decomposition algorithms with a sparseness constraint.
Binh Huy Le, Zhigang Deng 0001
ACM Trans. Graph.2
2013 Marker Optimization for Facial Motion Acquisition and Deformation
abstract
A long-standing problem in marker-based facial motion capture is what are the optimal facial mocap marker layouts. Despite its wide range of potential applications, this problem has not yet been systematically explored to date. This paper describes an approach to compute optimized marker layouts for facial motion acquisition as optimization of characteristic control points from a set of high-resolution, ground-truth facial mesh sequences. Specifically, the thin-shell linear deformation model is imposed onto the example pose reconstruction process via optional hard constraints such as symmetry and multiresolution constraints. Through our experiments and comparisons, we validate the effectiveness, robustness, and accuracy of our approach. Besides guiding minimal yet effective placement of facial mocap markers, we also describe and demonstrate its two selected applications: marker-based facial mesh skinning and multiresolution facial performance capture.
Binh Huy Le, Mingyang Zhu, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.3
2012 Intraoperative registration of preoperative 4D cardiac anatomy with real-time MR images
abstract
Co-registering pre- and intra- operative MR data is an important yet challenging problem due to different acquisition parameters, resolutions, and plane orientations. Despite its importance, previous approaches are often computationally intensive and thus cannot be employed in real-time. In this paper, a novel three-step approach is proposed to dynamically register pre-operative 4D MR data with intra-operative 2D RT-MRI to guide intracardiac procedures. Specifically, a novel preparatory step, executed in the pre-operative phase, is introduced to generate bridging information that can be used to significantly speed up the on-the-fly registration in the intraoperative procedure. Our experimental results demonstrate an accuracy of 0.42 mm and a processing speed of 26 FPS of the proposed approach on an off-the-shelf PC. This approach, is in particularly developed for performing intra-cardiac procedures with real-time MR guidance.
Xifeng Gao, Nikhil V. Navkar, Dipan J. Shah, Nikolaos V. Tsekos, Zhigang Deng 0001
BIBE5
2012 Visual and force-feedback guidance for robot-assisted interventions in the beating heart with real-time MRI
abstract
Robot-assisted surgical procedures are perpetually evolving due to potential improvement in patient treatment and healthcare cost reduction. Integration of an imaging modality intraoperatively further strengthens these procedures by incorporating the information pertaining to the area of intervention. Such information needs to be effectively rendered to the operator as a human-in-the-loop requirement. In this work, we propose a guidance approach that uses real-time MRI to assist the operator in performing robot-assisted procedure in a beating heart. Specifically, this approach provides both real-time visualization and force-feedback based guidance for maneuvering an interventional tool safely inside the dynamic environment of a heart's left ventricle. Experimental evaluation of the functionality of this approach was tested on a simulated scenario of transapical aortic valve replacement and it demonstrated improvement in control and manipulation by providing effective and accurate assistance to the operator in real-time.
Nikhil V. Navkar, Zhigang Deng 0001, Dipan J. Shah, Kostas E. Bekris, Nikolaos V. Tsekos
ICRA2
2012 A Surface-Based 3-D Dendritic Spine Detection Approach From Confocal Microscopy Images
abstract
Determining the relationship between the dendritic spine morphology and its functional properties is a fundamental challenge in neurobiology research. In particular, how to accurately and automatically analyse meaningful structural information from a large microscopy image data set is far away from being resolved. As pointed out in existing literature, one remaining challenge in spine detection and segmentation is how to automatically separate touching spines. In this paper, based on various global and local geometric features of the dendrite structure, we propose a novel approach to detect and segment neuronal spines, in particular, a breaking-down and stitching-up algorithm to accurately separate touching spines. Extensive performance comparisons show that our approach is more accurate and robust than two state-of-the-art spine detection and segmentation algorithms.
Qing Li 0008, Zhigang Deng 0001
IEEE Trans. Image Process.2
2012 Smooth skinning decomposition with rigid bones
abstract
This paper introduces the Smooth Skinning Decomposition with Rigid Bones (SSDR), an automated algorithm to extract the linear blend skinning (LBS) from a set of example poses. The SSDR model can effectively approximate the skin deformation of nearly articulated models as well as highly deformable models by a low number of rigid bones and a sparse, convex bone-vertex weight map. Formulated as a constrained optimization problem where the least squared error of the reconstructed vertices by LBS is minimized, the SSDR model can be solved by a block coordinate descent-based algorithm to iteratively update the weight map and the bone transformations. By employing the sparseness and convex constraints on the weight map, the SSDR model can be used for traditional skinning decomposition tasks such as animation compression and hardware-accelerated rendering. Moreover, by imposing the orthogonal constraints on the bone rotation matrices (rigid bones), the SSDR model can also be applied in motion editing, skeleton extraction, and collision detection tasks. Through qualitative and quantitative evaluations, we show the SSDR model can measurably outperform the state-of-the-art skinning decomposition schemes in terms of accuracy and applicability.
Binh Huy Le, Zhigang Deng 0001
ACM Trans. Graph.2
2012 A robust high-capacity affine-transformation-invariant scheme for watermarking 3D geometric models
abstract
In this article we propose a novel, robust, and high-capacity watermarking method for 3D meshes with arbitrary connectivities in the spatial domain based on affine invariants. Given a 3D mesh model, a watermark is embedded as affine-invariant length ratios of one diagonal segment to the residing diagonal intersected by the other one in a coplanar convex quadrilateral. In the extraction process, a watermark is recovered by combining all the watermark pieces embedded in length ratios through majority voting. Extensive experimental results demonstrate the robustness, high computational efficiency, high capacity, and affine-transformation-invariant characteristics of the proposed approach.
Xifeng Gao, Caiming Zhang 0001, Yan Huang 0003, Zhigang Deng 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2012 Live Speech Driven Head-and-Eye Motion Generators
abstract
This paper describes a fully automated framework to generate realistic head motion, eye gaze, and eyelid motion simultaneously based on live (or recorded) speech input. Its central idea is to learn separate yet interrelated statistical models for each component (head motion, gaze, or eyelid motion) from a prerecorded facial motion data set: 1) Gaussian Mixture Models and gradient descent optimization algorithm are employed to generate head motion from speech features; 2) Nonlinear Dynamic Canonical Correlation Analysis model is used to synthesize eye gaze from head motion and speech features, and 3) nonnegative linear regression is used to model voluntary eye lid motion and log-normal distribution is used to describe involuntary eye blinks. Several user studies are conducted to evaluate the effectiveness of the proposed speech-driven head and eye motion generator using the well-established paired comparison methodology. Our evaluation results clearly show that this approach can significantly outperform the state-of-the-art head and eye motion generation algorithms. In addition, a novel mocap+video hybrid data acquisition technique is introduced to record high-fidelity head movement, eye gaze, and eyelid motion simultaneously.
Binh Huy Le, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.3
2012 A Statistical Quality Model for Data-Driven Speech Animation
abstract
In recent years, data-driven speech animation approaches have achieved significant successes in terms of animation quality. However, how to automatically evaluate the realism of novel synthesized speech animations has been an important yet unsolved research problem. In this paper, we propose a novel statistical model (called SAQP) to automatically predict the quality of on-the-fly synthesized speech animations by various data-driven techniques. Its essential idea is to construct a phoneme-based, Speech Animation Trajectory Fitting (SATF) metric to describe speech animation synthesis errors and then build a statistical regression model to learn the association between the obtained SATF metric and the objective speech animation synthesis quality. Through delicately designed user studies, we evaluate the effectiveness and robustness of the proposed SAQP model. To the best of our knowledge, this work is the first-of-its-kind, quantitative quality model for data-driven speech animation. We believe it is the important first step to remove a critical technical barrier for applying data-driven speech animation techniques to numerous online or interactive talking avatar applications.
Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.2
2011 Perceptual analysis of talking avatar head movements: a quantitative perspective
abstract
Lifelike interface agents (e.g. talking avatars) have been increasingly used in human-computer interaction applications. In this work, we quantitatively analyze how human perception is affected by audio-head motion characteristics of talking avatars. Specifically, we quantify the correlation between perceptual user ratings (obtained via user study) and joint audio-head motion features as well as head motion patterns in the frequency-domain. Our quantitative analysis results clearly show that the correlation coefficient between the pitch of speech signals (but not the RMS energy of speech signals) and head motions is approximately linearly proportional to the perceptual user rating, and a larger proportion of high frequency signals in talking avatar head movements tends to degrade the user perception in terms of naturalness.
Binh Huy Le, Zhigang Deng 0001
CHI3
2011 Formation sketching: an approach to stylize groups in crowd simulation
Qin Gu, Zhigang Deng 0001
Graphics Interface2
2011 Magnetic resonance based control of a robotic manipulator for interventions in the beating heart
abstract
As a part of an ongoing project, in this paper we introduce the first version of a system which has a novel methodology for Cine (as in cinema) MRI based control of a cardiac robot for beating heart surgeries. The system uses the preoperative planning approach that we developed earlier, and integrates it to the intraoperative algorithms for controlling a robot and tracking some specific landmarks of a highly dynamical surgical field. In particular, our late studies presented herein aim to demonstrate the feasibility of integrating appropriate computational tools to achieve the volumetric image guidance for minimally invasive surgeries in the beating heart. We conceive of the system as practicable for in vitro experiments upon the completion of the first physical prototype, which may pave the way for expansion of the approach for other complex surgeries as well.
Erol Yeniaras, Johann Lamaury, Nikhil V. Navkar, Dipan J. Shah, Karen Chin, Zhigang Deng 0001, Nikolaos V. Tsekos
ICRA6
2011 Generation of 4D Access Corridors from Real-Time Multislice MRI for Guiding Transapical Aortic Valvuloplasties
Nikhil V. Navkar, Erol Yeniaras, Dipan J. Shah, Nikolaos V. Tsekos, Zhigang Deng 0001
MICCAI (1)5
2011 MR-Based Real Time Path Planning for Cardiac Operations with Transapical Access
Erol Yeniaras, Nikhil V. Navkar, Ahmet E. Sonmez, Dipan J. Shah, Zhigang Deng 0001, Nikolaos V. Tsekos
MICCAI (1)5
2011 A Global Spatial Similarity Optimization Scheme to Track Large Numbers of Dendritic Spines in Time-Lapse Confocal Microscopy
abstract
Dendritic spines form postsynaptic contact sites in the central nervous system. The rapid and spontaneous morphology changes of spines have been widely observed by neurobiologists. Determining the relationship between dendritic spine morphology change and its functional properties such as memory learning is a fundamental yet challenging problem in neurobiology research. In this paper, we propose a novel algorithm to track the morphology change of multiple spines simultaneously in time-lapse neuronal images based on nonrigid registration and integer programming. We also propose a robust scheme to link disappearing-and-reappearing spines. Performance comparisons with other state-of-the-art cell and spine tracking algorithms, and the ground truth show that our approach is more accurate and robust, and it is capable of tracking a large number of neuronal spines in time-lapse confocal microscopy images.
Qing Li 0008, Zhigang Deng 0001, Yong Zhang 0050, Xiaobo Zhou 0001, U. Valentin Nägerl, Stephen T. C. Wong
IEEE Trans. Medical Imaging2
2010 Crafting 3D faces using free form portrait sketching and plausible texture inference
Tanasai Sucontphunt, Borom Tunwattanapong, Zhigang Deng 0001, Ulrich Neumann
Graphics Interface3
2010 Perceiving motion transitions in pedestrian crowds
abstract
Perception of motion transitions in a pedestrian crowd is affected by many collective features such as crowd density, appearance variations, motion variations, and sub-group interaction patterns. We conducted a series of psychophysical experiments to investigate how these crowd features can influence human perception on walking motion transitions in a crowd when inexpensive motion blending algorithms are used. Our results provide useful implications and practical guidelines for performance-oriented crowd applications such as real-time games to improve the perceptual realism by effectively disguising motion transitions.
Qin Gu, Chang Yun, Zhigang Deng 0001
VRST3
2010 Image-based face illumination transferring using logarithmic total variation models
Qing Li 0008, Wotao Yin, Zhigang Deng 0001
Vis. Comput.3
2009 O' game, can you feel my frustration?: improving user's gaming experience via stresscam
abstract
One of the major challenges of video game design is to have appropriate difficulty levels for users in order to maximize the entertainment value of the game. Game players may lose interests if a game is either too easy or too difficult. This paper presents a novel methodology to improve user's experience in computer games by automatically adjusting the level of the game difficulty. The difficulty level is computed from measurements of the facial physiology of the players at a distance. The measurements are based on the assumption that the players' performance during the game-playing session alters blood flow in the supraorbital region, which is an indirect measurement of increased mental activities. This alters heat dissipation, which can be monitored in a contact-free manner through a thermal imaging-based stress monitoring and analysis system, known as StressCam.
Chang Yun, Dvijesh J. Shastri, Ioannis Pavlidis, Zhigang Deng 0001
CHI4
2009 Intermittency of slow arm movements increases in distal direction
abstract
When analyzed in the tangential speed domain, human movements exhibit a multi-peaked speed profile which is commonly interpreted as evidence for submovements. At slow speeds, the number of the peaks increases and the peaks also become more distinct, corresponding to non-smoothness or intermittency in the movement. In this study, we evaluate two potential sources proposed in the literature for the origins of movement intermittency and conclude that intermittency is more likely due to noise in the neuromuscular system as opposed to a central movement planner that generates intermittent plans. This conclusion is based on the assumption that the central planner would be expected to introduce similar levels of intermittency for different joints, while accumulating noise in the neuromuscular circuitry would be expected to exhibit itself as increase in noise in distal direction. We have used a 3D motion capture system to record trajectories of fingertip, wrist, elbow and shoulder as five participants completed a simple manual circular tracking task at various constant speed levels. Statistical analyses indicated that movement intermittency, quantified by a number of peaks metric, increased in distal direction, supporting the noise model for origins of intermittency. Movement speed was determined to have a significant effect on intermittency, while orientation of the task plane showed no significance.
Ozkan Celik, Qin Gu, Zhigang Deng 0001, Marcia Kilchenman O'Malley
IROS3
2009 An Interactive Geometric Technique for Upper and Lower Teeth Segmentation
Binh Huy Le, Zhigang Deng 0001, James J. Xia, Yu-Bing Chang, Xiaobo Zhou 0001
MICCAI (1)2
2009 Perceptually consistent example-based human motion retrieval
abstract
Large amount of human motion capture data have been increasingly recorded and used in animation and gaming applications. Efficient retrieval of logically similar motions from a large data repository thereby serves as a fundamental basis for these motion data based applications. In this paper we present a perceptually consistent, example-based human motion retrieval approach that is capable of efficiently searching for and ranking similar motion sequences given a query motion input. Our method employs a motion pattern discovery and matching scheme that breaks human motions into a part-based, hierarchical motion representation. Building upon this representation, a fast string match algorithm is used for efficient runtime motion query processing. Finally, we conducted comparative user studies to evaluate the accuracy and perceptual-consistency of our approach by comparing it with the state of the art example-based human motion search algorithms.
Zhigang Deng 0001, Qin Gu, Qing Li 0008
SI3D1
2009 Natural Eye Motion Synthesis by Modeling Gaze-Head Coupling
abstract
Due to the intrinsic subtlety and dynamics of eye movements, automated generation of natural and engaging eye motion has been a challenging task for decades. In this paper we present an effective technique to synthesize natural eye gazes given a head motion sequence as input, by statistically modeling the innate coupling between gazes and head movements. We first simultaneously recorded head motions and eye gazes of human subjects, using a novel hybrid data acquisition solution consisting of an optical motion capture system and off-the-shelf video cameras. Then, we statistically learn gaze-head coupling patterns using a dynamic coupled component analysis model. Finally, given a head motion sequence as input, we can synthesize its corresponding natural eye gazes based on the constructed gaze-head coupling model. Through comparative user studies and evaluations, we found that comparing with the state of the art algorithms in eye motion synthesis, our approach is more effective to generate natural gazes correlated with given head motions. We also showed the effectiveness of our approach for gaze simulation in two-party conversations.
Zhigang Deng 0001
VR2
2009 Crafting Personalized Facial Avatars Using Editable Portrait and Photograph Example
abstract
Computer-generated facial avatars have been increasingly used in a variety of virtual reality applications. Emulating the real-world face sculpting process, we present an interactive system to intuitively craft personalized 3D facial avatars by using 3D portrait editing and image example-based painting techniques. Starting from a default 3D face portrait, users can conveniently perform intuitive "pulling" operations on its 3D surface to sculpt the 3D face shape towards any individual. To automatically maintain the faceness of the 3D face being crafted, novel facial anthropometry constraints and a reduced face description space are incorporated into the crafting algorithms dynamically. Once the 3D face geometry is crafted, this system can automatically generate a face texture for the crafted model using an image example-based painting algorithm. Our user studies showed that with this system, users are able to craft a personalized 3D facial avatar efficiently on average within one minute.
Tanasai Sucontphunt, Zhigang Deng 0001, Ulrich Neumann
VR2
2009 Compression of Human Motion Capture Data Using Motion Pattern Indexing
abstract
Abstract In this work, a novel scheme is proposed to compress human motion capture data based on hierarchical structure construction and motion pattern indexing. For a given sequence of 3D motion capture data of human body, the 3D markers are first organized into a hierarchy where each node corresponds to a meaningful part of the human body. Then, the motion sequence corresponding to each body part is coded separately. Based on the observation that there is a high degree of spatial and temporal correlation among the 3D marker positions, we strive to identify motion patterns that form a database for each meaningful body part. Thereafter, a sequence of motion capture data can be efficiently represented as a series of motion pattern indices. As a result, higher compression ratio has been achieved when compared with the prior art, especially for long sequences of motion capture data with repetitive motion styles. Another distinction of this work is that it provides means for flexible and intuitive global and local distortion controls.
Qin Gu, Jingliang Peng, Zhigang Deng 0001
Comput. Graph. Forum3
2008 A Novel Visualization System for Expressive Facial Motion Data Exploration
abstract
Facial emotions and expressive facial motions have become an intrinsic part of many graphics systems and human computer interaction applications. The dynamics and high dimensionality of facial motion data make its exploration and processing challenging. In this paper, we propose a novel visualization system for expressive facial motion data exploration. Based on Principal Component Analysis (PCA) dimensionality reduction on anatomical facial sub regions, high dimensional facial motion data is mapped to 3D spaces. We further rendered it as colored 3D trajectories and color represents different emotion. We design an intuitive interface to allow users effectively explore and analyze high dimensional facial motion spaces. The applications of our visualization system on novel facial motion synthesis and emotion recognition are demonstrated.
Tanasai Sucontphunt, Xiaoru Yuan, Qing Li 0008, Zhigang Deng 0001
PacificVis4
2008 Interactive 3D facial expression posing through 2D portrait manipulation
Tanasai Sucontphunt, Zhenyao Mo, Ulrich Neumann, Zhigang Deng 0001
Graphics Interface4
2008 Expressive Speech Animation Synthesis with Phoneme-Level Controls
abstract
Abstract This paper presents a novel data‐driven expressive speech animation synthesis system with phoneme‐level controls. This system is based on a pre‐recorded facial motion capture database, where an actress was directed to recite a pre‐designed corpus with four facial expressions (neutral, happiness, anger and sadness). Given new phoneme‐aligned expressive speech and its emotion modifiers as inputs, a constrained dynamic programming algorithm is used to search for best‐matched captured motion clips from the processed facial motion database by minimizing a cost function. Users optionally specify ‘hard constraints’ (motion‐node constraints for expressing phoneme utterances) and ‘soft constraints’ (emotion modifiers) to guide this search process. We also introduce a phoneme–Isomap interface for visualizing and interacting phoneme clusters that are typically composed of thousands of facial motion capture frames. On top of this novel visualization interface, users can conveniently remove contaminated motion subsequences from a large facial motion dataset. Facial animation synthesis experiments and objective comparisons between synthesized facial motion and captured motion showed that this system is effective for producing realistic expressive speech animations.
Zhigang Deng 0001, Ulrich Neumann
Comput. Graph. Forum1
2007 Online Motion Capture Marker Labeling for Multiple Interacting Articulated Targets
abstract
Abstract In this paper, we propose an online motion capture marker labeling approach for multiple interacting articulated targets. Given hundreds of unlabeled motion capture markers from multiple articulated targets that are interacting each other, our approach automatically labels these markers frame by frame, by fitting rigid bodies and exploiting trained structure and motion models. Advantages of our approach include: 1) our method is an online algorithm, which requires no user interaction once the algorithm starts. 2) Our method is more robust than traditional the closest point‐based approaches by automatically imposing the structure and motion models. 3) Due to the use of the structure model which encodes the rigidity of each articulated body of captured targets, our method can recover missing markers robustly. Our approach is efficient and particularly suited for online computer animation and video game applications.
Qing Li 0008, Zhigang Deng 0001
Comput. Graph. Forum3
2007 Rigid Head Motion in Expressive Speech Animation: Analysis and Synthesis
abstract
Rigid head motion is a gesture that conveys important nonverbal information in human communication, and hence it needs to be appropriately modeled and included in realistic facial animations to effectively mimic human behaviors. In this paper, head motion sequences in expressive facial animations are analyzed in terms of their naturalness and emotional salience in perception. Statistical measures are derived from an audiovisual database, comprising synchronized facial gestures and speech, which revealed characteristic patterns in emotional head motion sequences. Head motion patterns with neutral speech significantly differ from head motion patterns with emotional speech in motion activation, range, and velocity. The results show that head motion provides discriminating information about emotional categories. An approach to synthesize emotional head motion sequences driven by prosodic features is presented, expanding upon our previous framework on head motion synthesis. This method naturally models the specific temporal dynamics of emotional head motion sequences by building hidden Markov models for each emotional category (sadness, happiness, anger, and neutral state). Human raters were asked to assess the naturalness and the emotional content of the facial animations. On average, the synthesized head motion sequences were perceived even more natural than the original head motion sequences. The results also show that head motion modifies the emotional perception of the facial animation especially in the valence and activation domain. These results suggest that appropriate head motion not only significantly improves the naturalness of the animation but can also be used to enhance the emotional content of the animation to effectively engage the users
Carlos Busso, Zhigang Deng 0001, Michael Grimm, Ulrich Neumann, Shri Narayanan
IEEE Trans. Speech Audio Process.2
2006 Perceiving Visual Emotions with Speech
Zhigang Deng 0001, Jeremy N. Bailenson, John P. Lewis, Ulrich Neumann
IVA1
2006 Animating blendshape faces by cross-mapping motion capture data
abstract
Animating 3D faces to achieve compelling realism is a challenging task in the entertainment industry. Previously proposed face transfer approaches generally require a high-quality animated source face in order to transfer its motion to new 3D faces. In this work, we present a semi-automatic technique to directly animate popularized 3D blendshape face models by mapping facial motion capture data spaces to 3D blendshape face spaces. After sparse markers on the face of a human subject are captured by motion capture systems while a video camera is simultaneously used to record his/her front face, then we carefully select a few motion capture frames and accompanying video frames as reference mocap-video pairs. Users manually tune blendshape weights to perceptually match the animated blendshape face models with reference facial images (the reference mocap-video pairs) in order to create reference mocap-weight pairs. Finally, the Radial Basis Function (RBF) regression technique is used to map any new facial motion capture frame to blendshape weights based on the reference mocap-weight pairs. Our results demonstrate that this technique is efficient to animate blendshape face models, while offering its generality and flexiblity.
Zhigang Deng 0001, Pei-Ying Chiang, Pamela Fox, Ulrich Neumann
SI3D1
2006 Expressive Facial Animation Synthesis by Learning Speech Coarticulation and Expression Spaces
abstract
Synthesizing expressive facial animation is a very challenging topic within the graphics community. In this paper, we present an expressive facial animation synthesis system enabled by automated learning from facial motion capture data. Accurate 3D motions of the markers on the face of a human subject are captured while he/she recites a predesigned corpus, with specific spoken and visual expressions. We present a novel motion capture mining technique that "learns" speech coarticulation models for diphones and triphones from the recorded data. A Phoneme-Independent Expression Eigenspace (PIEES) that encloses the dynamic expression signals is constructed by motion signal processing (phoneme-based time-warping and subtraction) and Principal Component Analysis (PCA) reduction. New expressive facial animations are synthesized as follows: First, the learned coarticulation models are concatenated to synthesize neutral visual speech according to novel speech input, then a texture-synthesis-based approach is used to generate a novel dynamic expression signal from the PIEES model, and finally the synthesized expression signal is blended with the synthesized neutral visual speech to create the final expressive facial animation. Our experiments demonstrate that the system can effectively synthesize realistic expressive facial animation.
Zhigang Deng 0001, Ulrich Neumann, John P. Lewis, Tae-Yong Kim 0002, Murtaza Bulut, Shri Narayanan
IEEE Trans. Vis. Comput. Graph.1
2005 Synthesizing speech animation by learning compact speech co-articulation models
abstract
While speech animation fundamentally consists of a sequence of phonemes over time, sophisticated animation requires smooth interpolation and co-articulation effects, where the preceding and following phonemes influence the shape of a phoneme. Co-articulation has been approached in speech animation research in several ways, most often by simply smoothing the mouth geometry motion over time. Data-driven approaches tend to generate realistic speech animation, but they need to store a large facial motion database, which is not feasible for real time gaming and interactive applications on platforms such as PDAs and cell phones. In this paper we show that accurate speech co-articulation model with compact size can be learned from facial motion capture data. An initial phoneme sequence is generated automatically from text-to-speech (TTS) systems. Then, our learned co-articulation model is applied to the resulting phoneme sequence, producing natural and detailed motion. The contribution of this work is that speech co-articulation models "learned" from real human motion data can be used to generate natural-looking speech motion while simultaneously preserving the expressiveness of the animation via keyframing control. Simultaneously, this approach can be effectively applied to interactive applications due to its compact size.
Zhigang Deng 0001, John P. Lewis, Ulrich Neumann
Computer Graphics International1
2005 Reducing blendshape interference by selected motion attenuation
abstract
Blendshapes (linear shape interpolation models) are perhaps the most commonly employed technique in facial animation practice. A major problem in creating blendshape animation is that of blendshape interference: the adjustment of a single blendshape "slider" may degrade the effects obtained with previous slider movements, because the blendshapes have overlapping, non-orthogonal effects. Because models used in commercial practice may have 100 or more individual blendshapes, the interference problem is the subject of considerable manual effort. Modelers iteratively resculpt models to reduce interference where possible, and animators must compensate for those interference effects that remain. In this short paper we consider the blendshape interference problem from a linear algebra point of view. We find that while full orthogonality is not desirable, the goal of preserving previous adjustments to the model can be effectively approached by allowing the user to temporarily designate a set of points as representative of the previous (desired) adjustments. We then simply solve for blendshape slider values that mimic desired new movement while moving these "tagged" points as little as possible. The resulting algorithm is easy to implement and demonstrably reduces cases of blendshape interference found in existing models.
John P. Lewis, Jonathan Mooser, Zhigang Deng 0001, Ulrich Neumann
SI3D3
2005 Natural head motion synthesis driven by acoustic prosodic features
abstract
Abstract Natural head motion is important to realistic facial animation and engaging human–computer interactions. In this paper, we present a novel data‐driven approach to synthesize appropriate head motion by sampling from trained hidden markov models (HMMs). First, while an actress recited a corpus specifically designed to elicit various emotions, her 3D head motion was captured and further processed to construct a head motion database that included synchronized speech information. Then, an HMM for each discrete head motion representation (derived directly from data using vector quantization) was created by using acoustic prosodic features derived from speech. Finally, first‐order Markov models and interpolation techniques were used to smooth the synthesized sequence. Our comparison experiments and novel synthesis results show that synthesized head motions follow the temporal dynamic behavior of real human subjects. Copyright © 2005 John Wiley & Sons, Ltd.
Carlos Busso, Zhigang Deng 0001, Ulrich Neumann, Shri Narayanan
Comput. Animat. Virtual Worlds2
2004 Analysis of emotion recognition using facial expressions, speech and multimodal information
abstract
The interaction between human beings and computers will be more natural if computers are able to perceive and respond to human non-verbal communication such as emotions. Although several approaches have been proposed to recognize human emotions based on facial expressions or speech, relatively limited work has been done to fuse these two, and other, modalities to improve the accuracy and robustness of the emotion recognition system. This paper analyzes the strengths and the limitations of systems based only on facial expressions or acoustic information. It also discusses two approaches used to fuse these two modalities: decision level and feature level integration. Using a database recorded from an actress, four emotions were classified: sadness, anger, happiness, and neutral state. By the use of markers on her face, detailed facial motions were captured with motion capture, in conjunction with simultaneous speech recordings. The results reveal that the system based on facial expression gave better performance than the system based on just acoustic information for the emotions considered. Results also show the complementarily of the two modalities and that when these two modalities are fused, the performance and the robustness of the emotion recognition system improve measurably.
Carlos Busso, Zhigang Deng 0001, Serdar Yildirim, Murtaza Bulut, Chul Min Lee, Abe Kazemzadeh, Sungbok Lee, Ulrich Neumann, Shri Narayanan
ICMI2
2004 Emotion recognition based on phoneme classes
abstract
Recognizing human emotions/attitudes from speech cues has gained increased attention recently. Most previous work has focused primarily on suprasegmental prosodic features calcu-lated at the utterance level for modeling against details at the segmental phoneme level. Based on the hypothesis that dif-ferent emotions have varying effects on the properties of the different speech sounds, this paper investigates the usefulness of phoneme-level modeling for the classification of emotional states from speech. Hidden Markov models (HMM) based on short-term spectral features are used for this purpose using data obtained from a recording of an actress ’ expressing 4 different emotional states- anger, happiness, neutral, and sadness. We designed and compared two sets of HMM classifiers: a generic set of “emotional speech ” HMMs (one for each emotion) and a set of broad phonetic-class based HMMs for each emotion type considered. Five broad phonetic classes were used to explore the effect of emotional coloring on different phoneme classes, and it was found that spectral properties of vowel sounds were the best indicator of emotions in terms of the classification per-formance. The experiments also showed that the better per-formance can be obtained by using phoneme-class classifiers than generic “emotional ” HMM classifier and classifiers based on global prosodic features. To see the complementary effect of the prosodic and spectral features, the two classifiers were combined at the decision level. The improvement was 0.55% in absolute (0.7 % relatively) compared with the result from phoneme-class based HMM classifier. 1.
Chul Min Lee, Serdar Yildirim, Murtaza Bulut, Abe Kazemzadeh, Carlos Busso, Zhigang Deng 0001, Sungbok Lee, Shri Narayanan
INTERSPEECH6
2004 An acoustic study of emotions expressed in speech
abstract
In this study, we investigate acoustic properties of speech associ-ated with four different emotions (sadness, anger, happiness, and neutral) intentionally expressed in speech by an actress. The aim is to obtain detailed acoustic knowledge on how speech is modulated when speaker’s emotion changes from neutral to a certain emotional state. It is based on measurements of acoustic parameters related to speech prosody, vowel articulation and spectral energy distribution. Acoustic similarities and differences among the emotions are then explored with mutual information computation, multidimensional scaling, and comparison of acoustic likelihoods relative to the neu-tral emotion. In addition, acoustic separability of the emotions is tested using the discriminant analysis at the utterance level and the result is compared with human evaluation. Results show that hap-piness/anger and neutral/sadness share similar acoustic properties in this speaker. Speech associated with anger and happiness are characterized by longer utterance duration, shorter inter-word si-lence, higher pitch and energy values with wider ranges, showing the characteristics of exaggerated or hyperarticulated speech. The discriminant analysis indicates that within-group acoustic separa-bility is relatively poor, suggesting that conventional acoustic pa-rameters examined in this study are not effective in describing the emotions along the valence (or pleasure) dimension. It is noted that RMS energy, inter-word silence and speaking rate are useful in dis-tinguishing sadness from others. Interestingly, the between-group difference in formant patterns seems better reflected in back vowels such as /a / (/father/) than in the front vowels. Larger lip opening and/or more tongue constriction at the mid or rear part of the vocal tract could be underlying reasons. 1.
Serdar Yildirim, Murtaza Bulut, Chul Min Lee, Abe Kazemzadeh, Zhigang Deng 0001, Sungbok Lee, Shri Narayanan, Carlos Busso
INTERSPEECH5
2003 Practical eye movement model using texture synthesis
abstract
As humans we are especially sensitive to the appearance of the face, and on the face, the eyes are particularly important. In fact, in attempts to animate photo-realistic CG humans, the eyes are very often what destroys the illusion [Williams 2003]. The state of art in eye movement synthesis is the Eyes Alive model [Lee et al. 2002] that develops a custom statistical model specifically for eye movement. While its results are the best to date, the model is complex and one wonders if it could be improved by using additional or different statistics. In fact the problem of generating novel animation that captures the “character” of given training data is the same problem as texture synthesis. In this sketch we describe a practical eye movement model using non-parametric texture synthesis techniques ([Efros and Leung 1999]), simulating the eye gaze motion and eye blink motion simultaneously. This approach uses the data directly and without an intervening humancrafted statistical model, yet it produces results that appear as good or better than the more complex statistical model. 2 Approach
Zhigang Deng 0001, John P. Lewis, Ulrich Neumann
SIGGRAPH1