Weidong Geng

dblp:55/6597 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 S2-Edit3DV: Diffusion-Guided Style Meets Structure for Consistent Multi-View 3D Video Generation
abstract
Consistently stylizing and editing 3D objects from multiple viewpoints is crucial for immersive applications such as virtual reality, augmented reality, and digital entertainment. Nevertheless, existing methods frequently face significant challenges, including inconsistent textures, pronounced drifting artifacts, and compromised geometric integrity when rendered from various perspectives. To effectively address these limitations, we introduce S2-Edit3DV, a novel diffusion-guided framework that reframes multi-view 3D objects editing as a temporally coherent video editing problem. By exploiting the robust single-view generative capabilities of SV3D, our approach reliably propagates initial style edits across different viewpoints, substantially mitigating drifting artifacts prevalent in current video-based editing methods. To further enhance semantic precision and structural preservation, we propose two innovative techniques: Attention-based Differential Style Injection (ADSI) and Adaptive Structural-aware Plug-and-Play (AS-PnP). ADSI utilizes attention-driven semantic embeddings for adaptive and precise style injection, effectively reducing semantic hallucinations. AS-PnP strategically modulates stylized latent features, balancing artistic expression with strict structural coherence. Comprehensive evaluations and ablation studies demonstrate that our proposed framework significantly enhances multi-view consistency, preserves fine-grained geometric details, and ensures accurate semantic alignment, showcasing superior performance and practical value for generating high-quality, creatively stylized, and structurally robust objects.
Xiubo Liang, Hongzhi Wang 0009, Weidong Geng
ACM Multimedia5
2025 SGM-Transformer: Rethinking Gradient Information Loss and Compensation in Spiking Neural Networks
Xiubo Liang, Hongzhi Wang 0009, Zigen Li, Jinxing Han, Weidong Geng
ACM Multimedia6
2025 Spike-RetinexFormer: Rethinking Low-light Image Enhancement with Spiking Neural Networks
abstract
Low-light image enhancement (LLIE) aims to improve the visibility and quality of images captured under poor illumination. However, existing deep enhancement methods often underemphasize computational efficiency, leading to high energy and memory costs. We propose \textbf{Spike-RetinexFormer}, a novel LLIE architecture that synergistically integrates Retinex theory, spiking neural networks (SNNs) and a Transformer-based design. Leveraging sparse spike-driven computation, the model reduces theoretical compute energy and memory traffic relative to ANN counterparts. Across standard benchmarks, the method matches or surpasses strong ANNs (25.50 dB on LOL-v1; 30.37 dB on SDSD-out) with comparable parameters and lower theoretical energy. Our work pioneers the synergistic integration of SNNs into Transformer architectures for LLIE, establishing a compelling pathway toward powerful, energy-efficient low-level vision on resource-constrained platforms.
Hongzhi Wang 0009, Xiubo Liang, Jinxing Han, Weidong Geng
NeurIPS4
2025 Effects of User Interface Orientation on Sense of Immersion in Augmented Reality
abstract
The sense of immersion provided by overlying computer-generated graphics over actual items is one of augmented reality’s distinguishing features. It is affected by multiple factors such as the user interface orientation in the virtual physical hybrid space. There have been many studies on realistic graphic rendering and multimodal interaction in augmented reality. Few studies, however, have systematically investigated the influence of the user interface orientations, more specifically the user-based orientation (UO) and object-based orientation (OO), on various aspects of immersion. We developed the UO and OO user interface prototypes and evaluated the sense of immersion in terms of engagement, engrossment, and overall immersion. The results showed that the user interface orientations in augmented reality had a significant influence on users’ perceived engagement and engrossment, whereas there was no significant difference in overall immersion. Particularly, compared to the OO user interfaces, the UO user interfaces exhibited higher perceived usability and slower focus of attention switching. Generalising implications for alternative augmented reality are provided.
Kailin Yin, Yifei Shan, Weidong Geng
Int. J. Hum. Comput. Interact.5
2024 MCM: Multi-condition Motion Synthesis Framework
Zeyu Ling, Bo Han 0003, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
IJCAI6
2024 PSSD-Transformer: Powerful Sparse Spike-Driven Transformer for Image Semantic Segmentation
abstract
Spiking Neural Networks (SNNs) have indeed shown remarkable promise in the field of computer vision, emerging as a low-energy alternative to traditional Artificial Neural Networks (ANNs). However, SNNs also face several challenges: i) Existing SNNs are not purely additive and involve a substantial amount of floating-point computations, which contradicts the original design intention of adapting to neuromorphic chips; ii) The incorrect positioning of convolutional and pooling layers relative to spiking layers leads to reduced accuracy; iii) Leaky Integrate-and-Fire (LIF) neurons have limited capability in representing local information, which is disadvantageous for downstream visual tasks like semantic segmentation.
Hongzhi Wang 0009, Xiubo Liang, Tao Zhang 0178, Weidong Geng
ACM Multimedia5
2024 Comparative Study on 2D and 3D User Interface for Eliminating Cognitive Loads in Augmented Reality Repetitive Tasks
abstract
The two-dimensional (2D) and three-dimensional (3D) user interfaces have been a prominent topic of augmented reality (AR) research, and their impact on the efficacy and usability of one-time tasks has been extensively examined. As AR is increasingly adopted in industry for repetitive tasks, there is an urgent need for research into the effect of the 2D and 3D user interfaces. In this study, we developed two prototypes of the user interfaces and conducted a comparison study with forty participants to assess their respective influence on cognitive load, perceived usability, and learnability. The results showed that the two user interfaces differed significantly in cognitive load and perceived usability. In particular, the 3D user interfaces exhibited substantially shorter eye blinking durations, shorter eye fixation durations, less dispersed eye gaze areas, lower subjective ratings of cognitive load, and higher overall scores of perceived usability than the 2D user interfaces. However, there was no difference in learnability between the two user interfaces throughout the repetitive tasks.
Cindy Zheng, Zhenghua Pan, Zhongnan Huang, Yuting Niu, Weidong Geng
Int. J. Hum. Comput. Interact.7
2024 Unsupervised Domain Adaptation by Causal Learning for Biometric Signal-based HCI
abstract
Biometric signal based human-computer interface (HCI) has attracted increasing attention due to its wide application in healthcare, entertainment, neurocomputing, and so on. In recent years, deep learning-based approaches have made great progress on biometric signal processing. However, the state-of-the-art (SOTA) approaches still suffer from model degradation across subjects or sessions. In this work, we propose a novel unsupervised domain adaptation approach for biometric signal-based HCI via causal representation learning. Specifically, three kinds of interventions on biometric signals (i.e., subjects, sessions, and trials) can be selected to generalize deep models across the selected intervention. In the proposed approach, a generative model is trained for producing intervened features that are subsequently used for learning transferable and causal relations with three modes. Experiments on the EEG-based emotion recognition task and sEMG-based gesture recognition task are conducted to confirm the superiority of our approach. An improvement of +0.21% on the task of inter-subject EEG-based emotion recognition is achieved using our approach. Besides, on the task of inter-session sEMG-based gesture recognition, our approach achieves improvements of +1.47%, +3.36%, +1.71%, and +1.01% on sEMG datasets including CSL-HDEMG, CapgMyo DB-b, 3DC, and Ninapro DB6, respectively. The proposed approach also works on the task of inter-trial sEMG-based gesture recognition and an average improvement of +0.66% on Ninapro databases is achieved. These experimental results show the superiority of the proposed approach compared with the SOTA unsupervised domain adaptation methods on HCIs based on biometric signal.
Qingfeng Dai, Yongkang Wong, Guofei Sun, Zhou Zhou 0012, Mohan Kankanhalli, Weidong Geng
ACM Trans. Multim. Comput. Commun. Appl.8
2024 Speech-Driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference
abstract
Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional, personality, and even subtle unknown features). Traditional approaches, reliant on various explicit feature inputs and complex multimodal processing, constrain the expressiveness of resulting gestures and limit their applicability. To address these challenges, we present Persona-Gestor, a novel end-to-end generative model designed to generate highly personalized 3D full-body gestures solely relying on raw speech audio. The model combines a fuzzy feature extractor and a non-autoregressive Adaptive Layer Normalization (AdaLN) transformer diffusion architecture (DiTs-based). The fuzzy feature extractor harnesses a fuzzy inference strategy that automatically infers implicit, continuous fuzzy features. These fuzzy features, represented as a unified latent feature, are fed into the AdaLN transformer. The AdaLN transformer introduces a conditional mechanism that applies a uniform function across all tokens, thereby effectively modeling the correlation between the fuzzy features and the gesture sequence. This module ensures a high level of gesture-speech synchronization while preserving naturalness. Finally, we employ the diffusion model to train and infer various gestures. Extensive subjective and objective evaluations on the Trinity, ZEGGS, and BEAT datasets confirm our model's superior performance to the current state-of-the-art approaches. Persona-Gestor improves the system's usability and generalization capabilities, setting a new benchmark in speech-driven gesture synthesis and broadening the horizon for virtual human technology.
Fan Zhang 0105, Zhaohan Wang, Xin Lyu 0004, Mengjian Li, Weidong Geng, Naye Ji, Fuxing Gao, Hao Wu 0141, Shunman Li
IEEE Trans. Vis. Comput. Graph.6
2022 Enhanced 3D Shape Reconstruction With Knowledge Graph of Category Concept
abstract
Reconstructing three-dimensional (3D) objects from images has attracted increasing attention due to its wide applications in computer vision and robotic tasks. Despite the promising progress of recent deep learning–based approaches, which directly reconstruct the full 3D shape without considering the conceptual knowledge of the object categories, existing models have limited usage and usually create unrealistic shapes. 3D objects have multiple forms of representation, such as 3D volume, conceptual knowledge, and so on. In this work, we show that the conceptual knowledge for a category of objects, which represents objects as prototype volumes and is structured by graph, can enhance the 3D reconstruction pipeline. We propose a novel multimodal framework that explicitly combines graph-based conceptual knowledge with deep neural networks for 3D shape reconstruction from a single RGB image. Our approach represents conceptual knowledge of a specific category as a structure-based knowledge graph. Specifically, conceptual knowledge acts as visual priors and spatial relationships to assist the 3D reconstruction framework to create realistic 3D shapes with enhanced details. Our 3D reconstruction framework takes an image as input. It first predicts the conceptual knowledge of the object in the image, then generates a 3D object based on the input image and the predicted conceptual knowledge. The generated 3D object satisfies the following requirements: (1) it is consistent with the predicted graph in concept, and (2) consistent with the input image in geometry. Extensive experiments on public datasets (i.e., ShapeNet, Pix3D, and Pascal3D+) with 13 object categories show that (1) our method outperforms the state-of-the-art methods, (2) our prototype volume-based conceptual knowledge representation is more effective, and (3) our pipeline-agnostic approach can enhance the reconstruction quality of various 3D shape reconstruction pipelines.
Guofei Sun, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventions
abstract
3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a step towards training robust models for cross-domain pose estimation task, which brings together ideas from causal representation learning and generative adversarial networks. Specifically, this paper introduces a novel framework for causal representation learning which explicitly exploits the causal structure of the task. We consider changing domain as interventions on images under the data-generation process and steer the generative model to produce counterfactual features. This help the model learn transferable and causal relations across different domains. Our framework is able to learn with various types of unlabeled datasets. We demonstrate the efficacy of our proposed method on both human and hand pose estimation task. The experiment results show the proposed approach achieves state-of-the-art performance on most datasets for both domain adaptation and domain generalization settings.
Xiheng Zhang, Yongkang Wong, Juwei Lu, Mohan Kankanhalli, Weidong Geng
ICCV7
2021 DeepDance: Music-to-Dance Motion Choreography With Adversarial Learning
abstract
The creation of improvised dancing choreographies is an important research field of cross-modal analysis. A key point of this task is how to effectively create and correlate music and dance with a probabilistic one-to-many mapping, which is essential to create realistic dances of various genres. To address this issue, we propose a GAN-based cross-modal association framework, DeepDance, which correlates two different modalities (dance motion and music) together, aiming at creating the desired dance sequence in terms of the input music. Its generator is to predictively produce the dance movements best-fit to current music piece by learning from examples. In another hand, its discriminator acts as an external evaluation from the audience and judges the whole performance. The generated dance movements and the corresponding input music are considered to be well-matched if the discriminator cannot distinguish the generated movements from the training samples according to the estimated probability. By adding motion consistency constraints in our loss function, the proposed framework is able to create long realistic dance sequences. To alleviate the problem of expensive and inefficient data collection, we propose an effective approach to create a large-scale dataset, YouTube-Dance3D, from open data source. Extensive experiments on currently available music-dance datasets and our YouTube-Dance3D dataset demonstrate that our approach effectively captures the correlation between music and dance and can be used to choreograph appropriate dance sequences.
Guofei Sun, Yongkang Wong, Zhiyong Cheng 0001, Mohan Kankanhalli, Weidong Geng
IEEE Trans. Multim.5
2019 Unsupervised Domain Adaptation for 3D Human Pose Estimation
abstract
Training an accurate 3D human pose estimator often requires a large amount of 3D ground-truth data which is inefficient and costly to collect. Previous methods have either resorted to weakly supervised methods to reduce the demand of ground-truth data for training, or using synthetically-generated but photo-realistic samples to enlarge the training data pool. Nevertheless, the former methods mainly require either additional supervision, such as unpaired 3D ground-truth data, or the camera parameters in multiview settings. On the other hand, the latter methods require accurately textured models, illumination configurations and background which need careful engineering. To address these problems, we propose a domain adaptation framework with unsupervised knowledge transfer, which aims at leveraging the knowledge in multi-modality data of the easy-to-get synthetic depth datasets to better train a pose estimator on the real-world datasets. Specifically, the framework first trains two pose estimators on synthetically-generated depth images and human body segmentation masks with full supervision, while jointly learning a human body segmentation module from the predicted 2D poses. Subsequently, the learned pose estimator and the segmentation module are applied to the real-world dataset to unsupervisedly learn a new RGB image based 2D/3D human pose estimator. Here, the knowledge encoded in the supervised learning modules are used to regularize a pose estimator without ground-truth annotations. Comprehensive experiments demonstrate significant improvements over weakly supervised methods when no ground-truth annotations are available. Further experiments with ground-truth annotations show that the proposed framework can outperform state-of-the-art fully supervised methods. In addition, we conducted ablation studies to examine the impact of each loss term, as well as with different amount of supervisions signal.
Xiheng Zhang, Yongkang Wong, Mohan Kankanhalli, Weidong Geng
ACM Multimedia4
2019 A multi-stream convolutional neural network for sEMG-based gesture recognition in muscle-computer interface
Yongkang Wong, Yu Du 0016, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng
Pattern Recognit. Lett.6
2018 Fine-Grained Grocery Product Recognition by One-Shot Learning
abstract
Fine-grained grocery product recognition via camera is a challenging task to identify the visually similar products with subtle differences by using single-shot training examples. To address this issue? we present a novel hybrid classification approach that combines feature-based matching and one-shot deep learning with a coarse-to-fine strategy. The candidate regions of product instances are first detected and coarsely labeled by recurring features in product images without any training. Then, attention maps are generated to guide the classifier to focus on fine discriminative details by magnifying the influences of the features in the candidate regions of interest (ROI) and suppressing the interferences of the features outside, improving the accuracy of fine-grained grocery products recognition effectively. Our framework also performs a good adaptability which allows existing classifier to be refined without retraining for new coming product classes. As an additional contribution, we collect a new grocery product database with 102 classes from 2 stores. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods.
Weidong Geng, Feilin Han, Jiangke Lin, Liuyi Zhu, Jieming Bai, Suzhen Wang 0001, Zhangjiong Lai
ACM Multimedia1
2018 Learning to sketch human facial portraits using personal styles by case-based reasoning
Bingwen Jin, Songhua Xu, Weidong Geng
Multim. Tools Appl.3
2018 Cross-objects user interfaces for video interaction in virtual reality museum context
Lingyun Sun, Yunzhan Zhou, Preben Hansen, Weidong Geng
Multim. Tools Appl.4
2017 Generation of view-dependent textures for an inaccurate model
abstract
Existing methods of texture generation from registered images only work well on accurate models. For inaccurate models, texture drifts may occur. In this paper, we propose a view-dependent seamless texture generation method for inaccurate models. Under a specific viewpoint, this method first assigns each mesh face with a label associated with a registered image to generate a primitive texture. Different from previous methods, our label assignment process is dependent on the current view direction to reduce projective displacements of texture due to imprecision of the geometry. Then, a gradient-domain editing method is used to eliminate seams between image segments. If more than one observation views are given, the seam-levelling method further ensures color consistency between views. Experiments show that our method endows more sense of reality to inaccurate models and succeeds in maintaining temporal color consistency throughout a pre-defined sequence of viewpoints.
Zhen Wang 0003, Weidong Geng
ICIS2
2017 Semi-Supervised Learning for Surface EMG-based Gesture Recognition
abstract
Conventionally, gesture recognition based on non-intrusive muscle-computer interfaces required a strongly-supervised learning algorithm and a large amount of labeled training signals of surface electromyography (sEMG). In this work, we show that temporal relationship of sEMG signals and data glove provides implicit supervisory signal for learning the gesture recognition model. To demonstrate this, we present a semi-supervised learning framework with a novel Siamese architecture for sEMG-based gesture recognition. Specifically, we employ auxiliary tasks to learn visual representation; predicting the temporal order of two consecutive sEMG frames; and, optionally, predicting the statistics of 3D hand pose with a sEMG frame. Experiments on the NinaPro, CapgMyo and csl-hdemg datasets validate the efficacy of our proposed approach, especially when the labeled samples are very scarce.
Yu Du 0016, Yongkang Wong, Wenguang Jin, Yu Hu 0005, Mohan Kankanhalli, Weidong Geng
IJCAI7
2017 Image mosaicking for oversized documents with a multi-camera rig
abstract
We provide a method to obtain high resolution images of oversized documents by mosaicking their images captured by a rig equipped with multiple consumer-level cameras. We only need to calibrate the system once, and then it gives accurate full-size document mosaicking results repeatedly and robustly afterwards. During calibration, we first determine global image registration information based on homography from camera calibration results; then we refine the alignment information locally with smoothed piecewise affine transformations based on feature matches. The resulting image transformation and warping information are stored for composition of new document images. We evaluate our method against the state-of-the-art image stitching algorithms and software to demonstrate its advantages on suppressing seams and artifacts.
Zhen Wang 0003, Feilin Han, Weidong Geng
SERA3
2016 Marker-Less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps
Yu Du 0016, Yongkang Wong, Feilin Han, Yilin Gui, Zhen Wang 0003, Mohan Kankanhalli, Weidong Geng
ECCV (4)8
2015 Correspondence specification learned from master frames for automatic inbetweening
Bingwen Jin, Weidong Geng
Multim. Tools Appl.2
2014 Accurate Intrinsic Calibration of Depth Camera with Cuboids
Bingwen Jin, Weidong Geng
ECCV (5)3
2014 A hybrid approach to animating the murals with Dunhuang style
abstract
In order to animate the valuable murals of Dunhuang Mogao Grottoes, we propose a hybrid approach to creating the animation with the artistic style in the murals. Its key point is the fusion of 2D and 3D animation assets, for which a hybrid model is constructed from a 2.5D model, a 3D model, and registration information. The 2.5D model, created from 2D multi-view drawings, is composed of 2.5D strokes. For each 2.5D stroke, we let the user draw corresponding strokes on the surface of the 3D model in multiple views. Then the method automatically generates registration information, which enables 3D animation assets to animate the 2.5D model. At last, the animated line drawings are produced from 2.5D and 3D models respectively and blended under the control of per-stroke weights. The user can manually modify the weights to get the desired animation style.
Bingwen Jin, Linglong Feng, Huaqing Luo, Weidong Geng
MMSP5
2013 An HOG-CT human detector with histogram-based search
Jianhao Ding, Yigang Wang, Weidong Geng
Multim. Tools Appl.3
2012 A transformer-based filtering technique to lower LC-oscillator phase noise
abstract
Phase noise of oscillators can be improved through noise filtering. This paper presents a transformer-based harmonic filtering technique to improve the phase noise performance of the fully differential CMOS LC oscillators. The primary and secondary of the transformer are placed at the sources of NMOS and PMOS differential pairs, respectively. Due to the extra degree of freedom introduced by the transformer, the filter can resonate at two different frequencies, which are designed to be the second and fourth harmonics. The analysis of the filtering method is presented and principles of operation are discussed in details. The performance is compared with other filtering techniques, and the simulation results show improved FoM compared to published literature.
Qing Jin, Kaiyuan Yang 0001, Chunyuan Zhou, Lei Zhang 0033, Yan Wang 0023, Zhiping Yu, Weidong Geng
ISCAS8
2012 Machine Learning Approach for Gesture Recognition Based on Automatic Feature Selection
Xiubo Liang, Franck Multon, Weidong Geng
MIG3
2012 Example-Based Automatic Music-Driven Conventional Dance Motion Synthesis
abstract
We introduce a novel method for synthesizing dance motions that follow the emotions and contents of a piece of music. Our method employs a learning-based approach to model the music to motion mapping relationship embodied in example dance motions along with those motions' accompanying background music. A key step in our method is to train a music to motion matching quality rating function through learning the music to motion mapping relationship exhibited in synchronized music and dance motion data, which were captured from professional human dance performance. To generate an optimal sequence of dance motion segments to match with a piece of music, we introduce a constraint-based dynamic programming procedure. This procedure considers both music to motion matching quality and visual smoothness of a resultant dance motion sequence. We also introduce a two-way evaluation strategy, coupled with a GPU-based implementation, through which we can execute the dynamic programming process in parallel, resulting in significant speedup. To evaluate the effectiveness of our method, we quantitatively compare the dance motions synthesized by our method with motion synthesis results by several peer methods using the motions captured from professional human dancers' performance as the gold standard. We also conducted several medium-scale user studies to explore how perceptually our dance motion synthesis method can outperform existing methods in synthesizing dance motions to match with a piece of music. These user studies produced very positive results on our music-driven dance motion synthesis experiments for several Asian dance genres, confirming the advantages of our method.
Rukun Fan, Songhua Xu, Weidong Geng
IEEE Trans. Vis. Comput. Graph.3
2011 Acquisition of time-varying 3D foot shapes from video
Weidong Geng, Yunhe Pan
Sci. China Inf. Sci.3
2011 Performance-driven animation of hand-drawn cartoon faces
abstract
We present a novel performance-driven approach to animating cartoon faces starting from pure 2D drawings. A 3D approximate facial model automatically built from front and side view master frames of character drawings is introduced to enable the animated cartoon faces to be viewed from angles different from that in the input video. The expressive mappings are built by artificial neural network (ANN) trained from the examples of the real face in the video and the cartoon facial drawings in the facial expression graph for a specific character. The learned mapping model makes the resultant facial animation to properly get the desired expressiveness, instead of a mere reproduction of the facial actions in the input video sequence. Furthermore, the lit sphere, capturing the lighting in the painting artwork of faces, is utilized to color the cartoon faces in terms of the 3D approximate facial model, reinforcing the hand-drawn appearance of the resulting facial animation. We made a series of comparative experiments to test the effectiveness of our method by recreating the facial expression in the commercial animation. The comparison results clearly demonstrate the superiority of our method not only in generating high quality cartoon-style facial expressions, but also in speeding up the animation production of cartoon faces. Copyright © 2011 John Wiley & Sons, Ltd.
Xiang Li 0013, Yangchun Ren, Weidong Geng
Comput. Animat. Virtual Worlds4
2010 Responsive Action Generation by Physically-Based Motion Retrieval and Adaptation
Xiubo Liang, Ludovic Hoyet, Weidong Geng, Franck Multon
MIG3
2010 Animating cartoon faces by multi-view drawings
abstract
Abstract In this paper, we present a novel framework for creating cartoon facial animation from multi‐view hand‐drawn sketches. The input sketches are first employed to construct a base mesh model by using a hybrid sketch‐based method. The model is then deformed for each key viewpoint, yielding a set of models that closely match the corresponding sketches. We introduce a view‐dependent facial expression space defined by the key viewpoints and the basic emotions to generate various facial expressions viewed from arbitrary angles. The output facial animation conforms to the input sketches and maintains frame‐to‐frame correspondence. We demonstrate the potential of our approach through an easy‐to‐use system, where the animating of cartoon faces is automated once the user accomplishes sketching and configuration. Copyright © 2010 John Wiley & Sons, Ltd.
Xiang Li 0013, Yangchun Ren, Weidong Geng
Comput. Animat. Virtual Worlds4
2009 Performance-driven motion choreographing with accelerometers
abstract
Abstract Live performance is an intuitive way to naturally draft the desired motion in the choreographer's mind. In this paper we present a novel approach to choreographing motions by live performance captured with degree of freedom (3‐DOF) accelerometers. The process begins by placing the accelerometers on the user's limbs according to the pre‐specified positions. The computer then recognizes the performed actions using Hidden Markov Model (HMM), which is pre‐trained by the acceleration data samples automatically generated from a pre‐segmented motion capture database. At last, the captured actions are further synthesized with motion retiming and exaggeration based on the acceleration signals from the accelerometers. This method can intuitively rapid‐prototype the choreographed motions for pre‐production of animation, the avatar control in virtual reality and game‐like scenarios, etc. The experimental results show that it can effectively recognize actions with spatial‐time variance, and is easy‐to‐use especially for a novice with little experience. Copyright © 2009 John Wiley & Sons, Ltd.
Xiubo Liang, Qilei Li, Weidong Geng
Comput. Animat. Virtual Worlds5
2008 Interactive Animation of Virtual Characters: Application to Virtual Kung-Fu Fighting
abstract
This paper aims at proposing a framework for animating virtual humans that can efficiently interact with real users in virtual reality (VR). If the user's order can be modeled as targets and commands, the system searches a database for the most convenient behavior. In order to avoid using a huge database that can deal with any kind of situation, we propose to associate this searching process to an adaptation module. Hence, even if the selected motion is not perfectly suited with the situation, it can be adapted in order to reach accurately the target specified by the user. This framework is illustrated with a kung-fu fighter example. Two people are involved in this example: the user and the supervisor. The user is displacing in the real environment while the position of his head is tracked in real-time thanks to reflective markers. The virtual opponent follows the displacement of the user to stay close to him. At any time, the supervisor can ask the virtual character to kick or punch the user. Our system automatically searches for the convenient motion in an average-size database (75 motions compared to hundreds of motions required for motion graphs) and adapts it to the current situation.
Nicolas Pronost, Franck Multon, Qilei Li, Weidong Geng, Richard Kulpa, Georges Dumont
CW4
2005 Mocap data editing via movement notations
abstract
In most motion editing approaches, the users often make changes on mocap data by specifying numerical parameters. However, it is not intuitive for a novice to edit motion sequences in motion planning tasks. In this paper, we present a method of notation-based motion editing, in which Labanotation, a well-developed notation language for human movements, is employed as the editing interface. It allows the user to specify the editing requirements via notation score, and the system will semi-automatically generate the desired motion data. The algorithmic steps of its pipeline and the core issues of converting Labanotation into motion sequences are discussed in detail. We show the results of our systems using martial arts motion as its testing data.
XiaoJie Shen, Qilei Li, Weidong Geng, Newman Lau
CAD/Graphics4
2005 Motion retrieval based on movement notation language
abstract
Abstract With the increased availability of motion capture data, the volume of motion library grows so large that it is difficult for animators to manually browse the dataset to search desired motions for reuse. To address this issue, we implement a framework, which allows the user to retrieve motions via Labanotation. For each motion clip in the library, we generate a corresponding Labanotation sequence as additional motion property. A similarity metric for Labanotation sequences is proposed and used to search the motions that have similar Laban descriptions. Our search algorithm is able to retrieve motion segments that only match part of the query Laban sequence. Then based on dynamic programming, these segments are stitched together to form a smooth output motion that is in an optimal sense of matching query Laban sequence. Experimental results demonstrate our method could effectively improve the utilization of motion data. Copyright © 2005 John Wiley & Sons, Ltd.
XiaoJie Shen, Qilei Li, Weidong Geng
Comput. Animat. Virtual Worlds4
2003 Reuse of Motion Capture Data in Animation: A Review
Weidong Geng, Gino Yu
ICCSA (3)1
2002 Embedding visual cognition in 3D reconstruction from multi-view engineering drawings
Weidong Geng, Jingbin Wang
Comput. Aided Des.1
2001 Technical Illustration Based on Human-Like Approach
abstract
Presents a human-like, non-photorealistic rendering approach. A typical process of how human engineers learn to paint a technical illustration is as follows. First, they are trained in how to paint separate primitives, such as cubes and spheres, and accordingly accumulate empirical drawing principles and skills during their continuous practice, and finally they can freely express complicated shapes by composing the related primitives' drawings together. We manage to mimic this human-like approach by embedding established illustration rules into primitives' lighting models and drawing algorithms, and implement it in an illustration system called RETOUCH.
Weidong Geng, Monika Fleischmann, Yunhe Pan
Computer Graphics International1
1996 An intelligent multi-blackboard CAD system
Yunhe Pan, Weidong Geng, Xin Tong 0001
Artif. Intell. Eng.2