VLDB 2026 Research / reviewers in the wild / expert
Taku Komura
dblp:97/6832
· DBLP profile ↗
149ranked-venue papers
11as first author
69since 2021 · last 2026
0000-0002-2729-5860ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 135 · 10 first-author · 62 since 2021Artificial intelligence and machine learning · 39 · 1 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prior-First, Condition-Second: Scalable and Controllable Hand Motion CompletionabstractSynthesizing hand motion that matches the full body motion and the semantic labels is a difficult task due to their high degrees of freedom and the lack of semantic labels. To cope with this issue, we propose a prior-first, condition-second framework for body-conditioned hand motion completion. Our framework first learns a generic body-hand kinematic prior from large-scale unstructured and unlabeled motion data, capturing the intrinsic coordination between global body dynamics and hand articulation. Semantic control is then introduced through lightweight adaptation on top of the frozen prior, avoiding the need to relearn kinematic structure for each control interface. Our framework centers on a streaming, autoregressive body-hand prior that generates coherent, kinematically consistent hand motion from body dynamics in real time, using structured kinematic modeling to maintain mechanical body-hand coupling. To enable practical controllability under limited supervision, we introduce semantically-layered adapters that inject conditioning signals at appropriate kinematic levels, supporting both self-supervised attribute control and weakly supervised text-driven control with only a few hours of labeled data. Extensive evaluations demonstrate that our framework improves kinematic plausibility, robustness, and controllability compared to end-to-end conditioned baselines, particularly in low-resource and cross-dataset settings. We further showcase real-time inference and an interactive authoring workflow, highlighting the applicability to production animation pipelines. Homepage: https://AIGAnimation.github.io/HandPrior/ Mingyi Shi, Xuelin Chen, Taku Komura |
Comput. Graph. Forum | 3 |
| 2026 | CHOICE: Coordinated Human-Object Interaction in Cluttered Environments for Pick-and-Place ActionsabstractAnimating human-scene interactions such as picking and placing a wide range of objects with different geometries is a challenging task, especially in a cluttered environment where interactions with complex articulated containers are involved. The main difficulty lies in the sparsity of the motion data compared to the wide variation of the objects and environments, as well as the poor availability of transition motions between different actions, increasing the complexity of the generalization to arbitrary conditions. To cope with this issue, we develop a system that tackles the interaction synthesis problem as a hierarchical goal-driven task. Firstly, we develop a bimanual scheduler that plans a set of keyframes for simultaneously controlling the two hands to efficiently achieve the pick-and-place task from an abstract goal signal such as the target object selected by the user. Next, we develop a neural implicit planner that generates hand trajectories to guide reaching and leaving motions across diverse object shapes/types and obstacle layouts. Finally, we propose a linear dynamic model for our DeepPhase controller that incorporates a Kalman filter to enable smooth transitions in the frequency domain, resulting in a more realistic and effective multi-objective control of the character. Our system can synthesize a rich variety of natural pick-and-place movements that adapt to different object geometries, container articulations, and scene layouts. Jintao Lu, Yuting Ye, Takaaki Shiratori, Sebastian Starke, Taku Komura |
ACM Trans. Graph. | 6 |
| 2026 | Efficient B-Spline Finite Elements for Cloth SimulationabstractWe present an efficient B-spline finite element method (FEM) for cloth simulation. While higher-order FEM has long promised higher accuracy, its adoption in cloth simulators has been limited by its larger computational costs while generating results with similar visual quality. Our contribution is a full algorithmic pipeline that makes cloth simulation using quadratic B-spline surfaces faster than standard linear FEM in practice while consistently improving accuracy and visual fidelity. Using quadratic B-spline basis functions, we obtain a globally C 1 -continuous displacement field that supports consistent discretization of both membrane and bending energies, effectively reducing locking artifacts and mesh dependence common to linear elements. To close the performance gap, we introduce a reduced integration scheme that separately optimizes quadrature rules for membrane and bending energies, an accelerated Hessian assembly procedure tailored to the spline structure, and an optimized linear solver based on partial factorization. Together, these optimizations make high-order, smooth cloth simulation competitive at scale, yielding an average 2× speedup over linear FEM in our tests. Extensive experiments demonstrate improved accuracy, wrinkle detail, and robustness, including contact-rich scenarios, relative to linear FEM and recent higher-order approaches. Our method enables realistic wrinkling dynamics across a wide range of material parameters and supports practical garment animation, providing a new promising spatial discretization for high-quality cloth simulation. Yuqi Meng, Yihao Shi, Kemeng Huang, Taku Komura, Yin Yang 0002, Minchen Li |
ACM Trans. Graph. | 6 |
| 2026 | Strips as Tokens: Artist Mesh Generation with Native UV SegmentationabstractRecent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token ordering strategies employed by existing methods typically fail to meet professional artist standards, where coordinate-based sorting yields inefficiently long sequences, and patch-based heuristics disrupt the continuous edge flow and structural regularity essential for high-quality modeling. To address these limitations, we propose Strips as Tokens ( SATO ), a novel framework with a token ordering strategy inspired by triangle strips. By constructing the sequence as a connected chain of faces that explicitly encodes UV boundaries, our method naturally preserves the organized edge flow and semantic layout characteristic of artist-created meshes. A key advantage of this formulation is its unified representation, enabling the same token sequence to be decoded into either a triangle or quadrilateral mesh. This flexibility facilitates joint training on both data types: large-scale triangle data provides fundamental structural priors, while high-quality quad data enhances the geometric regularity of the outputs. Extensive experiments demonstrate that SATO consistently outperforms prior methods in terms of geometric quality, structural coherence, and UV segmentation. Rui Xu 0016, Dafei Qin, Kaichun Qiao, Qiujie Dong, Huaijin Pi, Qixuan Zhang, Longwen Zhang, Lan Xu 0003, Jingyi Yu 0001, Wenping Wang 0001, Taku Komura |
ACM Trans. Graph. | 11 |
| 2026 | ComboStoc: Combinatorial Stochasticity for Diffusion Generative ModelsabstractIn this paper, we study an under-explored but important factor of diffusion generative models, i.e., the combinatorial complexity. Data samples are generally high-dimensional, and for various structured generation tasks, additional attributes are combined to associate with data samples. We show that the space spanned by the combination of dimensions and attributes can be insufficiently covered by existing training schemes of diffusion generative models, potentially limiting test time performance. We present a simple fix to this problem by constructing stochastic processes that fully exploit the combinatorial structures, hence the name ComboStoc. Using this simple strategy, we show that network training is significantly accelerated across diverse data modalities, including images and 3D structured shapes. Moreover, ComboStoc enables a new way of test time generation which uses asynchronous time steps for different dimensions and attributes, thus allowing for varying degrees of control over them. Our code is available at: https://github.com/Xrvitd/ComboStoc. Rui Xu 0016, Jiepeng Wang 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Shi-Qing Xin, Changhe Tu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 8 |
| 2026 | SAND: Spatially Adaptive Network Depth for Fast Sampling of Neural Implicit SurfacesabstractImplicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of network evaluations. We observe that implicit representations require progressively lower accuracy as query points move farther from the target surface, and that even within the same iso-surface, representation difficulty varies spatially with local geometric complexity. However, conventional neural implicit models evaluate all query points with the same network depth and computational cost, ignoring this spatial variation and thereby incurring substantial computational waste. Motivated by this observation, we propose an efficient neural implicit geometry representation framework with spatially adaptive network depth (SAND). SAND leverages a volumetric network-depth map together with a tailed multi-layer perceptron (T-MLP) to model implicit representation. The volumetric depth map records, for each spatial region, the network depth required to achieve sufficient accuracy, while the T-MLP is a modified MLP designed to learn implicit functions such as signed distance functions, where an output branch, referred to as a tail, is attached to each hidden layer. This design allows network evaluation to terminate adaptively without traversing the full network and directs computational resources to geometrically important and complex regions, improving efficiency while preserving high-fidelity representations. Extensive experimental results demonstrate that our approach can significantly improve the inference-time query speed of implicit neural representations. Chuanxiang Yang, Junhui Hou, Yuan Liu 0025, Guangshun Wei, Taku Komura, Yuanfeng Zhou, Wenping Wang 0001 |
ACM Trans. Graph. | 6 |
| 2026 | SDRS: Shape-Differentiable Robot SimulatorabstractRobot simulators are indispensable tools across many fields, and recent research has significantly improved their functionality by incorporating additional gradient information. However, existing differentiable robot simulators suffer from non-differentiable singularities, when robots undergo substantial shape changes. To address this, we present the Shape-Differentiable Robot Simulator (SDRS), designed to be differentiable under significant robot shape changes. The core innovation of SDRS lies in its representation of robot shapes using a set of convex polyhedrons. This approach allows us to generalize smooth, penalty-based contact mechanics for interactions between any pair of convex polyhedrons. Using the separating hyperplane theorem, SDRS introduces a separating plane for each pair of contacting convex polyhedrons. This separating plane functions as a zero-mass auxiliary entity, with its state determined by the principle of least action. This setup ensures global differentiability, even as robot shapes undergo significant geometric and topological changes. To demonstrate the practical value of SDRS, we provide examples of robot co-design scenarios, where both robot shapes and control movements are optimized simultaneously. Xiaohan Ye, Xifeng Gao, Kui Wu 0003, Zherong Pan, Taku Komura |
IEEE Trans. Robotics | 5 |
| 2026 | SVGS: Enhancing Gaussian Splatting Using Primitives With Spatially Varying ColorsabstractGaussian Splatting demonstrates impressive results in multi-view reconstruction based on Gaussian explicit representations. However, the current Gaussian primitives only have a single view-dependent color and an opacity to represent the appearance and geometry of the scene, resulting in a non-compact representation. In this paper, we introduce a new method called SVGS (Spatially Varying Gaussian Splatting) that utilizes spatially varying colors and opacity in a single Gaussian primitive to improve its representation ability. We have implemented bilinear interpolation, movable kernels, and tiny neural networks as spatially varying functions. SVGS employs 2D Gaussian surfels as primitives, which significantly enhances novel-view synthesis while maintaining high-quality geometric reconstruction. This approach is particularly effective in practical applications, as scenes combining complex textures with relatively simple geometry occur frequently in real-world environments. Quantitative and qualitative experimental results demonstrate that all three functions outperform the baseline, with the best movable kernels achieving superior novel view synthesis performance on multiple datasets, highlighting the strong potential of spatially varying functions. Rui Xu 0016, Wenyue Chen, Jiepeng Wang 0001, Yuan Liu 0025, Peng Wang 0099, Cheng Lin 0001, Shi-Qing Xin, Xin Li 0003, Wenping Wang 0001, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2025 | TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task TokenizationabstractSynthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate controllers, each specialized for a specific interaction task. This significantly hinders the ability to tackle a wide variety of challenging HSI tasks that require the integration of multiple skills, e.g. sitting down while carrying an object (see Fig. 1). To address this issue, we present TokenHSI, a single, unified transformer-based policy capable of multi-skill unification and flexible adaptation. The key insight is to model the humanoid proprioception as a separate shared token and combine it with distinct task tokens via a masking mechanism. Such a unified policy enables effective knowledge sharing across skills, thereby facilitating the multi-task training. Moreover, our policy architecture supports variable length inputs, enabling flexible adaptation of learned skills to new scenarios. By training additional task tokenizers, we can not only modify the geometries of interaction targets but also coordinate multiple skills to address complex tasks. The experiments demonstrate that our approach can significantly improve versatility, adaptability, and extensibility in various HSI tasks. Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang 0009, Buzhen Huang, Bo Dai 0002, Taku Komura, Jingbo Wang 0003 |
CVPR | 7 |
| 2025 | Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations
Ruoxi Guo, Huaijin Pi, Zehong Shen, Qing Shuai, Zechen Hu, Zhumei Wang, Yajiao Dong, Ruizhen Hu, Taku Komura, Sida Peng, Xiaowei Zhou 0001 |
ICCV | 9 |
| 2025 | SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Wenjia Wang 0009, Liang Pan, Zhiyang Dou, Jidong Mei, Zhouyingcheng Liao, Yuke Lou, Yifan Wu 0039, Lei Yang 0045, Jingbo Wang 0003, Taku Komura |
ICCV | 10 |
| 2025 | DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageabstractReconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, complex deformations, and the ambiguity of the single-view setting. The previous state-of-the-art, Decaf, employs a global fitting optimization guided by contact and deformation estimation networks trained on studio-collected data with 3D annotations. However, Decaf suffers from a time-consuming optimization process and limited generalization capability due to its reliance on 3D annotations of hand-face interaction data. To address these issues, we present DICE, the first end-to-end method for Deformation-aware hand-face Interaction reCovEry from a single image. DICE estimates the poses of hands and faces, contacts, and deformations simultaneously using a Transformer-based architecture. It features disentangling the regression of local deformation fields and global mesh vertex locations into two network branches, enhancing deformation and contact estimation for precise and robust hand-face mesh recovery. To improve generalizability, we propose a weakly-supervised training approach that augments the training set using in-the-wild images without 3D ground-truth annotations, employing the depths of 2D keypoints estimated by off-the-shelf models and adversarial priors of poses for supervision. Our experiments demonstrate that DICE achieves state-of-the-art performance on a standard benchmark and in-the- wild data in terms of accuracy and physical plausibility. Additionally, our method operates at an interactive rate (20 fps) on an Nvidia 4090 GPU, whereas Decaf requires more than 15 seconds for a single image. The code will be available at: https://github.com/Qingxuan-Wu/DICE. Qingxuan Wu, Zhiyang Dou, Sirui Xu 0002, Soshi Shimada, Chen Wang 0049, Zhengming Yu, Yuan Liu 0025, Cheng Lin 0001, Zeyu Cao, Taku Komura, Vladislav Golyanik, Christian Theobalt, Wenping Wang 0001, Lingjie Liu |
ICLR | 10 |
| 2025 | CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated ObjectsabstractSynthesizing whole-body manipulation of articulated objects, including body motion, hand motion, and object motion, is a critical yet challenging task with broad applications in virtual humans and robotics.
The core challenges are twofold.
First, achieving realistic whole-body motion requires tight coordination between the hands and the rest of the body, as their movements are interdependent during manipulation.
Second, articulated object manipulation typically involves high degrees of freedom and demands higher precision, often requiring the fingers to be placed at specific regions to actuate movable parts.
To address these challenges, we propose a novel coordinated diffusion noise optimization framework.
Specifically, we perform noise-space optimization over three specialized diffusion models for the body, left hand, and right hand, each trained on its own motion dataset to improve generalization.
Coordination naturally emerges through gradient flow along the human kinematic chain, allowing the global body posture to adapt in response to hand motion objectives with high fidelity.
To further enhance precision in hand-object interaction, we adopt a unified representation based on basis point sets (BPS), where end-effector positions are encoded as distances to the same BPS used for object geometry.
This unified representation captures fine-grained spatial relationships between the hand and articulated object parts, and the resulting trajectories serve as targets to guide the optimization of diffusion noise, producing highly accurate interaction motion.
We conduct extensive experiments demonstrating that our method outperforms existing approaches in motion quality and physical plausibility, and enables various capabilities such as object pose control, simultaneous walking and manipulation, and whole-body generation from hand-only data.
The code will be released for reproducibility. Huaijin Pi, Zhi Cen, Zhiyang Dou, Taku Komura |
NeurIPS | 4 |
| 2025 | 🎧MOSPA: Human Motion Generation Driven by Spatial AudioabstractEnabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive "Spatial Audio-Driven Human Motion" (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human "MOtion generation driven by SPatial Audio," termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation.git Shuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan, Leo Ho, Jingbo Wang 0003, Yuan Liu 0025, Cheng Lin 0001, Yuexin Ma, Wenping Wang 0001, Taku Komura |
NeurIPS | 11 |
| 2025 | Motion2Motion: Cross-topology Motion Transfer with Sparse CorrespondenceabstractThis work studies the challenge of transfer animations between characters whose skeletal topologies differ substantially. While many techniques have advanced retargeting techniques in decades, transfer motions across diverse topologies remains less-explored. The primary obstacle lies in the inherent topological inconsistency between source and target skeletons, which restricts the establishment of straightforward one-to-one bone correspondences. Besides, the current lack of large-scale paired motion datasets spanning different topological structures severely constrains the development of data-driven approaches. To address these limitations, we introduce Motion2Motion, a novel, training-free framework. Simply yet effectively, Motion2Motion works with only one or a few example motions on the target skeleton, by accessing a sparse set of bone correspondences between the source and target skeletons. Through comprehensive qualitative and quantitative evaluations, we demonstrate that Motion2Motion achieves efficient and reliable performance in both similar-skeleton and cross-species skeleton transfer scenarios. The practical utility of our approach is further evidenced by its successful integration in downstream applications and user interfaces, highlighting its potential for industrial applications. Code and data are available at https://lhchen.top/Motion2Motion. Zixin Yin, Zhiyang Dou, Xin Chen 0040, Jingbo Wang 0003, Taku Komura, Lei Zhang 0001 |
SIGGRAPH Asia | 7 |
| 2025 | MATStruct: High-quality Medial Mesh Computation via Structure-aware Variational OptimizationabstractWe propose a novel optimization framework for computing the medial axis transform that simultaneously preserves the medial structure and ensures high medial mesh quality. The medial structure, consisting of interconnected sheets, seams, and junctions, provides a natural volumetric decomposition of a 3D shape. Our method introduces a structure-aware, particle-based optimization pipeline guided by the restricted power diagram (RPD), which partitions the input volume into convex cells whose dual encodes the connectivity of the medial mesh. Structure-awareness is enforced through a spherical quadratic error metric (SQEM) projection that constrains the movement of medial spheres, while a Gaussian kernel energy encourages an even spatial distribution. Compared to feature-preserving methods such as MATFP [Wang et al. 2022] and MATTopo [Wang et al. 2024b], our approach produces cleaner medial structures with significantly improved mesh quality. In contrast to voxel-based, point-cloud-based, and variational methods, our framework is the first to integrate structural awareness into the optimization process, yielding medial meshes with explicit structural decomposition, topological correctness, and geometric fidelity. Our code is available at our project website. Ningna Wang, Rui Xu 0016, Yibo Yin, Zichun Zhong, Taku Komura, Wenping Wang 0001, Xiaohu Guo |
SIGGRAPH Asia | 5 |
| 2025 | P2Seg: Distance query from point to segments
Jiantao Song, Rui Xu 0016, Wensong Wang, Shi-Qing Xin, Shuang-Min Chen, Jiaye Wang, Taku Komura, Wenping Wang 0001, Changhe Tu |
Comput. Aided Des. | 7 |
| 2025 | Explicit topology and connectivity constraints for 3D model repair
Jiantao Song, Wensong Wang, Rui Xu 0016, Wenlong Meng, Shuang-Min Chen, Shi-Qing Xin, Taku Komura, Changhe Tu, Wenping Wang 0001 |
Comput. Graph. | 7 |
| 2025 | A Hybrid Lagrangian-Eulerian Formulation of Thin-Shell FractureabstractAbstract The hybrid Lagrangian/Eulerian formulation of continuum shells is highly effective for producing challenging simulations of thin materials like cloth with bending resistance and frictional contact. However, existing formulations are restricted to materials that do not undergo tearing nor fracture due to the difficulties associated with incorporating strong discontinuities of field quantities like velocity via basis enrichment while maintaining continuity or regularity. We propose an extension of this formulation to simulate dynamic tearing and fracturing of thin shells using Kirchhoff–Love continuum theory. Damage, which manifests as cracks or tears, is propagated by tracking the evolution of a time‐dependent phase‐field in the co‐dimensional manifold, where a moving least‐squares (MLS) approximation then captures the strong discontinuities of interpolated field quantities near the crack. Our approach is capable of simulating challenging scenarios of this tearing and fracture, all‐the‐while harnessing the existing benefits of the hybrid Lagrangian/Eulerian formulation to expand the domain of possible effects. The method is also amenable to user‐guided control, which serves to influence the propagation of cracks or tears such that they follow prescribed paths during simulation. Linxu Fan, Floyd M. Chitalu, Taku Komura |
Comput. Graph. Forum | 3 |
| 2025 | NeurCross: A Neural Approach to Computing Cross Fields for Quad Mesh GenerationabstractQuadrilateral mesh generation plays a crucial role in numerical simulations within Computer-Aided Design and Engineering (CAD/E). Producing high-quality quadrangulation typically requires satisfying four key criteria. First, the quadrilateral mesh should closely align with principal curvature directions. Second, singular points should be strategically placed and effectively minimized. Third, the mesh should accurately conform to sharp feature edges. Lastly, quadrangulation results should exhibit robustness against noise and minor geometric variations. Existing methods generally involve first computing a regular cross field to represent quad element orientations across the surface, followed by extracting a quadrilateral mesh aligned closely with this cross field. A primary challenge with this approach is balancing the smoothness of the cross field with its alignment to pre-computed principal curvature directions, which are sensitive to small surface perturbations and often ill-defined in spherical or planar regions. To tackle this challenge, we propose NeurCross , a novel framework that simultaneously optimizes a cross field and a neural signed distance function (SDF), whose zero-level set serves as a proxy of the input shape. Our joint optimization is guided by three factors: faithful approximation of the optimized SDF surface to the input surface, alignment between the cross field and the principal curvature field derived from the SDF surface, and smoothness of the cross field. Acting as an intermediary, the neural SDF contributes in two essential ways. First, it provides an alternative, optimizable base surface exhibiting more regular principal curvature directions for guiding the cross field. Second, we leverage the Hessian matrix of the neural SDF to implicitly enforce cross field alignment with principal curvature directions, thus eliminating the need for explicit curvature extraction. Extensive experiments demonstrate that NeurCross outperforms the state-of-the-art methods in terms of singular point placement, robustness against surface noise and surface undulations, and alignment with principal curvature directions and sharp feature curves. Qiujie Dong, Huibiao Wen, Rui Xu 0016, Shuang-Min Chen, Jiaran Zhou, Shi-Qing Xin, Changhe Tu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 8 |
| 2025 | CrossGen: Learning and Generating Cross Fields for Quad MeshingabstractCross fields play a critical role in various geometry processing tasks, especially for quad mesh generation. Existing methods for cross field generation often struggle to balance computational efficiency with generation quality, using slow per-shape optimization. We introduce CrossGen , a novel framework that supports both feed-forward prediction and latent generative modeling of cross fields for quad meshing by unifying geometry and cross field representations within a joint latent space. Our method enables extremely fast computation of high-quality cross fields of general input shapes, typically within one second without per-shape optimization. Our method assumes a point-sampled surface, also called a point-cloud surface , as input, so we can accommodate various surface representations by a straightforward point sampling process. Using an auto-encoder network architecture, we encode input point-cloud surfaces into a sparse voxel grid with fine-grained latent spaces, which are decoded into both SDF-based surface geometry and cross fields (see the teaser figure). We also contribute a dataset of models with both high-quality signed distance fields (SDFs) representations and their corresponding cross fields, and use it to train our network. Once trained, the network is capable of computing a cross field of an input surface in a feed-forward manner, ensuring high geometric fidelity, noise resilience, and rapid inference. Furthermore, leveraging the same unified latent representation, we incorporate a diffusion model for computing cross fields of new shapes generated from partial input, such as sketches. To demonstrate its practical applications, we validate CrossGen on the quad mesh generation task for a large variety of surface shapes. Experimental results demonstrate that CrossGen generalizes well across diverse shapes and consistently yields high-fidelity cross fields, thus facilitating the generation of high-quality quad meshes. Qiujie Dong, Jiepeng Wang 0001, Rui Xu 0016, Cheng Lin 0001, Yuan Liu 0025, Shi-Qing Xin, Zichun Zhong, Xin Li 0003, Changhe Tu, Taku Komura, Leif Kobbelt, Scott Schaefer, Wenping Wang 0001 |
ACM Trans. Graph. | 10 |
| 2025 | KISSColor: Kinetic and Intuitive Stroke Stretching for Vector Drawing ColorizationabstractHand-drawn vector sketches often contain implied lines, imprecise intersections, and unintended gaps, making it challenging to identify closed regions for colorization. These challenges become more pronounced as the number of strokes increases. In this paper, we present KISSColor, a novel method for inferring users' intended closed regions. Specifically, we propose intuitive stroke stretching by extending open strokes along tangent isolines of winding-number fields, which provably form geometrically aligned closed regions. Extending all open strokes can lead to overly fragmented regions due to redundant intersections. While a Mixed Integer Programming (MIP) formulation helps reduce redundancy, it is computationally expensive. To improve efficiency, we introduce kinetic stroke stretching, which grows all strokes simultaneously and prioritizes early intersections using a kinetic data structure. This approach preserves stylistic ambiguity for lines requiring long extensions. Based on the growth results, redundant regions are suppressed to minimize fragmentation. We conduct extensive experiments demonstrating the effectiveness of KISSColor, which generates more intuitive partitions, especially for imprecise sketches (see teaser figure). Our code and data will be released upon publication. Yiming Dong, Hongxu Xin, Zhiyang Dou, Rui Xu 0016, Yuan Liu 0025, Shuang-Min Chen, Shi-Qing Xin, Changhe Tu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 9 |
| 2025 | CFC: Simulating Character-Fluid Coupling using a Two-Level World ModelabstractHumans possess the ability to master a wide range of motor skills, enabling them to quickly and flexibly adapt to the surrounding environment. Despite recent progress in replicating such versatile human motor skills, existing research often oversimplifies or inadequately captures the complex interplay between human body movements and highly dynamic environments, such as interactions with fluids. In this paper, we present a world model for Character-Fluid Coupling (CFC) for simulating human-fluid interactions via two-way coupling. We introduce a two-level world model which consists of a Physics-Informed Neural Network (PINN)-based model for fluid dynamics and a character world model capturing body dynamics under various external forces. This two-level world model adeptly predicts the dynamics of fluid and its influence on rigid bodies via force prediction, sidestepping the computational burden of fluid simulation and providing policy gradients for efficient policy training. Once trained, our system can control characters to complete high-level tasks while adaptively responding to environmental changes. We also present that the fluid initiates emergent behaviors of the characters, enhancing motion diversity and interactivity. Extensive experiments underscore the effectiveness of CFC, demonstrating its ability to produce high-quality, realistic human-fluid interaction animations. Zhiyang Dou, Xiaohan Ye, Lixing Fang, Yuan Liu 0025, Wenping Wang 0001, Chuang Gan 0001, Lingjie Liu, Taku Komura |
ACM Trans. Graph. | 10 |
| 2025 | StiffGIPC: Advancing GPU IPC for Stiff Affine-Deformable SimulationabstractIncremental Potential Contact (IPC) is a widely used, robust, and accurate method for simulating complex frictional contact behaviors. However, achieving high efficiency remains a major challenge, particularly as material stiffness increases, which leads to slower Preconditioned Conjugate Gradient (PCG) convergence, even with state-of-the-art preconditioners. In this article, we propose a fully GPU-optimized IPC simulation framework capable of handling materials across a wide range of stiffnesses, delivering consistent high performance and scalability with up to 10× speedup over state-of-the-art GPU IPC methods. Our framework introduces three key innovations: (1) A novel connectivity-enhanced Multilevel Additive Schwarz (MAS) preconditioner on the GPU, designed to efficiently capture both stiff and soft elastodynamics and improve PCG convergence at a reduced preconditioning cost. (2) A C 2 -continuous cubic energy with an analytic eigensystem for inexact strain limiting, enabling more parallel-friendly simulations of stiff membranes, such as cloth, without membrane locking. (3) For extremely stiff behaviors where elastic waves are barely visible, we employ affine body dynamics (ABD) with a hash-based two-level reduction strategy for fast Hessian assembly and efficient affine-deformable coupling. We conduct extensive performance analyses and benchmark studies to compare our framework against state-of-the-art methods and alternative design choices. Our system consistently delivers the fastest performance across soft, stiff, and hybrid simulation scenarios, even in cases with high resolution, large deformations, and high-speed impacts. Kemeng Huang, Huancheng Lin, Taku Komura, Minchen Li |
ACM Trans. Graph. | 4 |
| 2025 | Patch-Grid: An Efficient and Feature-Preserving Neural Implicit Surface RepresentationabstractNeural implicit representations are increasingly used to depict three-dimensional (3D) shapes owing to their inherent smoothness and compactness, contrasting with traditional discrete representations. Yet, the multilayer perceptron–based neural representation, because of its smooth nature, rounds sharp corners or edges, rendering it unsuitable for representing objects with sharp features like computer-aided design (CAD) models. Moreover, neural implicit representations need long training times to fit 3D shapes. While previous works address these issues separately, we present a unified neural implicit representation called Patch-Grid , which efficiently fits complex shapes, preserves sharp features delineating different patches, and can also represent surfaces with open boundaries and thin geometric features. Patch-Grid learns a signed distance field (SDF) to approximate an encompassing surface patch of the shape with a learnable patch feature volume. To form sharp edges and corners in a CAD model, Patch-Grid merges the learned SDFs via the constructive solid geometry (CSG) approach. Core to the merging process is a novel merge grid design that organizes different patch feature volumes in a common octree structure. This design choice ensures robust merging of multiple learned SDFs by confining the CSG operations to localized regions. Additionally, it drastically reduces the complexity of the CSG operations in each merging cell, allowing the proposed method to be trained in seconds to fit a complex shape at high fidelity. Experimental results demonstrate that the proposed Patch-Grid representation is capable of accurately reconstructing shapes with complex sharp features, open boundaries, and thin geometric elements, achieving state-of-the-art reconstruction quality with high computational efficiency within seconds. Guying Lin, Lei Yang 0048, Congyi Zhang 0001, Hao Pan 0001, Yuhan Ping, Guodong Wei, Taku Komura, John Keyser, Wenping Wang 0001 |
ACM Trans. Graph. | 7 |
| 2025 | StructRe: Rewriting for Structured Shape ModelingabstractMan-made 3D shapes are naturally organized in parts and hierarchies; such structures provide important constraints for shape reconstruction and generation. Modeling shape structures is difficult, because there can be multiple hierarchies for a given shape, causing ambiguity, and across different categories, the shape structures are correlated with semantics, limiting generalization. We present StructRe , a structure rewriting system, as a novel approach to structured shape modeling. Given a 3D object represented by points and components, StructRe can rewrite it upward into more concise structures, or downward into more detailed structures; by iterating the rewriting process, hierarchies are obtained. Such a localized rewriting process enables probabilistic modeling of ambiguous structures and robust generalization across object categories. We train StructRe on PartNet data and show its generalization to cross-category and multiple object hierarchies, and test its extension to ShapeNet. We also demonstrate the benefits of probabilistic and generalizable structure modeling for shape reconstruction, generation and editing tasks. Jiepeng Wang 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 5 |
| 2025 | On Optimal Sampling for Learning SDF Using MLPs Equipped With Positional EncodingabstractNeural implicit fields, such as the neural signed distance field (SDF) of a shape, have emerged as a powerful representation for many applications, e.g., encoding a 3D shape and performing collision detection. Typically, implicit fields are encoded by Multi-layer Perceptrons (MLP) with positional encoding (PE) to capture high-frequency geometric details. However, a notable side effect of such PE-equipped MLPs is the noisy artifacts present in the learned implicit fields. While increasing the sampling rate could in general mitigate these artifacts, in this paper we aim to explain this adverse phenomenon through the lens of Fourier analysis. We devise a tool to determine the appropriate sampling rate for learning an accurate neural implicit field without undesirable side effects. Specifically, we propose a simple yet effective method to estimate the intrinsic frequency of a given network with randomized weights based on the Fourier analysis of the network's responses. It is observed that a PE-equipped MLP has an intrinsic frequency much higher than the highest frequency component in the PE layer. Sampling against this intrinsic frequency following the Nyquist-Sannon sampling theorem allows us to determine an appropriate training sampling rate. We empirically show in the setting of SDF fitting that this recommended sampling rate is sufficient to secure accurate fitting results, while further increasing the sampling rate would not further noticeably reduce the fitting error. Training PE-equipped MLPs simply with our sampling strategy leads to performances superior to the existing methods. Guying Lin, Lei Yang 0048, Yuan Liu 0025, Congyi Zhang 0001, Junhui Hou, Xiaogang Jin 0001, Taku Komura, John Keyser, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | A Potential Field Method for Tooth Motion Planning in Orthodontic TreatmentabstractInvisible orthodontics, commonly known as clear alignment treatment, offers a more comfortable and aesthetically pleasing alternative in orthodontic care, attracting considerable attention in the dental community in recent years. It replaces conventional metal braces with a series of removable, and transparent aligners. Each aligner is crafted to facilitate a gradual adjustment of the teeth, ensuring progressive stages of dental correction. This necessitates the design for teeth motion. Here we present an automatic method and a system for generating collision-free teeth motion planning while avoiding gaps between adjacent teeth, which is unacceptable in clinical practice. To tackle this task, we formulate it as a constrained optimization problem and utilize the interior point method for its solution. We also developed an interactive system that enables dentists to easily visualize and edit the paths. Our method significantly speeds up the clear aligner planning process, creating the desired motion paths for a full set of teeth in under five minutes-a task that typically requires several hours of manual work. Our experiments and user studies confirm the effectiveness of this method in planning teeth movement, showcasing its potential to streamline orthodontic procedures. Yuexin Ma, Lei Yang 0048, Congyi Zhang 0001, Guangshun Wei, Runnan Chen, Min Gu 0003, Jia Pan 0001, Zhengbao Yang, Taku Komura, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2024 | SENC: Handling Self-collision in Neural Cloth Simulation
Zhouyingcheng Liao, Taku Komura |
ECCV (9) | 3 |
| 2024 | TLControl: Trajectory and Language Control for Human Motion Synthesis
Weilin Wan 0001, Zhiyang Dou, Taku Komura, Wenping Wang 0001, Dinesh Jayaraman, Lingjie Liu |
ECCV (37) | 3 |
| 2024 | Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models
Zhengming Yu, Zhiyang Dou, Xiaoxiao Long, Cheng Lin 0001, Zekun Li 0002, Yuan Liu 0025, Norman Müller, Taku Komura, Marc Habermann, Christian Theobalt, Xin Li 0003, Wenping Wang 0001 |
ECCV (39) | 8 |
| 2024 | EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang 0003, Wenjia Wang 0009, Yuan Liu 0025, Taku Komura, Wenping Wang 0001, Lingjie Liu |
ECCV (2) | 8 |
| 2024 | SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageabstractIn this paper, we present a novel diffusion model called SyncDreamer that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability to generate plausible novel views from a single-view image of an object. However, maintaining consistency in geometry and colors for the generated images remains a challenge. To address this issue, we propose a synchronized multiview diffusion model that models the joint probability distribution of multiview images, enabling the generation of multiview-consistent images in a single reverse process. SyncDreamer synchronizes the intermediate states of all the generated images at every step of the reverse process through a 3D-aware feature attention mechanism that correlates the corresponding features across different views. Experiments show that SyncDreamer generates images with high consistency across different views, thus making it well-suited for various 3D generation tasks such as novel-view-synthesis, text-to-3D, and image-to-3D. Project page: https://liuyuan-pal.github.io/SyncDreamer/. Yuan Liu 0025, Cheng Lin 0001, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, Wenping Wang 0001 |
ICLR | 6 |
| 2024 | ProLiF: Progressively-connected Light Field network for efficient view synthesis
Peng Wang 0099, Yuan Liu 0025, Guying Lin, Jiatao Gu, Lingjie Liu, Taku Komura, Wenping Wang 0001 |
Comput. Graph. | 6 |
| 2024 | Coverage Axis++: Efficient Inner Point Selection for 3D Shape SkeletonizationabstractAbstract We introduce Coverage Axis++, a novel and efficient approach to 3D shape skeletonization. The current state‐of‐the‐art approaches for this task often rely on the watertightness of the input [LWS*15; PWG*19; PWG*19] or suffer from substantial computational costs [DLX*22; CD23], thereby limiting their practicality. To address this challenge, Coverage Axis++ proposes a heuristic algorithm to select skeletal points, offering a high‐accuracy approximation of the Medial Axis Transform (MAT) while significantly mitigating computational intensity for various shape representations. We introduce a simple yet effective strategy that considers shape coverage, uniformity, and centrality to derive skeletal points. The selection procedure enforces consistency with the shape structure while favoring the dominant medial balls, which thus introduces a compact underlying shape representation in terms of MAT. As a result, Coverage Axis++ allows for skeletonization for various shape representations (e.g., water‐tight meshes, triangle soups, point clouds), specification of the number of skeletal points, few hyperparameters, and highly efficient computation with improved reconstruction accuracy. Extensive experiments across a wide range of 3D shapes validate the efficiency and effectiveness of Coverage Axis++. Our codes are available at https://github.com/Frank-ZY-Dou/Coverage_Axis . Zhiyang Dou, Rui Xu 0016, Cheng Lin 0001, Yuan Liu 0025, Xiaoxiao Long, Shi-Qing Xin, Taku Komura, Xiaoming Yuan 0001, Wenping Wang 0001 |
Comput. Graph. Forum | 8 |
| 2024 | GIPC: Fast and Stable Gauss-Newton Optimization of IPC Barrier EnergyabstractBarrier functions are crucial for maintaining an intersection- and inversion-free simulation trajectory but existing methods, which directly use distance can restrict implementation design and performance. We present an approach to rewriting the barrier function for arriving at an efficient and robust approximation of its Hessian. The key idea is to formulate a simplicial geometric measure of contact using mesh boundary elements, from which analytic eigensystems are derived and enhanced with filtering and stiffening terms that ensure robustness with respect to the convergence of a Project-Newton solver. A further advantage of our rewriting of the barrier function is that it naturally caters to the notorious case of nearly parallel edge-edge contacts for which we also present a novel analytic eigensystem. Our approach is thus well suited for standard second-order unconstrained optimization strategies for resolving contacts, minimizing nonlinear nonconvex functions where the Hessian may be indefinite. The efficiency of our eigensystems alone yields a 3× speedup over the standard Incremental Potential Contact (IPC) barrier formulation. We further apply our analytic proxy eigensystems to produce an entirely GPU-based implementation of IPC with significant further acceleration. Kemeng Huang, Floyd M. Chitalu, Huancheng Lin, Taku Komura |
ACM Trans. Graph. | 4 |
| 2024 | Online Neural Path Guiding with Normalized Anisotropic Spherical GaussiansabstractImportance sampling techniques significantly reduce variance in physically based rendering. In this article, we propose a novel online framework to learn the spatial-varying distribution of the full product of the rendering equation, with a single small neural network using stochastic ray samples. The learned distributions can be used to efficiently sample the full product of incident light. To accomplish this, we introduce a novel closed-form density model, called the Normalized Anisotropic Spherical Gaussian mixture, that can model a complex light field with a small number of parameters and that can be directly sampled. Our framework progressively renders and learns the distribution, without requiring any warm-up phases. With the compact and expressive representation of our density model, our framework can be implemented entirely on the GPU, allowing it to produce high-quality images with limited computational resources. The results show that our framework outperforms existing neural path guiding approaches and achieves comparable or even better performance than state-of-the-art online statistical path guiding techniques. Jiawei Huang 0005, Akito Iizuka, Hajime Tanaka 0001, Taku Komura, Yoshifumi Kitamura |
ACM Trans. Graph. | 4 |
| 2024 | Analytic rotation-invariant modelling of anisotropic finite elementsabstractAnisotropic hyperelastic distortion energies are used to solve many problems in fields like computer graphics and engineering with applications in shape analysis, deformation, design, mesh parameterization, biomechanics, and more. However, formulating a robust anisotropic energy that is low order and yet sufficiently non-linear remains a challenging problem for achieving the convergence promised by Newton-type methods in numerical optimization. In this article, we propose a novel analytic formulation of an anisotropic energy that is smooth everywhere, low order, rotationally invariant, and at least twice differentiable. At its core, our approach utilizes implicit rotation factorizations with invariants of the Cauchy-Green tensor that arises from the deformation gradient. The versatility and generality of our analysis is demonstrated through a variety of examples, where we also show that the constitutive law suggested by the anisotropic version of the well-known As-Rigid-As-Possible energy is the foundational parametric description of both passive and active elastic materials. The generality of our approach means that we can systematically derive the force and force-Jacobian expressions for use in implicit and quasistatic numerical optimization schemes, and we can also use our analysis to rewrite, simplify, and speed up several existing anisotropic and isotropic distortion energies with guaranteed inversion safety. Huancheng Lin, Floyd M. Chitalu, Taku Komura |
ACM Trans. Graph. | 3 |
| 2024 | Categorical Codebook Matching for Embodied Character ControllersabstractTranslating motions from a real user onto a virtual embodied avatar is a key challenge for character animation in the metaverse. In this work, we present a novel generative framework that enables mapping from a set of sparse sensor signals to a full body avatar motion in real-time while faithfully preserving the motion context of the user. In contrast to existing techniques that require training a motion prior and its mapping from control to motion separately, our framework is able to learn the motion manifold as well as how to sample from it at the same time in an end-to-end manner. To achieve that, we introduce a technique called codebook matching which matches the probability distribution between two categorical codebooks for the inputs and outputs for synthesizing the character motions. We demonstrate this technique can successfully handle ambiguity in motion generation and produce high quality character controllers from unstructured motion capture data. Our method is especially useful for interactive applications like virtual reality or video games where high accuracy and responsiveness are needed. Sebastian Starke, Paul Starke, Nicky He, Taku Komura, Yuting Ye |
ACM Trans. Graph. | 4 |
| 2024 | An Eulerian Vortex Method on Flow MapsabstractWe present an Eulerian vortex method based on the theory of flow maps to simulate the complex vortical motions of incompressible fluids. Central to our method is the novel incorporation of the flow-map transport equations for line elements , which, in combination with a bi-directional marching scheme for flow maps, enables the high-fidelity Eulerian advection of vorticity variables. The fundamental motivation is that, compared to impulse m , which has been recently bridged with flow maps to encouraging results, vorticity ω promises to be preferable for its numerical stability and physical interpretability. To realize the full potential of this novel formulation, we develop a new Poisson solving scheme for vorticity-to-velocity reconstruction that is both efficient and able to accurately handle the coupling near solid boundaries. We demonstrate the efficacy of our approach with a range of vortex simulation examples, including leapfrog vortices, vortex collisions, cavity flow, and the formation of complex vortical structures due to solid-fluid interactions. Yitong Deng, Molin Deng, Hong-Xing Yu, Junwei Zhou 0001, Duowen Chen 0003, Taku Komura, Jiajun Wu 0001, Bo Zhu 0002 |
ACM Trans. Graph. | 7 |
| 2024 | CBIL: Collective Behavior Imitation Learning for Fish from Real VideosabstractReproducing realistic collective behaviors presents a captivating yet formidable challenge. Traditional rule-based methods rely on hand-crafted principles, limiting motion diversity and realism in generated collective behaviors. Recent imitation learning methods learn from data but often require ground-truth motion trajectories and struggle with authenticity, especially in high-density groups with erratic movements. In this paper, we present a scalable approach, Collective Behavior Imitation Learning (CBIL), for learning fish schooling behavior directly from videos , without relying on captured motion trajectories. Our method first leverages Video Representation Learning, in which a Masked Video AutoEncoder (MVAE) extracts implicit states from video inputs in a self-supervised manner. The MVAE effectively maps 2D observations to implicit states that are compact and expressive for following the imitation learning stage. Then, we propose a novel adversarial imitation learning method to effectively capture complex movements of the schools of fish, enabling efficient imitation of the distribution of motion patterns measured in the latent space. It also incorporates bio-inspired rewards alongside priors to regularize and stabilize training. Once trained, CBIL can be used for various animation tasks with the learned collective motion priors. We further show its effectiveness across different species. Finally, we demonstrate the application of our system in detecting abnormal fish behavior from in-the-wild videos. Yifan Wu 0039, Zhiyang Dou, Yuko Ishiwaka, Shun Ogawa, Yuke Lou, Wenping Wang 0001, Lingjie Liu, Taku Komura |
ACM Trans. Graph. | 8 |
| 2024 | CWF: Consolidating Weak Features in High-quality Mesh SimplificationabstractIn mesh simplification, common requirements like accuracy, triangle quality, and feature alignment are often considered as a trade-off. Existing algorithms concentrate on just one or a few specific aspects of these requirements. For example, the well-known Quadric Error Metrics (QEM) approach [Garland and Heckbert 1997] prioritizes accuracy and can preserve strong feature lines/points as well, but falls short in ensuring high triangle quality and may degrade weak features that are not as distinctive as strong ones. In this paper, we propose a smooth functional that simultaneously considers all of these requirements. The functional comprises a normal anisotropy term and a Centroidal Voronoi Tessellation (CVT) [Du et al. 1999] energy term, with the variables being a set of movable points lying on the surface. The former inherits the spirit of QEM but operates in a continuous setting, while the latter encourages even point distribution, allowing various surface metrics. We further introduce a decaying weight to automatically balance the two terms. We selected 100 CAD models from the ABC dataset [Koch et al. 2019], along with 21 organic models, to compare the existing mesh simplification algorithms with ours. Experimental results reveal an important observation: the introduction of a decaying weight effectively reduces the conflict between the two terms and enables the alignment of weak features. This distinctive feature sets our approach apart from most existing mesh simplification methods and demonstrates significant potential in shape understanding. Please refer to the teaser figure for illustration. Rui Xu 0016, Longdu Liu, Ningna Wang, Shuang-Min Chen, Shi-Qing Xin, Xiaohu Guo, Zichun Zhong, Taku Komura, Wenping Wang 0001, Changhe Tu |
ACM Trans. Graph. | 8 |
| 2023 | Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB VideosabstractUnderstanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices. Yilin Wen 0001, Hao Pan 0001, Lei Yang 0048, Jia Pan 0001, Taku Komura, Wenping Wang 0001 |
CVPR | 5 |
| 2023 | F2-NeRF: Fast Neural Radiance Field Training with Free Camera TrajectoriesabstractThis paper presents a novel grid-based NeRF called F2- NeRF (Fast-Free-NeRF) for novel view synthesis, which enables arbitrary input camera trajectories and only costs a few minutes for training. Existing fast grid-based NeRF training frameworks, like Instant-NGP, Plenoxels, DVGO, or TensoRF, are mainly designed for bounded scenes and rely on space warping to handle unbounded scenes. Existing two widely-used space-warping methods are only designed for the forward-facing trajectory or the 360° object-centric trajectory but cannot process arbitrary trajectories. In this paper, we delve deep into the mechanism of space warping to handle unbounded scenes. Based on our analysis, we further propose a novel space-warping method called perspective warping, which allows us to handle arbitrary trajectories in the grid-based NeRF framework. Extensive experiments demonstrate that F2-NeRF is able to use the same perspective warping to render high-quality images on two standard datasets and a new free trajectory dataset collected by us. Project page: totoro97.github.io/projects/f2-nerf. Peng Wang 0099, Yuan Liu 0025, Zhaoxi Chen 0009, Lingjie Liu, Ziwei Liu 0002, Taku Komura, Christian Theobalt, Wenping Wang 0001 |
CVPR | 6 |
| 2023 | NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces with Arbitrary TopologiesabstractWe present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based reconstruction have achieved compelling results. However, these methods are limited to objects with closed surfaces since they adopt Signed Distance Function (SDF) as surface representation which requires the target shape to be divided into inside and outside. In this paper, we propose to represent surfaces as the Unsigned Distance Function (UDF) and develop a new volume rendering scheme to learn the neural UDF representation. Specifically, a new density function that correlates the property of UDF with the volume rendering scheme is introduced for robust optimization of the UDF fields. Experiments on the DTU and DeepFashion3D datasets show that our method not only enables high-quality reconstruction of non-closed shapes with complex typologies, but also achieves comparable performance to the SDF based methods on the reconstruction of closed surfaces. Visit our project page at https://www.xxlong.site/NeuralUDF. Xiaoxiao Long, Cheng Lin 0001, Lingjie Liu, Yuan Liu 0025, Peng Wang 0099, Christian Theobalt, Taku Komura, Wenping Wang 0001 |
CVPR | 7 |
| 2023 | TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerabstractIn this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caused by redundant tokens. We propose token reduction strategies based on two important aspects, i.e., the 3D geometry structure and 2D image feature, where we hierarchically recover the mesh geometry with priors from body structure and conduct token clustering to pass fewer but more discriminative image feature tokens to the Transformer. Our method massively reduces the number of tokens involved in high-complexity interactions in the Transformer. This leads to a significantly reduced computational cost while still achieving competitive or even higher accuracy in shape recovery. Extensive experiments across a wide range of benchmarks validate the superior effectiveness of the proposed method. We further demonstrate the generalizability of our method on hand mesh recovery. Visit our project page at https://frank-zy-dou.github.io/projects/Tore/index.html. Zhiyang Dou, Qingxuan Wu, Cheng Lin 0001, Zeyu Cao, Qiangqiang Wu, Weilin Wan 0001, Taku Komura, Wenping Wang 0001 |
ICCV | 7 |
| 2023 | PhaseMP: Robust 3D Pose Estimation via Phase-conditioned Human Motion PriorabstractWe present a novel motion prior, called PhaseMP, modeling a probability distribution on pose transitions conditioned by a frequency domain feature extracted from a periodic autoencoder. The phase feature further enforces the pose transitions to be unidirectional (i.e. no backward movement in time), from which more stable and natural motions can be generated. Specifically, our motion prior can be useful for accurately estimating 3D human motions in the presence of challenging input data, including long periods of spatial and temporal occlusion, as well as noisy sensor measurements. Through a comprehensive evaluation, we demonstrate the efficacy of our novel motion prior, showcasing its superiority over existing state-of-the-art methods by a significant margin across various applications, including video-to-motion and motion estimation from sparse sensor data, and etc. Mingyi Shi, Sebastian Starke, Yuting Ye, Taku Komura, Jungdam Won |
ICCV | 4 |
| 2023 | Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh ReconstructionabstractAs it is hard to calibrate single-view RGB images in the wild, existing 3D human mesh reconstruction (3DHMR) methods either use a constant large focal length or estimate one based on the background environment context, which can not tackle the problem of the torso, limb, hand or face distortion caused by perspective camera projection when the camera is close to the human body. The naive focal length assumptions can harm this task with the incorrectly formulated projection matrices. To solve this, we propose Zolly, the first 3DHMR method focusing on perspective-distorted images. Our approach begins with analysing the reason for perspective distortion, which we find is mainly caused by the relative location of the human body to the camera center. We propose a new camera model and a novel 2D representation, termed distortion image, which describes the 2D dense distortion scale of the human body. We then estimate the distance from distortion scale features rather than environment context features. Afterwards, We integrate the distortion feature with image features to reconstruct the body mesh. To formulate the correct projection matrix and locate the human body position, we simultaneously use perspective and weak-perspective projection loss. Since existing datasets could not handle this task, we propose the first synthetic dataset PDHuman and extend two real-world datasets tailored for this task, all containing perspective-distorted human images. Extensive experiments show that Zolly outperforms existing state-of-the-art methods on both perspective-distorted datasets and the standard benchmark (3DPW). Code and dataset will be released at https://wenjiawang0312.github.io/projects/zolly/. Wenjia Wang 0009, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Chunhua Shen, Lei Yang 0059, Taku Komura |
ICCV | 9 |
| 2023 | Surface Extraction from Neural Unsigned Distance FieldsabstractWe propose a method, named DualMesh-UDF, to extract a surface from unsigned distance functions (UDFs), encoded by neural networks, or neural UDFs. Neural UDFs are becoming increasingly popular for surface representation because of their versatility in presenting surfaces with arbitrary topologies, as opposed to the signed distance function that is limited to representing a closed surface. However, the applications of neural UDFs are hindered by the notorious difficulty in extracting the target surfaces they represent. Recent methods for surface extraction from a neural UDF suffer from significant geometric errors or topological artifacts due to two main difficulties: (1) A UDF does not exhibit sign changes; and (2) A neural UDF typically has substantial approximation errors.DualMesh-UDF addresses these two difficulties. Specifically, given a neural UDF encoding a target surface $\bar S$ to be recovered, we first estimate the tangent planes of $\bar S$ at a set of sample points close to $\bar S$. Next, we organize these sample points into local clusters, and for each local cluster, solve a linear least squares problem to determine a final surface point. These surface points are then connected to create the output mesh surface, which approximates the target surface. The robust estimation of the tangent planes of the target surface and the subsequent minimization problem constitute our core strategy, which contributes to the favorable performance of DualMesh-UDF over other competing methods. To efficiently implement this strategy, we employ an adaptive Octree. Within this framework, we estimate the location of a surface point in each of the octree cells identified as containing part of the target surface. Extensive experiments show that our method outperforms existing methods in terms of surface reconstruction quality while maintaining comparable computational efficiency. Congyi Zhang 0001, Guying Lin, Lei Yang 0048, Xin Li 0003, Taku Komura, Scott Schaefer, John Keyser, Wenping Wang 0001 |
ICCV | 5 |
| 2023 | C·ASE: Learning Conditional Adversarial Skill Embeddings for Physics-based CharactersabstractWe present C · ASE, an efficient and effective framework that learns Conditional Adversarial Skill Embeddings for physics-based characters. C · ASE enables the physically simulated character to learn a diverse repertoire of skills while providing controllability in the form of direct manipulation of the skills to be performed. This is achieved by dividing the heterogeneous skill motions into distinct subsets containing homogeneous samples for training a low-level conditional model to learn the conditional behavior distribution. The skill-conditioned imitation learning naturally offers explicit control over the character’s skills after training. The training course incorporates the focal skill sampling, skeletal residual forces, and element-wise feature masking to balance diverse skills of varying complexities, mitigate dynamics mismatch to master agile motions and capture more general behavior characteristics, respectively. Once trained, the conditional model can produce highly diverse and realistic skills, outperforming state-of-the-art models, and can be repurposed in various downstream tasks. In particular, the explicit skill control handle allows a high-level policy or a user to direct the character with desired skill specifications, which we demonstrate is advantageous for interactive character animation. Zhiyang Dou, Xuelin Chen, Qingnan Fan, Taku Komura, Wenping Wang 0001 |
SIGGRAPH Asia | 4 |
| 2023 | NeRO: Neural Geometry and BRDF Reconstruction of Reflective Objects from Multiview ImagesabstractWe present a neural rendering-based method called NeRO for reconstructing the geometry and the BRDF of reflective objects from multiview images captured in an unknown environment. Multiview reconstruction of reflective objects is extremely challenging because specular reflections are view-dependent and thus violate the multiview consistency, which is the cornerstone for most multiview reconstruction methods. Recent neural rendering techniques can model the interaction between environment lights and the object surfaces to fit the view-dependent reflections, thus making it possible to reconstruct reflective objects from multiview images. However, accurately modeling environment lights in the neural rendering is intractable, especially when the geometry is unknown. Most existing neural rendering methods, which can model environment lights, only consider direct lights and rely on object masks to reconstruct objects with weak specular reflections. Therefore, these methods fail to reconstruct reflective objects, especially when the object mask is not available and the object is illuminated by indirect lights. We propose a two-step approach to tackle this problem. First, by applying the split-sum approximation and the integrated directional encoding to approximate the shading effects of both direct and indirect lights, we are able to accurately reconstruct the geometry of reflective objects without any object masks. Then, with the object geometry fixed, we use more accurate sampling to recover the environment lights and the BRDF of the object. Extensive experiments demonstrate that our method is capable of accurately reconstructing the geometry and the BRDF of reflective objects from only posed RGB images without knowing the environment lights and the object masks. Codes and datasets are available at https://github.com/liuyuan-pal/NeRO. Yuan Liu 0025, Peng Wang 0099, Cheng Lin 0001, Xiaoxiao Long, Jiepeng Wang 0001, Lingjie Liu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 7 |
| 2023 | BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerabstractAutomatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the stochastic nature of the problem and the lack of a rich cross-modal dataset that is needed for training. In this paper, we propose a novel transformer-based framework for automatic 3D body gesture synthesis from speech. To learn the stochastic nature of the body gesture during speech, we propose a variational transformer to effectively model a probabilistic distribution over gestures, which can produce diverse gestures during inference. Furthermore, we introduce a mode positional embedding layer to capture the different motion speeds in different speaking modes. To cope with the scarcity of data, we design an intra-modal pre-training scheme that can learn the complex mapping between the speech and the 3D gesture from a limited amount of data. Our system is trained with either the Trinity speech-gesture dataset or the Talking With Hands 16.2M dataset. The results show that our system can produce more realistic, appropriate, and diverse body gestures compared to existing state-of-the-art approaches. Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost, Takaaki Shiratori, Junichi Yamagishi, Taku Komura |
ACM Trans. Graph. | 7 |
| 2022 | FaceFormer: Speech-Driven 3D Facial Animation with TransformersabstractSpeech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on learning phoneme-level features of short audio windows with limited context, occasionally resulting in inaccurate lip movements. To tackle this limitation, we propose a Transformer-based autoregressive model, Face-Former, which encodes the long-term audio context and autoregressively predicts a sequence of animated 3D face meshes. To cope with the data scarcity issue, we integrate the self-supervised pre-trained speech representations. Also, we devise two biased attention mechanisms well suited to this specific task, including the biased cross-modal multi-head (MH) attention and the biased causal MH self-attention with a periodic positional encoding strategy. The former effectively aligns the audio-motion modalities, whereas the latter offers abilities to generalize to longer audio sequences. Extensive experiments and a perceptual user study show that our approach outperforms the existing state-of-the-arts. The code and the video are available at: https://evelynfan.github.io/audio2face/ Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 0001, Taku Komura |
CVPR | 5 |
| 2022 | Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images
Yuan Liu 0025, Yilin Wen 0001, Sida Peng, Cheng Lin 0001, Xiaoxiao Long, Taku Komura, Wenping Wang 0001 |
ECCV (32) | 6 |
| 2022 | SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views
Xiaoxiao Long, Cheng Lin 0001, Peng Wang 0099, Taku Komura, Wenping Wang 0001 |
ECCV (32) | 4 |
| 2022 | NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors
Jiepeng Wang 0001, Peng Wang 0099, Xiaoxiao Long, Christian Theobalt, Taku Komura, Lingjie Liu, Wenping Wang 0001 |
ECCV (32) | 5 |
| 2022 | DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation
Yilin Wen 0001, Hao Pan 0001, Lei Yang 0048, Zheng Wang 0002, Taku Komura, Wenping Wang 0001 |
ECCV (9) | 6 |
| 2022 | Coverage Axis: Inner Point Selection for 3D Shape SkeletonizationabstractAbstract In this paper, we present a simple yet effective formulation called Coverage Axis for 3D shape skeletonization. Inspired by the set cover problem, our key idea is to cover all the surface points using as few inside medial balls as possible. This formulation inherently induces a compact and expressive approximation of the Medial Axis Transform (MAT) of a given shape. Different from previous methods that rely on local approximation error, our method allows a global consideration of the overall shape structure, leading to an efficient high‐level abstraction and superior robustness to noise. Another appealing aspect of our method is its capability to handle more generalized input such as point clouds and poor‐quality meshes. Extensive comparisons and evaluations demonstrate the remarkable effectiveness of our method for generating compact and expressive skeletal representation to approximate the MAT. Zhiyang Dou, Cheng Lin 0001, Rui Xu 0016, Lei Yang 0048, Shi-Qing Xin, Taku Komura, Wenping Wang 0001 |
Comput. Graph. Forum | 6 |
| 2022 | Simulating Brittle Fracture with Material PointsabstractLarge-scale topological changes play a key role in capturing the fine debris of fracturing virtual brittle material. Real-world, tough brittle fractures have dynamic branching behaviour but numerical simulation of this phenomena is notoriously challenging. In order to robustly capture these visual characteristics, we simulate brittle fracture by combining elastodynamic continuum mechanical models with rigid-body methods: A continuum damage mechanics (CDM) problem is solved, following rigid-body impact, to simulate crack propagation by tracking a damage field. We combine the result of this elastostatic continuum model with a novel technique to approximate cracks as a non-manifold mid-surface, which enables accurate and robust modelling of material fragment volumes to compliment fast-and-rigid shatter effects. For enhanced realism, we add fracture detail, incorporating particle damage-time to inform localised perturbation of the crack surface with artistic control. We evaluate our method with numerous examples and comparisons, showing that it produces a breadth of brittle material fracture effects and with low simulation resolution to require much less time compared to fully elastodynamic simulations. Linxu Fan, Floyd M. Chitalu, Taku Komura |
ACM Trans. Graph. | 3 |
| 2022 | Isotropic ARAP Energy Using Cauchy-Green InvariantsabstractIsotropic As-Rigid-As-Possible (ARAP) energy has been popular for shape editing, mesh parametrisation and soft-body simulation for almost two decades. However, a formulation using Cauchy-Green (CG) invariants has always been unclear, due to a rotation-polluted trace term that cannot be directly expressed using these invariants. We show how this incongruent trace term can be understood via an implicit relationship to the CG invariants. Our analysis reveals this relationship to be a polynomial where the roots equate to the trace term, and where the derivatives also give rise to closed-form expressions of the Hessian to guarantee positive semi-definiteness for a fast and concise Newton-type implicit time integration. A consequence of this analysis is a novel analytical formulation to compute rotations and singular values of deformation-gradient tensors without explicit/numerical factorization which is significant, resulting in up-to 3.5× speedup and benefits energy function evaluation for reducing solver time. We validate our energy formulation by experiments and comparison, demonstrating that our resulting eigendecomposition using the CG invariants is equivalent to existing ARAP formulations. We thus reveal isotropic ARAP energy to be a member of the "Cauchy-Green club", meaning that it can indeed be defined using CG invariants and therefore that the closed-form expressions of the resulting Hessian are shared with other energies written in their terms. Huancheng Lin, Floyd M. Chitalu, Taku Komura |
ACM Trans. Graph. | 3 |
| 2022 | DeepPhase: periodic autoencoders for learning motion phase manifoldsabstractLearning the spatial-temporal structure of body movements is a fundamental problem for character motion synthesis. In this work, we propose a novel neural network architecture called the Periodic Autoencoder that can learn periodic features from large unstructured motion datasets in an unsupervised manner. The character movements are decomposed into multiple latent channels that capture the non-linear periodicity of different body segments while progressing forward in time. Our method extracts a multi-dimensional phase space from full-body motion data, which effectively clusters animations and produces a manifold in which computed feature distances provide a better similarity measure than in the original motion space to achieve better temporal and spatial alignment. We demonstrate that the learned periodic embedding can significantly help to improve neural motion synthesis in a number of tasks, including diverse locomotion skills, style-based movements, dance motion synthesis from music, synthesis of dribbling motions in football, and motion query for matching poses within large animation databases. Sebastian Starke, Ian Mason, Taku Komura |
ACM Trans. Graph. | 3 |
| 2022 | Reconstruction of Dexterous 3D Motion Data From a Flexible Magnetic Sensor With Deep Learning and Structure-Aware FilteringabstractWe propose IM3D+, a novel approach to reconstructing 3D motion data from a flexible magnetic flux sensor array using deep learning and a structure-aware temporal bilateral filter. Computing the 3D configuration of markers (inductor-capacitor (LC) coils) from flux sensor data is difficult because the existing numerical approaches suffer from system noise, dead angles, the need for initialization, and limitations in the sensor array's layout. We solve these issues with deep neural networks to learn the regression from the simulation flux values to the LC coils' 3D configuration, which can be applied to the actual LC coils at any location and orientation within the capture volume. To cope with the influence of system noise and the dead-angle limitation caused by the characteristics of the hardware and sensing principle, we propose a structure-aware temporal bilateral filter for reconstructing motion sequences. Our method can track various movements, including fingers that manipulate objects, beetles that move inside a vivarium with leaves and soil, and the flow of opaque fluid. Since no power supply is needed for the lightweight wireless markers, our method can robustly track movements for a very long time, making it suitable for various types of observations whose tracking is difficult with existing motion-tracking systems. Furthermore, the flexibility of the flux sensor layout allows users to reconfigure it based on their own applications, thus making our approach suitable for a variety of virtual reality applications. Jiawei Huang 0005, Ryo Sugawara, Kinfung Chu, Taku Komura, Yoshifumi Kitamura |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Relationship-Based Point Cloud CompletionabstractWe propose a partial point cloud completion approach for scenes that are composed of multiple objects. We focus on pairwise scenes where two objects are in close proximity and are contextually related to each other, such as a chair tucked in a desk, a fruit in a basket, a hat on a hook and a flower in a vase. Different from existing point cloud completion methods, which mainly focus on single objects, we design a network that encodes not only the geometry of the individual shapes, but also the spatial relations between different objects. More specifically, we complete missing parts of the objects in a conditional manner, where the partial or completed point cloud of the other object is used as an additional input to help predict missing parts. Based on the idea of conditional completion, we further propose a two-path network, which is guided by a consistency loss between different sequences of completion. Our method can handle difficult cases where the objects heavily occlude each other. Also, it only requires a small set of training data to reconstruct the interaction area compared to existing completion approaches. We evaluate our method qualitatively and quantitatively via ablation studies and in comparison to the state-of-the-art point cloud completion methods. Xi Zhao 0002, Jinji Wu, Ruizhen Hu, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionabstractWe present a novel neural surface reconstruction method, called NeuS, for reconstructing objects and scenes with high fidelity from 2D image inputs. Existing neural surface reconstruction approaches, such as DVR [Niemeyer et al., 2020] and IDR [Yariv et al., 2020], require foreground mask as supervision, easily get trapped in local minima, and therefore struggle with the reconstruction of objects with severe self-occlusion or thin structures. Meanwhile, recent neural methods for novel view synthesis, such as NeRF [Mildenhall et al., 2020] and its variants, use volume rendering to produce a neural scene representation with robustness of optimization, even for highly complex objects. However, extracting high-quality surfaces from this learned implicit representation is difficult because there are not sufficient surface constraints in the representation. In NeuS, we propose to represent a surface as the zero-level set of a signed distance function (SDF) and develop a new volume rendering method to train a neural SDF representation. We observe that the conventional volume rendering method causes inherent geometric errors (i.e. bias) for surface reconstruction, and therefore propose a new formulation that is free of bias in the first order of approximation, thus leading to more accurate surface reconstruction even without the mask supervision. Experiments on the DTU dataset and the BlendedMVS dataset show that NeuS outperforms the state-of-the-arts in high-quality surface reconstruction, especially for objects and scenes with complex structures and self-occlusion. Peng Wang 0099, Lingjie Liu, Yuan Liu 0025, Christian Theobalt, Taku Komura, Wenping Wang 0001 |
NeurIPS | 5 |
| 2021 | Bas-relief modelling from enriched detail and geometry with deep normal transfer
Meili Wang 0001, Li Wang 0105, Tao Jiang 0020, Juncong Lin, Mingqiang Wei, Xiaosong Yang, Taku Komura, Jian J. Zhang 0001 |
Neurocomputing | 8 |
| 2021 | MotioNet: 3D Human Motion Reconstruction from Monocular Video with Skeleton ConsistencyabstractWe introduce MotioNet , a deep neural network that directly reconstructs the motion of a 3D human skeleton from a monocular video. While previous methods rely on either rigging or inverse kinematics (IK) to associate a consistent skeleton with temporally coherent joint rotations, our method is the first data-driven approach that directly outputs a kinematic skeleton, which is a complete, commonly used motion representation. At the crux of our approach lies a deep neural network with embedded kinematic priors, which decomposes sequences of 2D joint positions into two separate attributes: a single, symmetric skeleton encoded by bone lengths, and a sequence of 3D joint rotations associated with global root positions and foot contact labels. These attributes are fed into an integrated forward kinematics (FK) layer that outputs 3D positions, which are compared to a ground truth. In addition, an adversarial loss is applied to the velocities of the recovered rotations to ensure that they lie on the manifold of natural joint rotations. The key advantage of our approach is that it learns to infer natural joint rotations directly from the training data rather than assuming an underlying model, or inferring them from joint positions using a data-agnostic IK solver. We show that enforcing a single consistent skeleton along with temporally coherent joint rotations constrains the solution space, leading to a more robust handling of self-occlusions and depth ambiguities. Mingyi Shi, Kfir Aberman, Andreas Aristidou, Taku Komura, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2021 | Neural animation layering for synthesizing martial arts movementsabstractInteractively synthesizing novel combinations and variations of character movements from different motion skills is a key problem in computer animation. In this paper, we propose a deep learning framework to produce a large variety of martial arts movements in a controllable manner from raw motion capture data. Our method imitates animation layering using neural networks with the aim to overcome typical challenges when mixing, blending and editing movements from unaligned motion sources. The framework can synthesize novel movements from given reference motions and simple user controls, and generate unseen sequences of locomotion, punching, kicking, avoiding and combinations thereof, but also reconstruct signature motions of different fighters, as well as close-character interactions such as clinching and carrying by learning the spatial joint relationships. To achieve this goal, we adopt a modular framework which is composed of the motion generator and a set of different control modules. The motion generator functions as a motion manifold that projects novel mixed/edited trajectories to natural full-body motions, and synthesizes realistic transitions between different motions. The control modules are task dependent and can be developed and trained separately by engineers to include novel motion tasks, which greatly reduces network iteration time when working with large-scale datasets. Our modular framework provides a transparent control interface for animators that allows modifying or combining movements after network training, and enables iterative adding of control modules for different motion tasks and behaviors. Our system can be used for offline and online motion generation alike, and is relevant for real-time applications such as computer games. Sebastian Starke, Fabio Zinno, Taku Komura |
ACM Trans. Graph. | 4 |
| 2021 | ManipNet: neural manipulation synthesis with a hand-object spatial representationabstractNatural hand manipulations exhibit complex finger maneuvers adaptive to object shapes and the tasks at hand. Learning dexterous manipulation from data in a brute force way would require a prohibitive amount of examples to effectively cover the combinatorial space of 3D shapes and activities. In this paper, we propose a hand-object spatial representation that can achieve generalization from limited data. Our representation combines the global object shape as voxel occupancies with local geometric details as samples of closest distances. This representation is used by a neural network to regress finger motions from input trajectories of wrists and objects. Specifically, we provide the network with the current finger pose, past and future trajectories, and the spatial representations extracted from these trajectories. The network then predicts a new finger pose for the next frame as an autoregressive model. With a carefully chosen hand-centric coordinate system, we can handle single-handed and two-handed motions in a unified framework. Learning from a small number of primitive shapes and kitchenware objects, the network is able to synthesize a variety of finger gaits for grasping, in-hand manipulation, and bimanual object handling on a rich set of novel shapes and functional tasks. We also demonstrate a live demo of manipulating virtual objects in real-time using a simple physical prop. Our system is useful for offline animation or real-time applications forgiving to a small delay. Yuting Ye, Takaaki Shiratori, Taku Komura |
ACM Trans. Graph. | 4 |
| 2021 | Multi-agent reinforcement learning for character controlabstractAbstract Simultaneous control of multiple characters has been a research topic that has been extensively pursued for applications in computer games and computer animations, for applications such as crowd simulation, controlling two characters carrying objects or fighting with one another and controlling a team of characters playing collective sports. With the advance in deep learning and reinforcement learning, there is a growing interest in applying multi-agent reinforcement learning for intelligently controlling the characters to produce realistic movements. In this paper we will survey the state-of-the-art MARL techniques that are applicable for character control. We will then survey papers that make use of MARL for multi-character control and then discuss about the possible future directions of research. Levi Fussell, Taku Komura |
Vis. Comput. | 3 |
| 2020 | Learning 3D Global Human Motion Estimation from Unpaired, Disjoint Datasets
Julian Habekost, Takaaki Shiratori, Yuting Ye, Taku Komura |
BMVC | 4 |
| 2020 | Binary Ostensibly-Implicit Trees for Fast Collision DetectionabstractAbstract We present a simple, efficient and low‐memory technique, targeting fast construction of bounding volume hierarchies (BVH) for broad‐phase collision detection. To achieve this, we devise a novel representation of BVH trees in memory. We develop a mapping of the implicit index representation to compact memory locations, based on simple bit‐shifts, to then construct and evaluate bounding volume test trees (BVTT) during collision detection with real‐time performance. We model the topology of the BVH tree implicitly as binary encodings which allows us to determine the nodes missing from a complete binary tree using the binary representation of the number of missing nodes. The simplicity of our technique allows for fast hierarchy construction achieving over 6× speedup over the state‐of‐the‐art. Making use of these characteristics, we show that not only it is feasible to rebuild the BVH at every frame, but that using our technique, it is actually faster than refitting and more memory efficient. Floyd M. Chitalu, Christophe Dubach, Taku Komura |
Comput. Graph. Forum | 3 |
| 2020 | Displacement-Correlated XFEM for Simulating Brittle FractureabstractAbstract We present a remeshing‐free brittle fracture simulation method under the assumption of quasi‐static linear elastic fracture mechanics (LEFM). To achieve this, we devise two algorithms. First, we develop an approximate volumetric simulation, based on the extended Finite Element Method (XFEM), to initialize and propagate Lagrangian crack‐fronts. We model the geometry of fracture explicitly as a surface mesh, which allows us to generate high‐resolution crack surfaces that are decoupled from the resolution of the deformation mesh. Our second contribution is a mesh cutting algorithm, which produces fragments of the input mesh using the fracture surface. We do this by directly operating on the half‐edge data structures of two surface meshes, which enables us to cut general surface meshes including those of concave polyhedra and meshes with abutting concave polygons. Since we avoid triangulation for cutting, the connectivity of the resulting fragments is identical to the (uncut) input mesh except at edges introduced by the cut. We evaluate our simulation and cutting algorithms and show that they outperform state‐of‐the‐art approaches both qualitatively and quantitatively. Floyd M. Chitalu, Qinghai Miao, Kartic Subr, Taku Komura |
Comput. Graph. Forum | 4 |
| 2020 | Automatic spatial estimation of white matter hyperintensities evolution in brain MRI using disease evolution predictor deep neural networksabstractPrevious studies have indicated that white matter hyperintensities (WMH), the main radiological feature of small vessel disease, may evolve (i.e., shrink, grow) or stay stable over a period of time. Predicting these changes are challenging because it involves some unknown clinical risk factors that leads to a non-deterministic prediction task. In this study, we propose a deep learning model to predict the evolution of WMH from baseline to follow-up (i.e., 1-year later), namely "Disease Evolution Predictor" (DEP) model, which can be adjusted to become a non-deterministic model. The DEP model receives a baseline image as input and produces a map called "Disease Evolution Map" (DEM), which represents the evolution of WMH from baseline to follow-up. Two DEP models are proposed, namely DEP-UResNet and DEP-GAN, which are representatives of the supervised (i.e., need expert-generated manual labels to generate the output) and unsupervised (i.e., do not require manual labels produced by experts) deep learning algorithms respectively. To simulate the non-deterministic and unknown parameters involved in WMH evolution, we modulate a Gaussian noise array to the DEP model as auxiliary input. This forces the DEP model to imitate a wider spectrum of alternatives in the prediction results. The alternatives of using other types of auxiliary input instead, such as baseline WMH and stroke lesion loads are also proposed and tested. Based on our experiments, the fully supervised machine learning scheme DEP-UResNet regularly performed better than the DEP-GAN which works in principle without using any expert-generated label (i.e., unsupervised). However, a semi-supervised DEP-GAN model, which uses probability maps produced by a supervised segmentation method in the learning process, yielded similar performances to the DEP-UResNet and performed best in the clinical evaluation. Furthermore, an ablation study showed that an auxiliary input, especially the Gaussian noise, improved the performance of DEP models compared to DEP models that lacked the auxiliary input regardless of the model's architecture. To the best of our knowledge, this is the first extensive study on modelling WMH evolution using deep learning algorithms, which deals with the non-deterministic nature of WMH evolution. Muhammad Febrian Rachmadi, Maria del C. Valdés Hernández, Stephen D. Makin, Joanna M. Wardlaw, Taku Komura |
Medical Image Anal. | 5 |
| 2020 | Skeleton Filter: A Self-Symmetric Filter for Skeletonization in Noisy Text ImagesabstractRobustly computing the skeletons of objects in natural images is difficult due to the large variations in shape boundaries and the large amount of noise in the images. Inspired by recent findings in neuroscience, we propose the Skeleton Filter, which is a novel model for skeleton extraction from natural images. The Skeleton Filter consists of a pair of oppositely oriented Gabor-like filters; by applying the Skeleton Filter in various orientations to an image at multiple resolutions and fusing the results, our system can robustly extract the skeleton even under highly noisy conditions. We evaluate the performance of our approach using challenging noisy text datasets and demonstrate that our pipeline realizes state-of-the-art performance for extracting the text skeleton. Moreover, the presence of Gabor filters in the human visual system and the simple architecture of the Skeleton Filter can help explain the strong capabilities of humans in perceiving skeletons of objects, even under dramatically noisy conditions. Xiuxiu Bai, Lele Ye, Jihua Zhu, Li Zhu 0003, Taku Komura |
IEEE Trans. Image Process. | 5 |
| 2020 | Local motion phases for learning multi-contact character movementsabstractTraining a bipedal character to play basketball and interact with objects, or a quadruped character to move in various locomotion modes, are difficult tasks due to the fast and complex contacts happening during the motion. In this paper, we propose a novel framework to learn fast and dynamic character interactions that involve multiple contacts between the body and an object, another character and the environment, from a rich, unstructured motion capture database. We use one-on-one basketball play and character interactions with the environment as examples. To achieve this task, we propose a novel feature called local motion phase, that can help neural networks to learn asynchronous movements of each bone and its interaction with external objects such as a ball or an environment. We also propose a novel generative scheme to reproduce a wide variation of movements from abstract control signals given by a gamepad, which can be useful for changing the style of the motion under the same context. Our scheme is useful for animating contact-rich, complex interactions for real-time applications such as computer games. Sebastian Starke, Taku Komura, Kazi A. Zaman |
ACM Trans. Graph. | 3 |
| 2020 | Localization and Completion for 3D Object InteractionsabstractFinding where and what objects to put into an existing scene is a common task for scene synthesis and robot/character motion planning. Existing frameworks require development of hand-crafted features suitable for the task, or full volumetric analysis that could be memory intensive and imprecise. In this paper, we propose a data-driven framework to discover a suitable location and then place the appropriate objects in a scene. Our approach is inspired by computer vision techniques for localizing objects in images: using an all directional depth image (ADD-image) that encodes the 360-degree field of view from samples in the scene, our system regresses the images to the positions where the new object can be located. Given several candidate areas around the host object in the scene, our system predicts the partner object whose geometry fits well to the host object. Our approach is highly parallel and memory efficient, and is especially suitable for handling interactions between large and small objects. We show examples where the system can hang bags on hooks, fit chairs in front of desks, put objects into shelves, insert flowers into vases, and put hangers onto laundry rack. Xi Zhao 0002, Ruizhen Hu, Haisong Liu, Taku Komura, Xinyu Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Building hierarchical structures for 3D scenes with repeated elements
Xi Zhao 0002, Zhenqiang Su, Taku Komura, Xinyu Yang 0001 |
Vis. Comput. | 3 |
| 2019 | Predicting the Evolution of White Matter Hyperintensities in Brain MRI Using Generative Adversarial Networks and Irregularity Map
Muhammad Febrian Rachmadi, Maria del C. Valdés Hernández, Stephen D. Makin, Joanna M. Wardlaw, Taku Komura |
MICCAI (3) | 5 |
| 2019 | Neural state machine for character-scene interactionsabstractWe propose Neural State Machine , a novel data-driven framework to guide characters to achieve goal-driven actions with precise scene interactions. Even a seemingly simple task such as sitting on a chair is notoriously hard to model with supervised learning. This difficulty is because such a task involves complex planning with periodic and non-periodic motions reacting to the scene geometry to precisely position and orient the character. Our proposed deep auto-regressive framework enables modeling of multi-modal scene interaction behaviors purely from data. Given high-level instructions such as the goal location and the action to be launched there, our system computes a series of movements and transitions to reach the goal in the desired state. To allow characters to adapt to a wide range of geometry such as different shapes of furniture and obstacles, we incorporate an efficient data augmentation scheme to randomly switch the 3D geometry while maintaining the context of the original motion. To increase the precision to reach the goal during runtime, we introduce a control scheme that combines egocentric inference and goal-centric inference. We demonstrate the versatility of our model with various scene interaction tasks such as sitting on a chair, avoiding obstacles, opening and entering through a door, and picking and carrying objects generated in real-time just from a single model. Sebastian Starke, Taku Komura, Jun Saito |
ACM Trans. Graph. | 3 |
| 2019 | A Sampling Approach to Generating Closely Interacting 3D Pose-Pairs from 2D AnnotationsabstractWe introduce a data-driven method to generate a large number of plausible, closely interacting 3D human pose-pairs, for a given motion category, e.g., wrestling or salsa dance. With much difficulty in acquiring close interactions using 3D sensors, our approach utilizes abundant existing video data which cover many human activities. Instead of treating the data generation problem as one of reconstruction, either through 3D acquisition or direct 2D-to-3D data lifting from video annotations, we present a solution based on Markov Chain Monte Carlo (MCMC) sampling. Given a motion category and a set of video frames depicting the motion with the 2D pose-pair in each frame annotated, we start the sampling with one or few seed 3D pose-pairs which are manually created based on the target motion category. The initial set is then augmented by MCMC sampling around the seeds, via the Metropolis-Hastings algorithm and guided by a probability density function (PDF) that is defined by two terms to bias the sampling towards 3D pose-pairs that are physically valid and plausible for the motion category. With a focus on efficient sampling over the space of close interactions, rather than pose spaces, we develop a novel representation called interaction coordinates (IC) to encode both poses and their interactions in an integrated manner. Plausibility of a 3D pose-pair is then defined based on the IC and with respect to the annotated 2D pose-pairs from video. We show that our sampling-based approach is able to efficiently synthesize a large volume of plausible, closely interacting 3D pose-pairs which provide a good coverage of the input 2D pose-pairs. Kangxue Yin, Hui Huang 0004, Edmond S. L. Ho, Hao Wang 0057, Taku Komura, Daniel Cohen-Or, Hao (Richard) Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Automatic Irregular Texture Detection in Brain MRI Without Human Supervision
Muhammad Febrian Rachmadi, Maria del C. Valdés Hernández, Taku Komura |
MICCAI (3) | 3 |
| 2018 | Bulk-synchronous parallel simultaneous BVH traversal for collision detection on GPUsabstractSimultaneous BVH traversal, as a dynamic task of pair-wise proximity tests, poses several challenges in terms of parallelization using GPUs. It is a highly dynamic and data-dependent problem which can induce control-flow divergence and inefficient data-access patterns. We present a simple solution using the bulk-synchronous parallel model to ensure a uniform mode of execution, and balanced workloads across GPU threads. The method is easy to implement, fast and operates entirely on the GPU by relying on a topology-centred work expansion scheme to ensure large concurrent workloads. We demonstrate speedups of upto 7.1x over the widely used "streams" model for GPU based parallel collision detection. Floyd M. Chitalu, Christophe Dubach, Taku Komura |
I3D | 3 |
| 2018 | Random-forest-based initializer for solving inverse problem in 3D motion tracking systemsabstractMany motion tracking systems require solving inverse problem to compute the tracking result from original sensor measurements. For real-time motion tracking, such typical solutions as the Gauss-Newton method for solving their inverse problems need an initial value to optimize the cost function through iterations. A powerful initializer is crucial to generate a proper initial value for every time instance and, for achieving continuous accurate tracking without errors and rapid tracking recovery even when it is temporally interrupted. An improper initial value easily causes optimization divergence, and cannot always lead to reasonable solutions. Therefore, we propose a new initializer based on random-forest to obtain proper initial values for efficient real-time inverse problem computation. Our method trains a random-forest model with varied massive inputs and corresponding outputs and uses it as an initializer for runtime optimization. As an instance, we apply our initializer to IM3D[1], which is a real-time magnetic 3D motion tracking system with multiple tiny, identifiable, wireless, occlusion-free passive markers (LC coils). Ryo Sugawara, Jiawei Huang 0005, Kazuki Takashima, Taku Komura, Yoshifumi Kitamura |
VRST | 4 |
| 2018 | Few-shot Learning of Homogeneous Human Locomotion StylesabstractAbstract Using neural networks for learning motion controllers from motion capture data is becoming popular due to the natural and smooth motions they can produce, the wide range of movements they can learn and their compactness once they are trained. Despite these advantages, these systems require large amounts of motion capture data for each new character or style of motion to be generated, and systems have to undergo lengthy retraining, and often reengineering, to get acceptable results. This can make the use of these systems impractical for animators and designers and solving this issue is an open and rather unexplored problem in computer graphics. In this paper we propose a transfer learning approach for adapting a learned neural network to characters that move in different styles from those on which the original neural network is trained. Given a pretrained character controller in the form of a Phase‐Functioned Neural Network for locomotion, our system can quickly adapt the locomotion to novel styles using only a short motion clip as an example. We introduce a canonical polyadic tensor decomposition to reduce the amount of parameters required for learning from each new style, which both reduces the memory burden at runtime and facilitates learning from smaller quantities of data. We show that our system is suitable for learning stylized motions with few clips of motion data and synthesizing smooth motions in real‐time. Ian Mason, Sebastian Starke, Hakan Bilen, Taku Komura |
Comput. Graph. Forum | 5 |
| 2018 | Data-Driven Crowd Motion Control With Multi-Touch GesturesabstractAbstract Controlling a crowd using multi‐touch devices appeals to the computer games and animation industries, as such devices provide a high‐dimensional control signal that can effectively define the crowd formation and movement. However, existing works relying on pre‐defined control schemes require the users to learn a scheme that may not be intuitive. We propose a data‐driven gesture‐based crowd control system, in which the control scheme is learned from example gestures provided by different users. In particular, we build a database with pairwise samples of gestures and crowd motions. To effectively generalize the gesture style of different users, such as the use of different numbers of fingers, we propose a set of gesture features for representing a set of hand gesture trajectories. Similarly, to represent crowd motion trajectories of different numbers of characters over time, we propose a set of crowd motion features that are extracted from a Gaussian mixture model. Given a run‐time gesture, our system extracts the K nearest gestures from the database and interpolates the corresponding crowd motions in order to generate the run‐time control. Our system is accurate and efficient, making it suitable for real‐time applications such as real‐time strategy games and interactive animation controls. Joseph Henry, He Wang 0002, Edmond S. L. Ho, Taku Komura, Hubert P. H. Shum |
Comput. Graph. Forum | 5 |
| 2018 | Mode-adaptive neural networks for quadruped motion controlabstractQuadruped motion includes a wide variation of gaits such as walk, pace, trot and canter, and actions such as jumping, sitting, turning and idling. Applying existing data-driven character control frameworks to such data requires a significant amount of data preprocessing such as motion labeling and alignment. In this paper, we propose a novel neural network architecture called Mode-Adaptive Neural Networks for controlling quadruped characters. The system is composed of the motion prediction network and the gating network. At each frame, the motion prediction network computes the character state in the current frame given the state in the previous frame and the user-provided control signals. The gating network dynamically updates the weights of the motion prediction network by selecting and blending what we call the expert weights, each of which specializes in a particular movement. Due to the increased flexibility, the system can learn consistent expert weights across a wide range of non-periodic/periodic actions, from unstructured motion capture data, in an end-to-end fashion. In addition, the users are released from performing complex labeling of phases in different gaits. We show that this architecture is suitable for encoding the multi-modality of quadruped locomotion and synthesizing responsive motion in real-time. Sebastian Starke, Taku Komura, Jun Saito |
ACM Trans. Graph. | 3 |
| 2018 | Widening Viewing Angles of Automultiscopic Displays Using Refractive InsertsabstractDisplays that can portray environments that are perceivable from multiple views are known as multiscopic displays. Some multiscopic displays enable realistic perception of 3D environments without the need for cumbersome mounts or fragile head-tracking algorithms. These automultiscopic displays carefully control the distribution of emitted light over space, direction (angle) and time so that even a static image displayed can encode parallax across viewing directions (Iightfield). This allows simultaneous observation by multiple viewers, each perceiving 3D from their own (correct) perspective. Currently, the illusion can only be effectively maintained over a narrow range of viewing angles. In this paper, we propose and analyze a simple solution to widen the range of viewing angles for automultiscopic displays that use parallax barriers. We propose the use of a refractive medium, with a high refractive index, between the display and parallax barriers. The inserted medium warps the exitant lightfield in a way that increases the potential viewing angle. We analyze the consequences of this warp and build a prototype with a 93% increase in the effective viewing angle. Geng Lyu, Xukun Shen, Taku Komura, Kartic Subr, Lijun Teng |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | A Recurrent Variational Autoencoder for Human Motion Synthesis
Ikhsanul Habibie, Daniel Holden, Jonathan Schwarz, Joseph Yearsley, Taku Komura |
BMVC | 5 |
| 2017 | Character-Object Interaction Retrieval using the Interaction Bisector SurfaceabstractIn this paper, we propose a novel approach for the classification and retrieval of interactions between human characters and objects. We propose to use the interaction bisector surface (IBS) between the body and the object as a feature of the interaction. We define a multi-resolution representation of the body structure, and compute a correspondence matrix hierarchy that describes which parts of the character's skeleton take part in the composition of the IBS and how much they contribute to the interaction. Key-frames of the interactions are extracted based on the evolution of the IBS and used to align the query interaction with the interaction in the database. Through the experimental results, we show that our approach outperforms existing techniques in motion classification and retrieval, which implies that the contextual information plays a significant role for scene and interaction description. Our method also shows better performance than other techniques that use features based on the spatial relations between the body parts, or the body parts and the object. Our method can be applied for character motion synthesis and robot motion planning. Myung Geol Choi, Taku Komura |
Comput. Graph. Forum | 3 |
| 2017 | Phase-functioned neural networks for character controlabstractWe present a real-time character control mechanism using a novel neural network architecture called a Phase-Functioned Neural Network. In this network structure, the weights are computed via a cyclic function which uses the phase as an input. Along with the phase, our system takes as input user controls, the previous state of the character, the geometry of the scene, and automatically produces high quality motions that achieve the desired user control. The entire network is trained in an end-to-end fashion on a large dataset composed of locomotion such as walking, running, jumping, and climbing movements fitted into virtual environments. Our system can therefore automatically produce motions where the character adapts to different geometric environments such as walking and running over rough terrain, climbing over large rocks, jumping over obstacles, and crouching under low ceilings. Our network architecture produces higher quality results than time-series autoregressive models such as LSTMs as it deals explicitly with the latent variable of motion relating to the phase. Once trained, our system is also extremely fast and compact, requiring only milliseconds of execution time and a few megabytes of memory, even when trained on gigabytes of motion data. Our work is most appropriate for controlling characters in interactive scenes such as computer games and virtual reality systems. Daniel Holden, Taku Komura, Jun Saito |
ACM Trans. Graph. | 2 |
| 2017 | Learning Inverse Rig Mappings by Nonlinear RegressionabstractWe present a framework to design inverse rig-functions-functions that map low level representations of a character's pose such as joint positions or surface geometry to the representation used by animators called the animation rig. Animators design scenes using an animation rig, a framework widely adopted in animation production which allows animators to design character poses and geometry via intuitive parameters and interfaces. Yet most state-of-the-art computer animation techniques control characters through raw, low level representations such as joint angles, joint positions, or vertex coordinates. This difference often stops the adoption of state-of-the-art techniques in animation production. Our framework solves this issue by learning a mapping between the low level representations of the pose and the animation rig. We use nonlinear regression techniques, learning from example animation sequences designed by the animators. When new motions are provided in the skeleton space, the learned mapping is used to estimate the rig controls that reproduce such a motion. We introduce two nonlinear functions for producing such a mapping: Gaussian process regression and feedforward neural networks. The appropriate solution depends on the nature of the rig and the amount of data available for training. We show our framework applied to various examples including articulated biped characters, quadruped characters, facial animation rigs, and deformable characters. With our system, animators have the freedom to apply any motion synthesis algorithm to arbitrary rigging and animation pipelines for immediate editing. This greatly improves the productivity of 3D animation, while retaining the flexibility and creativity of artistic input. Daniel Holden, Jun Saito, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Scanning and animating characters dressed in multiple-layer garments
Pengpeng Hu, Taku Komura, Daniel Holden, Yueqi Zhong |
Vis. Comput. | 2 |
| 2016 | Modelling the Usage of Discourse Connectives as Rational Speech ActsabstractDiscourse relations can either be implicit or explicitly expressed by markers, such as 'therefore' and 'but'.How a speaker makes this choice is a question that is not well understood.We propose a psycholinguistic model that predicts whether a speaker will produce an explicit marker given the discourse relation s/he wishes to express.Based on the framework of the Rational Speech Acts model, we quantify the utility of producing a marker based on the information-theoretic measure of surprisal, the cost of production, and a bias to maintain uniform information density throughout the utterance.Experiments based on the Penn Discourse Treebank show that our approach outperforms stateof-the-art approaches, while giving an explanatory account of the speaker's choice.1 'Speakers' and 'listeners' are interchangeably used with 'authors' and 'readers' in this article Frances Yung, Kevin Duh, Taku Komura, Yuji Matsumoto 0001 |
CoNLL | 3 |
| 2016 | SkillVis: a visualization tool for boxing skill assessmentabstractMotion analysis and visualization are crucial in sports science for sports training and performance evaluation. While primitive computational methods have been proposed for simple analysis such as postures and movements, few can evaluate the high-level quality of sports players such as their skill levels and strategies. We propose a visualization tool to help visualizing boxers' motions and assess their skill levels. Our system automatically builds a graph-based representation from motion capture data and reduces the dimension of the graph onto a 3D space so that it can be easily visualized and understood. In particular, our system allows easy understanding of the boxer's boxing behaviours, preferred actions, potential strength and weakness. We demonstrate the effectiveness of our system on different boxers' motions. Our system not only serves as a tool for visualization, it also provides intuitive motion analysis that can be further used beyond sports science. Hubert P. H. Shum, He Wang 0002, Edmond S. L. Ho, Taku Komura |
MIG | 4 |
| 2016 | Coordinated Crowd Simulation With Topological Scene AnalysisabstractAbstract This paper proposes a new algorithm to produce globally coordinated crowds in an environment with multiple paths and obstacles. Simple greedy crowd control methods easily lead to congestion at bottlenecks within scenes, as the characters do not cooperate with one another. In computer animation, this problem degrades crowd quality especially when ordered behaviour is needed, such as soldiers marching towards a castle. Similarly, in applications such as real‐time strategy games, this often causes player frustration, as the crowd will not move as efficiently as it should. Also, planning of building would usually require visualization of ordered evacuation to maximize the flow. Planning such globally coordinated crowd movement is usually labour intensive. Here, we propose a simple solution that is easy to use and efficient in computation. First, we compute the harmonic field of the environment, taking into account the starting points, goals and obstacles. Based on the field, we represent the topology of the environment using a Reeb Graph, and calculate the maximum capacity for each path in the graph. With the harmonic field and the Reeb Graph, path planning of crowd can be performed using a lightweight algorithm, such that any blocking of one another's paths is minimized. Comparing to previous methods, our system can synthesize globally coordinated crowd with smooth and efficient movement. It also enables control of the crowd with high‐level parameters such as the degree of cooperation and congestion. Finally, the method is scalable to thousands of characters with minimal impact to computation time. It is best applied in interactive crowd synthesis systems such as animation designs and real‐time strategy games. Adam Barnett, Hubert P. H. Shum, Taku Komura |
Comput. Graph. Forum | 3 |
| 2016 | Character contact re-positioning under large environment deformationabstractAbstract Character animation based on motion capture provides intrinsically plausible results, but lacks the flexibility of procedural methods. Motion editing methods partially address this limitation by adapting the animation to small deformations of the environment. We extend one such method, the so‐called relationship descriptors, to tackle the issue of motion editing under large environment deformations. Large deformations often result in joint limits violation, loss of balance, or collisions. Our method handles these situations by automatically detecting and re‐positioning invalidated contacts. The new contact configurations are chosen to preserve the mechanical properties of the original contacts in order to provide plausible support phases. When it is not possible to find an equivalent contact, a procedural animation is generated and blended with the original motion. Thanks to an optimization scheme, the resulting motions are continuous and preserve the style of the reference motions. The method is fully interactive and enables the motion to be adapted on‐line even in case of large changes of the environment. We demonstrate our method on several challenging scenarios, proving its immediate application to 3D animation softwares and video games. Steve Tonneau, Rami Ali Al-Ashqar, Julien Pettré, Taku Komura, Nicolas Mansard |
Comput. Graph. Forum | 4 |
| 2016 | A deep learning framework for character motion synthesis and editingabstractWe present a framework to synthesize character movements based on high level parameters, such that the produced movements respect the manifold of human motion, trained on a large motion capture dataset. The learned motion manifold, which is represented by the hidden units of a convolutional autoencoder, represents motion data in sparse components which can be combined to produce a wide range of complex movements. To map from high level parameters to the motion manifold, we stack a deep feedforward neural network on top of the trained autoencoder. This network is trained to produce realistic motion sequences from parameters such as a curve over the terrain that the character should follow, or a target location for punching and kicking. The feedforward control network and the motion manifold are trained independently, allowing the user to easily switch between feedforward networks according to the desired interface, without re-training the motion manifold. Once motion is generated it can be edited by performing optimization in the space of the motion manifold. This allows for imposing kinematic constraints, or transforming the style of the motion, while ensuring the edited motion remains natural. As a result, the system can produce smooth, high quality motion sequences without any manual pre-processing of the training data. Daniel Holden, Jun Saito, Taku Komura |
ACM Trans. Graph. | 3 |
| 2016 | Relationship templates for creating scene variationsabstractWe propose a novel example-based approach to synthesize scenes with complex relations, e.g., when one object is 'hooked', 'surrounded', 'contained' or 'tucked into' another object. Existing relationship descriptors used in automatic scene synthesis methods are based on contacts or relative vectors connecting the object centers. Such descriptors do not fully capture the geometry of spatial interactions, and therefore cannot describe complex relationships. Our idea is to enrich the description of spatial relations between object surfaces by encoding the geometry of the open space around objects, and use this as a template for fitting novel objects. To this end, we introduce relationship templates as descriptors of complex relationships; they are computed from an example scene and combine the interaction bisector surface (IBS) with a novel feature called the space coverage feature (SCF), which encodes the open space in the frequency domain. New variations of a scene can be synthesized efficiently by fitting novel objects to the template. Our method greatly enhances existing automatic scene synthesis approaches by allowing them to handle complex relationships, as validated by our user studies. The proposed method generalizes well, as it can form complex relationships with objects that have a topology and geometry very different from the example scene. Xi Zhao 0002, Ruizhen Hu, Paul Guerrero 0001, Niloy J. Mitra, Taku Komura |
ACM Trans. Graph. | 5 |
| 2015 | Carpet unrolling for character control on uneven terrainabstractWe propose a type of relationship descriptor based on carpet unrolling that computes the joint positions of a character based on the sum of relative vectors originating from a local coordinate system embedded on the surface of a carpet. Given a terrain that a character is to walk over, the carpet is unrolled over the surface of the terrain. The carpet adapts to the geometry of the terrain and curves according to the trajectory of the character. Because trajectories of the body parts are computed as a weighted sum of the relative vectors, the character can smoothly adapt to the elevation of the terrain and the horizontal curves of the carpet. The carpet relationship descriptors are easy to parallelize and hundreds of characters can be animated in real-time by making use of the GPUs. This makes it applicable to real-time applications such as computer games. Mark Miller 0002, Daniel Holden, Rami Ali Al-Ashqar, Christophe Dubach, Kenny Mitchell, Taku Komura |
MIG | 6 |
| 2015 | An Energy-Driven Motion Planning Method for Two Distant PosturesabstractIn this paper, we present a local motion planning algorithm for character animation. We focus on motion planning between two distant postures where linear interpolation leads to penetrations. Our framework has two stages. The motion planning problem is first solved as a Boundary Value Problem (BVP) on an energy graph which encodes penetrations, motion smoothness and user control. Having established a mapping from the configuration space to the energy graph, a fast and robust local motion planning algorithm is introduced to solve the BVP to generate motions that could only previously be computed by global planning methods. In the second stage, a projection of the solution motion onto a constraint manifold is proposed for more user control. Our method can be integrated into current keyframing techniques. It also has potential applications in motion planning problems in robotics. He Wang 0002, Edmond S. L. Ho, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | A multi-resolution approach for adapting close character interactionabstractSynthesizing close interactions such as dancing and fighting between characters is a challenging problem in computer animation. While encouraging results are presented in [Ho et al. 2010], the high computation cost makes the method unsuitable for interactive motion editing and synthesis. In this paper, we propose an efficient multiresolution approach in the temporal domain for editing and adapting close character interactions based on the Interaction Mesh framework. In particular, we divide the original large spacetime optimization problem into multiple smaller problems such that the user can observe the adapted motion while playing-back the movements during run-time. Our approach is highly parallelizable, and achieves high performance by making use of multi-core architectures. The method can be applied to a wide range of applications including motion editing systems for animators and motion retargeting systems for humanoid robots. Edmond S. L. Ho, He Wang 0002, Taku Komura |
VRST | 3 |
| 2014 | Model topology change with correspondence using electrostaticsabstractThis paper introduces a method for finding a dense correspondence between objects of varying topology or connectivity by using a proxy, genus zero mesh alongside the technique of Blended Intrinsic Maps. Harmonic space parameterisation is used to create a closed, genus zero shape that approximates the geometry of the original object. This allows for noisy or topologically different representations of objects to be mapped to one another, with seams in the mapping falling in generally hidden concave areas and tunnels. The paper presents example mapping between objects with and without holes, as well as objects that consist of a number of disconnected segments. Peter Sandilands, Taku Komura |
VRST | 2 |
| 2014 | Braiding hair by braid theoryabstractIn this paper, we propose a system based on braid theory that help users to generate customized hair braiding, which is a function that is lacking in most existing hair design software. Our user interface for braid design is built upon braid theory, which is a subarea of knot theory in mathematics. The user designs braid patterns using braid index, and specifies the amount of hair for each braid as well as the area over the head where the braid is to be made. Then, the system automatically braids the hair of the character and generates a realistic image of the designed hair style. Theoretically, our system can produce arbitrary kinds of braids. Our system can also judge if two braids are equivalent or not by making use of the transition rules of braid index, which helps to register designed braids to the database. The system is implemented as a Maya plugin, and can be combinedly used with various functions including physical simulation, hair rendering and hair texturing. Our user study shows that our toolkit is easy-to-use for novice users as well as experienced users. Gaoxiang Zeng, Taku Komura |
VRST | 2 |
| 2014 | Natural preparation behavior synthesisabstractABSTRACT Humans adjust their movements in advance to prepare for the forthcoming action, resulting in efficient and smooth transitions. However, traditional computer animation approaches such as motion graphs simply concatenate a series of actions without taking into account the following one. In this paper, we propose a new method to produce preparation behaviors using reinforcement learning. As an offline process, the system learns the optimal way to approach a target and to prepare for interaction. A scalar value called the level of preparation is introduced, which represents the degree of transition from the initial action to the interacting action. To synthesize the movements of preparation, we propose a customized motion blending scheme based on the level of preparation, which is followed by an optimization framework that adjusts the posture to keep the balance. During runtime, the trained controller drives the character to move to a target with the appropriate level of preparation, resulting in a humanlike behavior. We create scenes in which the character has to move in a complex environment and to interact with objects, such as crawling under and jumping over obstacles while walking. The method is useful not only for computer animation but also for real‐time applications such as computer games, in which the characters need to accomplish a series of tasks in a given environment. Copyright © 2013 John Wiley & Sons, Ltd. Hubert P. H. Shum, Ludovic Hoyet, Edmond S. L. Ho, Taku Komura, Franck Multon |
Comput. Animat. Virtual Worlds | 4 |
| 2014 | Indexing 3D Scenes Using the Interaction Bisector SurfaceabstractThe spatial relationship between different objects plays an important role in defining the context of scenes. Most previous 3D classification and retrieval methods take into account either the individual geometry of the objects or simple relationships between them such as the contacts or adjacencies. In this article we propose a new method for the classification and retrieval of 3D objects based on the Interaction Bisector Surface (IBS), a subset of the Voronoi diagram defined between objects. The IBS is a sophisticated representation that describes topological relationships such as whether an object is wrapped in, linked to, or tangled with others, as well as geometric relationships such as the distance between objects. We propose a hierarchical framework to index scenes by examining both the topological structure and the geometric attributes of the IBS. The topology-based indexing can compare spatial relations without being severely affected by local geometric details of the object. Geometric attributes can also be applied in comparing the precise way in which the objects are interacting with one another. Experimental results show that our method is effective at relationship classification and content-based relationship retrieval. Xi Zhao 0002, He Wang 0002, Taku Komura |
ACM Trans. Graph. | 3 |
| 2014 | Interactive Formation Control in Complex EnvironmentsabstractThe degrees of freedom of a crowd is much higher than that provided by a standard user input device. Typically, crowd-control systems require multiple passes to design crowd movements by specifying waypoints, and then defining character trajectories and crowd formation. Such multi-pass control would spoil the responsiveness and excitement of real-time control systems. In this paper, we propose a single-pass algorithm to control a crowd in complex environments. We observe that low-level details in crowd movement are related to interactions between characters and the environment, such as diverging/merging at cross points, or climbing over obstacles. Therefore, we simplify the problem by representing the crowd with a deformable mesh, and allow the user, via multitouch input, to specify high-level movements and formations that are important for context delivery. To help prevent congestion, our system dynamically reassigns characters in the formation by employing a mass transport solver to minimize their overall movement. The solver uses a cost function to evaluate the impact from the environment, including obstacles and areas affecting movement speed. Experimental results show realistic crowd movement created with minimal high-level user inputs. Our algorithm is particularly useful for real-time applications including strategy games and interactive animation creation. Joseph Henry, Hubert P. H. Shum, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Dynamic Comics for Hierarchical Abstraction of 3D Animation DataabstractAbstract Image storyboards of films and videos are useful for quick browsing and automatic video processing. A common approach for producing image storyboards is to display a set of selected key‐frames in temporal order, which has been widely used for 2D video data. However, such an approach cannot be applied for 3D animation data because different information is revealed by changing parameters such as the viewing angle and the duration of the animation. Also, the interests of the viewer may be different from person to person. As a result, it is difficult to draw a single image that perfectly abstracts the entire 3D animation data. In this paper, we propose a system that allows users to interactively browse an animation and produce a comic sequence out of it. Each snapshot in the comic optimally visualizes a duration of the original animation, taking into account the geometry and motion of the characters and objects in the scene. This is achieved by a novel algorithm that automatically produces a hierarchy of snapshots from the input animation. Our user interface allows users to arrange the snapshots according to the complexity of the movements by the characters and objects, the duration of the animation and the page area to visualize the comic sequence. Our system is useful for quickly browsing through a large amount of animation data and semi‐automatically synthesizing a storyboard from a long sequence of animation. Myung Geol Choi, Seung-tak Noh, Taku Komura, Takeo Igarashi |
Comput. Graph. Forum | 3 |
| 2013 | Interaction capture using magnetic sensorsabstractABSTRACT Capturing a close interaction between an actor and an object can be difficult as a result of occlusion and having to recreate the geometry of the scene accurately. In this paper, we propose a technique that allows us to capture the object's motion and geometry alongside the actor's movements and optionally the local environment, using a magnetic motion capture system and an RGB‐D sensor. This not only gives greater information when placing a character in a scene but enables us to digitally recreate the scene in motion without significant animator work after capture. The use of magnetic sensors prevents occlusion or marker confusion that is common in optical techniques when dealing with close interactions, as the magnetic sensors do not require direct line of sight to a camera. The geometry reconstruction ensures that the proportions of objects and surfaces the character interacts with are accurate and alleviates the need for an artist to model the object. We perform validation of the results by comparison with an optical system and show a variety of motions, such as using a screwdriver or removing a cap to drink from a bottle, that can be captured using our technique. Copyright © 2013 John Wiley & Sons, Ltd. Peter Sandilands, Myung Geol Choi, Taku Komura |
Comput. Animat. Virtual Worlds | 3 |
| 2013 | Harmonic parameterization by electrostaticsabstractIn this article, we introduce a method to apply ideas from electrostatics to parameterize the open space around an object. By simulating the object as a virtually charged conductor, we can define an object-centric coordinate system which we call Electric Coordinates. It parameterizes the outer space of a reference object in a way analogous to polar coordinates. We also introduce a measure that quantifies the extent to which an object is wrapped by a surface. This measure can be computed as the electric flux through the wrapping surface due to the electric field around the charged conductor. The electrostatic parameters, which comprise the Electric Coordinates and flux, have several applications in computer graphics, including: texturing, morphing, meshing, path planning relative to a target object, mesh parameterization, designing deformable objects, and computing coverage. Our method works for objects of arbitrary geometry and topology, and thus is applicable in a wide variety of scenarios. He Wang 0002, Kirill A. Sidorov, Peter Sandilands, Taku Komura |
ACM Trans. Graph. | 4 |
| 2013 | Interactive partner control in close interactions for real-time applicationsabstractThis article presents a new framework for synthesizing motion of a virtual character in response to the actions performed by a user-controlled character in real time. In particular, the proposed method can handle scenes in which the characters are closely interacting with each other such as those in partner dancing and fighting. In such interactions, coordinating the virtual characters with the human player automatically is extremely difficult because the system has to predict the intention of the player character. In addition, the style variations from different users affect the accuracy in recognizing the movements of the player character when determining the responses of the virtual character. To solve these problems, our framework makes use of the spatial relationship-based representation of the body parts called interaction mesh, which has been proven effective for motion adaptation. The method is computationally efficient, enabling real-time character control for interactive applications. We demonstrate its effectiveness and versatility in synthesizing a wide variety of motions with close interactions. Edmond S. L. Ho, Jacky C. P. Chan, Taku Komura, Howard Leung |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Capturing Close Interactions with Objects Using a Magnetic Motion Capture System and a RGBD Sensor
Peter Sandilands, Myung Geol Choi, Taku Komura |
MIG | 3 |
| 2012 | Interaction Retrieval by Spacetime Proximity GraphsabstractAbstract In this paper, we propose a new method to index and retrieve animation scenes in which multiple characters closely interact with one another. Such a technique can be an important tool for animators when they want to automatically extract the desired scene from a large database of animation sequence. Existing methods for single character movements do not scale well for multiple characters as they do not take into account the interaction of different body parts. In this paper, we propose a new distance function that computes the similarity of two‐character interations using the spatial relationship of the body parts. For each interaction, we produce a time‐varying graph structure based on the proximity of different joints, and compute the similarity of interactions by comparing the topology and Laplacian coordinates of the time‐varying graph. Experimental results show that the proposed method outperforms previous methods which are based on the kinematics of individual characters. The top retrieved samples are found similar in high level semantics while containing style variations. Jeff K. T. Tang, Jacky C. P. Chan, Howard Leung, Taku Komura |
Comput. Graph. Forum | 4 |
| 2012 | Manipulation of Flexible Objects by Geodesic ControlabstractAbstract We propose an effective and intuitive method for controlling flexible models such as ropes and cloth. Automating manipulation of such flexible objects is not an easy task due to the high dimensionality of the objects and the low dimensionality of the control. In order to cope with this problem, we introduce a method called Geodesic Control, which greatly helps to manipulate flexible objects. The core idea is to decrease the degrees of freedom of the flexible object by moving it along the geodesic line of the object that it is interacting with. By repeatedly applying this control, users can easily synthesize animations of twisting and knotting a piece of rope or wrapping a cloth around an object. We show examples of “furoshiki wrapping”, in which an object is wrapped by a cloth by a series of maneuvers based on Geodesic Control. As our representation can abstract such maneuvers well, the procedure designed by a user can be re‐applied for different combinations of cloth and an object. The method is applicable not only for computer animation but also for 3D computer games and virtual reality systems. He Wang 0002, Taku Komura |
Comput. Graph. Forum | 2 |
| 2012 | Guest Editors' Introduction: Special Section on ACM VRSTabstractThe articles in this special section contain selected papers from the 2010 ACM Virtual Reality Software and Technology Symposium. Taku Komura, Qunsheng Peng 0001, George Baciu, Rynson W. H. Lau |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | Simulating Multiple Character Interactions with Collaborative and Adversarial GoalsabstractThis paper proposes a new methodology for synthesizing animations of multiple characters, allowing them to intelligently compete with one another in dense environments, while still satisfying requirements set by an animator. To achieve these two conflicting objectives simultaneously, our method separately evaluates the competition and collaboration of the interactions, integrating the scores to select an action that maximizes both criteria. We extend the idea of min-max search, normally used for strategic games such as chess. Using our method, animators can efficiently produce scenes of dense character interactions such as those in collective sports or martial arts. The method is especially effective for producing animations along story lines, where the characters must follow multiple objectives, while still accommodating geometric and kinematic constraints from the environment. Hubert P. H. Shum, Taku Komura, Shuntaro Yamazaki |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Real-time controllable fire using textured forces
Jake Lever, Taku Komura |
Vis. Comput. | 2 |
| 2011 | Energy-Based Pose Unfolding and Interpolation for 3D Articulated Characters
He Wang 0002, Taku Komura |
MIG | 2 |
| 2011 | A finite state machine based on topology coordinates for wrestling gamesabstractThis paper proposes a new framework to simulate the real-time attack-and-defense interactions by two virtual wrestlers in 3D computer games. The characters are controlled individually by two different players—one player controls the attacker and the other controls the defender. A finite state machine of attacks and defenses based on topology coordinates is precomputed and used to control the virtual wrestlers during the game play. As the states are represented by topology coordinates, which is an abstract representation for the spatial relationship of the bodies, the players have much more degree of freedom to control the virtual characters even during attacks and defenses. Experimental results show the methodology can simulate realistic competitive interactions of wrestling in real time, which is difficult by previous methods. Copyright © 2010 John Wiley & Sons, Ltd. Edmond S. L. Ho, Taku Komura |
Comput. Animat. Virtual Worlds | 2 |
| 2010 | Controlling humanoid robots in topology coordinatesabstractThis paper presents an approach to the control of humanoid robot motion, e.g., holding another robot or tangled interactions involving multiple limbs, in a space defined by `topology coordinates'. The constraints of tangling can be linearized at every frame of motion synthesis, and can be used together with constraints such as defined by the Zero Moment Point, Center of Mass, inverse kinematics and angular momentum for computing the postures by a linear programming procedure. We demonstrate the utility of this approach using the simulator for the Nao humanoid robot. We show that this approach enables us to synthesize complex motion, such as tangling, very efficiently. Edmond S. L. Ho, Taku Komura, Subramanian Ramamoorthy, Sethu Vijayakumar |
IROS | 2 |
| 2010 | Perception Based Real-Time Dynamic Adaptation of Human Motions
Ludovic Hoyet, Franck Multon, Taku Komura, Anatole Lécuyer |
MIG | 3 |
| 2010 | Physically-Based Character Control in Low Dimensional Space
Hubert P. H. Shum, Taku Komura, Takaaki Shiratori, Shu Takagi |
MIG | 2 |
| 2010 | Can we distinguish biological motions of virtual humans?: perceptual study with captured motions of weight liftingabstractPerception of biological motions is a key issue in order to evaluate the quality and the credibility of motions of virtual humans. This paper presents a perceptual study to evaluate if human beings are able to accurately distinguish differences in natural lifting motions with various masses in virtual environments (VE), which is not the case. However, they reached very close levels of accuracy when watching to computer animations compared to videos. Still, quotes of participants suggest that the discrimination process is easier in videos of real motions which included muscles contractions, more degrees of freedom, etc. These results can be used to help animators to design efficient physically-based animations. Ludovic Hoyet, Franck Multon, Anatole Lécuyer, Taku Komura |
VRST | 4 |
| 2010 | Spatial relationship preserving character motion adaptationabstractThis paper presents a new method for editing and retargeting motions that involve close interactions between body parts of single or multiple articulated characters, such as dancing, wrestling, and sword fighting, or between characters and a restricted environment, such as getting into a car. In such motions, the implicit spatial relationships between body parts/objects are important for capturing the scene semantics. We introduce a simple structure called an interaction mesh to represent such spatial relationships. By minimizing the local deformation of the interaction meshes of animation frames, such relationships are preserved during motion editing while reducing the number of inappropriate interpenetrations. The interaction mesh representation is general and applicable to various kinds of close interactions. It also works well for interactions involving contacts and tangles as well as those without any contacts. The method is computationally efficient, allowing real-time character control. We demonstrate its effectiveness and versatility in synthesizing a wide variety of motions with close interactions. Edmond S. L. Ho, Taku Komura, Chiew-Lan Tai |
ACM Trans. Graph. | 2 |
| 2009 | Character Motion Synthesis by Topology CoordinatesabstractAbstract In this paper, we propose a new method to efficiently synthesize character motions that involve close contacts such as wearing a T‐shirt, passing the arms through the strings of a knapsack, or piggy‐back carrying an injured person. We introduce the concept of topology coordinates, in which the topological relationships of the segments are embedded into the attributes. As a result, the computation for collision avoidance can be greatly reduced for complex motions that require tangling the segments of the body. Our method can be combinedly used with other prevalent frame‐based optimization techniques such as inverse kinematics. Edmond S. L. Ho, Taku Komura |
Comput. Graph. Forum | 2 |
| 2009 | Interactive animation of virtual humans based on motion capture dataabstractAbstract This paper presents a novel, parameteric framework for synthesizing new character motions from existing motion capture data. Our framework can conduct morphological adaptation as well as kinematic and physically‐based corrections. All these solvers are organized in layers in order to be easily combined together. Given locomotion as an example, the system automatically adapts the motion data to the size of the synthetic figure and to its environment; the character will correctly step over complex ground shapes and counteract with external forces applied to the body. Our framework is based on a frame‐based solver. This ensures animating hundreds of humanoids with different morphologies in real‐time. It is particularly suitable for interactive applications such as video games and virtual reality where a user interacts in an unpredictable way. Copyright © 2009 John Wiley & Sons, Ltd. Franck Multon, Richard Kulpa, Ludovic Hoyet, Taku Komura |
Comput. Animat. Virtual Worlds | 4 |
| 2009 | Angular momentum guided motion concatenationabstractAbstract In this paper, we propose a new method to concatenate two dynamic full‐body motions such as punches, kicks, and flips by using the angular momentum as a cue. Through the observation of real humans, we have identified two patterns of angular momentum that make the transition of such motions efficient. Based on these observations, we propose a new method to concatenate two full‐body motions in a natural manner. Our method is useful for applications where dynamic, full‐body motions are required, such as 3D computer games and animations. Copyright © 2009 John Wiley & Sons, Ltd. Hubert P. H. Shum, Taku Komura, Pranjul Yadav |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Indexing and Retrieving Motions of Characters in Close ContactabstractHuman motion indexing and retrieval are important for animators due to the need to search for motions in the database which can be blended and concatenated. Most of the previous researches of human motion indexing and retrieval compute the Euclidean distance of joint angles or joint positions. Such approaches are difficult to apply for cases in which multiple characters are closely interacting with each other, as the relationships of the characters are not encoded in the representation. In this research, we propose a topology-based approach to index the motions of two human characters in close contact. We compute and encode how the two bodies are tangled based on the concept of rational tangles. The encoded relationships, which we define as TangleList, are used to determine the similarity of the pairs of postures. Using our method, we can index and retrieve motions such as one person piggy-backing another, one person assisting another in walking, and two persons dancing / wrestling. Our method is useful to manage a motion database of multiple characters. We can also produce motion graph structures of two characters closely interacting with each other by interpolating and concatenating topologically similar postures and motion clips, which are applicable to 3D computer games and computer animation. Edmond S. L. Ho, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2008 | Simulating interactions of avatars in high dimensional state spaceabstractEfficient computation of strategic movements is essential to control virtual avatars intelligently in computer games and 3D virtual environments. Such a module is needed to control non-player characters (NPCs) to fight, play team sports or move through a mass crowd. Reinforcement learning is an approach to achieve real-time optimal control. However, the huge state space of human interactions makes it difficult to apply existing learning methods to control avatars when they have dense interactions with other characters. In this research, we propose a new methodology to efficiently plan the movements of an avatar interacting with another. We make use of the fact that the subspace of meaningful interactions is much smaller than the whole state space of two avatars. We efficiently collect samples by exploring the subspace where dense interactions between the avatars occur and favor samples that have high connectivity with the other samples. Using the collected samples, a finite state machine (FSM) called Interaction Graph is composed. At run-time, we compute the optimal action of each avatar by minmax search or dynamic programming on the Interaction Graph. The methodology is applicable to control NPCs in fighting and ball-sports games. Hubert P. H. Shum, Taku Komura, Shuntaro Yamazaki |
SI3D | 2 |
| 2008 | Emulating human perception of motion similarityabstractAbstract Evaluating the similarity of motions is useful for motion retrieval, motion blending, and performance analysis of dancers and athletes. Euclidean distance between corresponding joints has been widely adopted in measuring similarity of postures and hence motions. However, such a measure does not necessarily conform to the human perception of motion similarity. In this paper, we propose a new similarity measure based on machine learning techniques. We make use of the results of questionnaires from subjects answering whether arbitrary pairs of motions appear similar or not. Using the relative distance between the joints as the basic features, we train the system to compute the similarity of arbitrary pair of motions. Experimental results show that our method outperforms methods based on Euclidean distance between corresponding joints. Our method is applicable to content‐based motion retrieval of human motion for large‐scale database systems. It is also applicable to e‐Learning systems which automatically evaluates the performance of dancers and athletes by comparing the subjects' motions with those by experts. Copyright © 2008 John Wiley & Sons, Ltd. Jeff K. T. Tang, Howard Leung, Taku Komura, Hubert P. H. Shum |
Comput. Animat. Virtual Worlds | 3 |
| 2008 | Interaction patches for multi-character animationabstractWe propose a data-driven approach to automatically generate a scene where tens to hundreds of characters densely interact with each other. During off-line processing, the close interactions between characters are precomputed by expanding a game tree, and these are stored as data structures called interaction patches. Then, during run-time, the system spatio-temporally concatenates the interaction patches to create scenes where a large number of characters closely interact with one another. Using our method, it is possible to automatically or interactively produce animations of crowds interacting with each other in a stylized way. The method can be used for a variety of applications including TV programs, advertisements and movies. Hubert P. H. Shum, Taku Komura, Masashi Shiraishi, Shuntaro Yamazaki |
ACM Trans. Graph. | 2 |
| 2007 | Wrestle Alone : Creating Tangled Motions of Multiple Avatars from Individually Captured MotionsabstractAnimations of two avatars tangled with each other often appear in battle or fighting scenes in movies or games. However, creating such scenes is difficult due to the limitations of the tracking devices and the complex interactions of the avatars during such motions. In this paper, we propose a new method to generate animations of two persons tangled with each other based on individually captured motions. We use wrestling as an example. The inputs to the system are two individually captured motions and the topological relationship of the two avatars computed using Gauss Linking Integral (GLI). Then the system edits the captured motions so that they satisfy the given topological relationship. Using our method, it is possible to create / edit close-contact motions with minimum effort by the animators. The method can be used not only for wrestling, but also for any movement that requires the body to be tangled with others, such as holding a shoulder of an elderly to walk or a soldier piggy-backing another injured soldier. Edmond S. L. Ho, Taku Komura |
PG | 2 |
| 2007 | Interactive control of physically-valid aerial motion: application to VR training system for gymnastsabstractThis paper aims at proposing a new method to animate aerial motions in interactive environments while taking dynamics into account. Classical approaches are based on spacetime constraints and require a complete knowledge of the motion. However, in Virtual Reality, the user's actions are unpredictable so that such techniques cannot be used. In this paper, we deal with the simulation of gymnastic aerial motions in virtual reality. A user can directly interact with the virtual gymnast thanks to a real-time motion capture system. The user's arm motions are blended to the original aerial motions in order to verify their consequences on the virtual gymnast's performance. Hence, a user can select an initial motion, an initial velocity vector, an initial angular momentum, and a virtual character. Each of these choices has a direct influence on mechanical values such as the linear and angular momentum. We thus have developed an original method to adapt the character's poses at each time step in order to make these values compatible with mechanical laws: the angular momentum is constant during the aerial phase and the linear one is determined at take-off. Our method enables to animate up to 16 characters at 30hz on a common PC. To sum-up, our method enables to solve kinematic constraints, to retarget motion and to correct it to satisfy mechanical laws. The virtual gymnast application described in this paper is very promising to help sports-men getting some ideas which postures are better during the aerial phase for better performance. Franck Multon, Ludovic Hoyet, Taku Komura, Richard Kulpa |
VRST | 3 |
| 2007 | Simulating competitive interactions using singly captured motionsabstractIt is difficult to create scenes where multiple avatars are fighting / competing with each other. Manually creating the motions of avatars is time consuming due to the correlation of the movements between the avatars. Capturing the motions of multiple avatars is also difficult as it requires a huge amount of post-processing. In this paper, we propose a new method to generate a realistic scene of avatars densely interacting in a competitive environment. The motions of the avatars are considered to be captured individually, which will increase the easiness of obtaining the data. We propose a new algorithm called the temporal expansion approach which maps the continuous time action plan to a discrete space such that turn-based evaluation methods can be used. As a result, many mature algorithms in game such as the min-max search and α---β pruning can be applied. Using our method, avatars will plan their strategies taking into account the reaction of the opponent. Fighting scenes with multiple avatars are generated to demonstrate the effectiveness of our algorithm. The proposed method can also be applied to other kinds of continuous activities that require strategy planning such as sport games. Hubert P. H. Shum, Taku Komura, Shuntaro Yamazaki |
VRST | 2 |
| 2006 | Stepping Motion for a Human-like Character to Maintain Balance against Large PerturbationsabstractWe propose a method of maintaining balance for a human-like character against large perturbations. The method enables a human-like model to maintain its balance with active whole-body motion, such as rotating its arms, bending down, and taking a step, if necessary. First, we capture the human motions of maintaining balance and abstract essential mechanisms from these motions. Next, we construct a model of maintaining balance that has a simple structure, such as an inverted pendulum. This model has two modes of maintaining balance: keeping the feet on the ground, and stepping. In this paper, the stepping mode is mainly described. Finally, we generate whole-body motion based on the model against several perturbations, and we discuss the validity of our method Shunsuke Kudoh, Taku Komura, Katsushi Ikeuchi |
ICRA | 2 |
| 2006 | Learning system for human motion characters of traditional artsabstractA useful learning system for human motion characters of traditional arts, such as Mai (Japanese classic dance), Kabuki (one of Japan's traditional stage arts), etc., is being developed. In such arts an effective system to pass the tradition down from a top artist to next generations is required. Video contents are generally used to pass the human motions in traditional arts down for non-experts. However, the video contents normally show the human motions as the views from a single direction. If the human motions are presented from orthogonal three-directions at the same time, it comes more useful. So, our learning system produces the three-dimensional human motion by the sequences of views from a single direction in the video contents, then, the motions from any directions can be presented simultaneously with the video contents.In addition, a learner can check the difference in motion between a top artist and him/herself by the producing his/her three-dimensional skeleton motion with our system and the overlapping it with one of the top artist. Here, in the comparison between two different physiques (a top artist and a learner) a simple adjustment method is suggested.In our study standard motion capture processes are employed, but our goal is to develop an original practical training system of performers' motion for beginners in traditional arts. In this paper the developing system is demonstrated for Kyo-mai (Mai originated in Kyoto) as example. The developing system can be useful not only in performing arts but also in industry or in sports. Yoshinori Maekawa, Takuya Oda, Taku Komura, Yoshihisa Shinagawa |
VRST | 3 |
| 2006 | Real-time locomotion control by sensing glovesabstractAbstract Sensing gloves are often used as an input device for virtual 3D games. We propose a new method to control characters such as humans or animals in real‐time by using sensing gloves. Based on existing motion data of the body, a new method to map the hand motion of the user to the locomotion of 3D characters in real‐time is proposed. The method was applied to control locomotion of characters such as humans or dogs. Various motions such as trotting, running, hopping, and turning could be produced. As the computational cost needed for our method is low, the response of the system is short enough to satisfy the real‐time requirements that are essential to be used for games. Using our method, users can directly control their characters intuitively and precisely than previous controlling devices such as mouse, keyboards or joysticks. Copyright © 2006 John Wiley & Sons, Ltd. Taku Komura, Wai-Chun Lam |
Comput. Animat. Virtual Worlds | 1 |
| 2005 | A Feedback Controller for Biped Humanoids that Can Counteract Large Perturbations During GaitabstractIn this paper, we propose a new method for biped humanoids to compensate for large amounts of angular momentum induced by strong external perturbations applied to the body during gait motion. Such angular momentum can easily cause the humanoid to fall down onto the ground. We use an Angular Momentum inducing inverted Pendulum Model (AMPM), which is an enhanced version of the 3D linear inverted pendulum model to model the robot dynamics. Because the AMPM allows us to explicitly calculate the angular momentum generated by the ground reaction force, it is possible to calculate a counteracting motion that compensates for the angular momentum generated by external perturbations in real-time. Taku Komura, Howard Leung, Shunsuke Kudoh, James J. Kuffner |
ICRA | 1 |
| 2005 | Computing inverse kinematics with linear programmingabstractInverse Kinematics (IK) is a popular technique for synthesizing motions of virtual characters. In this paper, we propose a Linear Programming based IK solver (LPIK) for interactive control of arbitrary multibody structures. There are several advantages of using LPIK. First, inequality constraints can be handled, and therefore the ranges of the DOFs and collisions of the body with other obstacles can be handled easily. Second, the performance of LPIK is comparable or sometimes better than the IK method based on Lagrange multipliers, which is known as the best IK solver today. The computation time by LPIK increases only linearly proportional to the number of constraints or DOFs. Hence, LPIK is a suitable approach for controlling articulated systems with large DOFs and constraints for real-time applications. Edmond S. L. Ho, Taku Komura, Rynson W. H. Lau |
VRST | 2 |
| 2005 | Animating reactive motion using momentum-based inverse kinematicsabstractAbstract Interactive generation of reactive motions for virtual humans as they are hit, pushed and pulled are very important to many applications, such as computer games. In this paper, we propose a new method to simulate reactive motions during arbitrary bipedal activities, such as standing, walking or running. It is based on momentum based inverse kinematics and motion blending. When generating the animation, the user first imports the primary motion to which the perturbation is to be applied to. According to the condition of the impact, the system selects a reactive motion from the database of pre‐captured stepping and reactive motions. It then blends the selected motion into the primary motion using momentum‐based inverse kinematics. Since the reactive motions can be edited in real‐time, the criteria for motion search can be much relaxed than previous methods, and therefore, the computational cost for motion search can be reduced. Using our method, it is possible to generate reactive motions by applying external perturbations to the characters at arbitrary moment while they are performing some actions. Copyright © 2005 John Wiley & Sons, Ltd. Taku Komura, Edmond S. L. Ho, Rynson W. H. Lau |
Comput. Animat. Virtual Worlds | 1 |
| 2004 | Animating reactive motions for biped locomotionabstractIn this paper, we propose a new method for simulating reactive motions for running or walking human figures. The goal is to generate realistic animations of how humans compensate for large external forces and maintain balance while running or walking. We simulate the reactive motions of adjusting the body configuration and altering footfall locations in response to sudden external disturbance forces on the body. With our proposed method, the user first imports captured motion data of a run or walk cycle to use as the primary motion. While executing the primary motion, an external force is applied to the body. The system automatically calculates a reactive motion for the center of mass and angular momentum around the center of mass using an enhanced version of the linear inverted pendulum model. Finally, the trajectories of the generalized coordinates that realize the precalculated trajectories of the center of mass, zero moment point, and angular momentum are obtained using constrained inverse kinematics. The advantage of our method is that it is possible to calculate reactive motions for bipeds that preserve dynamic balance during locomotion, which was difficult using previous techniques. We demonstrate our results on an application that allows a user to interactively apply external perturbations to a running or walking virtual human model. We expect this technique to be useful for human animations in interactive 3D systems such as games, virtual reality, and potentially even the control of actual biped robots. Taku Komura, Howard Leung, James J. Kuffner |
VRST | 1 |
| 2003 | An inverse kinematics method for 3D figures with motion dataabstractWe present a new inverse kinematics method that utilizes the motion data for realtime control and editing. The key idea is to extract parameters necessary for inverse kinematics from the motion data. These parameters are the weight matrix, which determines the motion of the redundant joints, and the transformation functions that define the motion of the end effectors. The user can control the motion by dragging a body segment using a mouse, and the method calculates the new motion using the precomputed parameters. The method enables interactive editing, warping, and retargeting character motions. Taku Komura, Atsushi Kuroda, Shunsuke Kudoh, Chiew-Lan Tai, Yoshihisa Shinagawa |
Computer Graphics International | 1 |
| 2003 | C2 continuous gait-pattern generation for biped robotsabstractIn this paper, we propose a new method to generate C/sup 2/ continuous gait motion for biped robots. The method is based on the enhanced inverted pendulum mode, which can easily handle angular momentum around the center of gravity. Using our method, it is possible to plan motion paths for biped robots without discontinuity in the acceleration even during switching from single support phase to double support phase, and vice versa. Shunsuke Kudoh, Taku Komura |
IROS | 2 |
| 2002 | The dynamic postural adjustment with the quadratic programming methodabstractThe postural balance system is one of the most fundamental functions for humanoid robot control. In this paper, we propose a new feedback balance control system for the human body. This system can manipulate large perturbations. It finds the optimal motion for maintaining balance in the 3D space without receiving any feed-forward input beforehand. Two different strategies are adopted for the optimization: the quadratic programming method and the PD control. Simulation results are compared with real human motion; many common features such as rotating arms are observed. Shunsuke Kudoh, Taku Komura, Katsushi Ikeuchi |
IROS | 2 |
| 2001 | An Inverse Kinematics Method Based on Muscle DynamicsabstractInverse kinematics is one of the most popular method in computer graphics to control 3D multi-joint characters. In this paper, we propose an inverse kinematics algorithm that takes the characteristics of human bodies into account. The mausculoskeletal model is used to solve the redundancy of the human body. Using our method, feasible human body motion can be obtained simply by specifying the motion of several end effectors or body segments. Since muscle dynamics is taken into account, the configuration space of the human body is automatically calculated, and unrealistic postures can be avoided. It is also possible to tune the motion by changing the external load applied to the muscles. Using our method, the amount of work by the animators is reduced to create natural human animation. Taku Komura, Yoshihisa Shinagawa, Tosiyasu L. Kunii |
Computer Graphics International | 1 |
| 2001 | Motion Conversion based on the Musculoskeletal System
Taku Komura, Yoshihisa Shinagawa |
Graphics Interface | 1 |
| 2001 | Topology matching for fully automatic similarity estimation of 3D shapesabstractThere is a growing need to be able to accurately and efficiently search visual data sets, and in particular, 3D shape data sets. This paper proposes a novel technique, called Topology Matching, in which similarity between polyhedral models is quickly, accurately, and automatically calculated by comparing Multiresolutional Reeb Graphs (MRGs). The MRG thus operates well as a search key for 3D shape data sets. In particular, the MRG represents the skeletal and topological structure of a 3D shape at various levels of resolution. The MRG is constructed using a continuous function on the 3D shape, which may preferably be a function of geodesic distance because this function is invariant to translation and rotation and is also robust against changes in connectivities caused by a mesh simplification or subdivision. The similarity calculation between 3D shapes is processed using a coarse-to-fine strategy while preserving the consistency of the graph structures, which results in establishing a correspondence between the parts of objects. The similarity calculation is fast and efficient because it is not necessary to determine the particular pose of a 3D shape, such as a rotation, in advance. Topology Matching is particularly useful for interactively searching for a 3D object because the results of the search fit human intuition well. Masaki Hilaga, Yoshihisa Shinagawa, Taku Komura, Tosiyasu L. Kunii |
SIGGRAPH | 3 |
| 2000 | Creating and retargetting motion by the musculoskeletal human body model
Taku Komura, Yoshihisa Shinagawa, Tosiyasu L. Kunii |
Vis. Comput. | 1 |
| 1999 | Calculation and visualization of the dynamic ability of the human bodyabstractThere is a great demand for data on the mobility and strength capability of the human body in many areas, such as ergonomics, medical engineering, biomechanical engineering, computer graphics (CG) and virtual reality (VR). This paper proposes a new method that enables the calculation of the maximal force exertable and acceleration performable by a human body during arbitrary motion. A musculoskeletal model of the legs is used for the calculation. Using our algorithm, it is possible to evaluate whether a given posture or motion is a feasible one. A tool to visualize the calculated maximal feasibility of each posture is developed. The obtained results can be used as criteria of manipulability or strength capability of the human body, important in ergonomics and human animation. Since our model is muscle-based, it is possible to simulate and visualize biomechanical effects such as fatigue and muscle training. The solution is based on linear programming and the results can be obtained in real time. Copyright © 1999 John Wiley & Sons, Ltd. Taku Komura, Yoshihisa Shinagawa, Tosiyasu L. Kunii |
Comput. Animat. Virtual Worlds | 1 |
| 1997 | A Muscle-based Feed-forward Controller of the Human BodyabstractThere is an increasing demand for human body motion data. Motion capture and physical animation have been used to generate such data. It is, however, apparent that such methods cannot automatically generate arbitrary human body motions. A human body is a redundant multi‐linked body controlled by a number of muscles. For this reason, the muscles must work appropriately and cooperatively for controlling the whole body. It is well‐known that the human body control system is composed of two parts: The open‐loop feed‐forward control system and the closed‐loop feedback control system. Many researchers have investigated the characteristics of the latter by analyzing the response of a human body to various external perturbations. However, for the former, very few studies have been done. This paper proposes an open‐loop feed‐forward model of the lower extremities which includes postural control for accurate animation of a human body. Assumptions are made here that the feed‐forward controller minimizes a certain objective value while keeping the balance of the body stable. The actual human motion data obtained using a motion capturing technique is compared with the trajectory calculated using our method for verification. The best criteria which is based on muscle dynamics is proposed. Using our method, dynamically correct human animation can be created by merely specifying a few key postures. Taku Komura, Yoshihisa Shinagawa, Tosiyasu L. Kunii |
Comput. Graph. Forum | 1 |