Takaaki Shiratori

dblp:17/5270 · DBLP profile ↗
← Back
51ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-1012-415XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 11 since 2021Systems, architecture and hardware · 8 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 CHOICE: Coordinated Human-Object Interaction in Cluttered Environments for Pick-and-Place Actions
abstract
Animating human-scene interactions such as picking and placing a wide range of objects with different geometries is a challenging task, especially in a cluttered environment where interactions with complex articulated containers are involved. The main difficulty lies in the sparsity of the motion data compared to the wide variation of the objects and environments, as well as the poor availability of transition motions between different actions, increasing the complexity of the generalization to arbitrary conditions. To cope with this issue, we develop a system that tackles the interaction synthesis problem as a hierarchical goal-driven task. Firstly, we develop a bimanual scheduler that plans a set of keyframes for simultaneously controlling the two hands to efficiently achieve the pick-and-place task from an abstract goal signal such as the target object selected by the user. Next, we develop a neural implicit planner that generates hand trajectories to guide reaching and leaving motions across diverse object shapes/types and obstacle layouts. Finally, we propose a linear dynamic model for our DeepPhase controller that incorporates a Kalman filter to enable smooth transitions in the frequency domain, resulting in a more realistic and effective multi-objective control of the character. Our system can synthesize a rich variety of natural pick-and-place movements that adapt to different object geometries, container articulations, and scene layouts.
Jintao Lu, Yuting Ye, Takaaki Shiratori, Sebastian Starke, Taku Komura
ACM Trans. Graph.4
2025 Generative Modeling of Shape-Dependent Self-Contact Human Poses
Takehiko Ohkawa, Shunsuke Saito, Jason M. Saragih, Fabian Prada, Shoou-I Yu, Ryosuke Furuta, Yoichi Sato 0001, Takaaki Shiratori
ICCV10
2025 ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling
abstract
Parametric body models offer expressive 3D representation of humans across a wide range of poses, shapes, and facial expressions, typically derived by learning a basis over registered 3D meshes. However, existing human mesh modeling approaches struggle to capture detailed variations across diverse body poses and shapes, largely due to limited training data diversity and restrictive modeling assumptions. Moreover, the common paradigm first optimizes the external body surface using a linear basis, then regresses internal skeletal joints from surface vertices. This approach introduces problematic dependencies between internal skeleton and outer soft tissue, limiting direct control over body height and bone lengths. To address these issues, we present ATLAS, a high-fidelity body model learned from 600k high-resolution scans captured using 240 synchronized cameras. Unlike previous methods, we explicitly decouple the shape and skeleton bases by grounding our mesh representation in the human skeleton. This decoupling enables enhanced shape expressivity, fine-grained customization of body attributes, and keypoint fitting independent of external soft-tissue characteristics. ATLAS outperforms existing methods by fitting unseen subjects in diverse poses more accurately, and quantitative evaluations show that our non-linear pose correctives more effectively capture complex poses compared to linear models.
Jinhyung Park, Javier Romero 0002, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Federica Bogo, Shoou-I Yu, Kris Makoto Kitani, Rawal Khirodkar
ICCV5
2024 Diffusion Shape Prior for Wrinkle-Accurate Cloth Registration
abstract
Registering clothes from 4D scans with vertex-accurate correspondence is challenging, yet important for dynamic appearance modeling and physics parameter estimation from real-world data. However, previous methods either rely on texture information, which is not always reliable, or achieve only coarse-level alignment. In this work, we present a novel approach to enabling accurate surface registration of texture-less clothes with large deformation. Our key idea is to effectively leverage a shape prior learned from pre-captured clothing using diffusion models. We also propose a multi-stage guidance scheme based on learned functional maps, which stabilizes registration for large-scale deformation even when they vary significantly from training data. Using high-fidelity real captured clothes, our experiments show that the proposed approach based on diffusion models generalizes better than surface registration with VAE or PCA-based priors, outperforming both optimization-based and learning-based non-rigid registration methods for both interpolation and extrapolation tests.
Jingfan Guo, Fabian Prada, Donglai Xiang, Javier Romero 0002, Chenglei Wu, Hyun Soo Park, Takaaki Shiratori, Shunsuke Saito
3DV7
2024 Authentic Hand Avatar from a Phone Scan via Universal Hand Model
abstract
The authentic 3D hand avatar with every identifiable information, such as hand shapes and textures, is necessary for immersive experiences in AR/VR. In this paper, we present a universal hand model (UHM), which 1) can universally represent high-fidelity 3D hand meshes of arbitrary identities (IDs) and 2) can be adapted to each person with a short phone scan for the authentic hand avatar. For effective universal hand modeling, we perform tracking and modeling at the same time, while previous 3D hand models perform them separately. The conventional separate pipeline suffers from the accumulated errors from the tracking stage, which cannot be recovered in the modeling stage. On the other hand, ours does not suffer from the accumulated errors while having a much more concise overall pipeline. We additionally introduce a novel image matching loss function to address a skin sliding during the tracking and modeling, while existing works have not focused on it much. Finally, using learned priors from our UHM, we effectively adapt our UHM to each person's short phone scan for the authentic hand avatar.
Gyeongsik Moon, Weipeng Xu, Rohan Joshi, Chenglei Wu, Takaaki Shiratori
CVPR5
2024 High-Fidelity Modeling of Generalizable Wrinkle Deformation
Jingfan Guo, Jae Shin Yoon, Shunsuke Saito, Takaaki Shiratori, Hyun Soo Park
ECCV (81)4
2024 Expressive Whole-Body 3D Gaussian Avatar
Gyeongsik Moon, Takaaki Shiratori, Shunsuke Saito
ECCV (41)2
2024 3D Hand Sequence Recovery from Real Blurry Images and Event Stream
Joonkyu Park, Gyeongsik Moon, Weipeng Xu, Evan Kaseman, Takaaki Shiratori, Kyoung Mu Lee
ECCV (59)5
2024 Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars
abstract
To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (i.e., easily adaptable to novel identities).Towards these goals, paired captures, that is, captures of the same subject obtained from systems of diverse quality and availability, are crucial.However, paired captures are rarely available to researchers outside of dedicated industrial labs: Codec Avatar Studio is our proposal to close this gap.Towards generalization and driveability, we introduce a dataset of 256 subjects captured in two modalities: high resolution multi-view scans of their heads, and video from the internal cameras of a headset.Towards completeness, we introduce a dataset of 4 subjects captured in eight modalities: high quality relightable multi-view captures of heads and hands, full body multi-view captures with minimal and regular clothes, and corresponding head, hands and body phone captures.Together with our data, we also provide code and pre-trained models for different state-of-the-art human generation models.Our datasets and code are available at https://github.com/facebookresearch/ava-256 and https://github.com/facebookresearch/goliath.
Julieta Martinez 0001, Emily Kim, Javier Romero 0002, Timur M. Bagautdinov, Shunsuke Saito, Shoou-I Yu, Michael Zollhöfer, Te-Li Wang, Shaojie Bai, Chenghui Li, Shih-En Wei, Rohan Joshi, Wyatt Borsos, Tomas Simon, Jason M. Saragih, Paul Theodosis, Alexander Greene, Anjani Josyula, Silvio Maeta, Andrew Jewett, Simion Venshtain, Christopher Heilman, Yueh-Tung Chen, Sidi Fu, Mohamed Elshaer, Tingfang Du, Longhua Wu, Shen-Chi Chen, Youssef Emad, Steven Longay, Ashley Brewer, Hitesh Shah, Taylor Koska, Kayla Haidle, Matthew Andromalos, Joanna Hsu, Thomas Dauer, Peter Selednik, Timothy Godisart, Scott Ardisson, Matthew Cipperly, Ben Humberston, Lon Farr, Bob Hansen, Peihong Guo, Dave Braun, Steven Krenn, He Wen 0001, Lucas Evans, Natalia Fadeeva, Matthew Stewart, Gabriel Schwartz, Divam Gupta, Gyeongsik Moon, Takaaki Shiratori, Fabian Prada, Bernardo Pires, Julia Buffalini, Autumn Trimble, Kevyn McPhail, Melissa Schoeller, Yaser Sheikh
NeurIPS62
2023 RelightableHands: Efficient Neural Relighting of Articulated Hand Models
abstract
We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a teacher-student framework, where the teacher learns appearance under a single point light from images captured in a light-stage, allowing us to synthesize hands in arbitrary illuminations but with heavy compute. Using images rendered by the teacher model as training data, an efficient student model directly predicts appearance under natural illuminations in real-time. To achieve generalization, we condition the student model with physics-inspired illumination features such as visibility, diffuse shading, and specular reflections computed on a coarse proxy geometry, maintaining a small computational overhead. Our key insight is that these features have strong correlation with subsequent global light transport effects, which proves sufficient as conditioning data for the neural relighting network. Moreover, in contrast to bottleneck illumination conditioning, these features are spatially aligned based on underlying geometry, leading to better generalization to unseen illuminations and poses. In our experiments, we demonstrate the efficacy of our illumination feature representations, outperforming baseline approaches. We also show that our approach can photorealistically relight two interacting hands at real-time speeds. https://sh8.io/#/relightable_hands
Shun Iwase, Shunsuke Saito, Tomas Simon, Stephen Lombardi, Timur M. Bagautdinov, Rohan Joshi, Fabian Prada, Takaaki Shiratori, Yaser Sheikh, Jason M. Saragih
CVPR8
2023 A Dataset of Relighted 3D Interacting Hands
abstract
The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image appearances and 2) diverse and large-scale groundtruth (GT) 3D poses at the same time. In this work, we propose Re:InterHand, a dataset of relighted 3D interacting hands that achieve the two goals. To this end, we employ a state-of-the-art hand relighting network with our accurately tracked two-hand 3D poses. We compare our Re:InterHand with existing 3D interacting hands datasets and show the benefit of it. Our Re:InterHand is available in https://mks0601.github.io/ReInterHand/
Gyeongsik Moon, Shunsuke Saito, Weipeng Xu, Rohan Joshi, Julia Buffalini, Harley Bellan, Nicholas Rosen, Jesse Richardson, Mallorie Mize, Philippe de Bree, Tomas Simon, Shubham Garg, Kevyn McPhail, Takaaki Shiratori
NeurIPS15
2023 BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
abstract
Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the stochastic nature of the problem and the lack of a rich cross-modal dataset that is needed for training. In this paper, we propose a novel transformer-based framework for automatic 3D body gesture synthesis from speech. To learn the stochastic nature of the body gesture during speech, we propose a variational transformer to effectively model a probabilistic distribution over gestures, which can produce diverse gestures during inference. Furthermore, we introduce a mode positional embedding layer to capture the different motion speeds in different speaking modes. To cope with the scarcity of data, we design an intra-modal pre-training scheme that can learn the complex mapping between the speech and the 3D gesture from a limited amount of data. Our system is trained with either the Trinity speech-gesture dataset or the Talking With Hands 16.2M dataset. The results show that our system can produce more realistic, appropriate, and diverse body gestures compared to existing state-of-the-art approaches.
Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost, Takaaki Shiratori, Junichi Yamagishi, Taku Komura
ACM Trans. Graph.5
2022 3D Clothed Human Reconstruction in the Wild
Gyeongsik Moon, Hyeongjin Nam, Takaaki Shiratori, Kyoung Mu Lee
ECCV (2)3
2022 N-Penetrate: Active Learning of Neural Collision Handler for Complex 3D Mesh Deformations
abstract
We present a robust learning algorithm to detect and handle collisions in 3D deforming meshes. We first train a neural network to detect collisions and then use a numerical optimization algorithm to resolve penetrations guided by the network. Our learned collision handler can resolve collisions for unseen, high-dimensional meshes with thousands of vertices. To obtain stable network performance in such large and unseen spaces, we apply active learning by progressively inserting new collision data based on the network inferences. We automatically label these new data using an analytical collision detector and progressively fine-tune our detection networks. We evaluate our method for collision handling of complex, 3D meshes coming from several datasets with different shapes and topologies, including datasets corresponding to dressed and undressed human poses, cloth simulations, and human hand poses acquired using multi-view capture systems.
Qingyang Tan, Zherong Pan, Breannan Smith, Takaaki Shiratori, Dinesh Manocha
ICML4
2022 Pattern-Based Cloth Registration and Sparse-View Animation
abstract
We propose a novel multi-view camera pipeline for the reconstruction and registration of dynamic clothing. Our proposed method relies on a specifically designed pattern that allows for precise video tracking in each camera view. We triangulate the tracked points and register the cloth surface in a fine-grained geometric resolution and low localization error. Compared to state-of-the-art methods, our registration exhibits stable correspondence, tracking the same points on the deforming cloth surface along the temporal sequence. As an application, we demonstrate how the use of our registration pipeline greatly improves state-of-the-art pose-based drivable cloth models. Furthermore, we propose a novel model, Garment Avatar , for driving cloth from a dense tracking signal which is obtained from two opposing camera views. The method produces realistic reconstructions which are faithful to the actual geometry of the deforming cloth. In this setting, the user wears a garment with our custom pattern which enables our driving model to reconstruct the geometry. Our code and data are available at https://github.com/HalimiOshri/Pattern-Based-Cloth-Registration-and-Sparse-View-Animation. The released data includes our pattern and registered mesh sequences containing four different subjects and 15k frames in total.
Oshri Halimi, Tuur Stuyck, Donglai Xiang, Timur M. Bagautdinov, He Wen 0001, Ron Kimmel, Takaaki Shiratori, Chenglei Wu, Yaser Sheikh, Fabian Prada
ACM Trans. Graph.7
2022 Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated Clothing
abstract
Despite recent progress in developing animatable full-body avatars, realistic modeling of clothing - one of the core aspects of human self-expression - remains an open challenge. State-of-the-art physical simulation methods can generate realistically behaving clothing geometry at interactive rates. Modeling photorealistic appearance, however, usually requires physically-based rendering which is too expensive for interactive applications. On the other hand, data-driven deep appearance models are capable of efficiently producing realistic appearance, but struggle at synthesizing geometry of highly dynamic clothing and handling challenging body-clothing configurations. To this end, we introduce pose-driven avatars with explicit modeling of clothing that exhibit both photorealistic appearance learned from real-world data and realistic clothing dynamics. The key idea is to introduce a neural clothing appearance model that operates on top of explicit geometry: at training time we use high-fidelity tracking, whereas at animation time we rely on physically simulated geometry. Our core contribution is a physically-inspired appearance network, capable of generating photorealistic appearance with view-dependent and dynamic shadowing effects even for unseen body-clothing configurations. We conduct a thorough evaluation of our model and demonstrate diverse animation results on several subjects and different types of clothing. Unlike previous work on photorealistic full-body avatars, our approach can produce much richer dynamics and more realistic deformations even for many examples of loose clothing. We also demonstrate that our formulation naturally allows clothing to be used with avatars of different people while staying fully animatable, thus enabling, for the first time, photorealistic avatars with novel clothing.
Donglai Xiang, Timur M. Bagautdinov, Tuur Stuyck, Fabian Prada, Javier Romero 0002, Weipeng Xu, Shunsuke Saito, Jingfan Guo, Breannan Smith, Takaaki Shiratori, Yaser Sheikh, Jessica K. Hodgins, Chenglei Wu
ACM Trans. Graph.10
2021 Driving-signal aware full-body avatars
abstract
We present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driving signals, such as human pose and facial keypoints, and produces a high-quality representation of human geometry and view-dependent appearance. The core intuition behind our method is that better drivability and generalization can be achieved by disentangling the driving signals and remaining generative factors, which are not available during animation. To this end, we explicitly account for information deficiency in the driving signal by introducing a latent space that exclusively captures the remaining information, thus enabling the imputation of the missing factors required during full-body animation, while remaining faithful to the driving signal. We also propose a learnable localized compression for the driving signal which promotes better generalization, and helps minimize the influence of global chance-correlations often found in real datasets. For a given driving signal, the resulting variational model produces a compact space of uncertainty for missing factors that allows for an imputation strategy best suited to a particular application. We demonstrate the efficacy of our approach on the challenging problem of full-body animation for virtual telepresence with driving signals acquired from minimal sensors placed in the environment and mounted on a VR-headset.
Timur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabian Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, Jason M. Saragih
ACM Trans. Graph.5
2021 ManipNet: neural manipulation synthesis with a hand-object spatial representation
abstract
Natural hand manipulations exhibit complex finger maneuvers adaptive to object shapes and the tasks at hand. Learning dexterous manipulation from data in a brute force way would require a prohibitive amount of examples to effectively cover the combinatorial space of 3D shapes and activities. In this paper, we propose a hand-object spatial representation that can achieve generalization from limited data. Our representation combines the global object shape as voxel occupancies with local geometric details as samples of closest distances. This representation is used by a neural network to regress finger motions from input trajectories of wrists and objects. Specifically, we provide the network with the current finger pose, past and future trajectories, and the spatial representations extracted from these trajectories. The network then predicts a new finger pose for the next frame as an autoregressive model. With a carefully chosen hand-centric coordinate system, we can handle single-handed and two-handed motions in a unified framework. Learning from a small number of primitive shapes and kitchenware objects, the network is able to synthesize a variety of finger gaits for grasping, in-hand manipulation, and bimanual object handling on a rich set of novel shapes and functional tasks. We also demonstrate a live demo of manipulating virtual objects in real-time using a simple physical prop. Our system is useful for offline animation or real-time applications forgiving to a small delay.
Yuting Ye, Takaaki Shiratori, Taku Komura
ACM Trans. Graph.3
2020 Learning 3D Global Human Motion Estimation from Unpaired, Disjoint Datasets
Julian Habekost, Takaaki Shiratori, Yuting Ye, Taku Komura
BMVC2
2020 DeepHandMesh: A Weakly-Supervised Deep Encoder-Decoder Framework for High-Fidelity Hand Mesh Modeling
Gyeongsik Moon, Takaaki Shiratori, Kyoung Mu Lee
ECCV (2)2
2020 InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image
Gyeongsik Moon, Shoou-I Yu, He Wen 0001, Takaaki Shiratori, Kyoung Mu Lee
ECCV (20)4
2020 Constraining dense hand surface tracking with elasticity
abstract
Many of the actions that we take with our hands involve self-contact and occlusion: shaking hands, making a fist, or interlacing our fingers while thinking. This use of of our hands illustrates the importance of tracking hands through self-contact and occlusion for many applications in computer vision and graphics, but existing methods for tracking hands and faces are not designed to treat the extreme amounts of self-contact and self-occlusion exhibited by common hand gestures. By extending recent advances in vision-based tracking and physically based animation, we present the first algorithm capable of tracking high-fidelity hand deformations through highly self-contacting and self-occluding hand gestures, for both single hands and two hands. By constraining a vision-based tracking algorithm with a physically based deformable model, we obtain an algorithm that is robust to the ubiquitous self-interactions and massive self-occlusions exhibited by common hand gestures, allowing us to track two hand interactions and some of the most difficult possible configurations of a human hand.
Breannan Smith, Chenglei Wu, He Wen 0001, Patrick Peluse, Yaser Sheikh, Jessica K. Hodgins, Takaaki Shiratori
ACM Trans. Graph.7
2019 Lost in Style: Gaze-driven Adaptive Aid for VR Navigation
abstract
A key challenge for virtual reality level designers is striking a balance between maintaining the immersiveness of VR and providing users with on-screen aids after designing a virtual experience. These aids are often necessary for wayfinding in virtual environments with complex paths. We introduce a novel adaptive aid that maintains the effectiveness of traditional aids, while equipping designers and users with the controls of how often help is displayed. Our adaptive aid uses gaze patterns in predicting user's need for navigation aid in VR and displays mini-maps or arrows accordingly. Using a dataset of gaze angle sequences of users navigating a VR environment and markers of when users requested aid, we trained an LSTM to classify user's gaze sequences as needing navigation help and display an aid. We validated the efficacy of the adaptive aid for wayfinding compared to other commonly-used wayfinding aids.
Rawan Alghofaili, Yasuhito Sawahata, Haikun Huang, Hsueh-Cheng Wang, Takaaki Shiratori, Lap-Fai Yu
CHI5
2019 Self-Supervised Adaptation of High-Fidelity Face Models for Monocular Performance Tracking
abstract
Improvements in data-capture and face modeling techniques have enabled us to create high-fidelity realistic face models. However, driving these realistic face models requires special input data, e.g., 3D meshes and unwrapped textures. Also, these face models expect clean input data taken under controlled lab environments, which is very different from data collected in the wild. All these constraints make it challenging to use the high-fidelity models in tracking for commodity cameras. In this paper, we propose a self-supervised domain adaptation approach to enable the animation of high-fidelity face models from a commodity camera. Our approach first circumvents the requirement for special input data by training a new network that can directly drive a face model just from a single 2D image. Then, we overcome the domain mismatch between lab and uncontrolled environments by performing self-supervised domain adaptation based on ``consecutive frame texture consistency'' based on the assumption that the appearance of the face is consistent over consecutive frames, avoiding the necessity of modeling the new environment such as lighting or background. Experiments show that we are able to drive a high-fidelity face model to perform complex facial motion from a cellphone camera without requiring any labeled data from the new domain.
Jae Shin Yoon, Takaaki Shiratori, Shoou-I Yu, Hyun Soo Park
CVPR2
2019 Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis
abstract
We present a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations. Our dataset features synchronized body and finger motion as well as audio data. To the best of our knowledge, it represents the largest motion capture and audio dataset of natural conversations to date. The statistical analysis verifies strong intraperson and interperson covariance of arm, hand, and speech features, potentially enabling new directions on data-driven social behavior analysis, prediction, and synthesis. As an illustration, we propose a novel real-time finger motion synthesis method: a temporal neural network innovatively trained with an inverse kinematics (IK) loss, which adds skeletal structural information to the generative model. Our qualitative user study shows that the finger motion generated by our method is perceived as natural and conversation enhancing, while the quantitative ablation study demonstrates the effectiveness of IK loss.
Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha S. Srinivasa, Yaser Sheikh
ICCV4
2018 Deep incremental learning for efficient high-fidelity face tracking
abstract
In this paper, we present an incremental learning framework for efficient and accurate facial performance tracking. Our approach is to alternate the modeling step, which takes tracked meshes and texture maps to train our deep learning-based statistical model, and the tracking step, which takes predictions of geometry and texture our model infers from measured images and optimize the predicted geometry by minimizing image, geometry and facial landmark errors. Our Geo-Tex VAE model extends the convolutional variational autoencoder for face tracking, and jointly learns and represents deformations and variations in geometry and texture from tracked meshes and texture maps. To accurately model variations in facial geometry and texture, we introduce the decomposition layer in the Geo-Tex VAE architecture which decomposes the facial deformation into global and local components. We train the global deformation with a fully-connected network and the local deformations with convolutional layers. Despite running this model on each frame independently - thereby enabling a high amount of parallelization - we validate that our framework achieves sub-millimeter accuracy on synthetic data and outperforms existing methods. We also qualitatively demonstrate high-fidelity, long-duration facial performance tracking on several actors.
Chenglei Wu, Takaaki Shiratori, Yaser Sheikh
ACM Trans. Graph.2
2017 Perception Meets Examination: Studying Deceptive Behaviors in VR
Carla Aravena, Mark Vo, Tao Gao 0004, Takaaki Shiratori, Lap-Fai Yu
CogSci4
2017 DrawFromDrawings: 2D Drawing Assistance via Stroke Interpolation with a Sketch Database
abstract
We present DrawFromDrawings, an interactive drawing system that provides users with visual feedback for assistance in 2D drawing using a database of sketch images. Following the traditional imitation and emulation training from art education, DrawFromDrawings enables users to retrieve and refer to a sketch image stored in a database and provides them with various novel strokes as suggestive or deformation feedback. Given regions of interest (ROIs) in the user and reference sketches, DrawFromDrawings detects as-long-as-possible (ALAP) stroke segments and the correspondences between user and reference sketches that are the key to computing seamless interpolations. The stroke-level interpolations are parametrized with the user strokes, the reference strokes, and new strokes created by warping the reference strokes based on the user and reference ROI shapes, and the user study indicated that the interpolation could produce various reasonable strokes varying in shapes and complexity. DrawFromDrawings allows users to either replace their strokes with interpolated strokes (deformation feedback) or overlays interpolated strokes onto their strokes (suggestive feedback). The other user studies on the feedback modes indicated that the suggestive feedback enabled drawers to develop and render their ideas using their own stroke style, whereas the deformation feedback enabled them to finish the sketch composition quickly.
Yusuke Matsui 0001, Takaaki Shiratori, Kiyoharu Aizawa
IEEE Trans. Vis. Comput. Graph.2
2015 Efficient Large-Scale Point Cloud Registration Using Loop Closures
abstract
Alignment of many 3D point clouds, possibly captured by multiple devices at different times, is a critical step for increasingly popular applications such as 3D model construction and augmented reality. For very large data sets, traditional methods such as ICP can become computationally intractable, or produce poor results. We present an efficient method for accurately aligning very large numbers of dense 3D point clouds, and apply it to a city-scale data set. The method relies on the novel combination of 1) partitioning the point clouds based on loop structures detected across a combined network of all device capture paths, and 2) making use of the loop closure property to accurately align point clouds within each sub-problem. Final global alignment of the loop-based results is formulated as a least squares optimization with closed form solution. Experimental results are shown for aligning 3D points across the entire city of San Francisco with centimeter-scale accuracy, via an efficient parallelized architecture.
Takaaki Shiratori, Jérôme Berclaz, Michael Harville, Chintan Shah, Taoyu Li, Yasuyuki Matsushita, Stephen Shiller
3DV1
2015 An efficient volumetric method for non-rigid registration
Xuejin Chen, Takaaki Shiratori, Xin Tong 0001, Ligang Liu 0001
Graph. Model.3
2015 3D Trajectory Reconstruction under Perspective Projection
Hyun Soo Park, Takaaki Shiratori, Iain A. Matthews, Yaser Sheikh
Int. J. Comput. Vis.2
2015 Autocomplete hand-drawn animations
abstract
Hand-drawn animation is a major art form and communication medium, but can be challenging to produce. We present a system to help people create frame-by-frame animations through manual sketches. We design our interface to be minimalistic: it contains only a canvas and a few controls. When users draw on the canvas, our system silently analyzes all past sketches and predicts what might be drawn in the future across spatial locations and temporal frames. The interface also offers suggestions to beautify existing drawings. Our system can reduce manual workload and improve output quality without compromising natural drawing flow and control: users can accept, ignore, or modify such predictions visualized on the canvas by simple gestures. Our key idea is to extend the local similarity method in [Xing et al. 2014], which handles only low-level spatial repetitions such as hatches within a single frame, to a global similarity that can capture high-level structures across multiple frames such as dynamic objects. We evaluate our system through a preliminary user study and confirm that it can enhance both users' objective performance and subjective satisfaction.
Jun Xing, Li-Yi Wei, Takaaki Shiratori, Koji Yatani
ACM Trans. Graph.3
2014 Automatic Extraction of Moving Objects from Image and LIDAR Sequences
abstract
Detecting and segmenting moving objects in an image sequence has always been a crucial task for many computer vision applications. This task becomes especially challenging for real-world image sequences of busy street scenes, where moving objects are ubiquitous. Although it remains technologically elusive to develop an effective and scalable image-based moving object detection, modern street side imagery are often augmented with sparse point clouds captured with depth sensors. This paper develops a simple but effective system for moving object detection that fully harnesses the complementary nature of 2D image and 3D LIDAR point clouds. We demonstrate how moving objects can be much more easily and reliably detected with sparse 3D measurements and how such information can significantly improve segmentation for moving objects in the image sequences. The results of our system are highly accurate "joint segmentation" of 2D images and 3D points for all moving objects in street scenes, which can serve many subsequent tasks such as object removal in images, 3D reconstruction and rendering.
Jizhou Yan, Heesoo Myeong, Takaaki Shiratori
3DV4
2014 Extraction of person-specific motion style based on a task model and imitation by humanoid robot
abstract
In this paper, we present a humanoid robot which extracts and imitates the person-specific differences in motions, which we will call style. Synthesizing human-like and stylistic motion variations according to specific scenarios is becoming important for entertainment robots, and imitation of styles is one variation which makes robots more amiable. Our approach extends a learning from observation (LFO) paradigm which enables robots to understand what a human is doing and to extract reusable essences to be learned. The focus is on styles in the domain of LFO and the representation of them using the reusable essences. In this paper, we design an abstract model of a target motion defined in LFO, observe human demonstrations through the model, and formulate the representation of styles in the context of LFO. Then we introduce a framework of generating robot motions that reflect styles which are automatically extracted from human demonstrations. To verify our proposed method we applied it to a ring toss game, and generated robot motions for a physical humanoid robot. Styles from each of three random players were extracted automatically from their demonstrations, and used for generating robot motions. The robot imitates the styles of each player without exceeding the limitation of its physical constraints, while tossing the rings to the goal.
Takahiro Okamoto, Takaaki Shiratori, M. Glisson, K. Yamane, Shunsuke Kudoh, Katsushi Ikeuchi
IROS2
2014 Toward a Dancing Robot With Listening Capability: Keypose-Based Integration of Lower-, Middle-, and Upper-Body Motions for Varying Music Tempos
abstract
This paper presents the development toward a dancing robot that can listen to and dance along with musical performances. One of the key components of this robot is the ability to modify its dance motions with varying tempos, without exceeding motor limitations, in the same way that human dancers modify their motions. In this paper, we first observe human performances with varying musical tempos of the same musical piece, and then analyze human modification strategies. The analysis is conducted in terms of three body components: lower, middle, and upper bodies. We assume that these body components have different purposes and different modification strategies, respectively, for the performance of a dance. For all of the motions of these three components, we have found that certain fixed postures, which we call keyposes, tend to be preserved. Thus, this paper presents a method to create motions for robots at a certain music tempo, from human motion at an original music tempo, by using these keyposes. We have implemented these algorithms as an automatic process and validated their effectiveness by using a physical humanoid robot HRP-2. This robot succeeded in performing the Aizu-bandaisan dance, one of the Japanese traditional folk dances, 1.2 and 1.5 times faster than the tempo originally learned, while maintaining its physical constraints. Although we are not achieving a dancing robot which autonomously interacts with varying music tempos, we think that our method has a vital role in the dancing-to-music capability.
Takahiro Okamoto, Takaaki Shiratori, Shunsuke Kudoh, Shinichiro Nakaoka, Katsushi Ikeuchi
IEEE Trans. Robotics2
2013 HideOut: mobile projector interaction with tangible objects and surfaces
abstract
HideOut is a mobile projector-based system that enables new applications and interaction techniques with tangible objects and surfaces. HideOut uses a device mounted camera to detect hidden markers applied with infrared-absorbing ink. The obtrusive appearance of fiducial markers is avoided and the hidden marker surface doubles as a functional projection surface. We present example applications that demonstrate a wide range of interaction scenarios, including media navigation tools, interactive storytelling applications, and mobile games. We explore the design space enabled by the HideOut system and describe the hidden marker prototyping process. HideOut brings tangible objects to life for interaction with the physical world around us.
Karl D. D. Willis, Takaaki Shiratori, Moshe Mahler
TEI2
2013 BodyAvatar: creating freeform 3D avatars using first-person body gestures
abstract
BodyAvatar is a Kinect-based interactive system that allows users without professional skills to create freeform 3D avatars using body gestures. Unlike existing gesture-based 3D modeling tools, BodyAvatar centers around a first-person "you're the avatar" metaphor, where the user treats their own body as a physical proxy of the virtual avatar. Based on an intuitive body-centric mapping, the user performs gestures to their own body as if wanting to modify it, which in turn results in corresponding modifications to the avatar. BodyAvatar provides an intuitive, immersive, and playful creation experience for the user. We present a formative study that leads to the design of BodyAvatar, the system's interactions and underlying algorithms, and results from initial user trials.
Teng Han, Zhimin Ren, Nobuyuki Umetani, Xin Tong 0001, Yang Liu 0014, Takaaki Shiratori
UIST7
2011 Motionbeam: a metaphor for character interaction with handheld projectors
abstract
We present the MotionBeam metaphor for character interaction with handheld projectors. Our work draws from the tradition of pre-cinema handheld projectors that use direct physical manipulation to control projected imagery. With our prototype system, users interact and control projected characters by moving and gesturing with the handheld projector itself. This creates a unified interaction style where input and output are tied together within a single device. We introduce a set of interaction principles and present prototype applications that provide clear examples of the MotionBeam metaphor in use. Finally we describe observations and insights from a preliminary user study with our system.
Karl D. D. Willis, Ivan Poupyrev, Takaaki Shiratori
CHI3
2011 Motion capture from body-mounted cameras
abstract
Motion capture technology generally requires that recordings be performed in a laboratory or closed stage setting with controlled lighting. This restriction precludes the capture of motions that require an outdoor setting or the traversal of large areas. In this paper, we present the theory and practice of using body-mounted cameras to reconstruct the motion of a subject. Outward-looking cameras are attached to the limbs of the subject, and the joint angles and root pose are estimated through non-linear optimization. The optimization objective function incorporates terms for image matching error and temporal continuity of motion. Structure-from-motion is used to estimate the skeleton structure and to provide initialization for the non-linear optimization procedure. Global motion is estimated and drift is controlled by matching the captured set of videos to reference imagery. We show results in settings where capture would be difficult or impossible with traditional motion capture systems, including walking outside and swinging on monkey bars. The quality of the motion reconstruction is evaluated by comparing our results against motion capture data produced by a commercially available optical system.
Takaaki Shiratori, Hyun Soo Park, Leonid Sigal, Yaser Sheikh, Jessica K. Hodgins
ACM Trans. Graph.1
2010 3D Reconstruction of a Moving Point from a Series of 2D Projections
Hyun Soo Park, Takaaki Shiratori, Iain A. Matthews, Yaser Sheikh
ECCV (3)2
2010 Temporal scaling of leg motion for music feedback system of a dancing humanoid robot
abstract
In this paper, we propose a method to achieve temporal scaling of leg motions as a fundamental technique for a music feedback system of a dancing humanoid robot. We asked dancers to perform dance motion at normal musical tempo and faster musical tempos and observed how dancers modified performance for given musical tempos. The obtained insights from the observation are 1) a dancer needs to preserve leg postures that are important to emphasize dance expression, 2) there is a priority to determine what features of leg motion can be adjusted, and 3) stylistic leg motion resembles normal step motion if dancers cannot follow fast musical tempo completely. Based on these insights, we generate leg motion appropriately adjusted for changing musical tempo while maintaining balance. We validated our method via simulation experiments with a humanoid robot HRP-2.
Takahiro Okamoto, Takaaki Shiratori, Shunsuke Kudoh, Katsushi Ikeuchi
IROS2
2010 Detecting dance motion structure using body components and turning motions
abstract
This paper presents a novel method for robust dance motion structure detection. In the japanese folk dance domain, teachers created illustrations of dance poses. These poses characterize the most important movements of a dance. So far there is no simple and reliable extraction method which can extract all poses as shown in these drawings. We use these poses for the Task Model (TM) in the context of Learning from Observation (LFO). LFO which is a well known technique for successful human to robot motion mapping, consists of tasks (what to do) and skills (how to do). We propose a novel approach, to extract special motions from a dance, called turning motions useful for skill mapping in the LFO paradigm. Furthermore, we use a modified version of this approach, to detect all poses as shown in the drawings, called turning poses. To achieve this we observe both forearms at the same time and analyze their movement in different 2-D coordinate planes. We evaluate the parameters with and without a weighting function where we minimize acceleration, velocity and power. We successfully demonstrate this novel method using two very different japanese folk dances and discuss further implications of this work in respect to the LFO paradigm and dances of other domains.
Bjoern Rennhak, Takaaki Shiratori, Shunsuke Kudoh, Phongtharin Vinayavekhin, Katsushi Ikeuchi
IROS2
2010 Physically-Based Character Control in Low Dimensional Space
Hubert P. H. Shum, Taku Komura, Takaaki Shiratori, Shu Takagi
MIG3
2008 Accelerometer-based user interfaces for the control of a physically simulated character
abstract
In late 2006, Nintendo released a new game controller, the Wiimote, which included a three-axis accelerometer. Since then, a large variety of novel applications for these controllers have been developed by both independent and commercial developers. We add to this growing library with three performance interfaces that allow the user to control the motion of a dynamically simulated, animated character through the motion of his or her arms, wrists, or legs. For comparison, we also implement a traditional joystick/button interface. We assess these interfaces by having users test them on a set of tracks containing turns and pits. Two of the interfaces (legs and wrists) were judged to be more immersive and were better liked than the joystick/button interface by our subjects. All three of the Wiimote interfaces provided better control than the joystick interface based on an analysis of the failures seen during the user study.
Takaaki Shiratori, Jessica K. Hodgins
ACM Trans. Graph.1
2007 Humanoid Robot Painter: Visual Perception and High-Level Planning
abstract
This paper presents visual perception discovered in high-level manipulator planning for a robot to reproduce the procedure involved in human painting. First, we apply a technique of 2D object segmentation that considers region similarity as an objective function and edge as a constraint with artificial intelligent used as a criterion function. The system can segment images more effectively than most of existing methods, even if the foreground is very similar to the background. Second, we propose a novel color perception model that shows similarity to human perception. The method outperforms many existing color reduction algorithms. Third, we propose a novel global orientation map perception using a radial basis function. Finally, we use the derived model along with the brush's position- and force-sensing to produce a visual feedback drawing. Experiments show that our system can generate good paintings including portraits.
Miti Ruchanurucks, Shunsuke Kudoh, Koichi Ogawara, Takaaki Shiratori, Katsushi Ikeuchi
ICRA4
2007 Multilinear analysis for task recognition and person identification
abstract
This paper introduces a Multi Factor Tensor(MFT) model to recognize motion styles and person identities in dance sequences. We apply a musical information analysis method in segmenting the motion sequence relevant to the key poses and the musical rhythm. We define a task model considering the repeated motion segments, where the motion is decomposed into person invariant factor task and person dependant factor style. We capture the motion data of different people for a few cycles, segment it using the musical analysis approach, normalize the segments using a vectorization method, and realize our MFT model. The experiments are conducted according to two approaches. Various experiments that we conduct to evaluate the potential of the recognition ability of our proposed approaches and the results demonstrate the high accuracy of our model. The recognition results and the motion decomposition will be used in further extending the motion generation process in various styles and for different tasks.
Manoj Perera, Takaaki Shiratori, Shunsuke Kudoh, Atsushi Nakazawa, Katsushi Ikeuchi
IROS2
2007 Robot painter: from object to trajectory
abstract
This paper presents visual perception discovered in high-level manipulator planning for a robot to reproduce the procedure involved in human painting. First, we propose a technique of 3D object segmentation that can work well even when the precision of the cameras is inadequate. Second, we apply a simple yet powerful fast color perception model that shows similarity to human perception. The method outperforms many existing interactive color perception algorithms. Third, we generate global orientation map perception using a radial basis function. Finally, we use the derived foreground, color segments, and orientation map to produce a visual feedback drawing. Our main contributions are 3D object segmentation and color perception schemes.
Miti Ruchanurucks, Shunsuke Kudoh, Koichi Ogawara, Takaaki Shiratori, Katsushi Ikeuchi
IROS4
2007 Temporal scaling of upper body motion for Sound feedback system of a dancing humanoid robot
abstract
This paper proposes a method to model the modification of upper body motion of dance performance based on the speed of played music. When we observed structured dance motion performed at a normal music playback speed and motion performed at a faster music playback speed, we found that the detail of each motion is slightly different while the whole of the dance motion is similar in both cases. This phenomenon is derived from the fact that dancers omit the details and perform the essential part of the dance in order to follow the faster speed of the music. To clarify this phenomenon, we analyzed the motion differences in the frequency domain, and obtained two insights on the omission of motion details: (1) High frequency components are gradually attenuated depending on the musical speed, and (2) important stop motions are preserved even when high frequency components are attenuated. Based on these insights, we modeled our motion modification considering musical speed and joint limitations that a humanoid robot has. We show the effectiveness of our method via some applications for humanoid robot motion generation.
Takaaki Shiratori, Shunsuke Kudoh, Shinichiro Nakaoka, Katsushi Ikeuchi
IROS1
2006 Video Completion by Motion Field Transfer
abstract
Existing methods for video completion typically rely on periodic color transitions, layer extraction, or temporally local motion. However, periodicity may be imperceptible or absent, layer extraction is difficult, and temporally local motion cannot handle large holes. This paper presents a new approach for video completion using motion field transfer to avoid such problems. Unlike prior methods, we fill in missing video parts by sampling spatio-temporal patches of local motion instead of directly sampling color. Once the local motion field has been computed within the missing parts of the video, color can then be propagated to produce a seamless hole-free video. We have validated our method on many videos spanning a variety of scenes. We can also use the same approach to perform frame interpolation using motion fields from different videos.
Takaaki Shiratori, Yasuyuki Matsushita, Xiaoou Tang, Sing Bing Kang
CVPR (1)1
2006 Synthesizing Dance Performance using Musical and Motion Features
abstract
This paper proposes a method for synthesizing dance performance synchronized to played music and our method presents a system that imitates dancers' skills in performing their motion while they listen to the music. Our method consists of a motion analysis, a music analysis, and a motion synthesis based on results of the analyses. In these analysis steps, motion and music features are acquired. These features are derived from motion keyframes, motion intensity, music intensity, musical beats, and chord changes. Our system also constructs a motion graph to search similar poses from given dance sequences and to connect them as possible transitions. In the synthesis step, the trajectory that provides the best correlation between music and motion features is selected from the motion graph, and the resulting motion is generated. Our experimental results indicate that our proposed method actually creates dance as the system "hears" the music
Takaaki Shiratori, Atsushi Nakazawa, Katsushi Ikeuchi
ICRA1
2006 Dancing-to-Music Character Animation
abstract
Abstract In computer graphics, considerable research has been conducted on realistic human motion synthesis. However, most research does not consider human emotional aspects, which often strongly affect human motion. This paper presents a new approach for synthesizing dance performance matched to input music, based on the emotional aspects of dance performance. Our method consists of a motion analysis, a music analysis, and a motion synthesis based on the extracted features. In the analysis steps, motion and music feature vectors are acquired. Motion vectors are derived from motion rhythm and intensity, while music vectors are derived from musical rhythm, structure, and intensity. For synthesizing dance performance, we first find candidate motion segments whose rhythm features are matched to those of each music segment, and then we find the motion segment set whose intensity is similar to that of music segments. Additionally, our system supports having animators control the synthesis process by assigning desired motion segments to the specified music segments. The experimental results indicate that our method actually creates dance performance as if a character was listening and expressively dancing to the music. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism Animation; J.5 [Arts and Humanities]: Performing Arts Music
Takaaki Shiratori, Atsushi Nakazawa, Katsushi Ikeuchi
Comput. Graph. Forum1