Yusuke Yoshiyasu

dblp:40/9988 · DBLP profile ↗
← Back
28ranked-venue papers
13as first author
11since 2021 · last 2025
0000-0002-0433-9832ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 13 first-author · 4 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 10 since 2021Systems, architecture and hardware · 7 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 MeshMamba: State Space Models for Articulated 3D Mesh Generation and Reconstruction
abstract
In this paper, we introduce MeshMamba, a neural network model for learning 3D articulated mesh models by employing the recently proposed Mamba State Space Models (Mamba-SSMs). MeshMamba is efficient and scalable in handling a large number of input tokens, enabling the generation and reconstruction of body mesh models with more than 10,000 vertices, capturing clothing and hand geometries. The key to effectively learning MeshMamba is the serialization technique of mesh vertices into orderings that are easily processed by Mamba. This is achieved by sorting the vertices based on body part annotations or the 3D vertex locations of a template mesh, such that the ordering respects the structure of articulated shapes. Based on MeshMamba, we design 1) MambaDiff3D, a denoising diffusion model for generating 3D articulated meshes and 2) Mamba-HMR, a 3D human mesh recovery model that reconstructs a human body shape and pose from a single image. Experimental results showed that MambaDiff3D can generate dense 3D human meshes in clothes, with grasping hands, etc., and outperforms previous approaches in the 3D human shape generation task. Additionally, Mamba-HMR extends the capabilities of previous non-parametric human mesh recovery approaches, which were limited to handling body-only poses using around 500 vertex tokens, to the whole-body setting with face and hands, while achieving competitive performance in (near) real-time.
Yusuke Yoshiyasu, Leyuan Sun, Ryusuke Sagawa
ICCV1
2025 Enhancing multimodal-input object goal navigation by leveraging large language models for inferring room-object relationship knowledge
Leyuan Sun, Asako Kanezaki, Guillaume Caron, Yusuke Yoshiyasu
Adv. Eng. Informatics4
2025 Memory-MambaNav: Enhancing object-goal navigation through integration of spatial-temporal scanning with state space models
Leyuan Sun, Yusuke Yoshiyasu
Image Vis. Comput.2
2024 DiffSurf: A Transformer-Based Diffusion Model for Generating and Reconstructing 3D Surfaces in Pose
Yusuke Yoshiyasu, Leyuan Sun
ECCV (82)1
2024 NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance Fields
abstract
We present NeuralLabeling, a labeling approach and toolset for annotating 3D scenes using either bounding boxes or meshes and generating segmentation masks, affordance maps, 2D bounding boxes, 3D bounding boxes, 6DOF object poses, depth maps, and object meshes. NeuralLabeling uses Neural Radiance Fields (NeRF) as a renderer, allowing labeling to be performed using 3D spatial tools while incorporating geometric clues such as occlusions, relying only on images captured from multiple viewpoints as input. To demonstrate the applicability of NeuralLabeling to a practical problem in robotics, we added ground truth depth maps to 30000 frames of transparent object RGB and noisy depth maps of glasses placed in a dishwasher captured using an RGBD sensor, yielding the Dishwasher30k dataset. We show that training a simple deep neural network with supervision using the annotated depth maps yields a higher reconstruction performance than training with the previously applied weakly supervised approach. We also show how instance segmentation and depth completion datasets generated using NeuralLabeling can be incorporated into a robot application for grasping transparent objects placed in a dishwasher with an accuracy of 83.3%, compared to 16.3% without depth completion. Supplementary URI: https://florise.github.io/neural_labeling_web/.
Floris Erich, Naoya Chiba, Abdullah Mustafa, Yusuke Yoshiyasu, Noriaki Ando, Ryo Hanai, Yukiyasu Domae
IROS4
2024 PEGASUS: Physically Enhanced Gaussian Splatting Simulation System for 6DoF Object Pose Dataset Generation
abstract
We introduce Physically Enhanced Gaussian Splatting Simulation System (PEGASUS) for 6DoF object pose dataset generation, a versatile dataset generator based on 3D Gaussian Splatting. Environment and object representations can be easily obtained using commodity cameras to reconstruct with Gaussian Splatting. PEGASUS allows the composition of new scenes by merging the respective underlying Gaussian Splatting point cloud of an environment with one or multiple objects. Leveraging a physics engine enables the simulation of natural object placement within a scene through interaction between meshes extracted for the objects and the environment. Consequently, an extensive amount of new scenes - static or dynamic - can be created by combining different environments and objects. By rendering scenes from various perspectives, diverse data points such as RGB images, depth maps, semantic masks, and 6DoF object poses can be extracted. Our study demonstrates that training on data generated by PEGASUS enables pose estimation networks to successfully transfer from synthetic data to real-world data. Moreover, we introduce the Ramen dataset, comprising 30 Japanese cup noodle items. This dataset includes spherical scans that capture images from both the object hemisphere and the Gaussian Splatting reconstruction, making them compatible with PEGASUS.
Lukas Meyer, Floris Erich, Yusuke Yoshiyasu, Marc Stamminger, Noriaki Ando, Yukiyasu Domae
IROS3
2023 Deformable Mesh Transformer for 3D Human Mesh Recovery
abstract
We present Deformable mesh transFormer (DeFormer), a novel vertex-based approach to monocular 3D human mesh recovery. DeFormer iteratively fits a body mesh model to an input image via a mesh alignment feedback loop formed within a transformer decoder that is equipped with efficient body mesh driven attention modules: 1) body sparse self-attention and 2) deformable mesh cross attention. As a result, DeFormer can effectively exploit high-resolution image feature maps and a dense mesh model which were computationally expensive to deal with in previous approaches using the standard transformer attention. Experimental results show that DeFormer achieves state-of-the-art performances on the Human3.6M and 3DPW benchmarks. Ablation study is also conducted to show the effectiveness of the DeFormer model designs for leveraging multi-scale feature maps. Code is available at https://github.com/yusukey03012/DeFormer.
Yusuke Yoshiyasu
CVPR1
2022 Direct visual servoing in the non-linear scale space of camera pose
abstract
This paper proposes to consider direct visual servoing (DVS) for object manipulation by a robot arm. The convergence domain limits of DVS are overcome by introducing the non-linear scale space related to camera pose. Its use in a new, yet little complex, direct cost can enlarge twice the convergence domain of one state-of-the-art DVS, as many experiments of symmetric object orientation control assess.
Guillaume Caron, Yusuke Yoshiyasu
ICPR2
2022 Object Memory Transformer for Object Goal Navigation
abstract
This paper presents a reinforcement learning method for object goal navigation (ObjNav) where an agent navigates in 3D indoor environments to reach a target object based on long-term observations of objects and scenes. To this end, we propose Object Memory Transformer (OMT) that consists of two key ideas: 1) Object-Scene Memory (OSM) that enables to store long-term scenes and object semantics, and 2) Transformer that attends to salient objects in the sequence of previously observed scenes and objects stored in OSM. This mechanism allows the agent to efficiently navigate in the indoor environment without prior knowledge about the environments, such as topological maps or 3D meshes. To the best of our knowledge, this is the first work that uses a long-term memory of object semantics in a goal-oriented navigation task. Experimental results conducted on the AI2-THOR dataset show that OMT outperforms previous approaches in navigating in unknown environments. In particular, we show that utilizing the long-term object semantics information improves the efficiency of navigation.
Rui Fukushima, Kei Ota, Asako Kanezaki, Yoko Sasaki, Yusuke Yoshiyasu
ICRA5
2021 Rapid Pose Label Generation through Sparse Representation of Unknown Objects
abstract
Deep Convolutional Neural Networks (CNNs) have been successfully deployed on robots for 6-DoF object pose estimation through visual perception. However, obtaining labeled data on a scale required for the supervised training of CNNs is a difficult task - exacerbated if the object is novel and a 3D model is unavailable. To this end, this work presents an approach for rapidly generating real-world, pose-annotated RGB-D data for unknown objects. Our method not only circumvents the need for a prior 3D object model (textured or otherwise) but also bypasses complicated setups of fiducial markers, turntables, and sensors. With the help of a human user, we first source minimalistic labelings of an ordered set of arbitrarily chosen keypoints over a set of RGB-D videos. Then, by solving an optimization problem, we combine these labels under a world frame to recover a sparse, keypoint-based representation of the object. The sparse representation leads to the development of a dense model and the pose labels for each image frame in the set of scenes. We show that the sparse model can also be efficiently used for scaling to a large number of new scenes. We demonstrate the practicality of the generated labeled dataset by training a CNN based 6-DoF object pose estimator.
Rohan P. Singh, Mehdi Benallegue, Yusuke Yoshiyasu, Fumio Kanehiro
ICRA3
2021 MV-FractalDB: Formula-driven Supervised Learning for Multi-view Image Recognition
abstract
The paper proposes a method for automatic multi-view dataset construction based on formula-driven supervised learning (FDSL). Although data collection and human annotation of 3D objects are labor-intensive, we automatically generate their training data and labels in the proposed multi-view dataset. To create a large-scale multi-view dataset, we employ fractal geometry, which is considered the background information of many objects in the real world. We project in a circle from the rendered 3D fractal models to construct the Multi-view Fractal DataBase (MV-FractalDB), which is then used to make a pre-trained CNN model. According to the experimental results, the MV-FractalDB pre-trained model surpasses the accuracies with self-supervised methods (e.g., SimCLR and MoCo) and is close to supervised methods (e.g., ImageNet) in terms of performance rates on multi-view image datasets. We demonstrate the potential of FDSL for multi-view image recognition.
Ryosuke Yamada, Ryota Suzuki 0006, Akio Nakamura, Yusuke Yoshiyasu, Ryusuke Sagawa, Hirokatsu Kataoka
IROS5
2020 APE: A More Practical Approach To 6-Dof Pose Estimation
abstract
Recent advances in deep learning have shown high success in obtaining the 6-DoF pose of rigid objects. However, most works rely on a pre-existing dataset and do not tackle the data gathering part. The time-consuming and tedious tasks required to build datasets are, to a large extent, what is keeping these techniques from being more widely used in practical applications. We present a whole pipeline from data gathering to pose recognition and an example application of robot grasping. For our data gathering method we require as minimum user intervention as possible and, even without using depth information or 3D models, by using a novel RGB-only Neural Network design we are able to obtain results very close to the state of the art. We call this method Affordable Pose Estimation (APE).
Antonio Gabas, Yusuke Yoshiyasu, Rohan P. Singh, Ryusuke Sagawa, Eiichi Yoshida
ICIP2
2020 Pathnet: Learning To Generate Trajectories Avoiding Obstacles
abstract
This paper presents a novel approach to solving 2D motion planning problems using deep neural networks, which we refer to as PathNet. PathNet first takes a 2D environment map composed of obstacle zone and free zone and compresses it to a latent vector. The latent vector is afterward concatenated with the start and goal positions to generate a trajectory connecting those positions. Our learning-based neural planner can solve motion planning problems in unseen environments and is computationally efficient as it only needs one single pass in our network to generate trajectories.
Alassane Watt, Yusuke Yoshiyasu
ICIP2
2020 Efficient Exploration in Constrained Environments with Goal-Oriented Reference Path
abstract
In this paper, we consider the problem of building learning agents that can efficiently learn to navigate in constrained environments. The main goal is to design agents that can efficiently learn to understand and generalize to different environments using high-dimensional inputs (a 2D map), while following feasible paths that avoid obstacles in obstacle-cluttered environment. To achieve this, we make use of traditional path planning algorithms, supervised learning, and reinforcement learning algorithms in a synergistic way. The key idea is to decouple the navigation problem into planning and control, the former of which is achieved by supervised learning whereas the latter is done by reinforcement learning. Specifically, we train a deep convolutional network that can predict collision-free paths based on a map of the environment- this is then used by an reinforcement learning algorithm to learn to closely follow the path. This allows the trained agent to achieve good generalization while learning faster. We test our proposed method in the recently proposed Safety Gym suite that allows testing of safety-constraints during training of learning agents. We compare our proposed method with existing work and show that our method consistently improves the sample efficiency and generalization capability to novel environments.
Kei Ota, Yoko Sasaki, Devesh K. Jha, Yusuke Yoshiyasu, Asako Kanezaki
IROS4
2018 Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image
Yusuke Yoshiyasu, Ryusuke Sagawa, Ko Ayusawa, Akihiko Murai
ACCV (4)1
2018 Interspecies Retargeting of Homologous Body Posture Based on Skeletal Morphing
abstract
The paper aims to develop a methodology of transferring the knowledge obtained from the experiments of laboratory animals to human musculoskeletal system. To achieve the goal, we propose a method for estimating the homologous posture of the mammalian skeletal system corresponding to the human body posture. We hypothesize the homology of bone geometry between mammalian species implies that of biomechanical functions. The method relies on this homology and determines the homologous postures according to the anatomical landmarks of bone geometry. This paper shows the results of the analysis on homologous postures between the human and mouse skeletal models to validate our hypothesis. A pilot study also introduces comparison of mechanical functions between the two models by using the homologous postures.
Ko Ayusawa, Yosuke Ikegami, Akihiko Murai, Yusuke Yoshiyasu, Eiichi Yoshida, Satoshi Oota, Yoshihiko Nakamura
IROS4
2017 3D convolutional neural networks by modal fusion
abstract
We propose multi-view and volumetric convolutional neural networks (ConvNets) for 3D shape recognition, which combines surface normal and height fields to capture local geometry and physical size of an object. This strategy helps distinguishing between objects with similar geometries but different sizes. This is especially useful for enhancing volumetric ConvNets and classifying 3D scans with insufficient surface details. Experimental results on CAD and real-world scan datasets showed that our technique outperforms previous approaches.
Yusuke Yoshiyasu, Eiichi Yoshida, Sören Pirk, Leonidas J. Guibas
ICIP1
2017 Toward a Human(oid) Motion Planner
Eiichi Yoshida, Ko Ayusawa, Yusuke Yoshiyasu, Adrien Escande, Abderrahmane Kheddar
ISRR3
2017 Understanding and Exploiting Object Interaction Landscapes
abstract
Interactions play a key role in understanding objects and scenes for both virtual and real-world agents. We introduce a new general representation for proximal interactions among physical objects that is agnostic to the type of objects or interaction involved. The representation is based on tracking particles on one of the participating objects and then observing them with sensors appropriately placed in the interaction volume or on the interaction surfaces. We show how to factorize these interaction descriptors and project them into a particular participating object so as to obtain a new functional descriptor for that object, its interaction landscape , capturing its observed use in a spatiotemporal framework. Interaction landscapes are independent of the particular interaction and capture subtle dynamic effects in how objects move and behave when in functional use. Our method relates objects based on their function, establishes correspondences between shapes based on functional key points and regions, and retrieves peer and partner objects with respect to an interaction.
Sören Pirk, Vojtech Krs, Kaimo Hu, Suren Deepak Rajasekaran, Hao Kang, Yusuke Yoshiyasu, Bedrich Benes, Leonidas J. Guibas
ACM Trans. Graph.6
2016 Nonlinear dimensionality reduction by curvature minimization
abstract
In this paper, we introduce a nonlinear dimensionality reduction (NLDR) technique that can construct a low-dimensional embedding efficiently and accurately with low embedding distortions. The key idea is to divide NLDR into nonlinearity reduction and linear dimensionality reduction, which simplifies the overall NLDR process. Nonlinearity reduction is based on the elastic shell model that measures the in-plane stretching and bending energy. With this model, we minimize the curvature of the data, which is the source of nonlinearity, while preserving the original intrinsic property (i.e., local lengths) as-much-as possible. We discretize and linearize our nonlinearity reduction model such that it leads to an iterative deformation technique that alternates between two steps in order to flatten a manifold: the curvature minimization step that solves a bi-Laplace system and the local length restoration step that solves a Poisson system. We propose an efficient optimization technique for the both steps using a direct solver based on Cholesky decomposition, which exploits the fact that the system matrices stay constant; during iterations, we reuse the factorizations that are obtained once at the beginning and perform back substitutions only. Since our algorithm relies only on local geometric properties, it can accurately embed the data with complicated topology. Experimental results show that our algorithm is faster than the most of other state-of-the-art algorithms and preserves local areas and angles better than previous approaches.
Yusuke Yoshiyasu, Eiichi Yoshida
ICPR1
2016 Symmetry aware embedding for shape correspondence
Yusuke Yoshiyasu, Eiichi Yoshida, Leonidas J. Guibas
Comput. Graph.1
2015 Analyzing Muscle Activity and Force with Skin Shape Captured by Non-contact Visual Sensor
Ryusuke Sagawa, Yusuke Yoshiyasu, Alexander Alspach, Ko Ayusawa, Katsu Yamane, Adrian Hilton 0001
PSIVT2
2014 Symmetry-Aware Nonrigid Matching of Incomplete 3D Surfaces
abstract
We present a nonrigid shape matching technique for establishing correspondences of incomplete 3D surfaces that exhibit intrinsic reflectional symmetry. The key for solving the symmetry ambiguity problem is to use a point-wise local mesh descriptor that has orientation and is thus sensitive to local reflectional symmetry, e.g. discriminating the left hand and the right hand. We devise a way to compute the descriptor orientation by taking the gradients of a scalar field called the average diffusion distance (ADD). Because ADD is smoothly defined on a surface, invariant under isometry/scale and robust to topological errors, the robustness of the descriptor to non-rigid deformations is improved. In addition, we propose a graph matching algorithm called iterative spectral relaxation which combines spectral embedding and spectral graph matching. This formulation allows us to define pairwise constraints in a scale-invariant manner from k-nearest neighbor local pairs such that non-isometric deformations can be robustly handled. Experimental results show that our method can match challenging surfaces with global intrinsic symmetry, data incompleteness and non-isometric deformations.
Yusuke Yoshiyasu, Eiichi Yoshida, Kazuhito Yokoi, Ryusuke Sagawa
CVPR1
2014 As-Conformal-As-Possible Surface Registration
abstract
Abstract We present a non‐rigid surface registration technique that can align surfaces with sizes and shapes that are different from each other, while avoiding mesh distortions during deformation. The registration is constrained locally as conformal as possible such that the angles of triangle meshes are preserved, yet local scales are allowed to change. Based on our conformal registration technique, we devise an automatic registration and interactive registration technique, which can reduce user interventions during template fitting. We demonstrate the versatility of our technique on a wide range of surfaces.
Yusuke Yoshiyasu, Wan-Chun Ma, Eiichi Yoshida, Fumio Kanehiro
Comput. Graph. Forum1
2012 Example-based inverse kinematics using cage
abstract
ABSTRACT This paper presents a cage‐based inverse kinematics (CageIK) that enables interactive posing of character models in a wide range of mesh representations using handle points. By providing a set of cage geometries as examples, CageIK optimizes deformations of the cage according to handle movements and reconstructs the model using a subspace deformation method based on Green Coordinates. CageIK therefore not only poses the model naturally but also preserves details even when leaving the example space. Furthermore, by blending deformations based on bounded biharmonic weights, CageIK can edit the pose of the model locally. Because example cages can be generated from existing models, we can reuse animation assets that were already created by artists or simulations, which avoids repeating the time‐consuming process of creating examples. We show that CageIK can edit a wide variety of models, including multicomponent meshes, polygon soups, and quadrilateral meshes. Copyright © 2012 John Wiley & Sons, Ltd.
Yusuke Yoshiyasu, Nobutoshi Yamazaki
Comput. Animat. Virtual Worlds1
2012 Detail-aware spatial deformation transfer
abstract
ABSTRACT In this paper, we propose a deformation transfer method that is applicable to multi‐component objects and is able to transfer fine‐scale deformations. To accomplish our goal, we combine surface‐based and space‐based deformation transfers. Our system enhances the result of space deformation transfer, which is applicable to multi‐component objects but loses details, by using surface‐based deformation transfer to add fine‐scale deformations. Self‐collisions are handled by approximating spatial relationships between surfaces, from the result of space deformation. Experimental results show that our method can transfer fine‐scale deformations such as motions of a skirt and muscle bulging. Copyright © 2012 John Wiley & Sons, Ltd.
Yusuke Yoshiyasu, Nobutoshi Yamazaki
Comput. Animat. Virtual Worlds1
2011 Topology-adaptive multi-view photometric stereo
abstract
In this paper, we present a novel technique that enables capturing of detailed 3D models from flash photographs integrating shading and silhouette cues. Our main contribution is an optimization framework which not only captures subtle surface details but also handles changes in topology. To incorporate normals estimated from shading, we employ a mesh-based deformable model using deformation gradient. This method is capable of manipulating precise geometry and, in fact, it outperforms previous methods in terms of both accuracy and efficiency. To adapt the topology of the mesh, we convert the mesh into an implicit surface representation and then back to a mesh representation. This simple procedure removes self-intersecting regions of the mesh and solves the topology problem effectively. In addition to the algorithm, we introduce a hand-held setup to achieve multi-view photometric stereo. The key idea is to acquire flash photographs from a wide range of positions in order to obtain a sufficient lighting variation even with a standard flash unit attached to the camera. Experimental results showed that our method can capture detailed shapes of various objects and cope with topology changes well.
Yusuke Yoshiyasu, Nobutoshi Yamazaki
CVPR1
2009 Pose space deformation with rotation-invariant details
abstract
Pose space deformation (PSD) is a method for creating a shape from example meshes [Lewis et al. 2000; Weber et al. 2007]. This is thought of as a combination of the skeletal-subspace deformation (SSD) and the shape interpolation (morphing). In PSD, a base mesh is first computed by SSD. This is then corrected by adding displacements computed from examples. To interpolate displacement vectors, example meshes must be transformed back to the rest pose by SSD to store the displacements in a rotation-invariant manner.
Yusuke Yoshiyasu, Nobutoshi Yamazaki
SIGGRAPH ASIA Sketches1