EDBT 2026 Demo / reviewers in the wild / expert
Yukiyasu Domae
dblp:151/9310
· DBLP profile ↗
14ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-1366-9657ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Systems, architecture and hardware · 11 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot ManipulationabstractImitation learning (IL) aims to enable robots to perform tasks autonomously by observing a few human demonstrations. Recently, a variant of IL, called In-Context IL, utilized off-the-shelf large language models (LLMs) as instant policies that understand the context from a few given demonstrations to perform a new task, rather than explicitly updating network models with large-scale demonstrations. However, its reliability in the robotics domain is undermined by hallucination issues such as LLM-based instant policy, which occasionally generates poor trajectories that deviate from the given demonstrations. To alleviate this problem, we propose a new robust in-context imitation learning algorithm called the robust instant policy (RIP), which utilizes a Student’s t-regression model to be robust against the hallucinated trajectories of instant policies to allow reliable trajectory generation. Specifically, RIP generates several candidate robot trajectories to complete a given task from an LLM and aggregates them using the Student’s t-distribution, which is beneficial for ignoring outliers (i.e., hallucinations); thereby, a robust trajectory against hallucinations is generated. Our experiments, conducted in both simulated and real-world environments, show that RIP significantly outperforms state-of-the-art IL methods, with at least 26% improvement in task success rates, particularly in low-data scenarios for everyday tasks. Video results available at https://sites.google.com/view/robustinstantpolicy Hanbit Oh, Andrea M. Salcedo-Vázquez, Ixchel G. Ramirez, Yukiyasu Domae |
IROS | 4 |
| 2024 | NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance FieldsabstractWe present NeuralLabeling, a labeling approach and toolset for annotating 3D scenes using either bounding boxes or meshes and generating segmentation masks, affordance maps, 2D bounding boxes, 3D bounding boxes, 6DOF object poses, depth maps, and object meshes. NeuralLabeling uses Neural Radiance Fields (NeRF) as a renderer, allowing labeling to be performed using 3D spatial tools while incorporating geometric clues such as occlusions, relying only on images captured from multiple viewpoints as input. To demonstrate the applicability of NeuralLabeling to a practical problem in robotics, we added ground truth depth maps to 30000 frames of transparent object RGB and noisy depth maps of glasses placed in a dishwasher captured using an RGBD sensor, yielding the Dishwasher30k dataset. We show that training a simple deep neural network with supervision using the annotated depth maps yields a higher reconstruction performance than training with the previously applied weakly supervised approach. We also show how instance segmentation and depth completion datasets generated using NeuralLabeling can be incorporated into a robot application for grasping transparent objects placed in a dishwasher with an accuracy of 83.3%, compared to 16.3% without depth completion. Supplementary URI: https://florise.github.io/neural_labeling_web/. Floris Erich, Naoya Chiba, Abdullah Mustafa, Yusuke Yoshiyasu, Noriaki Ando, Ryo Hanai, Yukiyasu Domae |
IROS | 7 |
| 2024 | PEGASUS: Physically Enhanced Gaussian Splatting Simulation System for 6DoF Object Pose Dataset GenerationabstractWe introduce Physically Enhanced Gaussian Splatting Simulation System (PEGASUS) for 6DoF object pose dataset generation, a versatile dataset generator based on 3D Gaussian Splatting. Environment and object representations can be easily obtained using commodity cameras to reconstruct with Gaussian Splatting. PEGASUS allows the composition of new scenes by merging the respective underlying Gaussian Splatting point cloud of an environment with one or multiple objects. Leveraging a physics engine enables the simulation of natural object placement within a scene through interaction between meshes extracted for the objects and the environment. Consequently, an extensive amount of new scenes - static or dynamic - can be created by combining different environments and objects. By rendering scenes from various perspectives, diverse data points such as RGB images, depth maps, semantic masks, and 6DoF object poses can be extracted. Our study demonstrates that training on data generated by PEGASUS enables pose estimation networks to successfully transfer from synthetic data to real-world data. Moreover, we introduce the Ramen dataset, comprising 30 Japanese cup noodle items. This dataset includes spherical scans that capture images from both the object hemisphere and the Gaussian Splatting reconstruction, making them compatible with PEGASUS. Lukas Meyer, Floris Erich, Yusuke Yoshiyasu, Marc Stamminger, Noriaki Ando, Yukiyasu Domae |
IROS | 6 |
| 2023 | Learning Depth Completion of Transparent Objects using Augmented Unpaired DataabstractWe propose a technique for depth completion of transparent objects using augmented data captured directly from real environments with complicated geometry. Using cyclic adversarial learning we train translators to convert between painted versions of the objects and their real transparent counterpart. The translators are trained on unpaired data, hence datasets can be created rapidly and without any manual labeling. Our technique does not make any assumptions about the geometry of the environment, unlike SOTA systems that assume easily observable occlusion and contact edges, such as ClearGrasp. We show how our technique outperforms ClearGrasp in a dishwasher environment, in which occlusion and contact edges are difficult to observe. We also show how the technique can be used to create an object manipulation application with a humanoid robot. Supplementary URI: https://ftorise.github.io/faking_depth_web/. Floris Erich, Bruno Leme, Noriaki Ando, Ryo Hanai, Yukiyasu Domae |
ICRA | 5 |
| 2023 | Force Map: Learning to Predict Contact Force Distribution from VisionabstractWhen humans see a scene, they can roughly imagine the forces applied to objects based on their expe-rience and use them to handle the objects properly. This paper considers transferring this “force-visualization” ability to robots. We hypothesize that a rough force distribution (named “force map”) can be utilized for object manipulation strategies even if accurate force estimation is impossible. Based on this hypothesis, we propose a training method to predict the force map from vision. To investigate this hypothesis, we generated scenes where objects were stacked in bulk through simulation and trained a model to predict the contact force from a single image. We further applied domain randomization to make the trained model function on real images. The experimental results showed that the model trained using only synthetic images could predict approximate patterns representing the contact areas of the objects even for real images. Then, we designed a simple algorithm to plan a lifting direction using the predicted force distribution. We confirmed that using the predicted force distribution contributes to finding natural lifting directions for typical real-world scenes. Furthermore, the evaluation through simulations showed that the disturbance caused to surrounding objects was reduced by 26 % (translation displacement) and by 39 % (angular displacement) for scenes where objects were overlapping. Ryo Hanai, Yukiyasu Domae, Ixchel G. Ramirez, Bruno Leme, Tetsuya Ogata |
IROS | 2 |
| 2022 | Point Cloud Pre-training with Natural 3D StructuresabstractThe construction of 3D point cloud datasets requires a great deal of human effort. Therefore, constructing a large-scale 3D point clouds dataset is difficult. In order to rem-edy this issue, we propose a newly developed point cloud fractal database (PC-FractalDB), which is a novel family of formula-driven supervised learning inspired by fractal geometry encountered in natural 3D structures. Our re-search is based on the hypothesis that we could learn rep-resentations from more real-world 3D patterns than con-ventional 3D datasets by learning fractal geometry. We show how the PC-FractalDB facilitates solving several re-cent dataset-related problems in 3D scene understanding, such as 3D model collection and labor-intensive annotation. The experimental section shows how we achieved the performance rate of up to 61.9% and 59.0% for the Scan-NetV2 and SUN RGB-D datasets, respectively, over the current highest scores obtained with the PointContrast, con-trastive scene contexts (CSC), and RandomRooms. More-over, the PC-FractalDB pre-trained model is especially ef-fective in training with limited data. For example, in 10% of training data on ScanNetV2, the PC-FractalDB pre-trained VoteNet performs at 38.3%, which is +14.8% higher accu-racy than CSC. Of particular note, we found that the pro-posed method achieves the highest results for 3D object de-tection pre-training in limited point cloud data.11Dataset release: https://ryosuke-yamada.github.io/PointCloud-FractalDataBase/ Ryosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae, Tetsuya Ogata |
CVPR | 4 |
| 2022 | Use of Action Label in Deep Predictive Learning for Robot ManipulationabstractVarious forms of human knowledge can be explicitly used to enhance deep robot learning from demonstrations. Annotation of subtasks from task segmentation is one type of human symbolism and knowledge. Annotated subtasks can be referred to as action labels, which are more primitive symbols that can be building blocks for more complex human reasoning, like language instructions. However, action labels are not widely used to boost learning processes because of problems that include (1) real-time annotation for online manipulation, (2) temporal inconsistency by annotators, (3) difference in data characteristics of motor commands and action labels, and (4) annotation cost. To address these problems, we propose the Gated Action Motor Predictive Learning (GAMPL) framework to leverage action labels for improved performance. GAMPL has two modules to obtain soft action labels compatible with motor commands and to generate motion. In this study, GAMPL is evaluated for towel-folding manipulation tasks in a real environment with a six degrees-of-freedom (6 DoF) robot and shows improved generalizability with action labels. Kei Kase, Chikara Utsumi, Yukiyasu Domae, Tetsuya Ogata |
IROS | 3 |
| 2022 | Interpretable Navigation Agents Using Attention-Augmented MemoryabstractDeep reinforcement learning (DRL) has achieved remarkable success in various domains, from games to complex tasks. In several DRL applications, the agents must perform long-horizon tasks in partially observable environments. However, owing to the high performance of DRL, the decision-making process for the long-horizon task is unclear and difficult to interpret. Although memory-based models and the Transformer model have been proposed to overcome the interpretability issue, challenges of limited flexibility, lack of stability, and high computational cost remain. To address these concerns, we propose a low-computational-complexity and scalable DRL model that uses attention-augmented memory (AAM) to interpret the long-horizon decision-making process. AAM adds long short-term memory states of past observations to the memory and uses a soft attention bottleneck to combine them into a single contextual vector. The agent is then trained to make decisions based on this AAM together with the current observation. The AAM model is applied to the navigation problem of the Labyrinth [1], and attention and saliency maps are generated to show the areas highlighted by the agent in the current observation and the areas it attends to from memory. The resiliency of the model was evaluated using saliency and it was discovered that the proposed method is more resilient to the noise of visual observations compared with the baseline model. In summary, the proposed method demonstrates an interpretable and noise-robust DRL approach for long-horizon tasks. Shotaro Miwa, Yukiyasu Domae |
SMC | 3 |
| 2020 | Robotic General Parts Feeder: Bin-picking, Regrasping, and KittingabstractThe automatic parts feeding of multiple objects is an unsolved problem in the manufacturing industry. In this paper, we tackle the problem by proposing a multi-robot system. The system comprises three sub-components which perform bin-picking, regrasping, and kitting. The three subcomponents divide and conquer the automatic multiple parts feeding problem by considering a coarse-to-fine manipulation process. Multiple robot arms are connected in series as a pipeline. The robots are separated into three groups to perform the roles of each sub-component. The accuracy of the state and manipulation are getting higher along with the changes of the sub-components in the pipeline. In the experimental section, the performance of the system is evaluated by using the Mean Picks Per Hour (MPPH) metric and success rate, which are compared to traditional parts feeder and manual labor. The results show that the Mean Picks Per Hour (MPPH) of the proposed system is 351 with eleven various-shaped industrial parts, which is faster than the state-of-the-art robotic bin-picking system. The lead time of the proposed system for new parts is less than that of a traditional parts feeders and/or manual labor. Yukiyasu Domae, Akio Noda, Tatsuya Nagatani, Weiwei Wan |
ICRA | 1 |
| 2020 | Planning an Efficient and Robust Base Sequence for a Mobile Manipulator Performing Multiple Pick-and-place TasksabstractIn this paper, we address efficiently and robustly collecting objects stored in different trays using a mobile manipulator. A resolution complete method, based on precomputed reachability database, is proposed to explore collision-free inverse kinematics (IK) solutions and then a resolution complete set of feasible base positions can be determined. This method approximates a set of representative IK solutions that are especially helpful when solving IK and checking collision are treated separately. For real world applications, we take into account the base positioning uncertainty and plan a sequence of base positions that reduce the number of necessary base movements for collecting the target objects, the base sequence is robust in that the mobile manipulator is able to complete the part-supply task even there is certain deviation from the planned base positions. Our experiments demonstrate both the efficiency compared to regular base sequence and the feasibility in real world applications. Jingren Xu, Kensuke Harada, Weiwei Wan, Toshio Ueshiba, Yukiyasu Domae |
ICRA | 5 |
| 2019 | Fast and Precise Detection of Object Grasping Positions with Eigenvalue TemplatesabstractFast Graspability Evaluation (FGE) has been proposed as a method for detecting grasping positions on objects and is now being used for industrial robots. FGE uses convolution of hand templates with regions on the target object to estimate the optimum grasping posture. However, the hand opening width and rotation angles must be set with high resolution to achieve highly accurate results and the computational load is high. To address that issue, we propose a method in which hand templates are represented in compact form for faster processing by using singular value decomposition. Applying singular value decomposition enables hand templates to be represented as linear combinations of a small number of eigenvalue templates and eigenfunctions. Eigenfunctions take discrete values, but response values can be calculated with arbitrary parameters by fitting a continuous function. Experimental results show that the proposed method reduces computation time by two thirds while maintaining the same detection accuracy as conventional FGE for both parallel hands and three-finger hands. Kousuke Mano, Takahiro Hasegawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Yukiyasu Domae |
ICRA | 5 |
| 2019 | Learning Based Robotic Bin-picking for Potentially Tangled ObjectsabstractIn this research, we tackle the challenge of picking only one object from a randomly stacked pile where the objects can potentially be tangled. No solution has been proposed to solve this challenge due to the complexity of picking one and only one object from the bin of tangled objects. Therefore, we propose a method for avoiding the situation where a robot picks multiple objects. In our proposed method, first, grasping pose candidates are computed by using the graspability index. Then, a Convolutional Neural Network (CNN) is trained to predict whether or not the robot can pick one and only one object from the bin. Additionally, since a physics simulator is used to collect data to train the CNN, an automatic picking system can be built. The effectiveness of the proposed method is confirmed through experiments on robot Nextage and compare with previous bin-picking methods. Ryo Matsumura, Yukiyasu Domae, Weiwei Wan, Kensuke Harada |
IROS | 2 |
| 2017 | 3D Object Discovery and Modeling Using Single RGB-D Images Containing Multiple Object InstancesabstractUnsupervised object modeling is important in robotics, especially for handling a large set of objects. We present a method for unsupervised 3D object discovery, reconstruction, and localization that exploits multiple instances of an identical object contained in a single RGB-D image. The proposed method does not rely on segmentation, scene knowledge, or user input, and thus is easily scalable. Our method aims to find recurrent patterns in a single RGB-D image by utilizing appearance and geometry of the salient regions. We extract keypoints and match them in pairs based on their descriptors. We then generate triplets of the keypoints matching with each other using several geometric criteria to minimize false matches. The relative poses of the matched triplets are computed and clustered to discover sets of triplet pairs with similar relative poses. Triplets belonging to the same set are likely to belong to the same object and are used to construct an initial object model. Detection of remaining instances with the initial object model using RANSAC allows to further expand and refine the model. The automatically generated object models are both compact and descriptive. We show quantitative and qualitative results on RGB-D images with various objects including some from the Amazon Picking Challenge. We also demonstrate the use of our method in an object picking scenario with a robotic arm. Wim Abbeloos, Esra Ataer Cansizoglu, Sergio Caccamo, Yuichi Taguchi, Yukiyasu Domae |
3DV | 5 |
| 2014 | Fast graspability evaluation on single depth maps for bin picking with general grippersabstractWe present a method that estimates graspability measures on a single depth map for grasping objects randomly placed in a bin. Our method represents a gripper model by using two mask images, one describing a contact region that should be filled by a target object for stable grasping, and the other describing a collision region that should not be filled by other objects to avoid collisions during grasping. The graspability measure is computed by convolving the mask images with binarized depth maps, which are thresholded differently in each region according to the minimum height of the 3D points in the region and the length of the gripper. Our method does not assume any 3-D model of objects, thus applicable to general objects. Our representation of the gripper model using the two mask images is also applicable to general grippers, such as multi-finger and vacuum grippers. We apply our method to bin picking of piled objects using a robot arm and demonstrate fast pick-and-place operations for various industrial objects. Yukiyasu Domae, Haruhisa Okuda, Yuichi Taguchi, Kazuhiko Sumi, Takashi Hirai |
ICRA | 1 |