Tomoaki Nakamura

dblp:63/4486 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 2 since 2021Systems, architecture and hardware · 29 · 9 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 IC-GP-HSMM: Unsupervised Segmentation of Behaviors with Individual Differences
abstract
Robots and intelligent systems that understand and support human behavior are increasingly used in various settings, including public facilities, homes, offices, and manufacturing sites. To effectively understand human behavior, such systems must accurately segment and classify behaviors. Unsupervised learning, which does not require labeled data, is crucial for this task, as manually labeling every behavior in advance is challenging. Conventional unsupervised segmentation methods typically assume that instances of the same behavior share uniform characteristics. However, this assumption can lead to inaccurate segmentation when individual differences are present. To address this limitation, we propose the individuality-conditioned Gaussian process-hidden semi-Markov model (IC-GP-HSMM), an extension of the GP-HSMM. The original GP-HSMM uses unsupervised learning to model behavioral patterns as Gaussian processes based on observed sequences. Our extension assumes that information regarding individuals performing the behaviors is observable and incorporates this information as an individuality vector. This enables the model to learn behavior-specific Gaussian processes that consider individual variation, resulting in more accurate segmentation. Experiments conducted on synthetic and motion capture datasets demonstrate that IC-GP-HSMM outperforms conventional methods in segmenting behaviors with individual differences. The proposed IC-GPHSMM enables intelligent systems to more accurately recognize and adapt to individual variations in human behavior, enhancing their reliability and effectiveness in real-world applications.
Toshiyuki Hatta, Issei Saito, Masatoshi Nagano, Tomoaki Nakamura
IECON4
2025 Semi-Autonomous Teleoperation for Mobile Manipulator via Action Chunking with Transformers
abstract
In this paper, we propose a semi-autonomous teleoperation method to reduce operator workload and enhance task execution efficiency. The proposed method uses action chunking with transformers (ACT), an imitation learning method, to predict long-duration motion trajectories from a robot’s egocentric images and pose information during teleoperation. Because the images and poses input into ACT are continuously updated during teleoperation, the predicted trajectory dynamically reflects the control history of the operator. These continuously updated trajectories are displayed in real time in a simulation environment, serving as visual feedback for the operator. When the operator determines that the visualized trajectory aligns with their intended trajectory, they can switch from manual to autonomous control. Subsequently, the robot autonomously executes motion based on the predicted trajectory, thereby reducing operator workload and improving task efficiency. To verify the effectiveness of the proposed method, an object-grasping task was conducted using a human support robot, comparing manual control, fully autonomous control, and the proposed semi-autonomous method. The results demonstrated that the proposed method effectively reduces task completion time and improves success rates, even with unknown objects and environments.
Shuntaro Itakura, Masatoshi Nagano, Tomoaki Nakamura
IECON3
2025 Scalable Unsupervised Segmentation via Random Fourier Feature-based Gaussian Process
abstract
In this paper, we propose RFF-GP-HSMM, a fast unsupervised time-series segmentation method that incorporates random Fourier features (RFF) to address the high computational cost of the Gaussian process hidden semi-Markov model (GP-HSMM). GP-HSMM models time-series data using Gaussian processes, requiring inversion of an N × N kernel matrix during training, where N is the number of data points. As the scale of the data increases, matrix inversion incurs a significant computational cost. To address this, the proposed method approximates the Gaussian process with linear regression using RFF, preserving expressive power while eliminating the need for inversion of the kernel matrix. Experiments on the Carnegie Mellon University (CMU) motion-capture dataset demonstrate that the proposed method achieves segmentation performance comparable to that of conventional methods, with approximately 278 times faster segmentation on time-series data comprising 39,200 frames.
Issei Saito, Masatoshi Nagano, Tomoaki Nakamura, Daichi Mochihashi, Koki Mimura
IECON3
2025 Prior Anatomical Knowledge-guided GAN for ICL surgery postoperative prediction based on AS-OCT image
Yinglin Zhang, Ruiling Xi, Risa Higashita, Keiichiro Okamoto, Kazutaka Kamiya, Kazunori Miyata, Akihito Igarashi, Seiichiro Hata, Tomoaki Nakamura, Jiang Liu 0001
Medical Image Anal.9
2023 Unsupervised Work Behavior Analysis Using Hierarchical Probabilistic Segmentation
abstract
Workers' behavior should be analyzed to improve their efficiency and that of cell production systems. However, traditional approaches and supervised learning methods are time-consuming and require abundant labeled data, respectively. Therefore, this study proposes a novel unsupervised behavior analysis model, namely the Gaussian process-hidden semi-Markov model-based behavior analyzer (GP-HSMM-BA), which is a two-layered model consisting of GP-HSMM and HSMM. This model can efficiently capture the motion element and unit motion, which are the smallest unit motions, and their compositional motions. The first layer is GP-HSMM, which divides and classifies continuous motion into motion elements. The second layer is HSMM, which divides and classifies the sequence of motion elements into unit motions. Furthermore, more accurate motion elements and unit motions can be obtained by training the two models mutually. The proposed model was applied to a cell production environment, indicating that motion elements can be extracted more accurately using GP-HSMM-BA. Additionally, an analysis using the trained parameters and motion elements indicated that habituation can be quantified and behavior changes in repeated tasks are easily understood.
Issei Saito, Tomoaki Nakamura, Toshiyuki Hatta, Wataru Fujita, Shintaro Watanabe, Shotaro Miwa
IECON2
2023 An Efficient Approach Based on Graph Neural Networks for Predicting Wait Time in Job Schedulers
Tomoe Kishimoto, Tomoaki Nakamura
JSSPP2
2022 Analysis of User Behavior and Workload During Simultaneous Tele-operation of Multiple Mobile Manipulators
abstract
This paper discusses the tele-operation system for multiple mobile manipulators. If a single person could freely tele-operate multiple mobile manipulators simultaneously, it would be a great step toward the goal of “avatar-symbiotic society” allowing people to live beyond the constraints of their bodies, space, and time. At present, however, such a tele-operation system has not been developed. Therefore, we built a prototype system to tele-operate two mobile manipulators and conducted a subject experiment on a pick-and-place task to investigate the tele-operator's task performance, workload and gaze behavior. By analyzing these results, we obtained guidelines to design tele-operation system for multiple robots.
Tatsuya Aoki, Tomoaki Nakamura, Takayuki Nagai
IROS2
2022 A whole brain probabilistic generative model: Toward realizing cognitive architectures for developmental robots
abstract
Building a human-like integrative artificial cognitive system, that is, an artificial general intelligence (AGI), is the holy grail of the artificial intelligence (AI) field. Furthermore, a computational model that enables an artificial system to achieve cognitive development will be an excellent reference for brain and cognitive science. This paper describes an approach to develop a cognitive architecture by integrating elemental cognitive modules to enable the training of the modules as a whole. This approach is based on two ideas: (1) brain-inspired AI, learning human brain architecture to build human-level intelligence, and (2) a probabilistic generative model (PGM)-based cognitive architecture to develop a cognitive system for developmental robots by integrating PGMs. The proposed development framework is called a whole brain PGM (WB-PGM), which differs fundamentally from existing cognitive architectures in that it can learn continuously through a system based on sensory-motor information. In this paper, we describe the rationale for WB-PGM, the current status of PGM-based elemental cognitive modules, their relationship with the human brain, the approach to the integration of the cognitive modules, and future challenges. Our findings can serve as a reference for brain studies. As PGMs describe explicit informational relationships between variables, WB-PGM provides interpretable guidance from computational sciences to brain science. By providing such information, researchers in neuroscience can provide feedback to researchers in AI and robotics on what the current models lack with reference to the brain. Further, it can facilitate collaboration among researchers in neuro-cognitive sciences as well as AI and robotics.
Tadahiro Taniguchi, Hiroshi Yamakawa, Takayuki Nagai, Kenji Doya, Masamichi Sakagami, Tomoaki Nakamura, Akira Taniguchi
Neural Networks7
2019 Adjusting Weight of Action Decision in Exploration for Logistics Warehouse Picking Learning
abstract
The purpose of this study is for a robot to learn picking motions in a logistics warehouse environment. The picking operation performed by a robot often fails owing to the inclination of items placed on a shelf, as well as the minimum clearance between the products and their vinyl packaging. Therefore, we considered acquiring a specific motion trajectory by reinforcement learning. However, because numerous types of items are handled in logistics warehouses, efficient learning is required. Therefore, in this research, we propose a method to efficiently exploration for learning picking an object by determining a focus exploration area for learning based on previous results of different objects.
Yusuke Kato, Tomoaki Nakamura, Takayuki Nagai, Natsuki Yamanobe, Kazuyuki Nagata, Jun Ozawa
IROS2
2019 High-dimensional Motion Segmentation by Variational Autoencoder and Gaussian Processes
abstract
Humans perceive continuous high-dimensional information by dividing it into significant segments such as words and units of motion. We believe that such unsupervised segmentation is also important for robots to learn topics such as language and motion. To this end, we previously proposed a hierarchical Dirichlet process-Gaussian process-hidden semi-Markov model (HDP-GP-HSMM). However, an important drawback to this model is that it cannot divide high-dimensional time-series data. Further, low-dimensional features must be extracted in advance. Segmentation largely depends on the design of features, and it is difficult to design effective features, especially in the case of high-dimensional data. To overcome this problem, this paper proposes a hierarchical Dirichlet process-variational autoencoder-Gaussian process-hidden semi-Markov model (HVGH). The parameters of the proposed HVGH are estimated through a mutual learning loop of the variational autoencoder and our previously proposed HDP-GP-HSMM. Hence, HVGH can extract features from high-dimensional time-series data, while simultaneously dividing it into segments in an unsupervised manner. In an experiment, we used various motion-capture data to show that our proposed model estimates the correct number of classes and more accurate segments than baseline methods. Moreover, we show that the proposed method can learn latent space suitable for segmentation.
Masatoshi Nagano, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Wataru Takano
IROS2
2018 Learning and Generation of Actions from Teleoperation for Domestic Service Robots*This work was supported by JST, CREST
abstract
In this paper, we propose a method for motion learning aimed at the execution of autonomous household chores by service robots in real environments. For robots to act autonomously in a real environment, it is necessary to define the appropriate actions for the environment. However, it is difficult to define these actions manually. Therefore, body motions that are common to multiple actions are defined as motion primitives. Complex actions can then be learned by combining these motion primitives. For learning motion primitives, we propose a reference-point and object-dependent Gaussian process hidden semi-Markov model (RPOD-GP-HSMM). For verification, a robot is teleoperated to perform the actions included in several domestic household chores. The robot then learns the associated motion primitives from the robot's body information and object information.
Kensuke Iwata, Tatsuya Aoki, Takato Horii, Tomoaki Nakamura, Takayuki Nagai
IROS4
2018 Sequence Pattern Extraction by Segmenting Time Series Data Using GP-HSMM with Hierarchical Dirichlet Process
abstract
Humans recognize perceived continuous information by dividing it into significant segments such as words and unit motions. We believe that such unsupervised segmentation is also an important ability that robots need to learn topics such as language and motions. Hence, in this paper, we propose a method for dividing continuous time-series data into segments in an unsupervised manner. To this end, we proposed a method based on a hidden semi-Markov model with Gaussian process (GP-HSMM). If Gaussian processes, which are nonparametric models, are used, unit motion patterns can be extracted from complicated continuous motion. However, this approach requires the number of classes of segments in the time-series data in advance. To overcome this problem, in this paper, we extend GP-HSMM to a nonparametric Bayesian model by introducing a hierarchical Dirichlet process (HDP) and propose the hierarchical Dirichlet processes-Gaussian process-hidden semi-Markov model (HDP-GP-HSMM). In the nonparametric Bayesian model, an infinite number of classes is assumed and it becomes difficult to estimate the parameters naively. Instead, the parameters of the proposed HDP-GP-HSMM are estimated by applying slice sampling. In the experiments, we use various synthetic and motion-capture data to show that our proposed model can estimate a more correct number of classes and achieve more accurate segmentation than baseline methods.
Masatoshi Nagano, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Masahide Kaneko
IROS2
2018 Interaction Modeling Based on Segmenting Two Persons Motions Using Coupled GP-HSMM
abstract
Humans interact with one another daily and learn from their experiences via observation and interaction. To create robots that can coexist with humans, it is important that they learn how to appropriately interact in the human community. In this paper, we propose a novel model, the coupled Gaussian process hidden semi-Markov model (GP-HSMM), which enables robots to learn rules of interaction between two persons by observing them in an unsupervised manner. The continuous motions of the persons are segmented into discrete actions based on GP-HSMM, and the relationships between the actions are extracted. Moreover, all corresponding actions are not simultaneously conducted by two persons during actual interaction. Thus, the coupled GP-HSMM accounts for these lags. We conduct experiments using the motion data of interaction games. Experimental results showed that the coupled GP-HSMM can estimate actions, lags between them and their relationships.
Satoru Oshikawa, Tomoaki Nakamura, Takayuki Nagai, Masahide Kaneko, Kotaro Funakoshi, Naoto Iwahashi, Mikio Nakano
RO-MAN2
2018 Generation of Gestures During Presentation for Humanoid Robots
abstract
For presentation purposes, gestures play an exceptionally important role in improving the information transmission effect. It has been demonstrated that the body language expressing the enthusiasm and intention of the presenter affects the success of the presentation and the impression on the audience. For these reasons, presentation robots are required to perform such movements; however, manual design of these movements is a difficult task. In this research, we propose a method to model the relationship between speech prosodic information and motion using a recurrent neural network, and directly generate appropriate motions using the prosodic information. This study also proposes a method for generating motions that convey the meaning of specific words. We implement the proposed method on the “Pepper” robot to evaluate its performance.
Akihito Shimazu, Chie Hieida, Takayuki Nagai, Tomoaki Nakamura, Yuki Takeda, Takenori Hara, Osamu Nakagawa, Tsuyoshi Maeda
RO-MAN4
2017 Symbol Grounding from Natural Conversation for Human-Robot Communication
abstract
This paper proposes a new approach for research on chat-like conversational systems that enable robots to acquire physically grounded knowledge through natural interaction with humans. The proposed approach combines research on chat-like conversational systems, language acquisition, and symbol grounding in order to realize physically situated and natural human-robot interaction. In contrast to previous approaches for chat-like conversation, the proposed approach focuses on utterances which are situated in physical environments surrounding humans and robots. Based on the proposed approach, we develop a concrete method that enables robots to learn object image concepts and the words describe them from object-teaching utterances made by humans. The method is composed of two processes:(1) the detection of object-teaching utterances from chat-like conversation and (2) the learning of object image concepts and the words describing them. It applies a linear support vector machine, multimodal hierarchical Dirichlet process, and term frequency-inverse document frequency process. The experimental results show that the method enabled robots to learn object image concepts and the words that describe them through multimodal chat-like interactions with humans.
Ye Kyaw Thu, Takuya Ishida, Naoto Iwahashi, Tomoaki Nakamura, Takayuki Nagai
HAI4
2016 Modeling of Honest Signals for Human Robot Interaction
abstract
Recent studies have shown that human beings unconsciously use signals that represent their thoughts and/or intentions when communicating with each other. These signals are known as “honest signals.” This study involves the use of a sociometer to capture multimodal data resulting from the interaction between humans. These data are then used to model the interaction using a multimodal hierarchical Dirichlet process hidden Markov model, which is then implemented in the robot. The model enables robots to generate “honest signals” and to interact in a natural manner.
Muhammad Attamimi, Yusuke Katakami, Kasumi Abe, Takayuki Nagai, Tomoaki Nakamura
HRI5
2016 Online joint learning of object concepts and language model using multimodal hierarchical Dirichlet process
abstract
One of the biggest challenges in intelligent robotics is to build robots that can learn to use language. To this end, we think that the practical long-term on-line concept/word learning algorithm for robots is a key issue to be addressed. In this paper, we develop an unsupervised on-line learning algorithm that uses Bayesian nonparametrics for categorizing multimodal sensory signals such as audio, visual, and haptic information for robots. The robot uses its physical body to grasp and observe an object from various viewpoints as well as listen to the sound during the observation. The most important property of the proposed framework is to learn multimodal concepts and the language model simultaneously. This mutual learning framework of concepts and language significantly improves both speech recognition and multimodal categorization performances. We conducted a long-term experiment where a human subject interacted with a real robot over 100 hours using 499 objects. Some interesting results of the experiment are discussed in this paper.
Tatsuya Aoki, Joe Nishihara, Tomoaki Nakamura, Takayuki Nagai
IROS3
2016 Small Talk Improves User Impressions of Interview Dialogue Systems
abstract
This paper addresses the problem of how to build interview systems that users are willing to use.Existing interview dialogue systems are mainly focused on obtaining information from users, thus they just repeatedly ask questions.We propose a method for improving user impressions by engaging in small talk during interviews.The system performs framebased dialogue management for interviewing and generates small talk utterances after the user answers the system's questions.Experimental results using a textbased interview dialogue system for diet recording showed the proposed method gives a better impression to users than interview dialogues without small talk.It is also found that generating too many small talk utterances makes user impressions worse because of the system's low capability of continuously generating appropriate small talk utterances.
Takahiro Kobori, Mikio Nakano, Tomoaki Nakamura
SIGDIAL Conference3
2015 Learning Word Meanings and Grammar for Describing Everyday Activities in Smart Environments
abstract
Muhammad Attamimi, Yuji Ando, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi, Hideki Asoh. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Muhammad Attamimi, Yuji Ando, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Hideki Asoh
EMNLP3
2015 Concept formation by robots using an infinite mixture of models
abstract
We propose a method for a robot to form various concepts. The robot uses its embodiment to obtain visual, auditory, and haptic information by grasping, shaking, and observing objects. At the same time, a user teaches the robot object features through speech. From these kinds of information, the robot can form object concepts. The information obtained by the robot is converted into Bag-of-Words representations, which are classified into categories. The proposed method is based on a stochastic model, and objects can be classified by estimating their parameters. We introduce the Chinese restaurant process into the multimodal hierarchical Dirichlet process. This model is an infinite mixture of models and enables the robot to form various concepts such as object type, color, and so on. Therefore, the robot can form not only object concepts but also various concepts that are not represented by object concepts (e.g., color). Also, because the proposed method is based on a stochastic model, it makes it possible for the robot to estimate category and unobserved information of unseen objects. We implement the proposed method on a robot and show that it can form various concepts and perform various estimations for unseen objects.
Tomoaki Nakamura, Yoshiki Ando, Takayuki Nagai, Masahide Kaneko
IROS1
2015 Probabilistic modeling of mental models of others
abstract
Intimacy is a very important factor not only for the communication between humans but also for the communication between humans and robots. In this research we propose an action decision model based on others' friendliness value, which can be estimated using the mental models of others. We examine the mutual adaptation process of two agents, each of which has its own model of others, through the interaction between them.
Takayuki Nagai, Kasumi Abe, Tomoaki Nakamura, Natsuki Oka, Takashi Omori
RO-MAN3
2014 Integration of various concepts and grounding of word meanings using multi-layered multimodal LDA for sentence generation
abstract
In the field of intelligent robotics, object handling by robots can be achieved by capturing not only the object concept through object categorization, but also other concepts (e.g., the movement while using the object), as well as the relationship between concepts. Moreover, capturing the concepts of places and people is also necessary to enable the robot to gain real-world understanding. In this study, we propose multi-layered multimodal latent Dirichlet allocation (mMLDA) to realize the formation of various concepts, and the integration of those concepts, by robots. Because concept formation and integration can be conducted by mMLDA, the formation of each concept affects others, resulting in a more appropriate formation. Another issue to be addressed in this paper is the language acquisition by the robots. We propose a method to infer which words are originally connected to a concept using mutual information between words and concepts. Moreover, the order of concepts in teaching utterances can be learned using a simple Markov model, which corresponds to grammar. This grammar can be used to generate sentences that represent the observed information. We report the results of experiments to evaluate the effectiveness of the proposed method.
Muhammad Attamimi, Muhammad Fadlil, Kasumi Abe, Tomoaki Nakamura, Kotaro Funakoshi, Takayuki Nagai
IROS4
2014 Mutual learning of an object concept and language model based on MLDA and NPYLM
abstract
Humans develop their concept of an object by classifying it into a category, and acquire language by interacting with others at the same time. Thus, the meaning of a word can be learnt by connecting the recognized word and concept. We consider such an ability to be important in allowing robots to flexibly develop their knowledge of language and concepts. Accordingly, we propose a method that enables robots to acquire such knowledge. The object concept is formed by classifying multimodal information acquired from objects, and the language model is acquired from human speech describing object features. We propose a stochastic model of language and concepts, and knowledge is learnt by estimating the model parameters. The important point is that language and concepts are interdependent. There is a high probability that the same words will be uttered to objects in the same category. Similarly, objects to which the same words are uttered are highly likely to have the same features. Using this relation, the accuracy of both speech recognition and object classification can be improved by the proposed method. However, it is difficult to directly estimate the parameters of the proposed model, because there are many parameters that are required. Therefore, we approximate the proposed model, and estimate its parameters using a nested Pitman-Yor language model and multimodal latent Dirichlet allocation to acquire the language and concept, respectively.
Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi
IROS1
2013 Formation of hierarchical object concept using hierarchical latent Dirichlet allocation
abstract
In recent studies, it has been revealed that robots can form concepts and understand the meanings of words through inference. The key idea underlying these studies is “multimodal categorization” of a robot's experience. However, previous studies considered only nonhierarchical categorization methods, which led to nonhierarchical concept structures. Our concepts have a hierarchical structure, thus ensuring that the resulting inferences are more efficient and accurate. In this paper, we propose a novel hierarchical categorization method. The method involves extending multimodal latent Dirichlet allocation (MLDA) to hierarchical MLDA using the nested Chinese restaurant process, which makes it possible for robots to acquire concepts in a hierarchical structure. We show that a robot can form a hierarchical concept structure based on self-obtained multimodal information. Moreover, by focusing on the common features of each category in the hierarchy, the robot is able to infer unobserved information including word meanings.
Yoshiki Ando, Tomoaki Nakamura, Takaya Araki, Takayuki Nagai
IROS2
2013 Long-term learning of concept and word by robots: Interactive learning framework and preliminary results
abstract
One of the biggest challenges in intelligent robotics is to build robots that can understand and use language. Such robots will be a part of our everyday life; at the same time, they can be of great help to investigate the complex mechanism of language acquisition by infants in constructive approach. To this end, we think that the practical long-term on-line concept/word learning algorithm for robots and the interactive learning framework are the key issues to be addressed. In this paper we develop a practical on-line learning algorithm that solves three remaining problems in our previous study. We also propose an interactive learning framework, in which the proposed on-line learning algorithm is embedded. The main contribution of this paper is to develop such a practical learning framework, and we test it on a real robot platform to show its potential toward the ultimate goal.
Takaya Araki, Tomoaki Nakamura, Takayuki Nagai
IROS2
2013 Integrated concept of objects and human motions based on multi-layered multimodal LDA
abstract
The human understanding of things is based on prediction which is made through concepts formed by the categorization of experience. To mimic this mechanism in robots, multimodal categorization, which enables the robot to form concepts, has been studied. On the other hand, segmentation and categorization of human motions have also been studied to recognize and predict future motions. This paper addresses the issue on how these different kinds of concepts are integrated to generate higher level concepts and, more importantly, on how the higher level concepts affect the formation of each lower level concept. To this end, we propose the multi-layered multimodal latent Dirichlet allocation (mMLDA), which is an expansion of the MLDA to learn and represent the hierarchical structure of concepts. We also examine a simple integration model and compare it with the mMLDA. The experimental results reveal that the mMLDA leads to a better inference performance and, indeed, forms higher level concepts which integrate motions and objects that are necessary for real-world understanding.
Muhammad Fadlil, Keisuke Ikeda, Kasumi Abe, Tomoaki Nakamura, Takayuki Nagai
IROS4
2013 Multimodal concept and word learning using phoneme sequences with errors
abstract
In this study, we propose a method for concept formation and word acquisition for robots. The proposed method is based on multimodal latent Dirichlet allocation (MLDA) and the nested Pitman-Yor language model (NPYLM). A robot obtains haptic, visual, and auditory information by grasping, observing, and shaking an object. At the same time, a user teaches object features to the robot through speech, which is recognized using only acoustic models and transformed into phoneme sequences. As the robot is supposed to have no language model in advance, the recognized phoneme sequences include many phoneme recognition errors. Moreover, the recognized phoneme sequences with errors are segmented into words in an unsupervised manner; however, not all words are necessarily segmented correctly. The words including these errors have a negative effect on the learning of word meanings. To overcome this problem, we propose a method to improve unsupervised word segmentation and to reduce phoneme recognition errors by using multimodal object concepts. In the proposed method, object concepts are used to enhance the accuracy of word segmentation, reduce phoneme recognition errors, and correct words so as to improve the categorization accuracy. We experimentally demonstrate that the proposed method can improve the accuracy of word segmentation and reduce the phoneme recognition error and that the obtained words enhance the categorization accuracy.
Tomoaki Nakamura, Takaya Araki, Takayuki Nagai, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi
IROS1
2012 Hierarchical multilevel object recognition using Markov model
Muhammad Attamimi, Tomoaki Nakamura, Takayuki Nagai
ICPR2
2012 Playmate robots that can act according to a child's mental state
abstract
We propose a playmate robot system that can play with a child. Unlike many therapeutic service robots, our proposed playmate system is implemented as a functionality of the domestic service robot with a high degree of freedom. This implies that the robot can play high-level games with children, i.e., beyond therapeutic play, using its physical features. The proposed system currently consists of ten play modules, including a chatbot with eye contact, card playing, and drawing. The algorithms of these modules are briefly discussed in this paper. To sustain the player's interest in the system, we also propose an action-selection strategy based on a transition model of the child's mental state. The robot can estimate the child's state and select an appropriate action in the course of play. A portion of the proposed algorithms was implemented on a real robot platform, and experiments were carried out to design and evaluate the proposed system.
Kasumi Abe, Akiko Iwasaki, Tomoaki Nakamura, Takayuki Nagai, Ayami Yokoyama, Takayuki Shimotomai, Takashi Omori
IROS3
2012 Online learning of concepts and words using multimodal LDA and hierarchical Pitman-Yor Language Model
abstract
In this paper, we propose an online algorithm for multimodal categorization based on the autonomously acquired multimodal information and partial words given by human users. For multimodal concept formation, multimodal latent Dirichlet allocation (MLDA) using Gibbs sampling is extended to an online version. We introduce a particle filter, which significantly improve the performance of the online MLDA, to keep tracking good models among various models with different parameters. We also introduce an unsupervised word segmentation method based on hierarchical Pitman-Yor Language Model (HPYLM). Since the HPYLM requires no predefined lexicon, we can make the robot system that learns concepts and words in completely unsupervised manner. The proposed algorithms are implemented on a real robot and tested using real everyday objects to show the validity of the proposed system.
Takaya Araki, Tomoaki Nakamura, Takayuki Nagai, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi
IROS2
2012 A planning method for efficient mobile manipulation considering ambiguity
abstract
In this study, we propose a system to enable robots to navigate more efficiently to target objects in order to carry out mobile manipulation tasks. Because a robot arm has limited reach, a target object cannot be grasped if the distance to the object exceeds the reach. At the same time, navigation errors may increase as the robot moves around in an environment. To mitigate these problems, we propose the use of maps that express the arm reachability and navigation reachability. Places with minimum probability of navigation errors and maximum tolerable navigation errors for the reach of the arm can be determined by integrating these two maps. We also consider the fact that the target object position is not always known. To cope with this problem, we introduce an object existence map that represents the ambiguity of the target object position, considering the past position. If the ambiguity is large, it is possible to select a place from where the object searching task can be performed with a wide angle of view by the robot. Efficient object searching can be achieved by changing the task objective from searching to grasping and vice versa. We conducted some mobile manipulation tasks to evaluate the proposed method. The results showed that the proposed system can perform such tasks effectively.
Muhammad Attamimi, Keisuke Ito, Tomoaki Nakamura, Takayuki Nagai
IROS3
2012 Bag of multimodal hierarchical dirichlet processes: Model of complex conceptual structure for intelligent robots
abstract
The formation of categories, which constitutes the basis of developing concepts, requires multimodal information with a complex structure. We propose a model called the bag of multimodal hierarchical Dirichlet processes (BoMHDP), which enables robots to form a variety of multimodal categories. The BoMHDP model is a collection of a large number of MHDP models, each of which has a different set of weights for sensory information. The weights work to realize selective attention and enable the formation of various types of categories (e.g., object, haptic, and color). The BoMHDP model is an extension of the HDP, and categorization is unsupervised. However, categories that are not natural for humans are also formed. Therefore, only the significant categories are selected through interaction between the user and the robot. At the same time, words obtained during the interaction are connected to the categories. Finally, categories, which are represented by words, are selected. The BoMHDP model was implemented on a robot platform and a preliminary experiment was conducted to validate it. The results revealed that various categories can be formed with the BoMHDP model. We also analyzed the formed conceptual structure by using multidimensional scaling. The results indicate that the complex conceptual structure was represented reasonably well with the BoMHDP model.
Tomoaki Nakamura, Takayuki Nagai, Naoto Iwahashi
IROS1
2011 Bag of multimodal LDA models for concept formation
abstract
In this paper a novel framework for multimodal categorization using Bag of multimodal LDA models is proposed. The main issue, which is tackled in this paper, is granularity of categories. The categories are not fixed but varied according to context. Selective attention is the key to model this granularity of categories. This fact motivates us to introduce various sets of weights to the perceptual information. Obviously, as the weights change, the categories vary. In the proposed model, various sets of weights and model structures are assumed. Then the multimodal LDA-based categorization is carried out many times that results in a variety of models. In order to make the categories (concepts) useful for inference, significant models should be selected. The selection process is carried out through the interaction between the robot and the user. These selected models enable the robot to infer unobserved properties of the object. For example, the robot can infer audio information only from its appearance. Furthermore, the robot can describe appearance of any objects using some suitable words, thanks to the connection between words and perceptual information. The proposed algorithm is implemented on a robot platform and preliminary experiment is carried out to validate the proposed algorithm.
Tomoaki Nakamura, Takayuki Nagai, Naoto Iwahashi
ICRA1
2011 Autonomous acquisition of multimodal information for online object concept formation by robots
abstract
This paper proposes a robot that acquires multi-modal information, i.e. auditory, visual, and haptic information, fully autonomous way using its embodiment. We also propose an online algorithm of multimodal categorization based on the acquired multimodal information and words, which are partially given by human users. The proposed framework makes it possible for the robot to learn object concepts naturally in everyday operation in conjunction with a small amount of linguistic information from human users. In order to obtain multimodal information, the robot detects an object on a fla surface. Then the robot grasps and shakes it for gaining haptic and auditory information. For obtaining visual information, the robot uses a hand held small observation table, so that the robot can control the viewpoints for observing the object. As for the multimodal concept formation, the multimodal LDA using Gibbs sampling is extended to the online version in this paper. The proposed algorithms are implemented on a real robot and tested using real everyday objects in order to show validity of the proposed system.
Takaya Araki, Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Mikio Nakano, Naoto Iwahashi
IROS2
2011 Multimodal categorization by hierarchical dirichlet process
abstract
In this paper, we propose a nonparametric Bayesian framework for categorizing multimodal sensory signals such as audio, visual, and haptic information by robots. The robot uses its physical embodiment to grasp and observe an object from various viewpoints as well as listen to the sound during the observation. The multimodal information enables the robot to form human-like object categories that are bases of intelligence. The proposed method is an extension of Hierarchical Dirichlet Process (HDP), which is a kind of nonparametric Bayesian models, to multimodal HDP (MHDP). MHDP can estimate the number of categories, while the parametric model, e.g. LDA-based categorization, requires to specify the number in advance. As this is an unsupervised learning method, a human user does not need to give any correct labels to the robot and it can classify objects autonomously. At the same time the proposed method provides a probabilistic framework for inferring object properties from limited observations. Validity of the proposed method is shown through some experimental results.
Tomoaki Nakamura, Takayuki Nagai, Naoto Iwahashi
IROS1
2010 Learning novel objects using out-of-vocabulary word segmentation and object extraction for home assistant robots
abstract
This paper presents a method for learning novel objects from audio-visual input. Objects are learned using out-of-vocabulary word segmentation and object extraction. The latter half of this paper is devoted to evaluations. We propose the use of a task adopted from the RoboCup@Home league as a standard evaluation for real world applications. We have implemented proposed method on a real humanoid robot and evaluated it through a task called “Supermarket”. The results reveal that our integrated system works well in the real application. In fact, our robot outperformed the maximum score obtained in RoboCup@Home 2009 competitions.
Muhammad Attamimi, Attamini Mizutani, Tomoaki Nakamura, Komei Sugiura, Takayuki Nagai, Naoto Iwahashi, Takashi Omori
ICRA3
2010 Real-time 3D visual sensor for robust object recognition
abstract
This paper presents a novel 3D measurement system, which yields both depth and color information in real time, by calibrating a time-of-flight and two CCD cameras. The problem of occlusions is solved by the proposed fast occluded-pixel detection algorithm. Since the system uses two CCD cameras, missing color information of occluded pixels is covered by one another. We also propose a robust object recognition using the 3D visual sensor. Multiple cues, such as color, texture and 3D (depth) information, are integrated in order to recognize various types of objects under varying lighting conditions. We have implemented the system on our autonomous robot and made the robot do recognition tasks (object learning, detection, and recognition) in various environments. The results revealed that the proposed recognition system provides far better performance than the previous system that is based only on color and texture information.
Muhammad Attamimi, Akira Mizutani, Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Mikio Nakano
IROS3
2010 Object concept modeling based on the relationship among appearance, usage and functions
abstract
In this paper, a novel object concept model, which encodes the relationship among appearance, functions and usage, is proposed. The essential attribute of an object (artifact) is its function that achieves a particular purpose. Therefore, the function model is constructed through observations from a camera at first. The function is defined as changes in the work object before and after tool use. At the same time, the usage model is constructed from observations of the hand shape, grasping parts, and contact points of the tool. And then, the proposed system learns the object concept that is based on the relationship among appearance, and learnt function and usage models. The object appearance is represented by SIFT (Scale Invariant Feature Transform). Since the proposed models are based on the graphical model, it is possible for the system to stochastically infer unobservable information from observed one. For example, the system can infer usage and/or functions of the tool visually through the proposed model. Some experimental results using the system, in which the proposed model is implemented, are shown to validate the proposed model.
Tomoaki Nakamura, Takayuki Nagai
IROS1
2009 Grounding of word meanings in multimodal concepts using LDA
abstract
In this paper we propose LDA-based framework for multimodal categorization and words grounding for robots. The robot uses its physical embodiment to grasp and observe an object from various view points as well as listen to the sound during the observing period. This multimodal information is used for categorizing and forming multimodal concepts. At the same time, the words acquired during the observing period are connected to the related concepts using multimodal LDA. We also provide a relevance measure that encodes the degree of connection between words and modalities. The proposed algorithm is implemented on a robot platform and some experiments are carried out to evaluate the algorithm. We also demonstrate a simple conversation between a user and the robot based on the learned model.
Tomoaki Nakamura, Takayuki Nagai, Naoto Iwahashi
IROS1
2007 Multimodal object categorization by a robot
abstract
In this paper unsupervised object categorization by robots is examined. We propose an unsupervised multimodal categorization based on audio-visual and haptic information. The robot uses its physical embodiment to grasp and observe an object from various view points as well as listen to the sound during the observation. The proposed categorization method is an extension of probabilistic Latent Semantic Analysis(pLSA), which is a statistical technique. At the same time the proposed method provides a probabilistic framework for inferring the object property from limited observations. Validity of the proposed method is shown through some experimental results.
Tomoaki Nakamura, Takayuki Nagai, Naoto Iwahashi
IROS1