EDBT 2026 Demo / reviewers in the wild / expert
Tadahiro Taniguchi
dblp:25/4674
· DBLP profile ↗
58ranked-venue papers
7as first author
26since 2021 · last 2025
0000-0002-5682-2076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 4 first-author · 24 since 2021Systems, architecture and hardware · 29 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interpretable Anomaly Detection in a Hippocampal Formation-Inspired Spatial Cognition Model
Takeshi Nakashima, Akira Taniguchi, Tadahiro Taniguchi, Hiroshi Yamakawa |
ICONIP (4) | 3 |
| 2025 | Reward-Independent Messaging for Decentralized Multi-agent Reinforcement Learning
Naoto Yoshida, Tadahiro Taniguchi |
ICONIP (1) | 2 |
| 2025 | Haptic-Informed ACT with a Soft Gripper and Recovery-Informed Training for Pseudo Oocyte ManipulationabstractIn this paper, we introduce Haptic-Informed ACT, an advanced robotic system for pseudo oocyte manipulation, integrating multimodal information and Action Chunking with Transformers (ACT). Traditional automation methods for oocyte transfer rely heavily on visual perception, often requiring human supervision due to biological variability and environmental disturbances. Haptic-Informed ACT enhances ACT by incorporating haptic feedback, enabling real-time grasp failure detection and adaptive correction. Additionally, we introduce a 3D-printed TPU soft gripper to facilitate delicate manipulations. Experimental results demonstrate that Haptic-Informed ACT improves the task success rate, robustness, and adaptability compared to conventional ACT, particularly in dynamic environments. These findings highlight the potential of multimodal learning in robotics for biomedical automation. Pedro Miguel Uriguen Eljuri, Hironobu Shibata, Katsuyoshi Maeyama, Tadahiro Taniguchi |
IROS | 5 |
| 2025 | Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View InstructionsabstractDaily life support robots must interpret ambiguous verbal instructions involving demonstratives such as "Bring me that cup," even when objects or users are out of the robot’s view. Existing approaches to exophora resolution primarily rely on visual data and thus fail in real-world scenarios where the object or user is not visible. We propose Multimodal Interactive Exophora resolution with user Localization (MIEL), which is a multimodal exophora resolution framework leveraging sound source localization (SSL), semantic mapping, visual-language models (VLMs), and interactive questioning with GPT-4o. Our approach first constructs a semantic map of the environment and estimates candidate objects from a linguistic query with the user’s skeletal data. SSL is utilized to orient the robot toward users who are initially outside its visual field, enabling accurate identification of user gestures and pointing directions. When ambiguities remain, the robot proactively interacts with the user, employing GPT-4o to formulate clarifying questions. Experiments in a real-world environment showed results that were approximately 1.3 times better when the user was visible to the robot and 2.0 times better when the user was not visible to the robot, compared to the methods without SSL and interactive questioning. The project website is https://emergentsystemlabstudent.github.io/MIEL/. Akira Oyama, Shoichi Hasegawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi |
RO-MAN | 5 |
| 2024 | Lewis's Signaling Game as beta-VAE For Natural Word Lengths and SegmentsabstractAs a sub-discipline of evolutionary and computational linguistics, emergent communication (EC) studies communication protocols, called emergent languages, arising in simulations where agents communicate. A key goal of EC is to give rise to languages that share statistical properties with natural languages. In this paper, we reinterpret Lewis's signaling game, a frequently used setting in EC, as beta-VAE and reformulate its objective function as ELBO. Consequently, we clarify the existence of prior distributions of emergent languages and show that the choice of the priors can influence their statistical properties. Specifically, we address the properties of word lengths and segmentation, known as Zipf's law of abbreviation (ZLA) and Harris's articulation scheme (HAS), respectively. It has been reported that the emergent languages do not follow them when using the conventional objective. We experimentally demonstrate that by selecting an appropriate prior distribution, more natural segments emerge, while suggesting that the conventional one prevents the languages from following ZLA and HAS. Ryo Ueda, Tadahiro Taniguchi |
ICLR | 2 |
| 2024 | Helical Control in Latent Space: Enhancing Robotic Craniotomy Precision in Uncertain EnvironmentsabstractIn this paper, we introduce a double-stage transfer learning framework based on expert data. It employs probabilistic graphical models to effectively capture helical periodic features in the latent space, integrating Bayesian variational inference and neural networks for implementation. Compared to traditional methods, it achieves high precision and stable control even in environments with limited observation signals and high noise levels. We have successfully applied this method to a biomedical task of a simulated cranial window procedure. Preliminary results show promising performance comparable to those of human experts with only image information, further validating the efficacy of the proposed method. Jessica Ziyu Qu, Tadahiro Taniguchi |
ICRA | 3 |
| 2024 | Goal Estimation-based Adaptive Shared Control for Brain-Machine Interfaces Remote Robot NavigationabstractIn this study, we propose a shared control method for teleoperated mobile robots using brain-machine interfaces (BMI). The control commands generated through BMI for robot operation face issues of low input frequency, discreteness, and uncertainty due to noise. To address these challenges, our method estimates the user’s intended goal from their commands and uses this goal to generate auxiliary commands through the autonomous system that are both at a higher input frequency and more continuous. Furthermore, by defining the confidence level of the estimation, we adaptively calculated the weights for combining user and autonomous commands, thus achieving shared control. We conducted navigation experiments in both simulated environments and participant experiments in real environments including user ratings, using a pseudo-BMI setup. As a result, the proposed method significantly reduced obstacle collisions in all experiments. It markedly shortened path lengths under almost all conditions in simulations and, in participant experiments, especially when user inputs become more discrete and noisy (p<0.01). Furthermore, under such challenging conditions, it was demonstrated that users could operate more easily, with greater confidence, and at a comfortable pace through this system. Tomoka Muraoka, Tatsuya Aoki, Masayuki Hirata, Tadahiro Taniguchi, Takato Horii, Takayuki Nagai |
IROS | 4 |
| 2024 | A Contact Model based on Denoising Diffusion to Learn Variable Impedance Control for Contact-rich ManipulationabstractIn this paper, a novel approach is proposed for learning robot control in contact-rich tasks such as wiping, by developing Diffusion Contact Model (DCM). Previous methods of learning such tasks relied on impedance control with time-varying stiffness tuning by performing Bayesian optimization by trial-and-error with robots. The proposed approach aims to reduce the cost of robot operation by predicting the robot contact trajectories from the variable stiffness inputs and using neural models. However, contact dynamics are inherently highly nonlinear, and their simulation requires iterative computations such as convex optimization. Moreover, approximating such computations by using finite-layer neural models is difficult. To overcome these limitations, the proposed DCM used the denoising diffusion models that could simulate the complex dynamics via iterative computations, thus improving the prediction accuracy. Stiffness tuning experiments conducted in simulated and real environments showed that the DCM achieved comparable performance to a conventional robot-based optimization method while reducing the number of robot trials. Masashi Okada, Mayumi Komatsu, Tadahiro Taniguchi |
IROS | 3 |
| 2024 | Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic MapabstractRobots that assist humans in their daily lives should be able to locate specific instances of objects in an environment that match a user’s desired objects. This task is known as instance-specific image goal navigation (InstanceImageNav), which requires a model that can distinguish different instances of an object within the same class. A significant challenge in robotics is that when a robot observes the same object from various 3D viewpoints, its appearance may differ significantly, making it difficult to recognize and locate accurately. In this paper, we introduce a method called SimView, which leverages multi-view images based on a 3D semantic map of an environment and self-supervised learning using SimSiam to train an instance-identification model on-site. The effectiveness of our approach was validated using a photorealistic simulator, Habitat Matterport 3D, created by scanning actual home environments. Our results demonstrate a 1.7-fold improvement in task accuracy compared with contrastive language-image pre-training (CLIP), a pre-trained multimodal contrastive learning method for object searching. This improvement highlights the benefits of our proposed fine-tuning method in enhancing the performance of assistive robots in InstanceImageNav tasks. The project website is https://emergentsystemlabstudent.github.io/MultiViewRetrieve/. Taichi Sakaguchi, Akira Taniguchi, Yoshinobu Hagiwara, Lotfi El Hafi, Shoichi Hasegawa, Tadahiro Taniguchi |
IROS | 6 |
| 2024 | Stable Object Placing using Curl and Diff Features of Vision-based Tactile SensorsabstractEnsuring stable object placement is crucial to prevent objects from toppling over, breaking, or causing spills. When an object makes initial contact to a surface, and some force is exerted, the moment of rotation caused by the instability of the object’s placing can cause the object to rotate in a certain direction (henceforth referred to as direction of corrective rotation). Existing methods often employ a Force/Torque (F/T) sensor to estimate the direction of corrective rotation by detecting the moment of rotation as a torque. However, its effectiveness may be hampered by sensor noise and the tension of the external wiring of robot cables. To address these issues, we propose a method for stable object placing using GelSights, vision-based tactile sensors, as an alternative to F/T sensors. Our method estimates the direction of corrective rotation of objects using the displacement of the black dot pattern on the elastomeric surface of GelSight. We calculate the Curl from vector analysis, indicative of the rotational field magnitude and direction of the displacement of the black dots pattern. Simultaneously, we calculate the difference (Diff) of displacement between the left and right fingers’ GelSight’s black dots. Then, the robot can manipulate the objects’ pose using Curl and Diff features, facilitating stable placing. Across experiments, handling 18 differently characterized objects, our method achieves precise placing accuracy (less than 1-degree error) in nearly 100% of cases. Kuniyuki Takahashi, Shimpei Masuda, Tadahiro Taniguchi |
IROS | 3 |
| 2024 | SPPP: Stochastic Penalty Path Planning for MAPF Under Conditions of Uncertain Travel Time
Atsuyoshi Kita, Hideki Aoyama, Tadahiro Taniguchi |
PRIMA | 3 |
| 2024 | Pointing Frame Estimation With Audio-Visual Time Series Data for Daily Life Service RobotsabstractDaily life support robots in the home environment interpret the user's pointing and understand the instructions, thereby increasing the number of instructions accomplished. This study aims to improve the estimation performance of pointing frames by using speech information when a person gives pointing or verbal instructions to the robot. The estimation of the pointing frame, which represents the moment when the user points, can help the user understand the instructions. Therefore, we perform pointing frame estimation using a time-series model, utilizing the user's speech, images, and speech-recognized text observed by the robot. In our experiments, we set up realistic communication conditions, such as speech containing everyday conversation, non-upright posture, actions other than pointing, and reference objects outside the robot's field of view. The results showed that adding speech information improved the estimation performance, especially the Transformer model with Mel-Spectrogram as a feature. This study will lead to be applied to object localization and action planning in 3D environments by robots in the future. The project website is https://emergentsystemlabstudent.github.io/PointingImgEst/. Hikaru Nakagawa, Shoichi Hasegawa, Yoshinobu Hagiwara, Akira Taniguchi, Tadahiro Taniguchi |
SMC | 5 |
| 2023 | Holographic CCG ParsingabstractWe propose a method for formulating CCG as a recursive composition in a continuous vector space.Recent CCG supertagging and parsing models generally demonstrate high performance, yet rely on black-box neural architectures to implicitly model phrase structure dependencies.Instead, we leverage the method of holographic embeddings (Nickel et al., 2016) as a compositional operator to explicitly model the dependencies between words and phrase structures in the embedding space.Experimental results revealed that holographic composition effectively improves the supertagging accuracy to achieve state-of-the-art parsing performance when using a C&C parser.The proposed span-based parsing algorithm using holographic composition achieves performance comparable to state-of-the-art neural parsing with Transformers.Furthermore, our model can semantically and syntactically infill text at the phrase level due to the decomposability of holographic composition. Ryosuke Yamaki, Tadahiro Taniguchi, Daichi Mochihashi |
ACL (1) | 2 |
| 2023 | Representation Uncertainty in Self-Supervised Learning as Variational InferenceabstractIn this study, a novel self-supervised earning (SSL) method is proposed, which considers SSL in terms of variational inference to learn not only representation but also representation uncertainties. SSL is a method of learning representations without labels by maximizing the similarity between image representations of different augmented views of an image. Meanwhile, variational autoencoder (VAE) is an unsupervised representation learning method that trains a probabilistic generative model with variational inference. Both VAE and SSL can learn representations without labels, but their relationship has not been investigated in the past. Herein, the theoretical relationship between SSL and variational inference has been clarified. Furthermore, a novel method, namely variational inference SimSiam (VI-SimSiam), has been proposed. VI-SimSiam can predict the representation uncertainty by interpreting SimSiam with variational inference and defining the latent space distribution. The present experiments qualitatively show that VI-SimSiam could learn uncertainty by comparing input images and predicted uncertainties. Additionally, we described a relationship between estimated uncertainty and classification accuracy. Hiroki Nakamura, Masashi Okada, Tadahiro Taniguchi |
ICCV | 3 |
| 2023 | Goal-Image Conditioned Dynamic Cable Manipulation through Bayesian Inference and Multi-Objective Black-Box OptimizationabstractTo perform dynamic cable manipulation to realize the configuration specified by a target image, we formulate dynamic cable manipulation as a stochastic forward model. Then, we propose a method to handle uncertainty by maximizing the expectation, which also considers estimation errors of the trained model. To avoid issues like multiple local minima and requirement of differentiability by gradient-based methods, we propose using a black-box optimization (BBO) to optimize joint angles to realize a goal image. Among BBO, we use the Tree-structured Parzen Estimator (TPE), a type of Bayesian optimization. By incorporating constraints into the TPE, the optimized joint angles are constrained within the range of motion. Since TPE is population-based, it is better able to detect multiple feasible configurations using the estimated inverse model. We evaluated image similarity between the target and cable images captured by executing the robot using optimal transport distance. The results show that the proposed method improves accuracy compared to conventional gradient-based approaches and methods that use deterministic models that do not consider uncertainty. Kuniyuki Takahashi, Tadahiro Taniguchi |
ICRA | 2 |
| 2023 | A Bayesian Reinforcement Learning Method for Periodic Robotic Control Under Significant UncertaintyabstractThis paper addresses the lack of research on periodic reinforcement learning for physical robot control by presenting a 3-phase periodic Bayesian reinforcement learning method for uncertain environments. Drawing on cognition theory, the proposed approach achieves effective convergence with fewer training episodes. The coach-based demonstration phase narrows the search space and establishes a foundation for a coarse-to-fine control strategy. The reconnaissance phase enhances adaptability by discovering a valuable global repre-sentation, and the operation phase produces accurate robotic control by applying the learned representation and periodically updating local information. Comparative analysis with state-of-the-art methods validates the efficacy of our approach on exemplar control tasks in simulation and a biomedical project involving a simulated cranial window task. Pedro Miguel Uriguen Eljuri, Tadahiro Taniguchi |
IROS | 3 |
| 2023 | Learning Compliant Stiffness by Impedance Control-Aware Task Segmentation and Multi-Objective Bayesian Optimization with PriorsabstractRather than traditional position control, impedance control is preferred to ensure the safe operation of industrial robots programmed from demonstrations. However, variable stiffness learning studies have focused on task performance rather than safety (or compliance). Thus, this paper proposes a novel stiffness learning method to satisfy both task performance and compliance requirements. The proposed method optimizes the task and compliance objectives ($T/C$objectives) simultaneously via multi-objective Bayesian optimization. We define the stiffness search space by segmenting a demonstration into task phases, each with constant responsible stiffness. The segmentation is performed by identifying impedance control-aware switching linear dynamics (IC-SLD) from the demonstration. We also utilize the stiffness obtained by proposed IC-SLD as priors for efficient optimization. Experiments on simulated tasks and a real robot demonstrate that IC-SLD-based segmentation and the use of priors improve the optimization efficiency compared to existing baseline methods. Masashi Okada, Mayumi Komatsu, Ryo Okumura, Tadahiro Taniguchi |
IROS | 4 |
| 2023 | Exophora Resolution of Linguistic Instructions with a Demonstrative based on Real-World Multimodal InformationabstractTo enable a robot to provide support in a home environment through human-robot interaction, exophora resolution is crucial for accurately identifying the target of ambiguous linguistic instructions, which may include a demonstrative, such as “Take that one”. Unlike endophora resolution, which involves predicting the corresponding word from given sentences, exophora resolution necessitates comprehensive utilization of external real-world information to identify and disambiguate the target from the on-site environment. This study aims to resolve ambiguity in language instructions containing a demonstrative through exophora resolution, utilizing real-world multimodal information. The robot accomplishes this by using three types of information: 1) object categories, 2) demonstratives, and 3) pointing, as well as knowledge about objects obtained from the robot’s pre-exploration of the environment. We evaluated the accuracy of object identification under multiple conditions by identifying a user-indicated object in a field that mimics a home environment. Our results demonstrate that our proposed method of exophora resolution using multimodal information can identify the target with two to three times higher accuracy than baseline methods in cases where information is missing. Akira Oyama, Shoichi Hasegawa, Hikaru Nakagawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi |
RO-MAN | 6 |
| 2023 | Active Semantic Mapping for Household Robots: Rapid Indoor Adaptation and Reduced User BurdenabstractActive semantic mapping is essential for service robots to quickly capture both the map of an environment and its spatial meaning, while also minimizing the burden on users during robot operation and data collection. SpCoSLAM, a method of semantic mapping with place categorization and simultaneous localization and mapping (SLAM), is well suited to environmental adaptation, as it is not limited to predefined labels. However, SpCoSLAM presents two issues that increase the burden on users: 1) users struggle to efficiently determine a destination for the robot's quick adaptation, and 2) providing instructions to the robot becomes repetitive and cumbersome. To address these challenges, we propose Active-SpCoSLAM, which enables the robot to actively explore uncharted areas and employs CLIP for image captioning to provide a flexible vocabulary that replaces human instructions. The robot determines its actions by calculating information gain integrated from both semantics and SLAM uncertainties. We conducted experiments in a simulated environment, comparing the proposed method to other methods in terms of efficiency and applicability to object discovery tasks. Additionally, we tested the proposed method, which combines user instruction and CLIP, in a real environment. Our results demonstrated that the robot explored its environment with approximately five fewer iterations and 11 minutes faster compared to the case of random exploration. Moreover, our method achieved a higher success rate in object discovery tasks during earlier stages of learning compared to other methods. In conclusion, the proposed method rapidly covers an environment while gathering useful data for object discovery tasks, thus reducing the burden on users and enhancing the robot's adaptability. The project website is https://tomochika-ishikawa.github.io/Active-SpCoSLAM/. Tomochika Ishikawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi |
SMC | 4 |
| 2022 | Multimodal Object Categorization with Reduced User Load through Human-Robot Interaction in Mixed RealityabstractEnabling robots to learn from interactions with users is essential to perform service tasks. However, as a robot categorizes objects from multimodal information obtained by its sensors during interactive onsite teaching, the inferred names of unknown objects do not always match the human user's expectation, especially when the robot is introduced to new environments. Confirming the learning results through natural speech interaction with the robot often puts an additional burden on the user who can only listen to the robot to validate the results. Therefore, we propose a human-robot interface to reduce the burden on the user by visualizing the inferred results in mixed reality (MR). In particular, we evaluate the proposed interface on the system usability scale (SUS) and the NASA task load index (NASA-TLX) with three experimental object categorization scenarios based on multimodal latent Dirichlet allocation (MLDA) in which the robot: 1) does not share the inferred results with the user at all, 2) shares the inferred results through speech interaction with the user (baseline), and 3) shares the inferred results with the user through an MR interface (proposed). We show that providing feedback through an MR interface significantly reduces the temporal, physical, and mental burden on the human user compared to speech interaction with the robot. Hitoshi Nakamura, Lotfi El Hafi, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi |
IROS | 5 |
| 2022 | DreamingV2: Reinforcement Learning with Discrete World Models without ReconstructionabstractThe present paper proposes a novel reinforce-ment learning method with world models, DreamingV2, a collaborative extension of DreamerV2 and Dreaming. Dream- erV2 is a cutting-edge model-based reinforcement learning from pixels that uses discrete world models to represent latent states with categorical variables. Dreaming is also a form of reinforcement learning from pixels that attempts to avoid the auto encoding process in general world model training by involving a reconstruction-free contrastive learning objective. The proposed DreamingV2 is a novel approach of adopting both the discrete representation of DreamingV2 and the reconstruction-free objective of Dreaming. Compared to DreamerV2 and other recent model-based methods without reconstruction, DreamingV2 achieves the best scores on five simulated challenging 3D robot arm tasks. We believe that DreamingV2 will be a reliable solution for robot learning since its discrete representation is suitable to describe discontinuous environments, and the reconstruction-free fashion well manages complex vision observations. Masashi Okada, Tadahiro Taniguchi |
IROS | 2 |
| 2022 | Tactile-Sensitive NewtonianVAE for High-Accuracy Industrial Connector InsertionabstractAn industrial connector insertion task requires submillimeter positioning and grasp pose compensation for a plug. Thus, highly accurate estimation of the relative pose between a plug and socket is fundamental for achieving the task. World models are promising technologies for visuomotor control because they obtain appropriate state representation to jointly optimize feature extraction and latent dynamics model. Recent studies show that the Newto-nianVAE, a type of the world model, acquires latent space equivalent to mapping from images to physical coordinates. Proportional control can be achieved in the latent space of NewtonianVAE. However, applying NewtonianVAE to high-accuracy industrial tasks in physical environments is an open problem. Moreover, the existing framework does not consider the grasp pose compensation in the obtained latent space. In this work, we proposed tactile-sensitive Newtonian-VAE and applied it to a USB connector insertion with grasp pose variation in the physical environments. We adopted a GelSight-type tactile sensor and estimated the insertion position compensated by the grasp pose of the plug. Our method trains the latent space in an end-to-end manner, and no additional engineering and annotation are required. Simple proportional control is available in the obtained latent space. Moreover, we showed that the original NewtonianVAE fails in some situations, and demonstrated that domain knowledge induction improves model accuracy. This domain knowledge can be easily obtained using robot specification and grasp pose error measurement. We demonstrated that our proposed method achieved a 100% success rate and 0.3 mm positioning accuracy in the USB connector insertion task in the physical environment. It outperformed SOTA CNN-based two-stage goal pose regression with grasp pose compensation using coordinate transformation. Ryo Okumura, Nobuki Nishio, Tadahiro Taniguchi |
IROS | 3 |
| 2022 | A whole brain probabilistic generative model: Toward realizing cognitive architectures for developmental robotsabstractBuilding a human-like integrative artificial cognitive system, that is, an artificial general intelligence (AGI), is the holy grail of the artificial intelligence (AI) field. Furthermore, a computational model that enables an artificial system to achieve cognitive development will be an excellent reference for brain and cognitive science. This paper describes an approach to develop a cognitive architecture by integrating elemental cognitive modules to enable the training of the modules as a whole. This approach is based on two ideas: (1) brain-inspired AI, learning human brain architecture to build human-level intelligence, and (2) a probabilistic generative model (PGM)-based cognitive architecture to develop a cognitive system for developmental robots by integrating PGMs. The proposed development framework is called a whole brain PGM (WB-PGM), which differs fundamentally from existing cognitive architectures in that it can learn continuously through a system based on sensory-motor information. In this paper, we describe the rationale for WB-PGM, the current status of PGM-based elemental cognitive modules, their relationship with the human brain, the approach to the integration of the cognitive modules, and future challenges. Our findings can serve as a reference for brain studies. As PGMs describe explicit informational relationships between variables, WB-PGM provides interpretable guidance from computational sciences to brain science. By providing such information, researchers in neuroscience can provide feedback to researchers in AI and robotics on what the current models lack with reference to the brain. Further, it can facilitate collaboration among researchers in neuro-cognitive sciences as well as AI and robotics. Tadahiro Taniguchi, Hiroshi Yamakawa, Takayuki Nagai, Kenji Doya, Masamichi Sakagami, Tomoaki Nakamura, Akira Taniguchi |
Neural Networks | 1 |
| 2021 | Dreaming: Model-based Reinforcement Learning by Latent Imagination without ReconstructionabstractIn the present paper, we propose a decoder-free extension of Dreamer, a leading model-based reinforcement learning (MBRL) method from pixels. Dreamer is a sample- and cost-efficient solution to robot learning, as it is used to train latent state-space models based on a variational autoencoder and to conduct policy optimization by latent trajectory imagination. However, this autoencoding based approach often causes object vanishing, in which the autoencoder fails to perceives key objects for solving control tasks, and thus significantly limiting Dreamer's potential. This work aims to relieve this Dreamer's bottleneck and enhance its performance by means of removing the decoder. For this purpose, we firstly derive a likelihood- free and InfoMax objective of contrastive learning from the evidence lower bound of Dreamer. Secondly, we incorporate two components, (i) independent linear dynamics and (ii) the random crop data augmentation, to the learning scheme so as to improve the training performance. In comparison to Dreamer and other recent model-free reinforcement learning methods, our newly devised Dreamer with InfoMax and without generative decoder (Dreaming) achieves the best scores on 5 difficult simulated robotics tasks, in which Dreamer suffers from object vanishing. Masashi Okada, Tadahiro Taniguchi |
ICRA | 2 |
| 2021 | StarGAN-VC+ASR: StarGAN-Based Non-Parallel Voice Conversion Regularized by Automatic Speech RecognitionabstractPreserving the linguistic content of input speech is essential during voice conversion (VC). The star generative adversarial network-based VC method (StarGAN-VC) is a recently developed method that allows non-parallel many-to-many VC. Although this method is powerful, it can fail to preserve the linguistic content of input speech when the number of available training samples is extremely small. To overcome this problem, we propose the use of automatic speech recognition to assist model training, to improve StarGAN-VC, especially in low-resource scenarios. Experimental results show that using our proposed method, StarGAN-VC can retain more linguistic information than vanilla StarGAN-VC. Shoki Sakamoto, Akira Taniguchi, Tadahiro Taniguchi, Hirokazu Kameoka |
Interspeech | 3 |
| 2021 | World model learning and inferenceabstractUnderstanding information processing in the brain-and creating general-purpose artificial intelligence-are long-standing aspirations of scientists and engineers worldwide. The distinctive features of human intelligence are high-level cognition and control in various interactions with the world including the self, which are not defined in advance and are vary over time. The challenge of building human-like intelligent machines, as well as progress in brain science and behavioural analyses, robotics, and their associated theoretical formalisations, speaks to the importance of the world-model learning and inference. In this article, after briefly surveying the history and challenges of internal model learning and probabilistic learning, we introduce the free energy principle, which provides a useful framework within which to consider neuronal computation and probabilistic world models. Next, we showcase examples of human behaviour and cognition explained under that principle. We then describe symbol emergence in the context of probabilistic modelling, as a topic at the frontiers of cognitive robotics. Lastly, we review recent progress in creating human-like intelligence by using novel probabilistic programming languages. The striking consensus that emerges from these studies is that probabilistic descriptions of learning and inference are powerful and effective ways to create human-like artificial intelligent machines and to understand intelligence in the context of how humans interact with their world. Karl J. Friston, Rosalyn J. Moran, Yukie Nagai, Tadahiro Taniguchi, Hiroaki Gomi, Josh Tenenbaum |
Neural Networks | 4 |
| 2020 | Multi-person Pose Tracking using Sequential Monte Carlo with Probabilistic Neural Pose PredictorabstractIt is an effective strategy for the multi-person pose tracking task in videos to employ prediction and pose matching in a frame-by-frame manner. For this type of approach, uncertainty-aware modeling is essential because precise prediction is impossible. However, previous studies have relied on only a single prediction without incorporating uncertainty, which can cause critical tracking errors if the prediction is unreliable. This paper proposes an extension to this approach with Sequential Monte Carlo (SMC). This naturally reformulates the tracking scheme to handle multiple predictions (or hypotheses) of poses, thereby mitigating the negative effect of prediction errors. An important component of SMC, i.e., a proposal distribution, is designed as a probabilistic neural pose predictor, which can propose diverse and plausible hypotheses by incorporating epistemic uncertainty and heteroscedastic aleatoric uncertainty. In addition, a recurrent architecture is introduced to our neural modeling to utilize time-sequence information of poses to manage difficult situations, such as the frequent disappearance and reappearances of poses. Compared to existing baselines, the proposed method achieves a state-of-the-art MOTA score on the PoseTrack2018 validation dataset by reducing approximately 50% of tracking errors from a state-of-the art baseline method. Masashi Okada, Shinji Takenaka, Tadahiro Taniguchi |
ICRA | 3 |
| 2020 | Blind Bin Picking of Small Screws Through In-finger Manipulation With Compliant Robotic FingersabstractAlthough picking up objects a few centimeters in size is a common task, achieving such ability in a robot manipulator remains challenging. We take a step toward solving this problem by focusing on the task of picking a 1.0-cm screw from a bulk bin using only tactile information to achieve the task. Inspired by how humans pick up small objects from a bin, we propose a "grasp-separate" strategy for robotic picking, which involves grasping many objects first and then separating a single object through manipulation in the fingers, for robotic picking. Based on this strategy, we developed a tactile-based screw bin-picking system. We trained a convolution neural network to estimate the number of screws in the fingers first and built a controller that generates manipulation behaviors to separate a screw using reinforcement learning. To compensate for the low resolution of off-the-shelf tactile sensor arrays, we adopted active sensing, which uses observations obtained during a predefined simple movement. We show that this approach enhances the estimation accuracy and manipulation performance. Furthermore, to enable flexible finger motion, such as between the thumb and the index finger in a human hand, we propose a soft robot finger structure that leverages compliant materials. A soft actor-critic algorithm successfully found dexterous screw separation behaviors in compliant soft robotic fingers. In the evaluation, the system obtained an average success rate of 80%, which was difficult to achieve without the grasp-separate manipulation technique. Matthew Ishige, Takuya Umedachi, Yoshihisa Ijiri, Tadahiro Taniguchi, Yoshihiro Kawahara |
IROS | 4 |
| 2020 | SpCoMapGAN: Spatial Concept Formation-based Semantic Mapping with Generative Adversarial NetworksabstractIn semantic mapping, which connects semantic information to an environment map, it is a challenging task for robots to deal with both local and global information of environments. In addition, it is important to estimate semantic information of unobserved areas from already acquired partial observations in a newly visited environment. On the other hand, previous studies on spatial concept formation enabled a robot to relate multiple words to places from bottom-up observations even when the vocabulary was not provided beforehand. However, the robot could not transfer global information related to the room arrangement between semantic maps from other environments. In this paper, we propose SpCoMapGAN, which generates the semantic map in a newly visited environment by training an inference model using previously estimated semantic maps. SpCoMapGAN uses generative adversarial networks (GANs) to transfer semantic information based on room arrangements to a newly visited environment. Our proposed method assigns semantics to the map of an unknown environment using the prior distribution of the map trained in known environments and the multimodal observations made in the unknown environment. We experimentally show in simulation that SpCoMapGAN can use global information for estimating the semantic map and is superior to previous methods. Finally, we also demonstrate in a real environment that SpCoMapGAN can accurately 1) deal with local information, and 2) acquire the semantic information of real places. Yuki Katsumata, Akira Taniguchi, Lotfi El Hafi, Yoshinobu Hagiwara, Tadahiro Taniguchi |
IROS | 5 |
| 2020 | PlaNet of the Bayesians: Reconsidering and Improving Deep Planning Network by Incorporating Bayesian InferenceabstractIn the present paper, we propose an extension of the Deep Planning Network (PlaNet), also referred to as PlaNet of the Bayesians (PlaNet-Bayes). There has been a growing demand in model predictive control (MPC) in partially observable environments in which complete information is unavailable because of, for example, lack of expensive sensors. PlaNet is a promising solution to realize such latent MPC, as it is used to train state-space models via model-based reinforcement learning (MBRL) and to conduct planning in the latent space. However, recent state-of-the-art strategies mentioned in MBRR literature, such as involving uncertainty into training and planning, have not been considered, significantly suppressing the training performance. The proposed extension is to make PlaNet uncertainty-aware on the basis of Bayesian inference, in which both model and action uncertainty are incorporated. Uncertainty in latent models is represented using a neural network ensemble to approximately infer model posteriors. The ensemble of optimal action candidates is also employed to capture multimodal uncertainty in the optimality. The concept of the action ensemble relies on a general variational inference MPC (VI-MPC) framework and its instance, probabilistic action ensemble with trajectory sampling (PaETS). In this paper, we extend VI-MPC and PaETS, which have been originally introduced in previous literature, to address partially observable cases. We experimentally compare the performances on continuous control tasks, and conclude that our method can consistently improve the asymptotic performance compared with PlaNet. Masashi Okada, Norio Kosaka, Tadahiro Taniguchi |
IROS | 3 |
| 2020 | Domain-Adversarial and -Conditional State Space Model for Imitation LearningabstractState representation learning (SRL) in partially observable Markov decision processes has been studied to learn abstract features of data useful for robot control tasks. For SRL, acquiring domain-agnostic states is essential for achieving efficient imitation learning. Without these states, imitation learning is hampered by domain-dependent information useless for control. However, existing methods fail to remove such disturbances from the states when the data from experts and agents show large domain shifts. To overcome this issue, we propose a domain-adversarial and -conditional state space model (DAC-SSM) that enables control systems to obtain domain-agnostic and task- and dynamics-aware states. DAC-SSM jointly optimizes the state inference, observation reconstruction, forward dynamics, and reward models. To remove domain-dependent information from the states, the model is trained with domain discriminators in an adversarial manner, and the reconstruction is conditioned on domain labels. We experimentally evaluated the model predictive control performance via imitation learning for continuous control of sparse reward tasks in simulators and compared it with the performance of the existing SRL method. The agents from DAC-SSM achieved performance comparable to experts and more than twice the baselines. We conclude domain-agnostic states are essential for imitation learning that has large domain shifts and can be obtained using DAC-SSM. Ryo Okumura, Masashi Okada, Tadahiro Taniguchi |
IROS | 3 |
| 2019 | Evaluation of Word Representations in Grounding Natural Language Instructions Through Computational Human-Robot InteractionabstractIn order to interact with people in a natural way, a robot must be able to link words to objects and actions. Although previous studies in the literature have investigated grounding, they did not consider grounding of unknown synonyms. In this paper, we introduce a probabilistic model for grounding unknown synonymous object and action names using cross-situational learning. The proposed Bayesian learning model uses four different word representations to determine synonymous words. Afterwards, they are grounded through geometric characteristics of objects and kinematic features of the robot joints during action execution. The proposed model is evaluated through an interaction experiment between a human tutor and HSR robot. The results show that semantic and syntactic information both enable grounding of unknown synonyms and that the combination of both achieves the best grounding. Oliver Roesler, Amir Aly, Tadahiro Taniguchi, Yoshikatsu Hayashi |
HRI | 3 |
| 2018 | Towards Understanding Object-Directed Actions: A Generative Model for Grounding Syntactic Categories of Speech Through Visual PerceptionabstractCreating successful human-robot collaboration requires robots to have high-level cognitive functions that could allow them to understand human language and actions in space. To meet this target, an elusive challenge that we address in this paper is to understand object-directed actions through grounding language based on visual cues representing the dynamics of human actions on objects, object characteristics (color and geometry), and spatial relationships between objects in a tabletop scene. The proposed probabilistic framework investigates unsupervised Part-of-Speech (POS) tagging to determine syntactic categories of words so as to infer grammatical structure of language. The dynamics of object-directed actions are characterized through the locations of the human arm joints - modeled on a Hidden Markov Model (HMM) - while manipulating objects, in addition to those of objects represented in 3D point clouds. These corresponding point clouds to segmented objects encode geometric features and spatial semantics of referents and landmarks in the environment. The proposed Bayesian learning model is successfully evaluated through interaction experiments between a human user and Toyota HSR robot in space. Amir Aly, Tadahiro Taniguchi |
ICRA | 2 |
| 2018 | Acceleration of Gradient-Based Path Integral Method for Efficient Optimal and Inverse Optimal ControlabstractThis paper deals with a new accelerated path integral method, which iteratively searches optimal controls with a small number of iterations. This study is based on the recent observations that a path integral method for reinforcement learning can be interpreted as gradient descent. This observation also applies to an iterative path integral method for optimal control, which sets a convincing argument for utilizing various optimization methods for gradient descent, such as momentum-based acceleration, step-size adaptation and their combination. We introduce these types of methods to the path integral and demonstrate that momentum-based methods, like Nesterov Accelerated Gradient and Adam, can significantly improve the convergence rate to search for optimal controls in simulated control systems. We also demonstrate that the accelerated path integral could improve the performance on model predictive control for various vehicle navigation tasks. Finally, we represent this accelerated path integral method as a recurrent network, which is the accelerated version of the previously proposed path integral networks (PI-Net). We can train the accelerated PI-Net more efficiently for inverse optimal control with less RAM than the original PI-Net. Masashi Okada, Tadahiro Taniguchi |
ICRA | 2 |
| 2018 | Learning Oscillator-Based Gait Controller for String-Form Soft Robots Using Parameter-Exploring Policy GradientsabstractThis paper presents a methodology to design mechanosensor feedback to oscillator-based controller for worm-like soft-bodied robots. A reinforcement learning technique, i.e., PEPG, is employed to embed appropriate mechanosensor feedback to harness global entrainment among the controller, the body dynamics, and the environment without explicitly designing the interaction between the oscillators. Another reinforcement learning, actor-critic, was applied to train the controller for the simulation models to analyze the effectiveness of PEPG in the system. Furthermore, the gait controller was trained under different body dynamics, i.e., the physical model of a caterpillar and an earthworm. We found that PEPG is suitable for the system probably because it does not add exploration noise to actions and it conducts episode based parameter updates. The simulation results show the proposed method can acquire distinct behavior, i.e., caterpillars' crawling, inching and earthworms' crawling, under different body dynamics. The outcome implies, that by utilizing appropriate learning method, desired functionality can be achieved in soft-bodied robots without explicitly designing their behavior. Matthew Ishige, Takuya Umedachi, Tadahiro Taniguchi, Yoshihiro Kawahara |
IROS | 3 |
| 2017 | Online spatial concept and lexical acquisition with simultaneous localization and mappingabstractIn this paper, we propose an online learning algorithm based on a Rao-Blackwellized particle filter for spatial concept acquisition and mapping. We have proposed a nonparametric Bayesian spatial concept acquisition model (SpCoA). We propose a novel method (SpCoSLAM) integrating SpCoA and FastSLAM in the theoretical framework of the Bayesian generative model. The proposed method can simultaneously learn place categories and lexicons while incrementally generating an environmental map. Furthermore, the proposed method has scene image features and a language model added to SpCoA. In the experiments, we tested online learning of spatial concepts and environmental maps in a novel environment of which the robot did not have a map. Then, we evaluated the results of online learning of spatial concepts and lexical acquisition. The experimental results demonstrated that the robot was able to more accurately learn the relationships between words and the place in the environmental map incrementally by using the proposed method. Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi, Tetsunari Inamura |
IROS | 3 |
| 2017 | Visualization of Driving Behavior Based on Hidden Feature Extraction by Using Deep LearningabstractIn this paper, we propose a visualization method for driving behavior that helps people to recognize distinctive driving behavior patterns in continuous driving behavior data. Driving behavior can be measured using various types of sensors connected to a control area network. The measured multi-dimensional time series data are called driving behavior data. In many cases, each dimension of the time series data is not independent of each other in a statistical sense. For example, accelerator opening rate and longitudinal acceleration are mutually dependent. We hypothesize that only a small number of hidden features that are essential for driving behavior are generating the multivariate driving behavior data. Thus, extracting essential hidden features from measured redundant driving behavior data is a problem to be solved to develop an effective visualization method for driving behavior. In this paper, we propose using deep sparse autoencoder (DSAE) to extract hidden features for visualization of driving behavior. Based on the DSAE, we propose a visualization method called adriving color mapby mapping the extracted 3-D hidden feature to the red green blue (RGB) color space. A driving color map is produced by placing the colors in the corresponding positions on the map. The subjective experiment shows that feature extraction method based on the DSAE is effective for visualization. In addition, its performance is also evaluated numerically by using pattern recognition method. We also provide examples of applications that use driving color maps in practical problems. In summary, it is shown the driving color map based on DSAE facilitates better visualization of driving behavior. Hailong Liu 0001, Tadahiro Taniguchi, Yusuke Tanaka 0003, Kazuhito Takenaka, Takashi Bando |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Driving word2vec: Distributed semantic vector representation for symbolized naturalistic driving dataabstractThis study describes driving word2vec (DW2V), a new method for forming semantic representations of naturalistic driving data (NDD). To use big NDD for developing driver assistance systems or other information services, it is important to compress large amounts of data into an abstract and compact representation without losing semantic information. For this purpose, this study uses a symbolization method using a double articulation analyzer (DAA) assuming that NDD and human speech signals share an analogous structure, called a double articulation structure. The DAA can encode driving behavior data into sequences of driving words. However, the amount of semantic information contained in these sequences has not been clarified. Very few attempts have been made to develop a method for obtaining an adequate semantic representation of driving words that explains the relationship between different driving words. DW2V uses word2vec, proposed by Mikolov et al., to make a system learn the distributed semantic vector representation of symbolized naturalistic driving data (SNDD). Through experiments, we show that DW2V can restore the semantic relationships between different driving scenes from only a set of sequences of driving words, i.e., SNDD. In addition to quantitative analysis, a qualitative analysis of DW2V and its potential applications are discussed. Yusuke Fuchida, Tadahiro Taniguchi, Toshiaki Takano, Takuma Mori, Kazuhito Takenaka, Takashi Bando |
Intelligent Vehicles Symposium | 2 |
| 2016 | Determining Utterance Timing of a Driving Agent With Double Articulation AnalyzerabstractIn-vehicle speech-based interaction between a driver and a driving agent should be performed without affecting the driving behavior. A driving agent provides information to the driver and helps his/her driving behavior and non-driving-related tasks, e.g., selecting music and giving weather information. In this paper, we focus on a method for determining utterance timings when a driving agent provides non-driving-related information. If a driving agent provides a driver with non-driving-related information at an inappropriate moment, it will distract his/her driving behavior and deteriorate his/her safety driving. To solve or to mitigate the problem, we propose a novel method for determining the utterance timing of a driving agent on the basis of a double articulation analyzer, which is an unsupervised nonparametric Bayesian machine learning method for detecting contextual change points. To verify the effectiveness of the method, we conduct two experiments. One is an experiment on a short circuit around a park in an urban area, and the other is an experiment on a long course in a town. The results show that the proposed method enables a driving agent to avoid inappropriate timing better than baseline methods. Tadahiro Taniguchi, Kai Furusawa, Hailong Liu 0001, Yusuke Tanaka 0003, Kazuhito Takenaka, Takashi Bando |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Sequence Prediction of Driving Behavior Using Double Articulation AnalyzerabstractA sequence prediction method for driving behavior data is proposed in this paper. The proposed method can predict a longer latent state sequence of driving behavior data than conventional sequence prediction methods. The proposed method is derived by focusing on the double articulation structure latently embedded in driving behavior data. The double articulation structure is a two-layer hierarchical structure originally found in spoken language, i.e., a sentence is a sequence of words and a word is a sequence of letters. Analogously, we assume that driving behavior data comprise a sequence of driving words and a driving word is a sequence of driving letters. The sequence prediction method is obtained by extending a nonparametric Bayesian unsupervised morphological analyzer using a nested Pitman-Yor language model (NPYLM), which was originally proposed in the natural language processing field. This extension allows the proposed method to analyze incomplete sequences of latent states of driving behavior and to predict subsequent latent states on the basis of a maximum a posteriori criterion. The extension requires a marginalization technique over an infinite number of possible driving words. We derived such a technique on the basis of several characteristics of the NPYLM. We evaluated this proposed sequence prediction method using three types of data: 1) synthetic data; 2) data from test drives around a driving course at a factory; and 3) data from drives on a public thoroughfare. The results showed that the proposed method made better long-term predictions than did the previous methods. Tadahiro Taniguchi, Shogo Nagasaka, Kentarou Hitomi, Naiwala P. Chandrasiri, Takashi Bando, Kazuhito Takenaka |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2015 | Statistical localization exploiting convolutional neural network for an autonomous vehicleabstractIn this paper, we propose a self-localization method that exploits object recognition results by using convolutional neural networks (CNNs) for autonomous vehicles. Monte-Carlo localization (MCL) is one of the most popular localization methods that use odometry and distance sensor data for determining vehicle position. Some errors are often observed in the localization tasks and MCL often suffers from global positional errors. A global positional error means that particles representing a vehicle's position are distributed in the form of a multimodal distribution, i.e., the distribution has several peaks. To overcome this problem, we propose a method in which an autonomous vehicle employs object recognition results, obtained using CNNs, as the measurement data with a Bag-of-Features representation in an integrative manner. The semantic information found in the recognition results obtained using the CNN reduces the global errors in localization. The experimental results show that the proposed method can converge the distribution of the vehicle positions and particle orientations and reduce the global positional errors. Satoshi Ishibushi, Akira Taniguchi, Toshiaki Takano, Yoshinobu Hagiwara, Tadahiro Taniguchi |
IECON | 5 |
| 2015 | Automatic generation of summarized driving video with music and captionsabstractThis paper provides a novel summarization method for driving data, including a driving video captured by car-mounted camera, vehicle behavior such as velocity, steering angle, and other information. While a large amount of the driving data have been gathered recently, looking back method, however, has not been well-considered. In this paper, we integrate various information from the driving data as a summarized video with a time stretched video based on the driving behavior, text annotations describing salient events and geo-location, and background music adapted with the driving behavior. This representation is generated automatically from driving data with symbolization process of driving behavior based on nonparametric Bayesian approach. Through subjective evaluations, we evaluated the efficiency of proposed method for understanding the driving data appropriately in short time. Kazuhito Takenaka, Takashi Bando, Tadahiro Taniguchi |
IECON | 3 |
| 2015 | Essential feature extraction of driving behavior using a deep learning methodabstractDriving behavior can be represented by many different types of measured sensor information obtained through a control area network. We assume that the measured sensor information is generated from several hidden time-series data through multiple nonlinear transformations. These hidden time-series data are statistically independent of each other and capture essential driving behavior. Driving behavior information is usually generated by multiple nonlinear transformations that fuse essential features, e.g., "Yaw rate" is generated by fusing the velocity of the vehicle and the change of driving direction. However, driving behavior data is often redundant because such data includes multivariate information and involves duplicated essential features. In this paper, we propose a feature extraction method to extract essential features from redundant driving behavior data using a deep sparse autoencoder (DSAE), which is a deep learning method. Two-dimensional features are extracted from seven-dimensional artificial data using a DSAE and are determined experimentally to be highly correlated with the prepared essential features. DSAEs are also used to extract features from an actual driving behavior data set. To verify a DSAE's ability to extract essential driving behavior features and filter out redundant information, we prepare twelve data sets that include some or all of the driving behavior information. Twelve DSAEs are used to independently extract features from the twelve prepared data sets, and canonical correlation analysis is used to analyze the canonical correlation coefficients between extracted features. Furthermore, we verify DSAEs' ability to extract essential driving behavior features from the redundant driving behavior data sets. Hailong Liu 0001, Tadahiro Taniguchi, Yusuke Tanaka 0003, Kazuhito Takenaka, Takashi Bando |
Intelligent Vehicles Symposium | 2 |
| 2015 | Automatic lane change extraction based on temporal patterns of symbolized driving behavioral dataabstractThis paper proposes a method of automatically extracting lane change situations from large-scale driving corpora. Naturalistic driving data stored in large-scale corpora has a potential of contributing for developing novel advanced driver-assistance systems based on estimated information about driver's intent and/or potential risk of accidents. However, direct estimation of such kind of information from stream data is difficult. To address the issue, we apply an unsupervised symbolization method and topic representation to driving data. Driving stream data is converted to sequences of discrete symbols by a non-parametric symbolization method, and then the symbols are characterized by topics which represent typical distribution of driving behavior observed during the symbols. Because these symbols are separated on changing points of driving behavior, similar driving situations are effectively retrieved from sequences of the symbols. For evaluating effectiveness of the symbolization approach, we extract lane change situations based on the topic proportions and their temporal patterns. Distinctive elements of topic proportions and their temporal patterns for lane change situations are extracted by AdaBoost classifier. As a result, proposed approach outperforms baselines with neither topic proportions nor their temporal patterns in terms of extracting lane change situations. This result shows effectiveness of symbols with topic proportions for representing characteristics of driving situations. Masataka Mori, Kazuhito Takenaka, Takashi Bando, Tadahiro Taniguchi, Chiyomi Miyajima, Kazuya Takeda |
Intelligent Vehicles Symposium | 4 |
| 2015 | Unsupervised Hierarchical Modeling of Driving Behavior and Prediction of Contextual Changing PointsabstractAn unsupervised learning method, called double articulation analyzer with temporal prediction (DAA-TP), is proposed on the basis of the original DAA model. The method will enable future advanced driving assistance systems to determine driving context and predict possible scenarios of driving behavior by segmenting and modeling incoming driving-behavior time series data. In previous studies, we applied the DAA model to driving-behavior data and argued that contextual changing points can be estimated as changing points of chunks. A sequence prediction method, which predicts the next hidden state sequence, was also proposed in a previous study. However, the original DAA model does not model the duration of chunks of driving behavior and is not able to do a temporal prediction of the scenarios. Our DAA-TP method explicitly models the duration of chunks of driving behavior on the assumption that driving-behavior data have a two-layered hierarchical structure, i.e., double articulation structure. For this purpose, the hierarchical Dirichlet process hidden semi-Markov model is used for explicitly modeling the duration of segments of driving-behavior data. A Poisson distribution is also used to model the duration distribution of driving-behavior segments. The duration distribution of chunks of driving-behavior data is also theoretically calculated using the reproductive property of the Poisson distribution. We also propose a calculation method for obtaining the probability distribution of the remaining duration of current driving words as a mixture of Poisson distribution with a theoretical approximation for unobserved driving words. This method can calculate the posterior probability distribution of the next termination time of chunks by explicitly modeling all probable chunking results for observed data. The DAA-TP was applied to a synthetic data set having a double articulation structure to evaluate its model consistency. To evaluate the effectiveness of DAA-TP, we applied it to a driving-behavior data set recorded at actual factory circuits. The DAA-TP could predict the next termination time of chunks more accurately than the compared methods. We also report the qualitative results for understanding the potential capability of DAA-TP. Tadahiro Taniguchi, Shogo Nagasaka, Kentarou Hitomi, Kazuhito Takenaka, Takashi Bando |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | Mutual learning of an object concept and language model based on MLDA and NPYLMabstractHumans develop their concept of an object by classifying it into a category, and acquire language by interacting with others at the same time. Thus, the meaning of a word can be learnt by connecting the recognized word and concept. We consider such an ability to be important in allowing robots to flexibly develop their knowledge of language and concepts. Accordingly, we propose a method that enables robots to acquire such knowledge. The object concept is formed by classifying multimodal information acquired from objects, and the language model is acquired from human speech describing object features. We propose a stochastic model of language and concepts, and knowledge is learnt by estimating the model parameters. The important point is that language and concepts are interdependent. There is a high probability that the same words will be uttered to objects in the same category. Similarly, objects to which the same words are uttered are highly likely to have the same features. Using this relation, the accuracy of both speech recognition and object classification can be improved by the proposed method. However, it is difficult to directly estimate the parameters of the proposed model, because there are many parameters that are required. Therefore, we approximate the proposed model, and estimate its parameters using a nested Pitman-Yor language model and multimodal latent Dirichlet allocation to acquire the language and concept, respectively. Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi |
IROS | 5 |
| 2014 | Generating contextual description from driving behavioral dataabstractThis paper presents an automatic translation method from time-series driving behavior into natural language with contextual information. Nowadays, various advanced driver-assistance systems (ADASs) have been developed to reduce the number of traffic accidents and multiple ADASs are required to reduce further accidents. For such multiple ADASs, considering the context of driving and selecting appropriate assistance is key because the systems have to handle extremely complicated driving situations consisting of drivers (and their intents and maneuvers), environments (including other traffic participants such as vehicles and pedestrians), and vehicles dynamics. In this paper, time-series driving behavior is segmented into typical driving situation symbols, and the natural language expression of each situation is generated via the behavioral feature distribution observed in each situation. Owing to the symbolization of the driving behavior, the generated behavioral descriptions can be associated with their causes and results not on an actual time axis but on a situation-symbol axis as contextual descriptions, e.g., “letting up on gas pedal to pass tollgate.” The effectiveness of the proposed method was evaluated by using an actual data set of more than eight hours over a distance of 300 km in total. Although contextual expressions are very diverse even among human drivers, the proposed method obtained an agreement of more than 70%. Takashi Bando, Kazuhito Takenaka, Shogo Nagasaka, Tadahiro Taniguchi |
Intelligent Vehicles Symposium | 4 |
| 2014 | Visualization of driving behavior using deep sparse autoencoderabstractDriving behavioral data is too high-dimensional for people to review their driving behavior. It includes accelerator opening rate, steering angle, brake Master-Cylinder pressure and other various information. The high-dimensional data is not very intuitive for drivers to understand their driving behavior when they take a look back on their recorded driving behavior. We used a deep sparse autoencoder to extract the low-dimensional high-level representation from high-dimensional raw driving behavioral data obtained from a control area network. Based on this low-dimensional representation, we propose two visualization methods called Driving Cube and Driving Color Map. Driving Cube is a cubic representation displaying extracted three-dimensional features. Driving Color Map is a colored trajectory shown on a road map representing the extracted features. The trajectory is colored using the RGB color space, which corresponds to the extracted three-dimensional features. To evaluate the proposed method for extracting low-dimensional feature, we conducted an experiment and found several differences with recorded driving behavior by viewing the visualized Driving Color Map and that our visualization methods can help people to recognize different driving behavior. To evaluate the effectiveness of low-dimensional representation, we compared deep sparse autoencoder with other conventional methods from the viewpoint of linear separability of elemental driving behavior. As a result, our methods outperformed other conventional methods. Hailong Liu 0001, Tadahiro Taniguchi, Toshiaki Takano, Yusuke Tanaka 0003, Kazuhito Takenaka, Takashi Bando |
Intelligent Vehicles Symposium | 2 |
| 2014 | Prediction of Next Contextual Changing Point of Driving Behavior Using Unsupervised Bayesian Double Articulation AnalyzerabstractFuture advanced driver assistance systems (ADASs) should observe a driving behavior and detect contextual changing points of driving behaviors. In this paper, we propose a novel method for predicting the next contextual changing point of driving behavior on the basis of a Bayesian double articulation analyzer. To develop the method, we extended a previously proposed semiotic predictor using an unsupervised double articulation analyzer that can extract a two-layered hierarchical structure from driving-behavior data. We employ the hierarchical Dirichlet process hidden semi-Markov model [4] to model duration time of a segment of driving behavior explicitly instead of the sticky hierarchical Dirichlet process hidden Markov model (HDP-HMM) employed in the previous model [13]. Then, to recover the hierarchical structure of contextual driving behavior as a sequence of chunks, we use the Nested Pitman-Yor Language model [6], which can extract latent words from sequences of latent letters. On the basis of the extension, we develop a method for calculating posterior probability distribution of the next contextual changing point by marginalizing potentially possible results of the chunking method and potentially successive words theoretically. To evaluate the proposed method, we applied the method to synthetic data and driving behavior data that was recorded in a real environment. The results showed that the proposed method can predict the next contextual changing point more accurately and in a longer-term manner than the compared methods: linear regression and Recurrent Neural Networks, which were trained through a supervised learning scheme. Shogo Nagasaka, Tadahiro Taniguchi, Kentarou Hitomi, Kazuhito Takenaka, Takashi Bando |
Intelligent Vehicles Symposium | 2 |
| 2013 | Automatic drive annotation via multimodal latent topic modelabstractTime-series driving behavioral data and image sequences captured with car-mounted video cameras can be annotated automatically in natural language, for example, “in a traffic jam,” “leading vehicle is a truck,” or “there are three and more lanes.” Various driving support systems nowadays have been developed for safe and comfortable driving. To develop more effective driving assist systems, abstractive recognition of driving situation performed just like a human driver is important in order to achieve fully cooperative driving between the driver and vehicle. To achieve human-like annotation of driving behavioral data and image sequences, we first divided continuous driving behavioral data into discrete symbols that represent driving situations. Then, using multimodal latent Dirichlet allocation, latent driving topics laid on each driving situation were estimated as a relation model among driving behavioral data, image sequences, and human-annotated tags. Finally, automatic annotation of the behavioral data and image sequences can be achieved by calculating the predictive distribution of the annotations via estimated latent-driving topics. The proposed method intuitively annotated more than 50,000 pieces of frame data, including urban road and expressway data. The effectiveness of the estimated drive topics was also evaluated by analyzing the performances of driving-situation classification. The topics represented the drive context efficiently, i.e., the drive topics lead to a 95% lower-dimensional feature space and 6% higher accuracy compared with a high-dimensional raw-feature space. Moreover, the drive topics achieved performance almost equivalent performance to human annotators, especially in classifying traffic jams and the number of lanes. Takashi Bando, Kazuhito Takenaka, Shogo Nagasaka, Tadahiro Taniguchi |
IROS | 4 |
| 2013 | Multimodal concept and word learning using phoneme sequences with errorsabstractIn this study, we propose a method for concept formation and word acquisition for robots. The proposed method is based on multimodal latent Dirichlet allocation (MLDA) and the nested Pitman-Yor language model (NPYLM). A robot obtains haptic, visual, and auditory information by grasping, observing, and shaking an object. At the same time, a user teaches object features to the robot through speech, which is recognized using only acoustic models and transformed into phoneme sequences. As the robot is supposed to have no language model in advance, the recognized phoneme sequences include many phoneme recognition errors. Moreover, the recognized phoneme sequences with errors are segmented into words in an unsupervised manner; however, not all words are necessarily segmented correctly. The words including these errors have a negative effect on the learning of word meanings. To overcome this problem, we propose a method to improve unsupervised word segmentation and to reduce phoneme recognition errors by using multimodal object concepts. In the proposed method, object concepts are used to enhance the accuracy of word segmentation, reduce phoneme recognition errors, and correct words so as to improve the categorization accuracy. We experimentally demonstrate that the proposed method can improve the accuracy of word segmentation and reduce the phoneme recognition error and that the obtained words enhance the categorization accuracy. Tomoaki Nakamura, Takaya Araki, Takayuki Nagai, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi |
IROS | 5 |
| 2013 | Unsupervised drive topic finding from driving behavioral dataabstractContinuous driving-behavioral data can be converted automatically into sequences of “drive topics” in natural language; for example, “gas pedal operating,” “high-speed cruise,” then “stopping and standing still with brakes on.” In regard to developing advanced driver-assistance systems (ADASs), various methods for recognizing driver behavior have been proposed. Most of these methods employ a supervised approach based on human tags. Unfortunately, preparing complete annotations is practically impossible with massive driving-behavioral data because of the great variety of driving scenes. To overcome that difficulty, in this study, a double articulation analyzer (DAA) is used to segment continuous driving-behavioral data into sequences of discrete driving scenes. Thereafter, latent Dirichlet allocation (LDA) is used for clustering the driving scenes into a small number of so-called “drive topics” according to emergence frequency of physical features observed in the scenes. Because both DAA and LDA are unsupervised methods, they achieve data-driven scene segmentation and drive topic estimation without human tags. Labels of the extracted drive topics are also determined automatically by using distributions of the physical behavioral features included in each drive topic. The proposed framework therefore translates the output of sensors monitoring the driver and the driving environment into natural language. Efficiency of proposed method is evaluated by using a massive data set of driving behavior, including 90 drives for more than 78 hours over 3700km in total. Takashi Bando, Kazuhito Takenaka, Shogo Nagasaka, Tadahiro Taniguchi |
Intelligent Vehicles Symposium | 4 |
| 2012 | Online learning of concepts and words using multimodal LDA and hierarchical Pitman-Yor Language ModelabstractIn this paper, we propose an online algorithm for multimodal categorization based on the autonomously acquired multimodal information and partial words given by human users. For multimodal concept formation, multimodal latent Dirichlet allocation (MLDA) using Gibbs sampling is extended to an online version. We introduce a particle filter, which significantly improve the performance of the online MLDA, to keep tracking good models among various models with different parameters. We also introduce an unsupervised word segmentation method based on hierarchical Pitman-Yor Language Model (HPYLM). Since the HPYLM requires no predefined lexicon, we can make the robot system that learns concepts and words in completely unsupervised manner. The proposed algorithms are implemented on a real robot and tested using real everyday objects to show the validity of the proposed system. Takaya Araki, Tomoaki Nakamura, Takayuki Nagai, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi |
IROS | 5 |
| 2012 | Contextual scene segmentation of driving behavior based on double articulation analyzerabstractVarious advanced driver assistance systems (ADASs) have recently been developed, such as Adaptive Cruise Control and Precrash Safety System. However, most ADASs can operate in only some driving situations because of the difficulty of recognizing contextual information. For closer cooperation between a driver and vehicle, the vehicle should recognize a wider range of situations, similar to that recognized by the driver, and assist the driver with appropriate timing. In this paper, we assumed a double articulation structure in driving behavior data and segmented driving behavior into meaningful chunks for driving scene recognition in a similar manner to natural language processing (NLP). A double articulation analyzer translated the driving behavior into meaningless manemes, which are the smallest units of the driving behavior just like phonemes in NLP, and from them it constructed navemes, which are meaningful chunks of driving behavior just like morphemes. As a result of this two-phase analysis, we found that driving chunks equivalent to language words were closer to the complicated or contextual driving scene segmentation produced by human recognition. Kazuhito Takenaka, Takashi Bando, Shogo Nagasaka, Tadahiro Taniguchi, Kentarou Hitomi |
IROS | 4 |
| 2012 | Semiotic prediction of driving behavior using unsupervised double articulation analyzerabstractIn this paper, we propose a novel semiotic prediction method for driving behavior based on double articulation structure. It has been reported that predicting driving behavior from its multivariate time series behavior data by using machine learning methods, e.g., hybrid dynamical system, hidden Markov model and Gaussian mixture model, is difficult because a driver's behavior is affected by various contextual information. To overcome this problem, we assume that contextual information has a double articulation structure and develop a novel semiotic prediction method by extending nonparametric Bayesian unsupervised morphological analyzer. Effectiveness of our prediction method was evaluated using synthetic data and real driving data. In these experiments, the proposed method achieved long-term prediction 2-6 times longer than some conventional methods. Tadahiro Taniguchi, Shogo Nagasaka, Kentarou Hitomi, Naiwala P. Chandrasiri, Takashi Bando |
Intelligent Vehicles Symposium | 1 |
| 2012 | Drive video summarization based on double articulation structure of driving behaviorabstractThis paper provides a novel summarization method for drive videos using driving behavior, such as driver maneuvers and vehicle reaction, recorded simultaneously alongside video. We segmented the driving behavior into chunks via an unsupervised manner and summarized the drive videos using the chunks, i.e., the switching points of the chunks were emphasized and the middle of the chunks were compressed. As the result of subjective evaluation, we found that the chunks were more consistent with human-recognized driving context than the image-based method and that the summarized video was more suitable for reviewing entire driving scenes, i.e., our method achieved an efficient summarization. Kazuhito Takenaka, Takashi Bando, Shogo Nagasaka, Tadahiro Taniguchi |
ACM Multimedia | 4 |
| 2008 | Imitation Learning from Unsegmented Human Motion Using Switching Autoregressive Model and Singular Vector Decomposition
Tadahiro Taniguchi, Naoto Iwahashi |
ICONIP (1) | 1 |
| 2008 | Incremental acquisition of multiple nonlinear forward models based on differentiation process of schema model
Tadahiro Taniguchi, Tetsuo Sawaragi |
Neural Networks | 1 |