Takeshi Shibuya

dblp:45/2208 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0003-4645-5898ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Understanding Behavioral Differences Between Machine Agents and Human Participants Based on How They Play the Energy Transition Game
Kengo Suzuki, Yuta Nakadegawa, Kento Miura, Takeshi Shibuya, Susumu Ohnuma
ISAGA4
2024 Action Robust Reinforcement Learning with Highly Expressive Policy
abstract
In traditional reinforcement learning, there can be a degradation in the control performance of the policy when the environmental parameters differ between the training and application phase. The policy that minimizes this degradation is referred to as a robust policy. A framework called Noisy action Robust Markov Decision Process (NR-MDP) was proposed for training robust policies, and the Action Robust Deep Determin-istic Policy Gradient (AR-DDPG) algorithm was introduced as a method for solving NR-MDP. The optimal policy in NR-MDP includes policies following various probability distributions, whereas AR-DDPG is restricted to deterministic policies. We propose a new robust reinforcement learning method called Action Robust Q-Learning (AR-QL) that enables the training of optimal policies in NR-MDP by leveraging various sampling techniques to extend the representational capacity of policies, targeting an improvement in policy robustness. To validate this, we confirmed that AR-QL can acquire the optimal policy for a simple NR-MDP problem, for which AR-DDPG fails to obtain the optimal policy. Furthermore, we confirmed that the robust performance of policy trained by AR-QL in the OpenAI's InvertedPendulum environment surpasses that of policy trained by AR-DDPG.
Seong-in Kim, Takeshi Shibuya
SMC2
2024 Explainable Reinforcement Learning via Causal Model Considering Agent's Intention
abstract
Explaining agent's decision can offer valuable insights for designers and end-users. One proposed method for explaining an agent's decision-making involves representing a causal relation between state components and action as a causal model and providing explanations for the decisions made using causal model. However, traditional causal model often faces structural limitation, restricting the range of representable control problems. Additionally, providing accurate explanation becomes challenging in control problems with various types of rewards because agent's intention of an action is unknown. In this study, we introduce a causal model capable of representing a broader range of control problems and a method to provide accurate explanations in control problems with various types of reward structures. Through redefining the relationships between nodes in the causal model, we have enabled a broader representation of control problems. Also, by incorporating agent's intention into the explanation, we have achieved to provide a more precise explanation. To validate the effectiveness of our proposed method, we conducted experiments using OpenAI's LunarLander environment. Using a proposed causal model, we defined the causal model of LunarLander, which could not be represented by conventional causal models. Furthermore, by incorporating the intentions of an agent into the explanation, novel interpretations previously inaccessible have become feasible.
Seong-in Kim, Takeshi Shibuya
SMC2
2022 Reinforcement Learning to Efficiently Recover Control Performance of Robots Using Imitation Learning After Failure
abstract
Extreme environments, such as space and underwater, are difficult for humans to enter because they involve risks, hence it is necessary to employ autonomous robots instead of humans. When robots fail in extreme environments, it is essential for the robot to automatically recover control following the rules of failure because humans cannot repair the robot directly. Reinforcement learning is expected to automatically acquire the control rule; however, the retrieval of the control rule requires significant trial-and-error. Imitation learning cannot acquire the control rule if no suitable expert data exist. Methods combining imitation learning and reinforcement learning reduce the number of trial-and-errors; however, they are still not effective against robot failure because these methods cannot utilize expert data effectively. This paper proposes a reinforcement learning method that efficiently recovers control performance from failures by utilizing both the control rules prepared by the designer and multiple discriminators to calculate measures similar to expert data. Experimental results show that the proposed method recovers the control performance with fewer episodes than the conventional method. The main contribution of the proposed method is its efficiency against robot failure through utilizing expart data prepared for failure by the designers for imitation learning.
Shoki Kobayashi, Takeshi Shibuya
SMC2
2020 Topological Visualization Method for Understanding the Landscape of Value Functions and Structure of the State Space in Reinforcement Learning
Yuki Nakamura, Takeshi Shibuya
ICAART (2)2
2020 Inferring Underlying Manifold of Low Density Data using Adaptive Interpolation
Noritaka Yamada, Takeshi Shibuya
ICAART (2)2
2020 Reinforcement Learning Compensator Robust to the Time Constants of First Order Delay Elements
abstract
Reinforcement learning is a learning paradigm in which a control is learned automatically based on rewards through trial and error based on rewards. When reinforcement learning is employed for robot control, the action that is output by reinforcement learning and the input of the actuator are often the same. A robot's actuator has a time constant of a first-order delay element between input and output. Delays result in the deterioration of the reinforcement learning performance because the environments that contain them lack the Markov property. Although there have been studies of such environments, they are problematic in that performance deteriorates when the time constant of a first-order time-delay element greater than the control cycle. The principal contribution of this paper is to propose a compensator for reinforcement learning that is more effective than conventional methods for environments with a time constant of a first-order time-delay element greater than the control cycle. The purpose of the compensator is to minimize the difference between actions in delayed environments and those not in delayed environments. Experiments reveal that the compensator increases rewards within wider ranges than conventional methods.
Shoki Kobayashi, Takeshi Shibuya
SMC2
2011 Reinforcement learning with nonstationary reward depending on the episode
abstract
A model which represents nonstationary reward is proposed for reinforcement learning(RL). RL is a framework that the agent learns by the interaction with an environment. The agent receives the reward, and learns its behavior. The reward is determined by the designer. It is not necessary to design the behavior of the agent so that RL is expected to be applied to various applications. However, conventional RL algorithms work under the assumption that the environment is stationary. In other words, conventional RL can not accept unstationary rewards and the change of the objective. From the point of view of real world applications, it is necessary for the agent to deal with a change of the objective. In this paper, a learning technique to deal with a temporal change of the reward is proposed. In the proposed reward representation, the reward is divided into two parts: episode-dependent part and episode-independent part. The simulation experiments show the effectiveness of the proposed method.
Takeshi Shibuya, Seiji Yasunobu
SMC1
2010 Reinforcement learning in continuous state space with perceptual aliasing by using complex-valued RBF network
abstract
Reinforcement learning for continuous state space with perceptual aliasing is proposed. Complex-valued reinforcement learning is effective for perceptual aliasing. In continuous state space, the conventional complex-valued reinforcement learning demands the discretization of continuous state. However, it is difficult to discretize continuous state suitably. In this paper, complex-valued reinforcement learning using complex-valued RBF network is proposed. An experiment shows that proposed method is effective for continuous state space with perceptual aliasing.
Takeshi Shibuya, Hideaki Arita, Tomoki Hamagami
SMC1
2010 Complex-valued reinforcement learning with hierarchical architecture
abstract
Hierarchical complex-valued reinforcement learning is proposed in order to solve the perceptual aliasing problem. The perceptual aliasing problem is encountered when an incomplete set of sensors is used in an actual environment, and this problem makes learning difficult for an agent. Hierarchical Q-learning (HQ-learning) and complex-valued reinforcement learning are proposed in order to solve this problem. HQ-learning is a hierarchical extension of Q-learning. In HQ-learning, tasks are divided into sequences of simpler sub-tasks that can be solved by adopting memory-less policies, but a considerable amount of time is required for learning. In complex-valued reinforcement learning, the dependence of contexts can be represented by using complex-valued action-value functions. It enables the agent to adaptively perform actions, but may not deal problems because of the cycle of perceptual aliasing. In this paper, complex-valued reinforcement learning based on HQ-learning with a hierarchical design is proposed. Experimental results show the effectiveness of the proposed method.
Atsuhiro Yamazaki, Tomoki Hamagami, Takeshi Shibuya
SMC3
2007 Experimental study of the eligibility traces in complex valued reinforcement learning
abstract
Effectiveness of eligibility traces in complex valued reinforcement learning is studied. Complex valued reinforcement learning is a new method inspired by complex valued nerual networks. In this study, it is desired that various approaches in the ordinally real valued reinforcement learning are applied to the complex valued reinforcement learning. This paper focuses attention on an experimental study of the eligibility traces. Simulation results infer that there is a possibility of overcoming tight perceptual aliasing with long trace back up.
Takeshi Shibuya, Shingo Shimada, Tomoki Hamagami
SMC1
2006 Complex-Valued Reinforcement Learning
abstract
A new reinforcement learning algorithm with complex-valued functions is proposed. The algorithm is inspired by complex-valued neural networks introducing complex numbers representing phase and amplitude into a conventional neural network. The strong advantage of using complex values in reinforcement learning is that the state-action function in a time series can be easily extended. In particular, considering the coherence of each complex value, the proposed learning algorithm can represent the context of agent behavior. This extension allows compensating for the perceptual aliasing problem and provides for the intelligent behavior of mobile robots in the real world. The complex-valued functions are applied to the conventional reinforcement learning algorithms: Q-learning and profit sharing. These algorithms are evaluated by simple maze problems and a bar-carrying task involving perceptual aliasings. Simulation experiments show that the new algorithm can efficiently solve perceptual aliasing.
Tomoki Hamagami, Takeshi Shibuya, Shingo Shimada
SMC2